Next-Gen Sequencing: Genome Coverage from BAM file

Jan 31, 2018

Genome Coverage from BAM file

There are many excellent tools for analysis of Next Gen Sequencing data in the standard BAM alignment format so I was surprised how difficult it was for me to get a nice graph of genome coverage. This will be trivial for a lot of hard core bioinformatics coders, so just move along if you are bored/annoyed.

I needed to check the evenness of coverage across intervals of a bacterial genome that we were re-sequencing for various experimental reasons. I aligned my FASTQ to a reference genome from GenBank using Bowtie2. There are several nice tools in the SAMTools and BEDTools kits that produce either a base by base coverage count or a histogram of coverage showing how many bases are covered how deeply. I wanted a map at 1 Kb resolution. It took a while to figure out that I first need to make a BED file of intervals from my genome - with correct names for the 'chromosomes' that match the SAM header, and then use 'samtools bedcov' to get my intervals. Then a simple graph from Excel or R shows the coverage per interval along the genome.

Here are the steps (as much for me to remember as for usefulness to anyone else)

1) Create a Bowtie2 index of the reference genome (can be from GenBank, or can be a de novo assembly of contigs created locally from this FASTQ data). Reference_in must be FASTA format.
bt2_base is the name you will call the index.

bowtie2-build

2) Align the FASTQ file(s) to the Reference. bt2-idx is the name of the index created in the previous step. There are a ton of options for stringency of alignment, format of input data, etc. I set this to use 16 CPU threads. I usually like to leave out the unaligned reads, which can reduce the output file size somewhat. [Of course, the unalinged reads are the goal when you use Bowtie as a filter to remove human from microbiome data or any other sort of contaminant screen.]

bowtie2 -p 16 --no-unal -x -1 -2 -S

3) Is is annoying that Bowtie produces output in SAM format, and 99% of the time, the very first thing you have to do is convert to sorted BAM. Note that samtools sort puts its own .bam on the end, so if your are not careful you will get files named output.bam.bam

samtools view -bS output.sam | samtools sort - file_sorted

4) Create a 'genome' file for your reference genome. This is just a tab delimited file that names the chromosomes (or individudal sequences that may be in a multi-FASTA file, such as contig names).
It is super easy to mess this up. The best way is to view the top of your output.sam file created by bowtie2. The lines that start with @ are your chromosome headers, and they very helpfully already show the length of each one. This is a bit of a pain if you have a genome with lots of contigs, but a little 'cut' and 'paste' in bash or Excel will get you there.

Here is what mine looked like:

@HD VN:1.0 SO:unsorted
@SQ SN:MKZW02000001.1 LN:5520555
@SQ SN:MKZW02000002.1 LN:248293
@PG ID:bowtie2 PN:bowtie2 VN:2.2.7 CL:"/local/apps/bowtie2/2.2.7/bowtie2-align-s --wrapper basic-0 -p 8 -x Kluy

and here is the genome file I made:

MKZW02000001.1 5520555
MKZW02000002.1 248293

5) Make a set of intervals with bedtools makewindows. I wanted 1 Kb intervals, so I use -w 1000.

The result is a simple BED file with one line for each 1Kb window of the genome.

bedtools makewindows -g genome.txt -w 1000 > genome_1k.bed

$ more genome_1k.bed

MKZW02000001.1 0 1000

MKZW02000001.1 1000 2000

MKZW02000001.1 2000 3000

MKZW02000001.1 3000 4000

6) Use samtools bedcov to count the total number of bases in the BAM file that are located in each of the intervals ('sum of per base read depths per BED region'). This works much faster than any other coverage tool that I have tested.

samtools bedcov genome_1k.bed kv_sorted.bam > kv_1k.cov

7) Divide the sum of coverage by the window size (/1000 in my case), and plot the average coverage per window as a scatter plot, using the end of each interval as the X axis and the coverage as the Y.

Histogram will only work nicely if you have very few intervals. In my case, high and low coverage outlier intervals are easily visible.

8 comments:

Tim said...: Seen deeptools bamCoverage? https://deeptools.readthedocs.io/en/latest/content/tools/bamCoverage.html; Jan 31, 2018, 9:09:00 PM
Unknown said...: With a grateful hearth , I want to give my sincere appreciation to Dr Sambo He is the best , he's so wonderful and helping. Few weeks ago , I contacted him for a lotto lucky number, without delay , he gave me the real lucky winning numbers I played and won $100,000,000. Contact him for lucky winning numbers and your story will change for good . Thank you Dr Sambo I will forever be grateful for your kind gesture. Contact him today through his email he can help you.
divinespellhome@gmail.com
WhatsApp him now +1(267)527-9481; Oct 14, 2018, 5:53:00 AM
Anacyte Laboratories said...: Very nice posting. Your article is quite informative. Thanks for the same. Our service also helps you to market your products with various marketing.

Click Here:- single cell rna sequencing; Jan 26, 2021, 12:09:00 AM
Donna said...: I am Doctor Paul I got affected with HIV in the process of attending to my HIV patient I tried all I can to get cured but all to no avail, until I saw a post in a health forum about a herbalist man who prepare herbal medication to cure all kind of diseases including HIV virus, at first I doubted if it was real but decided to give it a try, when I contact this herbalist via his email Blessedlovetemple@gmail.com and he prepared a HIV herbal cure and sent it to me via fed-ex delivery company service, when I received this herbal cure, he gave me step by directions on how to apply it, when I applied it as instructed, I was totally cured of this deadly disease within 5 days of usage, I am now free from the deadly disease called HIV, all thanks to Dr Mark. Contact this great herbal spell caster. Kindly contact him. Blessedlovetemple@gmail.com
He cures all kinds of sickness or diseases such as: 1. HERPES VIRUS 2. LASSA FEVER 3. GONORRHEA 4. HIV/AID 5. EX BACK.
Thanks Dr Mark for saving my life.; Apr 7, 2021, 4:18:00 PM
Ric Clayton said...: I really want to thank Dr Emu for saving my marriage. My wife really treated me badly and left home for almost 3 month this got me sick and confused. Then I told my friend about how my wife has changed towards me. Then she told me to contact Dr Emu that he will help me bring back my wife and change her back to a good woman. I never believed in all this but I gave it a try. Dr Emu casted a spell of return of love on her, and my wife came back home for forgiveness and today we are happy again. If you are going through any relationship stress or you want back your Ex or Divorce husband you can contact his whats app +2347012841542 or email emutemple@gmail.com website: Https://emutemple.wordpress.com/; Sep 2, 2021, 11:02:00 AM
Nadia Adams said...: This is such an informative post about genome coverage and BAM files! Your clear explanations make a complex topic much easier to understand. Just Amazing Discounts; Dec 3, 2024, 4:00:00 PM
Barbara Nimmo said...: As someone who appreciates high-quality resources, I’ve found Saving Cents Together to be helpful in finding tools and materials for scientific research at great prices. Thanks for sharing this detailed and educational guide!; Dec 3, 2024, 4:00:00 PM
URGENT LOAN said...: My name is Mrs Aisha Mohamed, am a Citizen Of Qatar.Have you been looking for a loan?Do you need an urgent personal loan or business loan?contact Adam Ibrahim Finance Home he help me with a loan of $80,000 some days ago after been scammed of $6,800 from a woman claiming to been a loan lender but i thank God today that i got my loan worth $80,000.Feel free to contact the company for a genuine financial service. Email:adamibrahimfinanceltd1976@gmail.com call/whats-App Contact Number +918119841594 Adam Ibrahim Finance Pvt Ltd; Mar 24, 2025, 9:07:00 AM

Next-Gen Sequencing

Jan 31, 2018

Genome Coverage from BAM file

8 comments:

Stuart Brown

Resources

Blog Archive

List of Blogs relevant to NG Seq

Popular Posts