May 19, 2011

$10K bioinformatics on thousand dollar genome

It is now possible to get 100x coverage of the exome sequence for a cancer sample (or any other type of human genomic sample) on one lane of an Illumina HiSeq machine. With the Sure Select 50 MB exome kit, it still costs quite a bit more than one thousand dollars to get this data, but it is getting close. At maximum yield, it might currently be possible to multiplex 4 samples into a singe lane and still get 100x coverage of each. This will certainly be true when planned upgrades to the HiSeq machine are available.

Illumina provides some nice software (called CASAVA) that is typically run at the default settings by Core labs and sequencing outsourcing companies. This software gives high-quality genome alignments and pretty good SNP calls - useful for many purposes. However, real-world research needs are often not satisfied with default automated bioinformatics analysis. Narrowing down hundreds of thousands of SNP calls to the few real disease-related mutations is difficult hands-on work for skilled bioinformaticians. Today in my lab group, we are fighting with false-negatives: SNPs that were present but not called in the germ line sample, leading to false identification of mutations unique to the tumor. It looks like we will have to re-run the SNP detection software many times with small changes in various parameters to optimize specificity vs. sensitivity in each sample. Investigators may sub-contract this type of work to the lab that does the sequencing, they may have skilled bioinformaticians in their lab group, or they may hire bioinformatics consultants. In any case, $1K of sequence data may cost more than $10K for analysis.

Feb 15, 2011

Archiving NGS data

Anyone who has worked with NextGen sequence data quickly gains an appreciation for the difficulties associated with long term data storage. The current 'state of the art,' at least for Illumina machines, involves saving some fairly raw data files such as fastq text to the NCBI Short Read Archive (SRA).


Our GAIIx is producing about 30 million reads per lane, which gives files of 8-10 GB (72 cycles) per lane in either qseq (completely unfiltered) or fastq (quality scored) format. If we max out two runs per week, that is about 140 GB of raw sequence data per Illumina machine per week.

There has been some recent discussion about the possibility of phasing out the SRA at NCBI.
[see this post which claims to be a memo from NCBI director David Lipman: "The Sequence Read Archive (SRA) will also be phased out over the next 12 months."]
If cost cutting is truly necessary for our national biomedical research infrastructure, I can see why the raw SRA data might be growing at an awkwardly rapid rate and have less value than the higly used databases of GenBank non-redundant nucleotide, GEO, etc.

I think that it is interesting to turn this discussion around and ask why are we archiving all of this raw sequence data? The trivial argument is that: Journals require open access to raw data as a condition of publication." But that argument ignores the more interesting question: What is the 'raw data' for a sequencing project? No one is loading Illumina (or SOLID or 454) image data into public archives. The impracticality of saving multiple terabytes of image data for each run made that approach moot a couple of years ago. We are saving raw qseq or fastq files right now because our methods for basecalling and SNP calling (and indel/translocation/copy number calling) are imprecise. I have seen data analysts go back into primary sequence reads for a single sample and find a SNP that was not called because a few reads had below threshold quality scores.

If we consider the actual "useful" data content of a NGS run on a single sample, the landscape looks quite different. ChIP-seq is our most common NGS application. The useful data from a ChIP-seq run is actually just a set of genome positions where read starts are mapped. At most, this is 20-30 million positions. In actuality, 30% of reads are not mapped, and another 10-50% are duplicates (multiple reads that map to the exact same position), so the final data set might be compressed to about 10 million genomic loci with a read count at each spot. After sorting and indexing, this information could be efficiently stored in a very compact file.

RNA sequencing is becoming increasingly popular. Our clients are typically not interested in the sequence data itself, only in gene expression counts - essentially the same data as produced by a microarray. However, there are some cool new applications that look at alternative splicing, so we may have to keep the actual sequence reads on hand for a while longer.

Human (and mouse) SNP/indel/cnv detection is another popular NGS application. We are only really interested in the variants. However, SNP calling software requires both numbers of reads with reference vs. variant bases and quality scores for each basecall. Some software also uses context dependent quality metrics, such as distance from other SNPs, distance from indels, etc. Given the highly diverse collection of existing SNP detection software, and the likelihood of new software development, it seems impossible to compress this class of data to a set of variant calls and discard the raw reads. This is very unfortunate, since typical variant detection projects use anything from 20x to 50x coverage of the genome. So we are storing 150 GB of raw sequence data in order to track a few million bytes worth of actual variation in the genome of each research sample.

Other applications, such as de novo genome sequencing of new organisms, or metagenomic sequencing of environmental or medical samples will not be easily compressed. Fortunately, these data are currently archived in places other than the SRA.




Jan 24, 2011

Genetic Disease Diagnostics with NGS

A couple of recent papers demonstrate a significant opportunity for the use of NextGen Sequencing in the diagnosis of genetic disease. Dennis Lo et al, at The Chinese University of Hong Kong, have published results for a NGS fetal genetic diagnostic test based on recovery of fragments of fetal DNA from the mother's blood plasma. Preliminary results show that complete coverage of the fetal diploid genome is possible [Science Translational Medicine] at a resolution that allows for differentiation of heterozygous vs. homozygous mutations in disease genes; and also that aneuploidy, such as trisomy 21 can be detected with high specificity and sensitivity [British Medical Journal]. The key benefit of this approach is that it can be done non-invasively from simple blood draw from the mother, so it avoids the relatively high incidence of pregnancy complications created by amniocentesis or chorionic villus sampling procedures.

Meanwhile, the lab of Stephen Kingsmore at the US National Center for Genome Resources reported results of a targeted sequencing carrier screen for a total of 448 severe (rare) recessive genetic diseases [Science Translational Medicine]. This work is particularly significant because the screen is designed to work in multiplex, allowing for a potential total cost per patient of below $500 (less than $1 per disease screened). While each gene is rare in isolation, the combined screen shows an average of 2.8 mutations per individual tested in the proof-of-concept phase of the study.

Taken together, these advances suggest that routine clinical applications of NGS will soon be practical, attractive, and economically feasible for large numbers of healthy people (pregnant women and marriage minded couples). This is great news for NGS equipment vendors, and also suggests a software engineering opportunity for the development of much more robust bioinformatics pipelines for processing this data and including it in electronic medical records. At the same time, I am worried that the lab folks may be progressing much more rapidly than the thinking in the ELSI community. What kind of databases will be created when every pregnancy and every marriage license is associated with gigabyte files of deep sequencing data? This issue is all the more problematic because disease carrier testing and Down syndrome screening are already so widely accepted. Changing prenatal tests to use sequencing in order to reduce complications in pregnancy, and adding pre-conception tests for diseases that were previously thought to be too rare to merit widespread screening are non-controversial medical advances. The downside might come from the unintentional discovery of other genetic information, the availability to law enforcement and other organizations of large files of genetic information on every person, etc.

Dec 6, 2010

"NGS without bioinformatics expertise"

This showed up in my inbox today:

Integromics launches SeqSolve, a Next Generation Sequencing functional ...
PharmaLive.com (press release)
SeqSolve is the first NGS analysis software on the market specifically developed for data interpretation without requiring bioinformatics expertise. ...

Might be a bit presumptious, but I would guess that a lot of scientists would like to have someone with "bioinformatics expertise" do data analysis for their Next-Gen Sequencing projects. But hey, if someone wants to spend $10-$50K on sequencing, but doesn't want an expert to look at the data, good luck with that. Our "Sequencing Informatics Group" at NYU is now taking outside consulting work for all types of NGS bioinformatics projects. It is sort of like fixing your Porche - you can go to the Foreign Car specialist mechanic, or you can go to PepBoys and buy some spark plugs and a wrench kit.

We (Ross Smith) have been developing our own visualization toolkit for NGS. Latest version allows us to integrate RNAseq and ChIPseq with RefSeq or other annotation. Net result is much more accurate assignment of TF or histone modification sites to genes, and the ability to clearly see which of multiple TSS are actually being used in a particular sample/cell type. It is very beautiful.


Sep 25, 2010

ChIPseq QC

I am speaking at the CHI Next-Gen Sequencing conf in Providence 9.26.2010 (Sunday short course). My topic is going to be about the role of bioinformatics in QC for NG sequencing, with examnples from ChIPseq, where I have the most experience.

My main point is that the informatics team works hardest on experiments that produce poor data - or where the data contradict the investigator's expectations. When the experiment is beautiful, then you can use your automated (or semi-automated) pipeline, and hand over the analyzed data with a standard report. For a transcription factor type ChIPseq, the standard result is a set of peaks with p-value and fold change vs. an input DNA, annotated by distance to the nearest gene's Transcription Start Site. If pressed, we can deliver this about 2 days after the sequencing run is complete. For an epigenomics type ChIPseq (histone methylation, acetylation, etc) we deliver both peaks vs input DNA and some type of fold-change for each peak comparing one biological condition vs. another.

However, we spend a lot more time squabbling about runs with high PCR duplication, weird artifacts, low yield, peaks in the input DNA lane, etc. To deal with this, we have been developing a variety of tools to quantify overall data quality in a ChIPseq run. We are looking as the overall clustering of mapped reads on the Reference genome (average spacing of adjacent/overlapping reads), as well as coverage at various depths. Some of these metrics make intersting graphs, but we have not completely pinned down their predictive power for understanding the data.

We have recently been playing with selecting sets of genes based on external data such as gene exprssion values from microarray or RNAseq experiments, and looking at the aggregate profile of reads mapped near the TSS of groups of genes that are upregulated, downregulated, unchanged, etc. By combining reads for a bunch of genes, we get smooter curves and you can actually say fairly clearly that upregualted genes have (or do NOT have) a change in histone methylation near the TSS as compared with downreg or unchanged genes.