Bioinformatics for the Clinical Lab
From raw reads to a report: FASTQ, FASTA, BAM, and VCF in order, Phred quality (Q30 = 99.9%), trimming, alignment, variant calling, annotation, and filtering; depth and variant allele fraction with a worked pileup; and the lab-operations side - genome builds, pipeline validation, and version control.
- 5 min
- 6 steps
- 4 questions
- Lesson 32 of 60
In this lesson
- The files, in order
- Phred quality
- The pipeline
- Depth and VAF
- Running a pipeline in a lab
- What to take from this
Picking up where you left off.
The files, in order
| File | Holds | Made by |
|---|---|---|
| FASTQ | raw reads, four lines each: name, bases, separator, a quality character per base | the sequencer (after demultiplexing) |
| FASTA | the reference genome sequence | a curated public reference |
| SAM / BAM | alignments: where each read maps and how well (BAM is the compressed binary form) | the aligner |
| VCF | variant calls: one per line with evidence and quality | the variant caller |
@RUN1:lane1:read00042
GATCACAGGTCTATCACCCTATTAACCAC
+
IIIIIIIIIFFFFFAAAA<<<<7777,,,
The fourth line encodes a quality score for each base: here high at the start and falling toward the end, as reads typically do.
Quick check
FASTQ = raw reads with qualities, FASTA = reference, BAM = alignments, VCF = variant calls.
Phred quality
Q = -10 × log₁₀(P), where P is the chance the base call is wrong 1:
| Q | Error rate | Accuracy |
|---|---|---|
| 10 | 1 in 10 | 90% |
| 20 | 1 in 100 | 99% |
| 30 | 1 in 1,000 | 99.9% |
| 40 | 1 in 10,000 | 99.99% |
Run QC reports the percentage of bases at Q30 or above; most runs aim for 80% or more.
Quick check
Q = -10 log10(P): Q20 = 1/100, Q30 = 1/1,000, Q40 = 1/10,000.
The pipeline
- QC and trimming: remove adapter sequence and low-quality tails; discard poor reads 1.
- Alignment: place each read at its best match on the reference, tolerating real variants and errors; output BAM. Mark PCR duplicates.
- Variant calling: at each position, decide statistically whether reads that disagree with the reference reflect a real variant - weighing depth, base and mapping quality, and strand balance 1. Separate callers handle small variants, copy number, and structural variants.
- Annotation: add gene, transcript, HGVS name, predicted effect, population frequency (gnomAD), and database entries (ClinVar, COSMIC) 2.
- Filtering and interpretation: drop low-quality, common, or irrelevant calls; a qualified analyst classifies the rest (ACMG/AMP for germline, AMP/ASCO/CAP tiers for somatic) and decides what is reported.
Depth and VAF
- Depth (coverage) at a position: how many reads span it after duplicate removal. Reports state the minimum depth and the fraction of target bases covered (for example, 99% at 100x or more).
- Variant allele fraction (VAF): variant reads ÷ total reads.
Interpretation depends on the sample:
| VAF | Germline sample | Tumor sample |
|---|---|---|
| about 50% | heterozygous | possibly germline - confirm in normal tissue |
| about 100% | homozygous or hemizygous | germline, or a variant with loss of the other allele |
| 10-40% | suspicious for mosaicism or artifact | typical somatic, scaled by tumor purity |
| 1-5% | usually artifact | subclone; needs deep coverage and often UMIs |
A tumor at 40% purity with a heterozygous clonal mutation would show a VAF near 20%. At low depth, VAF is meaningless: 1 of 4 reads is “25%” with huge uncertainty.
Quick check
25/500 = 0.05. In a tumor, that could be a subclone; in germline blood it would be suspicious.
Running a pipeline in a lab
- Reference build: GRCh37 (hg19) and GRCh38 (hg38) give the same variant different coordinates. Record the build, and name variants against versioned transcripts.
- Validation: a pipeline is validated like any assay - sensitivity and specificity for each variant type against reference materials - and revalidated when software, versions, or settings change.
- Version control and locked configurations, so every report can be traced to exactly the code that produced it.
- Storage and privacy: FASTQ and BAM files are huge and contain identifiable genomes; retention and security policies apply.
- Known blind spots: low-complexity regions, pseudogenes, large indels, and GC-extreme regions; reports should state them.
Quick check
Mixing builds misplaces variants; HGVS names should cite transcripts and build.
What to take from this
FASTQ (raw reads with qualities) → align to the FASTA reference → BAM → call → VCF → annotate and interpret. Phred Q30 means a 1-in-1,000 error. Depth counts reads at a base; VAF is variant reads over total, read in light of germline versus tumor and tumor purity. A clinical pipeline names its genome build, is validated and version-controlled, and states its blind spots.
Lesson complete
Nice work.
Sources for this lesson
- 1Lela Buckingham. Molecular Diagnostics: Fundamentals, Methods, and Clinical Applications. 3rd ed. F.A. Davis Company. 2019. verifiedThe standard clinical molecular-diagnostics textbook for MLS/MB programs; author holds MB DLM(ASCP). Covers nucleic-acid chemistry, techniques, lab operations, and applications across infectious disease, oncology, genetics, and identity. Primary topic reference for the ASCP MB program.
- 2Bruce Alberts, Rebecca Heald, Alexander Johnson, David Morgan, Martin Raff, Keith Roberts, Peter Walter. Molecular Biology of the Cell. 7th ed. W. W. Norton & Company. 2022. verifiedThe canonical cell/molecular biology textbook; used for nucleic-acid chemistry and the central dogma.
Further reading
- Michael R. Green, Joseph Sambrook. Molecular Cloning: A Laboratory Manual. 4th ed. Cold Spring Harbor Laboratory Press. 2012. verifiedThe classic three-volume molecular-biology methods manual — authoritative for nucleic-acid isolation, electrophoresis, restriction digestion, labeling, and hybridization techniques. Standard-tier topic reference for the techniques courses.