I have a recurring lab dream where someone casually says, “We should just do sequencing,” and then ten minutes later we’re arguing about stranded vs non-stranded, rRNA depletion, adapter ligation bias, and whether ATAC-seq is DNA-seq (it is… but emotionally it’s chromatin therapy). 😝
So here’s my attempt to categorize the entire “-seq” universe in a way that is scientifically accurate, actually usable, and only mildly chaotic.🙂
How to categorize sequencing without losing your will to live (the only system that matters)
• Every “-seq” is basically answering one of these questions:
1) What DNA is present? (genome, variants, copy number)
⭐️DNA-seq (WGS/WES)
⭐️targeted panels
⭐️amplicon sequencing
2) What RNA is present and how much? (expression, isoforms, small RNAs)
⭐️bulk RNA-seq (poly(A) or total RNA; stranded/non-stranded)
⭐️scRNA-seq / snRNA-seq (RNA-seq, but per cell/nucleus)
⭐️small RNA-seq (miRNA etc.)
3) What parts of chromatin are open? (accessibility)
⭐️ATAC-seq
4) Where does a protein bind DNA / where are histone marks? (protein–DNA)
⭐️ChIP-seq (modern cousins: CUT&RUN / CUT&Tag)
5) What RNAs are bound to a protein? (protein–RNA)
⭐️RIP-seq (higher-resolution cousins: CLIP-seq / eCLIP / iCLIP)
• Inside each category, two “library reality checks” explain 90% of the confusion:
Enrichment / depletion: What are you keeping, and what are you removing? (poly(A) capture, rRNA depletion, size selection, target capture…)
Adapters & directionality: How did adapters get added, and do you preserve strand? (ligation vs tagmentation, stranded vs non-stranded, UMIs, barcodes…) Adapters are basically the tiny handles that let the sequencer grab your fragments. Without adapters, your sample is just screaming into the void. 📣
Part 1 — DNA sequencing (aka “what’s in the genome?”) 🧬
The question:
“What DNA sequence is present, and what variants are there?”
Common flavors:
WGS (Whole Genome Sequencing): all genomic DNA. Expensive, comprehensive, honest. 💸
WES (Whole Exome Sequencing): mostly coding exons. Cheaper, but misses regulatory/noncoding regions.
Targeted / amplicon sequencing: specific genes/regions. Great when you already know what you care about.
Typical library logic: (simplified but accurate)
Genomic DNA → fragment → end repair / A-tailing (platform dependent) → adapter ligation → PCR (often) → sequencing.
DNA-seq is usually the most straightforward in the “-seq zoo” because DNA mostly behaves like… DNA. (RNA could never.) 😭

Part 2 — RNA-seq (aka “what’s being expressed?”) 🧫🧬
RNA-seq sounds simple until you remember RNA comes with built-in problems: Total RNA is dominated by rRNA (often ~80–90% of reads if you don’t remove it) RNA can degrade strand matters for overlapping genes and antisense transcription small RNAs are tiny divas with adapter bias. So RNA-seq is really “RNA-seq + decisions.”
2A) Poly(A) RNA-seq (mRNA-seq) 🧲
Best for: typical eukaryotic mRNA expression profiling
How it selects: capture polyadenylated RNA using oligo-dT (poly(A) selection)
What it misses: many non-poly(A) RNAs (some noncoding RNAs, many histone mRNAs), and degraded samples can perform poorly because poly(A) tails may be compromised.
2B) Total RNA-seq + rRNA depletion 🧹
Best for: noncoding RNAs, messy transcriptomes, and often degraded samples
Key step: rRNA depletion (probe-based capture or RNase H–based depletion, depending on kit/approach)
Why it matters: otherwise your sequencing run becomes Ribosome Fan Club: Greatest Hits. 🎶

2C) Stranded vs non-stranded RNA-seq 🧭
Non-stranded: you don’t know which DNA strand the RNA originated from.
Stranded: preserves directionality so you can distinguish overlapping transcripts, antisense RNA, and sense/antisense signals.
How strandedness is commonly achieved (conceptually): the chemistry “marks” one cDNA strand during library prep (often via incorporating dUTP in second-strand cDNA so it won’t amplify), resulting in reads that preserve original RNA orientation.
If you have overlapping genes, antisense transcription, or compact genomes: stranded is worth it. ✅
Part 2D — scRNA-seq (single-cell RNA-seq): RNA-seq, but every cell wears a name tag 😭🏷️
Where it belongs: Under RNA-seq. Same biological question (“what’s expressed?”), but now the unit is one cell (or one nucleus) instead of a mixed population.
The question: “What is each individual cell expressing?”
Bulk RNA-seq averages everything together. scRNA-seq lets you see cell types, cell states, and rare populations that bulk would blur into a single smoothie. 🥤
scRNA-seq introduces identity labels during capture so reads can be assigned to: which cell they came from (cell barcode), and which original RNA molecule they came from (UMI, Unique Molecular Identifier). UMIs help prevent PCR from turning one molecule into “100 identical molecules” in your data. (PCR is helpful, but also a liar.) 😭
Most high-throughput scRNA-seq captures RNA using barcoded oligo-dT primers, meaning it primarily targets poly(A)+ RNA (mostly mRNA). Then it performs end-counting rather than full-length transcript sequencing.
3′ scRNA-seq: sequences near the 3′ end; very common for expression profiling
5′ scRNA-seq: sequences near the 5′ end; often used when you also want immune repertoire (TCR/BCR) information in paired workflows
Droplet scRNA-seq is usually tag-based end counting, and the design is inherently directional around the captured end (3′ or 5′). So people don’t talk about it exactly like bulk “stranded vs non-stranded” whole-transcript RNA-seq. If you do full-length scRNA-seq (plate-based approaches like SMART-seq-style workflows), strand/isoform questions become more relevant — but throughput is lower and cost per cell is higher. 💸

snRNA-seq (single-nucleus RNA-seq): the close cousin! 🧠🧫 You profile nuclei rather than intact cells. This is especially useful when: tissue is hard to dissociate, intact (brain, fibrotic tissue) samples are frozen, or you want less dissociation-induced stress artifacts. You often see more intronic reads (pre-mRNA) in snRNA-seq because nuclear RNA includes unspliced transcripts. This is expected, not a mistake.🙂
Part 3 — Small RNA-seq (aka “tiny RNAs, huge attitude”) 🧬🪙
Best for: miRNA and other small RNAs (often ~18–30 nt, protocol dependent)
Why it’s different: Small RNAs are too short for standard workflows.
Typical steps: size selection to enrich small fragments → adapter ligation (commonly 3′ adapter first, then 5′ adapter) → reverse transcription → PCR → sequencing
The big gotcha (real science): Ligation bias: certain sequences ligate adapters more efficiently, which can skew quantification. Some workflows use UMIs or optimized ligation conditions to reduce bias, but it never fully disappears…… you just learn to live with it politely. 😭
Part 4 — ATAC-seq (aka “show me open chromatin”) ⚡️
The question: “Where is chromatin accessible (open)?”
ATAC-seq uses Tn5 transposase to simultaneously cut DNA and insert adapters (“tagmentation”).
So the workflow is basically: accessible chromatin + Tn5 = sequencing-ready fragments
Why it’s cool: 👀relatively fast library prep; 👀accessibility maps regulatory regions and can support motif analysis; 👀fragment size patterns carry biological signal (nucleosome-free vs nucleosome-associated)
What it’s not: ATAC-seq measures accessibility potential, not expression. Open ≠ expressed, but they can correlate.

Part 5 — ChIP-seq (protein–DNA binding and histone marks) 🧲🧬
The question: “Where does this protein bind DNA?” or “Where is this histone modification located?”
Classic ChIP-seq logic: (often) crosslink protein–DNA fragment chromatin (sonication or enzymatic) → immunoprecipitate using an antibody to the protein or histone mark reverse crosslinks (if crosslinked) → purify DNA adapter ligation → PCR → sequencing
General patterns: Transcription factor ChIP-seq: sharper peaks, often harder (low abundance) Histone mark ChIP-seq: broader domains (mark dependent)
ChIP-seq is where your experimental success can feel like: 50% biology + 50% antibody + 80% sample quality. Yes, the math is broken. That’s the point. 😭

Part 6 — RIP-seq (protein–RNA interactions: “who’s hanging out with whom?”) 🧲🧫
The question: “What RNAs are associated with this RNA-binding protein (RBP)?”
RIP-seq logic: lyse under conditions preserving RBP–RNA complexes → immunoprecipitate the RBP isolate associated RNA → build RNA-seq library → sequence
Key nuance (important): RIP-seq tells you which RNAs associate with your protein, but not necessarily the exact nucleotide-level binding site. If you want binding sites, the CLIP family (eCLIP/iCLIP/etc.) is often used because it can provide higher resolution via crosslinking-based strategies.
The cheat-sheet: pick your -seq in four questions ✅
😃Do you care about DNA sequence/variants or RNA expression?
variants → DNA-seq expression → RNA-seq / scRNA-seq
😃Are you studying regulation?
accessibility → ATAC-seq binding/marks → ChIP-seq (or CUT&Tag/CUT&RUN)
😃Are you studying protein–RNA interactions?
association → RIP-seq binding sites → CLIP-like methods
😃What will dominate your library if you don’t control it?
total RNA → rRNA will eat your budget small RNA → ligation bias gremlins scRNA-seq → cell quality and handling matter a lot ATAC → tagmentation and sample quality can make or break you
“Let’s just do sequencing” is not a plan 😭✨
Sequencing isn’t confusing, it’s just aggressively diverse. Every method is a different way of asking: What’s here? How much? Where is it? Who is it interacting with?
And the hardest part isn’t the sequencing. It’s the moment someone says: “Let’s just do some sequencing,” and you realize “some” is not a plan. 🧬🥲 CHOOSE WISELY.
