If you’ve ever looked at a tissue sample and thought, “There are so many cells in here… and they’re all doing different things… and I would like to know their secrets 🥲🔬”… welcome. This is the exact emotional origin story of single cell RNA sequencing.
I wanted to write this blog because 10x droplet scRNA-seq sounds like science fiction until you learn the core idea, and then it becomes… oddly simple… and also kind of genius. It’s basically tiny droplets + barcodes + sequencing + computational therapy 😂💻✨.
Today I’m walking through the classic 10x style workflow (especially 3’ Gene Expression) from “cell soup” 🥣🧪 to a gene by cell matrix 📊… with extra detail on library preparation and a section on different chemistries and what they’re best for.
The whole trick in one sentence: 🧠✨
10x droplet scRNA-seq is basically tiny oil bubbles playing “one cell per droplet” (ideally): each droplet captures a single cell together with a barcoded bead, the cell lyses, and its RNAs get converted to cDNA that’s tagged with a cell barcode (who this cell is) plus a UMI (which original RNA molecule this was). After that, everything is pooled and sequenced together, but those tags let you sort reads back to individual cells and count molecules more honestly instead of being fooled by PCR duplicates.
Step 1… sample prep is where success is decided (quietly) 🧊🧪😅
Everything starts with a single cell suspension (or nuclei, depending on the assay). This step looks innocent… but it’s where most “why does my dataset feel cursed 😭” stories begin.
The reason is simple… droplet scRNA-seq is sensitive to cell viability and RNA leakage. When cells die, RNA leaks into the solution and becomes ambient RNA. That ambient RNA can get captured too, which adds background signal and can blur cell type boundaries. It’s not “wrong”… it’s just biology being biology.
So the practical goal is… get a clean suspension, keep cells happy-ish, remove debris, and avoid turning the tube into a transcriptome soup buffet 🍜🧬.
Step 2… limiting dilution means probability is in charge 🎲
Here’s a thing that feels rude at first… the system does not guarantee one cell per droplet. Instead, cells are loaded at concentrations where many droplets end up empty on purpose. The logic is… if you try to stuff the system so every droplet has a cell, you will drastically increase the number of droplets that contain two cells.
Those are doublets or multiplets, and later they look like one “cell” expressing markers from two different cell types… because it literally was two cells sharing a barcode. Plot twist, but in a bad way 🙃.
So yes… empty droplets are not a failure. Empty droplets are part of the design… like emotional boundaries, but microfluidic. 😃
Step 3… GEMs form and each droplet becomes a tiny lab 🏠🧬
In the Chromium droplet workflow, cells, reagents, oil, and a special barcoded gel bead come together to form GEMs… Gel Beads in Emulsion. These droplets are tiny reaction chambers where chemistry happens locally.
The ideal droplet contains one cell + one gel bead. The gel bead carries many copies of the same barcoded oligo design so that everything captured in that droplet can be assigned to the same cell barcode later. That is the entire “how do we keep cells separate” strategy… but implemented as thousands of microscopic bubbles 😭✨.
Step 4… lysis and capture happen inside the droplet 🧬🧪🔬
Once a cell is inside a droplet, it lyses and releases RNA. In the classic 3’ gene expression workflow, mRNA is captured through poly(A) selection… the gel bead oligos include a poly(dT) stretch that binds poly(A) tails. That means the workflow is intentionally tuned to capture polyadenylated mRNA rather than everything in the cell.
This is why the assay is called 3’ Gene Expression… capture is anchored near the poly(A) tail, so the resulting reads are biased toward transcript 3’ ends. This design is perfect for quantifying gene expression across many cells, even if it’s not meant for full length isoform discovery.
Step 5… reverse transcription adds identity tags (cell barcode + UMI) 🏷️🧬✨
This is the part where scRNA-seq becomes scRNA-seq and not just “RNA but chaotic.”
During reverse transcription inside droplets, each captured mRNA molecule gets converted to cDNA while picking up two labels.
First is the cell barcode… every droplet has a bead with one cell barcode, so all molecules in that droplet share the same cell identity.
Second is the UMI… a unique molecular identifier attached so you can tell whether multiple reads came from the same original molecule or from PCR duplication. In the Next GEM 3’ v3.1 library design, Read 1 is structured specifically to encode the 16 bp 10x Barcode and 12 bp UMI, and Read 2 sequences the cDNA fragment.
Why UMIs matter is simple… PCR is powerful, but it is also dramatic 😂📈. UMIs let you collapse PCR duplicates so counts better reflect “how many molecules were captured,” not “how loud PCR screamed that day.”
Step 6… break the emulsion, clean up, amplify cDNA 🧼📦😌
After reverse transcription finishes, the emulsion is broken and everything is pooled into one tube. This is the moment beginners often panic… “Wait, aren’t we mixing all cells together???”
Yes… and it’s fine… because by now, single cell identity is embedded into the cDNA through the barcode and UMI. Pooling does not erase cell identity… it simply collects all labeled molecules in one place.
Then cDNA is amplified by PCR to generate enough material for library construction. At this stage you have barcoded cDNA that is ready to be turned into sequencing libraries.
Library preparation… how barcoded cDNA becomes a sequencer ready library 🧬🔧✨
Library prep is where your cDNA gets transformed into an Illumina compatible library that sequencing instruments can read efficiently. A 10x 3’ gene expression library is built as a standard Illumina paired end construct that begins and ends with P5 and P7, and uses index reads for sample indices.
1) Fragmentation… making cDNA a sequenceable size ✂️🧬
Amplified cDNA is typically too long and heterogeneous for efficient sequencing, so it is fragmented down to a target size range. This makes the insert lengths compatible with sequencing and improves cluster generation and read usability.
2) End repair and A tailing… making the ends ligation friendly 🧪🙂
After fragmentation, DNA ends are processed so they have the right structure for adapter addition. In modern Illumina style library prep, this often includes end repair to generate blunt or compatible ends and addition of a single A overhang to support adapter ligation. The details are kit dependent, but the principle is consistent across library construction workflows.
3) Adapter addition and indexing… adding the sequencer handles 🏷️✨
Adapters provide the sequences needed for sequencing primers and flow cell binding. Sample indices are incorporated so multiple libraries can be pooled and later demultiplexed. In the Next GEM 3’ v3.1 library design, the construct uses standard Illumina paired end structure and incorporates both i7 and i5 index reads as sample indices.
4) Sample Index PCR… finalizing the library 📈🧬
A PCR step adds the full adapter sequences and amplifies the final libraries so there is enough material for sequencing. This step is also where the final index structure is completed.
5) Sequencing read layout… who reads what 👀🧬
For 3’ gene expression, the reads are designed so that barcode and UMI information is read separately from the transcript sequence. In v3.1, Read 1 encodes the 16 bp cell barcode and 12 bp UMI, while Read 2 sequences the cDNA fragment, and i7 and i5 index reads carry sample indices.
That structure is the reason the pipeline can later say… “these reads belong to cell 000123… and these UMIs indicate 412 original molecules for gene X”… instead of “everything is mixed, good luck” 😭🙏.

After sequencing… how reads become a gene by cell matrix 📊🧠✨
Once you have FASTQs, software assigns reads to cells using barcodes, deduplicates using UMIs, and quantifies expression by aligning or mapping the transcript sequence.
A key step is cell calling… deciding which barcodes correspond to real cell containing droplets and which represent empty droplets or background. The gene expression pipeline uses a two step approach… first a cutoff based on total UMI counts, and then RNA profile information to distinguish low RNA cells from empty droplets using an EmptyDrops style method.
This matters because many droplets were intentionally empty, and some real cells have low RNA content. Good cell calling is basically the difference between “clean atlas” and “why do I have 40,000 cells but 30,000 of them look like air.” 😅
Different 10x chemistries… what they are and what they’re best for 🧪🧬🧭
Now for the part that sounds like phone models… “v3.1,” “v4,” “GEM-X,” “Flex,” “5’,” and your labmate saying “we upgraded” like it’s an iOS update 😂📱.
In 10x terms, “chemistry” means the generation of reagents and workflow design that impacts recovery, sensitivity, and what types of samples you can realistically run.
1) Next GEM 3’ v3.1… the classic workhorse 📊
This is the widely used 3’ gene expression workflow many published datasets are built on. It uses droplet barcoding, poly(A) capture, and a read structure where Read 1 contains the 16 bp cell barcode and 12 bp UMI and Read 2 reads the cDNA fragment.
👀Use it when you want scalable single cell expression profiling across many cell types and states and your samples are compatible with fresh cells or nuclei workflows.
2) GEM-X Universal 3’ v4… newer generation, higher recovery and sensitivity 📈🧬✨
GEM-X is the newer generation of the same general droplet idea. 10x reports improved cell recovery efficiency for GEM-X compared to Next GEM in internal benchmarking, including a higher median recovery for 3’ gene expression runs.
They also market Universal 3’ as offering up to 80% cell recovery and higher sensitivity versus prior Next GEM technology, plus more usable reads and potential sequencing cost savings.
Use it when you are starting a new project and want the current generation performance benefits, especially if cell recovery and sensitivity are key constraints.
3) Universal 5’ Gene Expression and Immune Profiling… when you want TCR or BCR too 🛡️🧬🧫
If your question involves adaptive immunity, 5’ workflows are the main character. 10x describes Universal 5’ as enabling analysis of full length paired BCR or TCR sequences, surface protein expression, and 5’ gene expression from a single cell.
This is what you choose when you want to connect “what the cell is doing” with “what receptor it has,” meaning clonotypes, antigen receptor diversity, immune responses, tumor immunology, and vaccine studies.
👀If you are running multi modality immune experiments, the multi analysis approach can analyze gene expression, V(D)J, and feature barcode data together and helps keep cell calling consistent across modalities.
4) Flex Fixed RNA Profiling… when you need fixation and logistics freedom 🧊📦😌
Sometimes your biggest limitation is not biology… it’s timing. Flex is designed for fixed samples. Instead of relying on poly(A) capture of mRNA, Fixed RNA Profiling measures RNA levels using probe pairs that hybridize to targets and are then ligated, after which the barcoded ligation products are used to generate libraries for sequencing.
The user guide also describes that in GEMs, probes are ligated and a GEM barcode is added so that ligated probes within a GEM share a common barcode.
Use Flex when you need the ability to fix samples and run later, want easier batching across time, or have sample types where fixation is part of feasibility.
5) Feature Barcoding… when RNA alone is not enough 😤🧬➕🧫
Feature Barcoding is how you measure extra modalities alongside RNA, most commonly cell surface proteins with oligo tagged antibodies, and sometimes other features depending on the application. 10x describes cell surface protein labeling as using antibodies conjugated to a Feature Barcode oligonucleotide for use with single cell RNA sequencing protocols.
There are also workflows that extend feature barcoding to fixed samples, including cell surface and intracellular protein labeling with Feature Barcode oligos in GEM-X Flex contexts.
Use Feature Barcoding when you want cleaner cell type annotation, better resolution of immune subsets, multiplexing strategies, or protein level confirmation of markers that RNA is lying about today 😅.
How to choose a chemistry… without turning this blog into a grant 😭📄
If you want broad discovery gene expression across many cells and you have fresh cells or nuclei… start with 3’ Gene Expression, and consider GEM-X if available for newer performance.
If you care about immune receptors and clonotypes… use 5’ immune workflows.
If you need fixation and batching flexibility… Flex Fixed RNA Profiling is built for that.
If you keep saying “RNA is not telling the whole truth”… add Feature Barcoding for proteins.
And that’s the 10x droplet scRNA-seq story… If you take only one thing from this… it’s that single-cell isn’t “one perfect transcriptome per cell”… it’s probability, chemistry, and computation teaming up to give you a surprisingly powerful snapshot of what your cells were up to. So be kind to your sample prep, respect the doublets 🙃, feed your sequencing depth responsibly 💸, and remember… every beautiful UMAP started as a chaotic cell soup 🥣✨.

And… don’t forget… single-cell RNA-seq is expensive 💸. The budget can disappear faster than your cells after a harsh dissociation 😭🧪. So before you run “just one more sample” (famous last words 😂), plan your experimental design carefully… decide how many samples you actually need, what read depth answers your question, and whether you can multiplex to stretch your dollars a little further 🧬📊✨.
