RNA-seq Analysis
for Absolute Beginners
Seven recorded sessions that start at zero and don’t skip. The Linux command line, your raw reads turned into a count matrix, then R, the statistics that find your differentially expressed genes, publication figures, pathway analysis, batch effects, and getting your data into GEO. On your own laptop, at your own pace.
Three minutes on what the course is, who built it, and who it’s for.
Your data came back.
Now what?
The data comes back. You open the folder, and it’s files you can’t read, in formats nobody explained β and everyone around you is too busy to walk you through it.
You ask how to start, and you get handed a reading list
Learn Bash. Python, R, AWK, SED. Statistics, matrix algebra β all of it, before you’re allowed to touch your own data. I took that road. It cost me about ten years.
The Linux command line, before a single analysis begins
Almost every quantification tool runs in a terminal. Most tutorials skip it entirely and assume you already know it. Most beginners don’t, and get stuck on the first line.
A maze of file formats with no map
FASTQ, SAM, BAM, GTF β knowing which tool wants which format, and what’s actually inside each one, is genuinely confusing when nobody has ever opened one in front of you.
Written tutorials have no “start here” arrow
Even good ones. If you don’t already know which section to read first, what to skip, and how the pieces connect into a working pipeline, a library of tutorials is just more things you haven’t read.
I wasn’t a bioinformatician either.
I didn’t come from computing. I was a wet-lab researcher studying temperature-dependent sex determination in turtles, and when our first RNA-seq data arrived I had no idea what to do with it. Nobody in the department did. I ended up asking someone in the medical school for help; they took a long time, and eventually told me they couldn’t do it.
Everything I know, I taught myself the hard way. I spent years in the dark, piecing together fragments from tutorials that all assumed prerequisites I didn’t have. The advice I was given β learn the whole language first, then the statistics, then you may touch your data β is the reason it took as long as it did.
Since then I’ve worked on hundreds of projects with more than two hundred collaborators, and I’ve watched the same thing happen to other people every week. Researchers waiting months for results. Students stuck before they start. So I built this the way I wish someone had built it for me β starting from the terminal you’ve never opened, and not skipping the parts that are supposed to be obvious.
So I teach it
the other way round.
It has a name β lean learning. It is the opposite of the reading list, and it is why a bench scientist gets through this course when they didn’t get through the others.
Only what the analysis actually needs
Not the whole of R. Not the whole of Linux. The commands and the concepts this pipeline uses, learned at the moment you need them β which is far less than that list, and it is enough to do the work.
Take the action first, aim as you go
You run real commands on real data in the first session, before you understand everything about them. Understanding arrives from having done it, not the other way round. Aiming forever and never firing is how beginners lose years.
You are a junior analyst, from day one
You don’t invent methods. You understand the workflow and point existing tools at your own data. That’s the job, that’s what this teaches, and settling it on day one is what stops the impostor feeling that makes people quit.
Seven sessions. One complete pipeline.
In order, nothing skipped. Sessions 1 and 2 are in Linux, because that’s where quantification happens. From Session 3 on, everything runs on your own laptop, in R. Each session assumes only the ones before it, and there’s nothing you need before Session 1.
Linux from scratch, on your own machine
The session that assumes the least of any in the course. The Linux file system as one upside-down tree, and a path as the route down it. What an HPC actually is and whether your project needs one. Then the habit that isn’t glamorous and decides everything: raw data you never touch, numbered results folders, scripts kept separate from data, a README that says what you did and why.
From raw reads to your first count matrix
Where you meet the data. A FASTQ record is four lines per read, and once you can read those four lines it stops being frightening. Then SAM and BAM, so you know where every read landed, and GTF, so you know where the genes are. Then the tools in the order you run them: FastQC, fastp, STAR, featureCounts, MultiQC. The example data is one chromosome rather than a whole genome, so an alignment that would take hours finishes while you’re still watching.
The longest session, and deliberately so
This is where most beginners are quietly lost, so it gets the room it needs. What programming actually is. RStudio’s four panels. Vectors, data frames, matrices and lists β the four shapes your data can take. Reading files in, writing them out, saving objects so they survive between sessions. Conditions, loops, functions. Bioconductor. Then normalization: why raw counts lie to you, and which method to use for what.
The session you’ve been working toward
It starts by explaining why you cannot just run a t-test: counts aren’t normally distributed, the variance depends on the mean, and testing twenty thousand genes at once means false positives by the hundred. That last problem has a name, and the fix for it is why your results table has an adjusted p-value column. Then the eight steps in order β load, filter, normalize, QC with a correlation heatmap and your first PCA, design matrix, voom and limma, read the table, export. And at the end, the same data through DESeq2 and edgeR, so you see where three good tools agree and where they don’t.
The table becomes pictures, then biology
The biggest deck in the course, because figures take room to teach properly. First ggplot2 β the grammar of it, so you’re not memorising a different command per plot type. Then volcano, MA, heatmap and PCA, each built up step by step from your own results. Then the part reviewers actually notice: colours that work for colourblind readers, type sizes that survive a printed column, multi-panel figures, 300 DPI export. Then GO, KEGG and GSEA β a gene list into pathways, and pathways into a paragraph you can put in a paper.
The session that asks whether any of it was true
A batch effect is technical noise that looks exactly like biology, and if you don’t go looking for it, it ends up in your results and nobody catches it until review. Detect it with PCA, then three different ways to correct it, with all four states side by side so you can see what each method did to the data. Then the part people get wrong: for hypothesis testing you don’t correct the matrix at all β you put batch into the design. Correct the matrix for pictures, model the batch for tests.
Where the two pipelines are handed over
The first is the downstream analysis, walked through block by block, with the real expected output printed beside every call so you can check your own console against it. There’s a slide on making it yours, too β it comes down to five edits, and everything in between stays exactly as it is. The second is quantification on a cluster: SLURM, one sample first, then a job array across all of them. And in between, the whole GEO submission in seven steps, because most journals want an accession number before they’ll accept your paper, and almost nobody teaches it.
You don’t just learn the pipeline.
You leave with it.
Two folders that sit outside the sessions, because this is the part you copy into your next project. The numbering is the instructions.
RNAseq_Analysis_Pipeline/
The downstream half. A README, a script that installs every package once, the notebook that runs the analysis end to end, the function library behind it, and the example dataset β counts, metadata and the annotation β so the whole thing runs before you ever touch your own data.
ββ README + one-time package setup
ββ notebook Β· QC β DE β figures β pathways
ββ function library + example dataset
RNAseq_Quantification_SLURM/
The cluster half. Setup for Linux and for Mac, then build the index, run one sample, run the job array across all of them, collect the QC. Built for SLURM, the scheduler used by most university and institute clusters.
ββ setup Β· Linux and Mac
ββ index β single sample β job array
ββ QC collection Β· FASTQ β counts table
Start from raw reads or from a counts table β the pipeline meets your data wherever it is.
Every session folder
has the same things in it.
So you always know where to look. All of it works offline, and all of it is yours to keep.
The videos themselvesSeven sessions, 6 h 26 m, yours for as long as you want them
Every deck as a PDFTo keep, to search, and to print
The scripts and notebooksR notebooks for the analysis sessions, shell scripts for the Linux ones
The example data, bundled in the folderNothing depends on a server staying up
Cheatsheets and handoutsWhere they earn their place, not padded into every session
Both pipeline foldersThe downstream analysis and the cluster quantification
A setup guide for Windows and MacOne pass, 45 minutes to an hour, and you’re ready for Session 1
Downloaded, not streamedEverything sits on your own disk, so nothing depends on a server staying up
Here’s how I’d actually do it.
This is a course you run, not a course you watch. The difference is entirely in how you use the pause button.
Set up once β 45 minutes to an hour
Follow the setup guide for Windows or Mac. You do this one time, before Session 1, and then it’s done for the rest of the course.
Then the sessions in order
Each one assumes the ones before it and nothing else. Every session opens by telling you what you should have finished and what file you need in hand, so picking it back up after two weeks away costs you nothing.
Watch a section, pause, run the code yourself
Then carry on. The sessions are built with the stopping points already in them β this is the difference between finishing the course able to do the analysis and finishing it having seen someone else do it.
Don’t try to memorise any of it
The notebooks keep every line for you. What you want to be watching is the process β what each step is for, what goes in, what comes out, and how to tell when something has gone wrong.
When something breaks β and it will
Copy the exact error, paste it into an AI, and ask it how to fix it. That’s what I do, and it’s faster than searching. You should know before you buy: this course doesn’t come with me answering your questions. That’s the live workshop, and it’s a different thing.
See exactly what’s inside.
Twelve minutes, all seven sessions, what each one teaches and what you actually walk away with. Everything in it is a real slide out of the real deck β not a summary made for the video. If you’re still deciding, this should be enough to decide on.
Free, no email required. Watch it before you spend anything.
Who this is for β
and who it isn’t.
There are no prerequisites. If you’ve never opened a terminal, you are exactly who I built this for.
This is for you ifβ¦
- βYou’re a bench scientist and your own RNA-seq data has come back, or is about to.
- βYou’ve never opened a terminal, or you’ve opened one a couple of times and nothing more.
- βYou’ve tried to learn this before and stalled on a tutorial that assumed something you didn’t have.
- βYour core facility hands you a counts table and you’d like to know what’s inside it β and what to do with it.
- βYou’re heading for single-cell, ATAC-seq or variant calling later. All of them assume the foundation this course teaches.
- βYou want to work at your own pace, at your own hours, without a cohort schedule to keep up with.
This isn’t for you ifβ¦
- βYou already write your own pipelines. You’ll find this slow β it’s built for someone starting at zero.
- βYou want an advanced or specialist course. This is the foundation, taught thoroughly, not a survey of frontier methods.
- βYou want someone to answer your questions as you go, or to look at your own dataset with you. That’s the live workshop.
- βYou want to skip the Linux half. You can β Sessions 3 onward stand on their own from a counts table β but then you’re buying about two-thirds of what’s here.
Dr. Lei Guo
Computational biologist Β· Founder, NGS101.com
A computational biologist with more than a decade analysing next-generation sequencing data. Hundreds of projects, more than two hundred collaborators, and a tutorial library that draws around thirty thousand visits a month.
He has trained individual researchers, whole labs and entire departments β most of them wet-lab scientists who had never opened a terminal before they started.
And he started there himself. Every course claims a decade of experience. Almost none of them can claim the person teaching it began exactly where you’re standing, got the standard advice, and lost ten years to it.
Same material.
Two different things.
The live workshop teaches this same curriculum, with me in the room. Here is exactly what changes, including the parts that don’t flatter the recorded course.
| Recorded course$497 | Live workshop$997 | |
|---|---|---|
| Format | Seven recorded sessions, 6 h 26 m. Watch any time. | Seven live sessions with me, ~2 hours each, Tue Β· Thu Β· Sat. |
| When you start | The moment you buy. | The next cohort. Three a year, 15 seats each. |
| Pace | Yours. Pause, rewind, stop for a month, start again. | The cohort’s. You keep up, or catch up on the recordings. |
| Questions answered | No. It’s self-serve β you get the material, not me. | Yes, live, in every session. This is what the live workshop sells. |
| Your own data | No. You learn on the example dataset, then run the pipeline yourself. | A private 1:1 after Session 7, where we run the pipeline on your counts table. |
| Setup | Your own laptop. A 45β60 minute setup guide, then everything runs offline. | A pre-configured cloud environment. Log in and start β nothing to install. |
| After it ends | Nothing further. The material is yours permanently. | One month of email support while you apply it to your data. |
| Certificate | β | Yes, documenting your training hours. |
| Materials | Videos, decks, scripts, notebooks, example data, both pipelines. Yours to keep. | All of the same β plus this recorded course, on enrolment, to work through before Session 1. |
Your $497 comes off the price of a live seat. You pay the $500 difference for a place in any cohort within 12 months of buying the recorded course. You never pay more for having started with the recording β the total is the same $997 either way.
Which makes this the low-risk way in: take the course, work through it at your own pace, and if you get to the end wanting someone to look at your own data with you, the live workshop is still there at no penalty.
How it works: your welcome email contains a single-use code worth $500 off a live seat. Apply it at checkout when you enrol in a cohort. Nothing to claim and nothing to ask for β just keep the email.
Want the live version instead? See the live workshop β
Stop waiting for a bioinformatician.
Become one.
Seven sessions, six and a half hours, two pipelines, and nothing you need to know before you start. If you were looking for the catch β there isn’t one. It’s just long.
All sales are final β this course does not come with refunds. Every part of it, the videos included, is downloadable and yours to keep the moment you have access, so a refund would mean returning nothing.
That is exactly why the full twelve-minute walkthrough is free and sits above. It shows you every session, every deliverable and the real slides before you spend anything. Watch all of it, then decide.
And what it doesn’t include: this course is self-serve. It does not come with email support, a forum, office hours, or any other way to reach me with questions. If that is what you want, it is the live workshop.
What βyours to keepβ means. Everything downloads to your own machine. Once you have the files, nothing you paid for depends on a server staying up β the videos play, the scripts run and the data is already there. Your download area on this site stays open for at least twelve months after you buy.
If you move up to the live workshop later. The full $497 comes off a seat in any cohort. Nothing to arrange now and nothing to keep track of β the details arrive with your purchase.