Welcome to Pseudocount
Pseudocount is about the statistics inside bioinformatics workflows. A typical tutorial shows a pipeline running from raw counts to a list of genes; the posts here stop at one step of that pipeline and ask what it does to the answer.
The format
Each post picks a single decision: a normalisation method, a design formula, whether cells or samples are the unit of replication, how a batch is handled, which multiple-testing correction is used. It then simulates data where the truth is known, runs the analysis with and without the problem, and measures the difference. Code comes in R, with Python where that is where the reader is likely to be working.
Why simulation
With real data you rarely know which genes truly changed, so you cannot tell whether a method got it right. With simulated data you can, which turns a debate about methods into a measurement. The cost is that a simulation is only as realistic as its assumptions, and each post says which assumptions it makes.
Where to start
New posts appear on the front page. The about page explains how corrections are handled.