Welcome to Pseudocount

meta
What Pseudocount covers and how its bioinformatics statistics tutorials work: one analysis decision per post, tested on simulated omics data with a known answer.
Author

Pseudocount

Published

24 August 2026

Pseudocount is about the statistics inside bioinformatics workflows. A typical tutorial shows a pipeline running from raw counts to a list of genes; the posts here stop at one step of that pipeline and ask what it does to the answer.

The format

Each post picks a single decision: a normalisation method, a design formula, whether cells or samples are the unit of replication, how a batch is handled, which multiple-testing correction is used. It then simulates data where the truth is known, runs the analysis with and without the problem, and measures the difference. Code comes in R, with Python where that is where the reader is likely to be working.

Why simulation

With real data you rarely know which genes truly changed, so you cannot tell whether a method got it right. With simulated data you can, which turns a debate about methods into a measurement. The cost is that a simulation is only as realistic as its assumptions, and each post says which assumptions it makes.

Where to start

New posts appear on the front page. The about page explains how corrections are handled.