Reading Your Blood Type From Your DNA
How to genotype ABO and RhD from whole-genome sequencing data, which variants matter, where short reads fail, and why a genomic call is not a transfusion-grade answer.
How to genotype ABO and RhD from whole-genome sequencing data, which variants matter, where short reads fail, and why a genomic call is not a transfusion-grade answer.
A working guide to running DESeq2 on bulk RNA-seq counts: building the count matrix, filtering, size factors, dispersion, Wald versus likelihood ratio tests, log fold change shrinkage, and the sanity checks that catch most mistakes.
A worked walkthrough of what real DNA results look like: array text files, VCF lines field by field, coverage and callability checks, annotation with VEP, and what a 'positive' result does and does not mean.
A practical protocol for estimating functional age from physical performance tests, blood chemistry algorithms, and DNA methylation clocks, with the reliability problems that make most single-timepoint results uninterpretable.
A breakdown of the genomic testing market by category — clinical labs, direct-to-consumer arrays, research-grade sequencing providers, and tumor profiling — with the file formats, coverage depths, and failure modes that decide whether the data is useful to you.
There is no single lifespan blood test. There is a small set of analytes that, combined, predict all-cause mortality better than any one of them, and a way to order, store, and recompute them so the numbers mean something over time.
What direct-access lab testing is, how self-ordered blood work is priced, and how to store and interpret your own results as structured data rather than PDFs.
A step-by-step pipeline for small RNA sequencing data, from adapter chemistry and UMI handling through hierarchical annotation, isomiR-level quantification, and differential expression, with the biofluid-specific failure modes that break naive runs.
A working guide to reading, normalizing, filtering, annotating, and querying a Variant Call Format file from whole-genome sequencing, including how to get a slice of it into a spreadsheet without breaking it.
A breakdown of WES pricing: research-grade versus CLIA-clinical, what drives the per-sample cost, how trios are billed, and why the sequencing is the cheap part of the project.
Consumer 30x WGS runs $300-600, clinical CLIA-reported WGS runs roughly $1,000-3,000 out of pocket, and the difference is depth, deliverables, and interpretation rather than sequencing chemistry. Here is how to read a price and verify what you received.
A breakdown of consumer, research-core, and clinical whole genome sequencing prices, what each tier gets you in files and coverage, and where the hidden costs sit.
An essay prompted by Alan Kay's 2011 talk on why the Internet and living systems scale while most software does not, with his numbers checked against real sources, and what the argument implies for a personal molecular data system built from a genome, a blood panel, and a glucose trace.
A working guide to diagnosing and handling batch effects in bulk RNA-seq, from design and TMM normalization through limma, ComBat-seq, and RUV, with the checks that tell you whether the correction helped.
What biological age tests measure, which clocks are worth running, how to compute them yourself from IDATs and a blood panel, and how much of the number is noise.
An essay prompted by Carolyn Bertozzi's lecture on bioorthogonal chemistry: what the reactions are, the second-order rate constants that decide what can be imaged in a living animal, and why the glycome is still a missing data layer in most molecular profiles.
How to use R and Bioconductor as the annotation, statistics, and integration layer for a personal molecular dataset: VCFs, RNA-seq counts, proteomics, and continuous glucose data, with the parts you should not do in R.
Why the data structure chosen for a genome, a variant set, an expression matrix, or a phylogeny decides which questions can be asked and what runs in reasonable time, prompted by a Boundary talk on coding agents.
A technical account of the privacy properties of consumer and clinical DNA testing: what data exists, who holds it, what re-identification attacks work, and a concrete setup for keeping your own sequence data under your control.
In the US, FreeStyle Libre 2 and 3 still require a prescription; Libre Rio, Dexcom Stelo, and Abbott Lingo do not. What each one costs, how they differ, and how to get the raw glucose data out.