Research
Spatial molecular technologies report where every measurement came from, and that context changes what the measurement is. Effects that are safely treated as nuisance in dissociated data — library size, background fluorescence, the identity of neighbouring cells — become spatially structured and entangled with the biology. My work builds measurement models that take this seriously, so that what comes out of an analysis is biology rather than an artefact of how the tissue was sampled.
Normalising counts
In spatial transcriptomics the library size at each location is not scattered at random: it varies smoothly across the tissue and tracks the tissue's own structure. Dividing counts by it, as single-cell pipelines do, therefore removes real biological signal along with the technical effect — the result my group demonstrated in Library size confounds biology in spatial transcriptomics data (Genome Biology, 2024).
SpaNorm is the alternative: a gene-wise generalised linear model in which spatially smooth functions — thin-plate splines — represent location-specific size factors, letting spatial variation be decomposed into a library-size-associated (technical) part and a library-size-independent (biological) part. Only the technical part is adjusted away, using percentile-adjusted counts.
SpaNorm: spatially-aware normalization for spatial transcriptomics data (Genome Biology, 2025) · source
Differential expression that depends on the neighbourhood
A cell's behaviour is not a property of the cell alone. The same cell type can respond to a treatment one way when surrounded by immune cells and another way at the edge of a tumour nest, and a conventional differential expression test — which asks only whether a gene changes with the condition — averages those opposite responses into nothing. spiDE asks the sharper question: does the condition effect itself depend on the local neighbourhood?
The diagram on the home page shows how the niche covariates, the model, and the testing procedure fit together.
Background and staining quality in multiplex proteomics
The same principle applies beyond transcriptomics. In multiplex spatial proteomics — Akoya PhenoCycler-Fusion, Cell DIVE, IMC, CosMx — a pixel's fluorescence is a mixture of true antibody binding, non-specific binding, and autofluorescent background. Thresholding to remove the background forces a hard yes-or-no call on every pixel and discards the uncertainty that is actually present.
bgnorm models the whole distribution instead, which turns background correction into an inference problem and yields a quality metric for free.
Kharbanda, Tubelleza et al., 2025 · available for R (bgnormR) and Python (bgnormPy), including a Java-free reader for Akoya QPTIFF, OME-TIFF, and OME-Zarr images.
Interested in this work?
I'm always glad to talk about spatial methods, collaborations, or graduate research projects.