Research

Spatial molecular technologies report where every measurement came from, and that context changes what the measurement is. Effects that are safely treated as nuisance in dissociated data — library size, background fluorescence, the identity of neighbouring cells — become spatially structured and entangled with the biology. My work builds measurement models that take this seriously, so that what comes out of an analysis is biology rather than an artefact of how the tissue was sampled.

Normalising counts

In spatial transcriptomics the library size at each location is not scattered at random: it varies smoothly across the tissue and tracks the tissue's own structure. Dividing counts by it, as single-cell pipelines do, therefore removes real biological signal along with the technical effect — the result my group demonstrated in Library size confounds biology in spatial transcriptomics data (Genome Biology, 2024).

SpaNorm is the alternative: a gene-wise generalised linear model in which spatially smooth functions — thin-plate splines — represent location-specific size factors, letting spatial variation be decomposed into a library-size-associated (technical) part and a library-size-independent (biological) part. Only the technical part is adjusted away, using percentile-adjusted counts.

Observed counts library size varies with tissue biology Standard normalisation divide counts by library size Biology removed too SpaNorm spatially-smooth GLM technical biology remove technical, keep biology Biology retained
In spatial data the library size at each location is spatially smooth and correlated with the tissue's biology, so scaling counts by it — the single-cell default — strips real biological signal along with the technical effect. SpaNorm instead fits a generalised linear model in which each gene's counts are decomposed into a library-size-dependent (technical) component and a library-size-independent (biological) component, both varying smoothly across space via thin-plate splines, and adjusts away only the technical part.

SpaNorm: spatially-aware normalization for spatial transcriptomics data (Genome Biology, 2025) · source

Differential expression that depends on the neighbourhood

A cell's behaviour is not a property of the cell alone. The same cell type can respond to a treatment one way when surrounded by immune cells and another way at the edge of a tumour nest, and a conventional differential expression test — which asks only whether a gene changes with the condition — averages those opposite responses into nothing. spiDE asks the sharper question: does the condition effect itself depend on the local neighbourhood?

The diagram on the home page shows how the niche covariates, the model, and the testing procedure fit together.

Background and staining quality in multiplex proteomics

The same principle applies beyond transcriptomics. In multiplex spatial proteomics — Akoya PhenoCycler-Fusion, Cell DIVE, IMC, CosMx — a pixel's fluorescence is a mixture of true antibody binding, non-specific binding, and autofluorescent background. Thresholding to remove the background forces a hard yes-or-no call on every pixel and discards the uncertainty that is actually present.

bgnorm models the whole distribution instead, which turns background correction into an inference problem and yields a quality metric for free.

log₂ pixel intensity density background non-specific binding biological signal observed JSD → staining quality bgnorm deconvolution · keeps the signal component only · soft per-pixel correction · no hard thresholding · flags weak staining via JSD
bgnorm treats the observed distribution of log-transformed pixel intensities as a three-component mixture. Deconvolving it isolates the biological signal component, so each pixel is corrected in proportion to how likely it is to carry real signal rather than by a hard threshold. The Jensen-Shannon divergence between the non-specific and signal components measures how well a marker separated from its background, giving an automated staining-quality check across markers, samples, and tissue slices.

Kharbanda, Tubelleza et al., 2025 · available for R (bgnormR) and Python (bgnormPy), including a Java-free reader for Akoya QPTIFF, OME-TIFF, and OME-Zarr images.

Interested in this work?

I'm always glad to talk about spatial methods, collaborations, or graduate research projects.

Email me