Research

We work on the edge cases in bioinformatics — problems that cannot be satisfactorily solved by ‘standard’ workflows. Fortunately for us, bioinformatics has a rather large edge, large enough for us to dance on.

New questions

Can we reconstruct cell lineages from sparse multiomics data?

Cells that descend from the same ancestor share mutations, so in principle genetic variation can be used to reconstruct cellular ‘family trees’. In practice, single-cell experiments only reveal a tiny and highly non-random fraction of this information: one measurement may observe a mutation that another completely misses.

This creates an interesting statistical problem. Can we infer hidden clones from many sparse and biased observations, especially when different data types reveal different parts of the same underlying genotype? We are interested in models that combine partially overlapping information from single-cell RNA and chromatin accessibility (ATAC) data to recover clonal structure, quantify uncertainty and determine when such reconstruction is actually possible.

Biologically this project may lead to a method that infers cell lineages, such as cancer clones, from public scRNA-seq, scATAC-seq and multiome (scRNA+ATAC-seq) data.

What’s so special about spatial? Effective design of spatial omics experiments

The field of spatial biology has entered a new phase: researchers are beginning to ask what spatial omics offers beyond single-cell omics and low-throughput imaging. This leads to two sets of questions. First, when do we genuinely benefit from spatial information? What do spatial markers look like, and how can we make full use of tissue location, cellular neighbourhoods and tissue architecture? Second, when do we genuinely need high-throughput measurements? Which questions can be answered only with high-dimensional spatial omics, rather than conventional lower-dimensional approaches such as H&E imaging and immunostaining? These questions are central to experimental design. Spatial omics remains expensive, so understanding when spatial information and high dimensionality add value will help us use limited budgets more effectively, potentially by combining complementary technologies.

Sample size calculation for single-cell? What a horrifying idea!

What counts as a ‘sample’ in a single-cell study? Suppose we recruit six patients with lung cancer, collect tumour and adjacent normal tissue from each patient, and measure three million cells. The dataset contains six patients, twelve specimens and three million cells — but which number matters depends on the biological question and the level at which we want to draw conclusions.

These observational units are not independent: cells from the same specimen are usually more similar than cells from different patients. The effective sample size may therefore be much smaller than the number of cells, and ignoring this dependence can lead to pseudo-replication and exaggerated statistical significance. We are developing principled, computationally efficient ways to model this dependence, including meta-cells and spatial autocorrelation—to support sample-size and power calculations for single-cell and spatial omics studies.

Ongoing projects (methodology)

How many genes should be used for Gene Ontology analysis?

Gene Ontology analysis is widely used in bioinformatics to interpret lists of genes by identifying biological processes that are over-represented. For example, after comparing two biological conditions, researchers may take a list of differentially expressed genes and ask whether the list is enriched for terms such as immune response, cell cycle or metabolism.

A key practical question is how many genes should be included in the analysis. In many studies, this choice is made arbitrarily, such as using the top 50. However, the number of selected genes can strongly affect the resulting Gene Ontology terms and therefore the biological interpretation.

This project will explore how Gene Ontology results change as the number of selected genes varies. The student will use public gene-expression data, generate ranked gene lists, run enrichment analyses across a range of gene-list sizes, and compare the stability, specificity and interpretability of the resulting GO terms. The aim is to develop practical guidance for choosing gene-list sizes in exploratory bioinformatics analyses.

Cell states on a continuum

Many biological systems cannot be described well by discrete cell-type labels alone. We develop methods for representing continuous cell states and identifying state variation in single-cell and spatial omics studies.

Representative work:

  • Φ-Space: continuous phenotyping of single-cell multi-omics data.
  • Φ-Space ST: a platform-agnostic method to identify cell states in spatial transcriptomics studies.

Φ-Space is essentially supervised dimension reduction (thousands of genes –> dozens of cell types), hence Φ-Space is very flexible. We are interested in extending this framework to solve problems where mixture of cell types is an issue (e.g. doublet deconvolution).

What’s in a doublet? Doublet deconvolution

Doublets are a common by-product of single-cell protocols: the tissue sample is expected to be dissolved and separated into individual cells, but some cells may fail to be separted, resulting in doublets (multiplets are also possible but they are much rarer than doublets). As a result, some observational units in the single-cell matrix might not be single cells, but mixture of cells.

Φ-Space provides a solution to finding out what types of cells are present in doublets. Existing tools such as scDblFinder only classify doublets, differentiating them from singlets (single cells), but do not identify what cell types are present in doublets. Our Φ-Space-based doublet detection and deconvolution provides additional layer of information to doublet detection, helping researchers to make informed decisions about doublet processing.

Harmony in diversity: shared, partially shared and individual variations in multiomics

Biomedical studies increasingly combine gene expression, histology, proteomics, epigenomics, microbiome profiles, clinical records, and other data types. Most multiomics integration tools tell you what factors are shared by different omics layer some also tell you what factors are individual to each omics layer; but they do not tell you what is partially shared. Biologically jointly shared variations are likely to represent biological pathways that involve all omics layers.

However, as we increase the number of omics layers, we do not expect to see many factors that are common to all of them. What is partially shared and what is not shared (individual to omics layers) become more important.

Collaborating with Prof J S Marron’s group at UNC Chapel Hill, we have developed the DIVAS R package, implementing Marron’s DIVAS method. DIVAS is a powerful matrix factorisation tool for sort all the signal components in multiomics data into jontly shared, partially shared and individual ones.

Representative work:

  • DIVAS Software (RA Yinuo Sun, collaboration with Prof J S Marron at UNC Chapel Hill)
  • Integration of histological images and gene expression data (PhD student Chengyi Ma).

Gene regulation and causal discovery

Representative work:

  • NeighbourNet for scalable cell-specific coexpression networks (Yidi Deng).
  • StableMate (Yidi Deng).

Ongoing projects (biology)

Clonal tracing in cancer and blood cell formation (haematopoiesis)

Cancer is an evolutionary process. Cancer cells acquire different mutations, compete for resources and leave descendants, allowing some populations to undergo ‘clonal expansion’. A clone is a group of cells descended from the same ancestral cell and carrying a shared genetic marker or lineage label. Because treatment may eliminate some clones while others survive and expand, the clonal composition of a cancer can change over time. This ongoing evolution is one reason why cancer can be difficult to cure.

Blood formation, or haematopoiesis, is also a clonal process: a relatively small number of stem and progenitor cells continually produce the many mature cells in our blood. We use clonal tracing data and dynamic models to reconstruct these lineages and quantify how clones emerge, persist, compete and respond to intervention. Our projects apply this framework to questions such as detecting and understanding minimal residual disease after cancer treatment, and tracking how donor stem-cell clones establish and maintain blood production following stem-cell transplantation.

Mapping the spatio-temporal metabolic atlas of aging and obesity

This KU Leuven-Melbourne Joint PhD Program supports two PhD students on the spatio-temporal metabolic changes that accompany aging and obesity. The project connects expertise in metabolism, spatial biology, and computational analysis across the Fendt lab at KU Leuven and the Lê Cao lab at the University of Melbourne. As a co-PI at Melbourne, Jiadong contributes to the Melbourne supervision team, helping guide the statistical and computational side of the project.

Finding the missing causes of early-onset breast cancer

This Cancer Australia Research Initiative project, led by A/Prof Shuai Li at the School of Population and Global Health, University of Melbourne, investigates why breast cancer is increasingly diagnosed in women before age 50 and why much of the risk remains unexplained. As an Associate Investigator, Jiadong contributes expertise in multi-omics integration, with a focus on discovering disease mechanisms and risk markers.