Multi-omics
The integration of data from multiple 'omic' layers — genomics, transcriptomics, proteomics, metabolomics — to build a more complete picture of biological state than any single layer provides.
What it means
Multi-omics refers to the combined analysis of two or more “omic” data types — large-scale molecular measurements that characterize a biological system at a particular level:
| Layer | What is measured | Key technology |
|---|---|---|
| Genomics | DNA sequence, structural variants | Whole-genome sequencing |
| Transcriptomics | RNA expression levels (which genes are active) | RNA-seq |
| Proteomics | Protein abundance and modification state | Mass spectrometry |
| Metabolomics | Small molecule metabolite levels | LC-MS, NMR |
| Epigenomics | DNA methylation, histone modification | ATAC-seq, ChIP-seq |
| Spatial omics | Gene/protein expression with spatial location | Visium, MERFISH |
Each layer captures a different aspect of biology. Genomics tells you what is possible; transcriptomics tells you what is being expressed; proteomics tells you what is actually present and active; metabolomics tells you what the cell is doing. Integration across layers gives a more complete picture of disease, development, or treatment response.
Why integration is hard
Different scales and distributions. Genomic data is binary or categorical (variant present/absent); transcript data is count-based with heavy zero-inflation; proteomic data is continuous with heavy right skew. Combining them requires careful normalization.
Different sample requirements. Collecting transcriptomics, proteomics, and metabolomics from the same patient sample at the same time is technically challenging — different protocols require different sample preparation, and some require tissue amounts that conflict with clinical constraints.
Dimensionality. A single patient’s genome contributes ~4 million common variants; their transcriptome ~20,000 gene expression values; their proteome ~7,000 quantified proteins; their metabolome ~1,000 metabolites. The total feature space vastly exceeds the number of patients in most studies, requiring dimensionality reduction and regularization.
AI methods for multi-omics
Factor analysis and matrix factorization (MOFA, NMF) decompose multi-omics data into shared latent factors that represent biological axes of variation (disease progression, cell type composition, treatment response).
Graph neural networks on multi-omics data represent each sample as a node in a patient similarity network, with edges weighted by molecular similarity.
Deep learning integration — multi-input neural networks with separate encoders for each omic type, combined at a fusion layer. These require large sample sizes to avoid overfitting.
Practical relevance
Multi-omics data from large biobanks (UK Biobank, GTEx, TCGA) is publicly available for secondary analysis. For researchers entering this space, the Bioconductor ecosystem (R) and tools like MOFA+ provide well-documented starting points.