Modern sequencing technologies have made it possible to characterize individual cells along multiple molecular axes, such as gene expression, chromatin accessibility, surface protein abundance, at once, and recently to do so while retaining each cell's position within a tissue. As a result, the limiting factor in biological discovery is no longer data acquisition but interpretation: the resulting measurements are sparse, high-dimensional, technically heterogeneous, and frequently incomplete, with different modalities observed for different subsets of cells. Variational Autoencoders (VAEs) have become a standard tool for this setting, as they combine flexible neural-network likelihoods with a probabilistic treatment of uncertainty and can scale, through amortized inference and mini-batching, to atlases of hundreds of thousands of cells; multimodal VAEs are a natural extension to the multiomics case. This thesis revisits two problematic aspects of such algorithms. The first concerns how multimodal VAEs combine modalities. Most existing architectures, whether based on Product-of-Experts, Mixture-of-Experts, or a joint encoder, summarize all modalities into a single shared latent representation, from which every modality is then decoded. We show that this collapses the joint statistical structure of the data: generated modalities become deterministically related, carrying maximal mutual information regardless of the correlation actually present, and uncertainty is systematically underestimated when a modality is missing or predicted from the others. This behavior is a structural consequence of the aggregation strategy, and it implies that these models fail to provide faithful estimates of uncertainties, especially for the conditional distributions. We introduce Correlated VAEs (CoVAE), which encode each modality in its own latent space and represent cross-modality dependence explicitly through a learned, non-diagonal prior over the concatenated latent variables. CoVAE is the only architecture to properly generalize a probabilistic version of Canonical Correlation Analysis (CCA); on synthetic benchmarks with controlled correlation, CoVAE is the only model tested that recovers the true correlation and propagates it correctly into predictive uncertainty; on a pan-cancer mRNA/miRNA dataset it remains competitive with established multimodal VAEs across joint and conditional tasks. The second issue concerns the decoder and its interpretability. We developed TopOmics, a modular framework that extends topic modeling to essentially any combination of single-cell and spatial 'omics modalities, with several options pre-implemented for spatial integration, multiomic integration, observation models and priors over topics. Because inference is amortized, TopOmics is a multimodal VAE whose decoder is constrained to an interpretable topic structure rather than a black-box multilayer perceptron. We show that this constraint has little effect on the model performance: TopOmics matches or exceeds fully black-box competitors such as MultiVI, SpatialGLUE and COSMOS, measured by agreement between latent space clusters and ground truth annotations, across three structurally distinct datasets: spatial transcriptome–chromatin co-profiling of mouse brain, a subcellular-resolution VisiumHD colorectal cancer sample, and a three-modality TEA-seq PBMC dataset. In each of these tests, TopOmics recovers directly interpretable topics, and its mini-batched inference engine scales to dataset sizes at which competing implementations fail outright. Both contributions share a common thread: they modify a standard component of multimodal VAEs, specifically the aggregation strategy in one case, the decoder in the other, to recover a relevant property for scientific use, namely correct correlation and interpretability, respectively, at limited cost to the model's other capabilities.
Multimodal Variational Autoencoders: from Correlated Latent Spaces to Interpretable Multi-Omic Models / Caretti, F.. - (2026 Sep 30).
Multimodal Variational Autoencoders: from Correlated Latent Spaces to Interpretable Multi-Omic Models
CARETTI, FEDERICO
2026-09-30
Abstract
Modern sequencing technologies have made it possible to characterize individual cells along multiple molecular axes, such as gene expression, chromatin accessibility, surface protein abundance, at once, and recently to do so while retaining each cell's position within a tissue. As a result, the limiting factor in biological discovery is no longer data acquisition but interpretation: the resulting measurements are sparse, high-dimensional, technically heterogeneous, and frequently incomplete, with different modalities observed for different subsets of cells. Variational Autoencoders (VAEs) have become a standard tool for this setting, as they combine flexible neural-network likelihoods with a probabilistic treatment of uncertainty and can scale, through amortized inference and mini-batching, to atlases of hundreds of thousands of cells; multimodal VAEs are a natural extension to the multiomics case. This thesis revisits two problematic aspects of such algorithms. The first concerns how multimodal VAEs combine modalities. Most existing architectures, whether based on Product-of-Experts, Mixture-of-Experts, or a joint encoder, summarize all modalities into a single shared latent representation, from which every modality is then decoded. We show that this collapses the joint statistical structure of the data: generated modalities become deterministically related, carrying maximal mutual information regardless of the correlation actually present, and uncertainty is systematically underestimated when a modality is missing or predicted from the others. This behavior is a structural consequence of the aggregation strategy, and it implies that these models fail to provide faithful estimates of uncertainties, especially for the conditional distributions. We introduce Correlated VAEs (CoVAE), which encode each modality in its own latent space and represent cross-modality dependence explicitly through a learned, non-diagonal prior over the concatenated latent variables. CoVAE is the only architecture to properly generalize a probabilistic version of Canonical Correlation Analysis (CCA); on synthetic benchmarks with controlled correlation, CoVAE is the only model tested that recovers the true correlation and propagates it correctly into predictive uncertainty; on a pan-cancer mRNA/miRNA dataset it remains competitive with established multimodal VAEs across joint and conditional tasks. The second issue concerns the decoder and its interpretability. We developed TopOmics, a modular framework that extends topic modeling to essentially any combination of single-cell and spatial 'omics modalities, with several options pre-implemented for spatial integration, multiomic integration, observation models and priors over topics. Because inference is amortized, TopOmics is a multimodal VAE whose decoder is constrained to an interpretable topic structure rather than a black-box multilayer perceptron. We show that this constraint has little effect on the model performance: TopOmics matches or exceeds fully black-box competitors such as MultiVI, SpatialGLUE and COSMOS, measured by agreement between latent space clusters and ground truth annotations, across three structurally distinct datasets: spatial transcriptome–chromatin co-profiling of mouse brain, a subcellular-resolution VisiumHD colorectal cancer sample, and a three-modality TEA-seq PBMC dataset. In each of these tests, TopOmics recovers directly interpretable topics, and its mini-batched inference engine scales to dataset sizes at which competing implementations fail outright. Both contributions share a common thread: they modify a standard component of multimodal VAEs, specifically the aggregation strategy in one case, the decoder in the other, to recover a relevant property for scientific use, namely correct correlation and interpretability, respectively, at limited cost to the model's other capabilities.| File | Dimensione | Formato | |
|---|---|---|---|
|
caretti_thesis.pdf
accesso aperto
Tipologia:
Tesi
Licenza:
Non specificato
Dimensione
22.1 MB
Formato
Adobe PDF
|
22.1 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


