RNA structures and interactions are often accessible only through indirect, noisy measurements. Principled models of how observations arise allow data-driven regression to infer not only measured labels but also latent biological quantities. This thesis develops such physical and statistical frameworks for RNA secondary structure, RNA--protein interactions, and siRNA potency. MERGE-RNA describes RNA structural ensembles by modelling the physics of dimethyl sulfate (DMS) probing. It learns transferable, interpretable parameters shared across molecules, probe concentrations, and replicates, while maximum entropy minimally adjusts thermodynamic populations to the data. Across systems and probe concentrations, inferred ensembles recapitulate measured DMS reactivity better than the traditional pseudo-free-energy model. On an adenine riboswitch, it recovers NMR-resolved conformations and ligand-induced rearrangement, with a structural midpoint matching the NMR-derived Kd; on a designed RNA, it resolves strand-displacement intermediate populations. For RNA-binding proteins, a generative model uses frozen foundation-model representations to predict eCLIP read counts and quantify fold-enrichment uncertainty under a shared input control. IP replicate tests support near-Poisson noise, while held-out SMInput genes validate the control rate. This separates reproducible signal from IP and control noise and estimates shared-control covariance. For Staufen-2 in HepG2 at 125-nucleotide resolution, counting noise accounts for 78% of fold-enrichment variance; control correction reduces replicate correlation from 0.54 to 0.20 and yields an honest prediction ceiling of 0.46. Though experiment- and resolution-dependent, this decomposition and correction apply whenever eCLIP replicates share a control. Silscore predicts siRNA potency from a position-resolved nearest-neighbour stacking-energy profile, a quadratic duplex-stability preference, and sparse identity terms selected from training residuals. Its 26-parameter main fit uses fewer than half the parameters of any trainable comparator. It has the highest Pearson correlation point estimate in three transfer evaluations, with a statistically supported advantage over every applicable comparator on the Katoh--Suzuki panel. When evaluated on a fixed test set after a change in the provenance of the training sequences, Silscore retains the highest correlation and shows the smallest decrease among fitted models. Its transferable, interpretable parameters formalise established rules within a principled data-driven framework, recovering an intermediate stability optimum, terminal asymmetry, and regional preferences for weaker seed and stronger central pairing. Together, these studies show how principled data-driven frameworks make latent quantities explicit and testable, with transfer and independent measurements probing whether they capture RNA biology.
Data-driven models for RNA: physical and statistical approaches to secondary structure, RNA–protein interactions, and siRNA potency / Sacco, G.. - (2026 Sep 28).
Data-driven models for RNA: physical and statistical approaches to secondary structure, RNA–protein interactions, and siRNA potency
SACCO, GIUSEPPE
2026-09-28
Abstract
RNA structures and interactions are often accessible only through indirect, noisy measurements. Principled models of how observations arise allow data-driven regression to infer not only measured labels but also latent biological quantities. This thesis develops such physical and statistical frameworks for RNA secondary structure, RNA--protein interactions, and siRNA potency. MERGE-RNA describes RNA structural ensembles by modelling the physics of dimethyl sulfate (DMS) probing. It learns transferable, interpretable parameters shared across molecules, probe concentrations, and replicates, while maximum entropy minimally adjusts thermodynamic populations to the data. Across systems and probe concentrations, inferred ensembles recapitulate measured DMS reactivity better than the traditional pseudo-free-energy model. On an adenine riboswitch, it recovers NMR-resolved conformations and ligand-induced rearrangement, with a structural midpoint matching the NMR-derived Kd; on a designed RNA, it resolves strand-displacement intermediate populations. For RNA-binding proteins, a generative model uses frozen foundation-model representations to predict eCLIP read counts and quantify fold-enrichment uncertainty under a shared input control. IP replicate tests support near-Poisson noise, while held-out SMInput genes validate the control rate. This separates reproducible signal from IP and control noise and estimates shared-control covariance. For Staufen-2 in HepG2 at 125-nucleotide resolution, counting noise accounts for 78% of fold-enrichment variance; control correction reduces replicate correlation from 0.54 to 0.20 and yields an honest prediction ceiling of 0.46. Though experiment- and resolution-dependent, this decomposition and correction apply whenever eCLIP replicates share a control. Silscore predicts siRNA potency from a position-resolved nearest-neighbour stacking-energy profile, a quadratic duplex-stability preference, and sparse identity terms selected from training residuals. Its 26-parameter main fit uses fewer than half the parameters of any trainable comparator. It has the highest Pearson correlation point estimate in three transfer evaluations, with a statistically supported advantage over every applicable comparator on the Katoh--Suzuki panel. When evaluated on a fixed test set after a change in the provenance of the training sequences, Silscore retains the highest correlation and shows the smallest decrease among fitted models. Its transferable, interpretable parameters formalise established rules within a principled data-driven framework, recovering an intermediate stability optimum, terminal asymmetry, and regional preferences for weaker seed and stronger central pairing. Together, these studies show how principled data-driven frameworks make latent quantities explicit and testable, with transfer and independent measurements probing whether they capture RNA biology.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_PhD.pdf
embargo fino al 27/09/2027
Descrizione: PDF tesi
Tipologia:
Tesi
Licenza:
Non specificato
Dimensione
3.65 MB
Formato
Adobe PDF
|
3.65 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


