The distribution of neutral hydrogen and its connection to the underlying dark matter field is a powerful cosmological probe, but predicting the observable signal is hindered by the large computational cost of hydrodynamical simulations, particularly on the small, non-linear scales that upcoming intensity-mapping and imaging surveys are most sensitive to. This thesis develops a suite of machine learning methods, built around conditional generative diffusion models, to replace or augment these simulations directly at the field level, across the chain from field generation to field-level inference. We first introduce \LODI, a pipeline combining an attention-based ResUNet (\HALOgen) with a conditional variational diffusion model to generate high-fidelity 21\,cm brightness temperature maps from dark matter-only simulations, recovering the power spectrum to within 10\% for $k\leq10\hMpci$. We then present \FOLD, which removes \LODI's explicit sub-volume overlap prescription in favor of a sliding window diffusion trajectory, generating large, spatially coherent 3D fields from small trained sub-volumes without the boundary artifacts, or the added overlap compute of tiling-based approaches. Trained on $25^3\Mpch$ \CAMELS\ volumes, \FOLD\ generalizes to the $205\Mpch$ \TNG\ box with the power spectrum accurate to within 10\% up to $k\lesssim5\hMpci$, and a positional encoding scheme further improves the recovered bispectrum across several configurations. We also adapt \FOLD\ to generate field level dark matter posteriors conditioned on sparse stellar tracers and an external Planck prior on $\Om$ and $\sige$, injected at inference time via Monte Carlo marginalization with no retraining. This is presented as a first, single field proof of concept demonstration rather than a fully validated result. Next, we develop \FOREST, that rests on the framework of Diffusion Posterior Sampling (DPS), using a latent Diffusion model as a generative prior, for application on the field level line of sight reconstruction of the full physical state of the intergalactic medium from Lyman-$\alpha$ forest sightlines, using a differentiable physics based forward model, and only a single observation of Lyman-$\alpha$ flux. Finally, we develop a fast neural emulator mapping dark matter fields to HI and galaxy fields, conditioned on six astrophysical and cosmological parameters, and embed it within a simulation-based inference pipeline. Moving from power-spectrum to full 3D field-level summaries increases the median figure of merit on $\Om$ and $\sige$ by a factor of $\sim3$, and combining both tracers rather than either alone increases it by a further factor of $2$--$7$ depending on configuration, with relative parameter bias staying below $1.5\%$ throughout. Together, these results show that diffusion based generative models can reproduce cosmological and astrophysical fields with high fidelity at otherwise computationally prohibitive volumes, and, in their most developed form, be turned into practical posterior samplers, recovering cosmological information that is otherwise discarded when a field is compressed to a summary statistic.
Generating the Universe: Field-Level Diffusion Models for Emulation, Reconstruction, and Posterior Inference / Mishra, S.. - (2026 Sep 25).
Generating the Universe: Field-Level Diffusion Models for Emulation, Reconstruction, and Posterior Inference
MISHRA, SATVIK
2026-09-25
Abstract
The distribution of neutral hydrogen and its connection to the underlying dark matter field is a powerful cosmological probe, but predicting the observable signal is hindered by the large computational cost of hydrodynamical simulations, particularly on the small, non-linear scales that upcoming intensity-mapping and imaging surveys are most sensitive to. This thesis develops a suite of machine learning methods, built around conditional generative diffusion models, to replace or augment these simulations directly at the field level, across the chain from field generation to field-level inference. We first introduce \LODI, a pipeline combining an attention-based ResUNet (\HALOgen) with a conditional variational diffusion model to generate high-fidelity 21\,cm brightness temperature maps from dark matter-only simulations, recovering the power spectrum to within 10\% for $k\leq10\hMpci$. We then present \FOLD, which removes \LODI's explicit sub-volume overlap prescription in favor of a sliding window diffusion trajectory, generating large, spatially coherent 3D fields from small trained sub-volumes without the boundary artifacts, or the added overlap compute of tiling-based approaches. Trained on $25^3\Mpch$ \CAMELS\ volumes, \FOLD\ generalizes to the $205\Mpch$ \TNG\ box with the power spectrum accurate to within 10\% up to $k\lesssim5\hMpci$, and a positional encoding scheme further improves the recovered bispectrum across several configurations. We also adapt \FOLD\ to generate field level dark matter posteriors conditioned on sparse stellar tracers and an external Planck prior on $\Om$ and $\sige$, injected at inference time via Monte Carlo marginalization with no retraining. This is presented as a first, single field proof of concept demonstration rather than a fully validated result. Next, we develop \FOREST, that rests on the framework of Diffusion Posterior Sampling (DPS), using a latent Diffusion model as a generative prior, for application on the field level line of sight reconstruction of the full physical state of the intergalactic medium from Lyman-$\alpha$ forest sightlines, using a differentiable physics based forward model, and only a single observation of Lyman-$\alpha$ flux. Finally, we develop a fast neural emulator mapping dark matter fields to HI and galaxy fields, conditioned on six astrophysical and cosmological parameters, and embed it within a simulation-based inference pipeline. Moving from power-spectrum to full 3D field-level summaries increases the median figure of merit on $\Om$ and $\sige$ by a factor of $\sim3$, and combining both tracers rather than either alone increases it by a further factor of $2$--$7$ depending on configuration, with relative parameter bias staying below $1.5\%$ throughout. Together, these results show that diffusion based generative models can reproduce cosmological and astrophysical fields with high fidelity at otherwise computationally prohibitive volumes, and, in their most developed form, be turned into practical posterior samplers, recovering cosmological information that is otherwise discarded when a field is compressed to a summary statistic.| File | Dimensione | Formato | |
|---|---|---|---|
|
PhD Thesis Satvik.pdf
accesso aperto
Tipologia:
Tesi
Licenza:
Non specificato
Dimensione
69.55 MB
Formato
Adobe PDF
|
69.55 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


