This thesis develops predictive theories for neural networks in high-dimensional settings, focusing on three problems of particular relevance to modern machine learning: generalization, transfer learning, and neural scaling laws. We first study the generalization error of two-layer neural networks with generic activation functions in a setting closer to empirical risk minimization, rather than restricting the analysis to an optimal statistical benchmark, and we characterize how the regularization strength and the ratio between the number of training samples and the number of trainable parameters shape the equilibrium solutions and the generalization error they achieve. For networks with Erf and ReLU activations, we identify phase transitions at which an equilibrium solution carrying non-trivial correlations with the ground-truth directions becomes optimal, and we show that the predicted generalization error is recovered by gradient-based learning algorithms. We then investigate transfer learning through Low-Rank Adaptation (LoRA), with particular emphasis on two of its essential ingredients, namely the pretrained weights and the rank of the low-rank update. Our analysis reveals a separation between two stages of the learning dynamics: an initial search phase, whose duration is primarily controlled by the pretrained weights and which ends once the network develops non-trivial correlations with the directions relevant to the downstream task, followed by a convergence phase in which the final generalization performance is controlled by the available adaptation rank. Finally, we study neural scaling laws in a controlled solvable setting and construct a model that reproduces the power-law decay of the error with model size observed empirically in large language models. We show that, when the model is optimally calibrated, the way the optimal norm of the learned directions varies with model width is by itself sufficient to generate the observed power law, and, beyond the exponent, we identify low-order geometric properties of the learned directions as the quantities controlling the scaling coefficient. This provides a mechanism for understanding how different learning algorithms can exhibit the same scaling exponent while reaching systematically different levels of performance through different scaling coefficients.
Theory of Learning in High-Dimensional Neural Networks: Generalization, Transfer Learning, and Neural Scaling Laws / Nwemadji Tiako, A.G.. - (2026 Sep 30).
Theory of Learning in High-Dimensional Neural Networks: Generalization, Transfer Learning, and Neural Scaling Laws
NWEMADJI TIAKO, ARSENE GIBBS
2026-09-30
Abstract
This thesis develops predictive theories for neural networks in high-dimensional settings, focusing on three problems of particular relevance to modern machine learning: generalization, transfer learning, and neural scaling laws. We first study the generalization error of two-layer neural networks with generic activation functions in a setting closer to empirical risk minimization, rather than restricting the analysis to an optimal statistical benchmark, and we characterize how the regularization strength and the ratio between the number of training samples and the number of trainable parameters shape the equilibrium solutions and the generalization error they achieve. For networks with Erf and ReLU activations, we identify phase transitions at which an equilibrium solution carrying non-trivial correlations with the ground-truth directions becomes optimal, and we show that the predicted generalization error is recovered by gradient-based learning algorithms. We then investigate transfer learning through Low-Rank Adaptation (LoRA), with particular emphasis on two of its essential ingredients, namely the pretrained weights and the rank of the low-rank update. Our analysis reveals a separation between two stages of the learning dynamics: an initial search phase, whose duration is primarily controlled by the pretrained weights and which ends once the network develops non-trivial correlations with the directions relevant to the downstream task, followed by a convergence phase in which the final generalization performance is controlled by the available adaptation rank. Finally, we study neural scaling laws in a controlled solvable setting and construct a model that reproduces the power-law decay of the error with model size observed empirically in large language models. We show that, when the model is optimally calibrated, the way the optimal norm of the learned directions varies with model width is by itself sufficient to generate the observed power law, and, beyond the exponent, we identify low-order geometric properties of the learned directions as the quantities controlling the scaling coefficient. This provides a mechanism for understanding how different learning algorithms can exhibit the same scaling exponent while reaching systematically different levels of performance through different scaling coefficients.| File | Dimensione | Formato | |
|---|---|---|---|
|
Gibbs_Nwemadji_PhD_thesis_2026.pdf
embargo fino al 28/02/2027
Tipologia:
Tesi
Licenza:
Non specificato
Dimensione
6.97 MB
Formato
Adobe PDF
|
6.97 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


