Speaker
Description
Scaling laws have become a central topic in modern machine learning, providing a quantitative understanding of how model performance improves with increasing model size, training data, and compute. They also offer insights into whether learning is approaching the information limits of a given dataset. In this contribution, we present the first study of neural scaling laws for generative jet modeling. Beyond the conventional next-token prediction validation loss, we also investigate the scaling behavior of the sliced Wasserstein distance computed for five high-level jet observables, providing a physics-motivated measure of generative performance. As a function of model size, we observe that both metrics exhibit the expected logarithmic scaling behavior. In contrast, scaling with dataset size and training compute is significantly weaker for both the validation loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss several possible explanations for this observation, including the intrinsic stochasticity of QCD jet formation and the fundamental differences between supervised prediction and generative modeling.