24–28 Aug 2026
Kirchhoff Institute for Physics (KIP)
Europe/Berlin timezone

A Recipe for Scale: Hyperparameter-Optimal Scaling Laws End-to-End, from Analytic Toy Problems to Jet Tagging

25 Aug 2026, 11:00
8m
1.404

1.404

Foundation Models 🔀 Foundation Models

Speaker

Matthias Vigl (TUM)

Description

Recent progress in machine learning has come as much from scale as from architecture, and high-energy physics is starting to follow: work on transformer scaling for flavour tagging in ATLAS [https://inspirehep.net/literature/3114069] has shown that tagger performance follows predictable power laws over orders of magnitude in compute on a multi-billion-jet dataset. This matters particularly in a field built on domain-specific inductive bias: physics-motivated features, symmetry-aware architectures, auxiliary objectives, whose benefits are almost always established at a modest scale. Which of them still pay off as compute grows is therefore an open question, and answering it requires scaling laws that reflect genuine method differences rather than mistuned training dynamics.
We present a recipe that scales every tunable axis of the system optimally. We first validate the full life of a scaling law: onset, power-law regime, and irreducible floor, on domain-agnostic toy problems where the floor is known analytically, and finally apply the complete procedure to multi-task transformers for jet flavour tagging.
Tuned this way, design choices can be compared fairly: we quantify how auxiliary objectives and input representations affect both the scaling of the loss with compute and the optimal scaling of the hyperparameters themselves. We further show that the onset of the power-law regime is itself set by scale: below a threshold in model size and training tokens the loss carries no information about high-compute scaling, though fits in that region will still return a biased exponent, showing the need for large open datasets of high-quality full simulation.

Authors

Alexander Froch (Université de Genève) Dan Guest (Humboldt University of Berlin) Jackson Barr (UCL) Lukas Heinrich (Technical University of Munich) Matthias Vigl (TUM) Michael Kagan (SLAC National Accelerator Laboratory) Nicole Hartman (TUM) Nikita Pond (UCL)

Presentation materials