Speaker
Description
The high-energy physics pipeline already resembles a multi-modal foundation model: a general-purpose representation built once and reused across downstream tasks - but it lacks scale and differentiability. Already, we are seeing the gains from large-scale pre-training in individual pieces of the ATLAS pipeline, such as jet flavour identification. Flavour taggers, however, are normally trained in isolation - a separate model from scratch for each task such as small-R and large-R classification - leaving the benefits of scale untapped. In this work, using ATLAS simulation, we show that the representations learned by a small-R tagger, pre-trained on billions of jets, transfer to the large-R H→bb tagging task, demonstrating that the tagger learns general jet features that generalise across radii. This goes beyond H→bb: the same transfer that carries across radii motivates a unification of taggers - a step toward the foundation-model view of the high-energy physics pipeline.