24–28 Aug 2026
Kirchhoff Institute for Physics (KIP)
Europe/Berlin timezone

Foundation Models for Anomaly Detection in HEP

26 Aug 2026, 14:20
8m
1.404

1.404

Foundation Models 🔀 Foundation Models

Speaker

Desheng Yang (Sapienza Università di Roma and INFN sezione di Roma)

Description

Foundation models have recently emerged as powerful tools for learning general-purpose representations that can outperform task-specific models while being generalizable to multiple downstream tasks. In this project, we explore their application to particle physics by training a foundation model on simulated jet data from hadronic collisions. The goal is to build a model that can learn physically meaningful representations of collision events and to use the model as a general tool for different tasks such as anomaly detection, jet reconstruction, and jet tagging. To this end, a foundation model based on the Perceiver architecture was developed and trained on the CoLLiDe-2V dataset, which provides simulations of hadronic background processes in Large Hadron Collider (LHC) detectors. By utilizing self-attention and cross-attention mechanisms, the model processes all jets and their constituents within an event simultaneously to capture inter-jet correlations. Crucially, it maps these tokens onto a fixed-size latent space, maintaining a linear computational complexity relative to the number of input objects, which is essential for scaling to High-Luminosity LHC (HL-LHC) pile-up conditions. The model was pretrained without labeled data via self-supervised learning, using multiple parallel decoding heads tasked with masked-object reconstruction, event-level property prediction, and consistency constraints. After the model was trained on QCD, W and Z into hadronic events, it was tested for anomaly detection using tt̅into hadronic as the signal events. To evaluate the learned representation, a Variational Autoencoder and a Logistic Regression were applied both to the Perceiver's latent space and directly to the raw data. Probing the latent space yielded a competitive performance compared to their raw-data baselines. Transfer learning experiments further validated the representation, showing that a frozen pretrained Perceiver converges faster and achieves comparative performance to training an anomaly detection model from scratch. Finally, an explainability study using Linear Probes and Gaussian Mixture Models (GMM) confirmed that the latent space does not rely on trivial correlations; instead, it geometrically separates anomalies from the background and explicitly encodes physically meaningful quantities, including jet pT, jet mass, and complex inter-object kinematic relationships.
These results demonstrate that a Perceiver-based foundation model can successfully learn robust, transferable, and physically interpretable representations of collision events without supervision, outperforming conventional task-specific baselines. If extended to the full Standard Model, this approach could provide a model-agnostic way to search for new physics over a wider range than conventional methods.

Authors

Desheng Yang (Sapienza Università di Roma and INFN sezione di Roma) Giuliano Gustavino (Sapienza Università di Roma and INFN Sezione di Roma) Stefano Giagu (Sapienza Università di Roma and Istituto Nazionale di Fisica Nucleare) Valerio Ippolito (INFN Sezione di Roma)

Presentation materials