Speaker
Description
The Tokenized Representations for Energy-frontier AI Searches via Understanding and Reasoning (TREASURE) project aims to make high-energy physics collider data usable by modern AI methods through standardized tokenized representations. Collider experiments produce rich datasets in different, experiment-specific formats, which limits cross-experiment analysis and the reuse of legacy data. TREASURE addresses this challenge by developing tokenized representations of collider data at multiple levels, including events, jets, particles, tracks, clusters, and hits, and by using these representations to build foundation models that can learn across experiments, detectors, and eras of data. This will enable multi-experiment training, improve sensitivity to rare signals, and help extract new information from current and legacy high-energy physics datasets.
We present a first TREASURE pipeline for LHC open data, which converts experiment-native formats into a common event schema and then into tokenized sequences suitable for transformer-based models. We show initial checks of the tokenized representation, including reconstruction-quality studies and event-level classification benchmarks based on Higgs physics signatures. The goal is to build reusable tokenized collider datasets and shared benchmarks, as a first step toward cross-experiment foundation models for current, legacy, and future high-energy physics experiments.