EuCAIFCon 2026

Europe/Berlin
Foyer (Kirchhoff Institute for Physics (KIP))

Foyer

Kirchhoff Institute for Physics (KIP)

Im Neuenheimer Feld 227 69120 Heidelberg Germany
Description

The third “European AI for Fundamental Physics Conference” (EuCAIFCon) will be held in Heidelberg, from 24th to 28th of August 2026. The aim of this event is to provide a platform for establishing new connections between AI activities across various branches of fundamental physics, by bringing together researchers that face similar challenges and/or use similar AI solutions. The conference will be organized “horizontally”: sessions are centered on specific AI methods and themes, while being cross-disciplinary regarding the scientific questions.


The conference fee amounts to 100€. It covers all the coffee breaks, 3 lunches (Tuesday, Wednesday, Thursday) and the reception on Monday evening. The fee is to be paid either in cash when arriving or via wire transfer (detail can be found in the registration form)

Thanks to the generous support of The European Physical Journal, we are able to waive the conference fee for up to 10 PhD or Undergrad Students. If funding is a barrier to your participation, please get in touch with the organizers via email. 

EuCAIFCon 2026 is organised by EuCAIF (https://eucaif.org/).

Posters
We will have two dedicated poster sessions and the posters will remain up during the whole conference. We can accommodate poster sizes up to A0 portrait, please do not exceed 84 cm width, 119 cm height. Additionally, we will hold a vote for the best posters, the winners of which will be invited to give a plenary highlight talk about their work on Friday. Copy shops are available in Heidelberg, in walking distance from the venue
 

Plenary Speakers:

  • Adrian Oeftiger (University of Oxford)
  • Dario Hügel (Barra Labs)
  • Eilam Gross (Weizmann Institute of Science)
  • Emille Ishida (Clermont Auvergne)
  • Jan Kieseler (KIT)
  • Johann Brehmer (CuspAI)
  • Lucie Flek  (University of Bonn)
  • Luigi DelDebbio (University of Edinburgh)
  • Mario Krenn (University of Tübingen)
  • Maximilian Dax (ELLIS Institute Tübingen)
  • Michelle Kuchera (Florida State University)
  • Nicole Hartman (TUM)
  • Ramon Winterhalder (University of Milan)
  • Roberto Trotta (SISSA)
  • Tristan Bereau (University of Heidelberg)

 

Local organizing committee:

  • Henning Bahl (bahl@thphys.uni-heidelberg.de)
  • Ranit Das (das@thphys.uni-heidelberg.de)
  • Sascha Diefenbacher (diefenbacher@thphys.uni-heidelberg.de)
  • Caroline Heneka (heneka@thphys.uni-heidelberg.de)
  • Marvin Kohls (m.kohls@gsi.de)
  • Johan Messchendorp (j.messchendorp@gsi.de)
  • Ayodele Ore (ore@thphys.uni-heidelberg.de)
  • Tilman Plehn (plehn@uni-heidelberg.de)

 

 

Registration
Registration
Participants
    • 13:00
      Registration Goldbox

      Goldbox

    • 1
      👋 Welcome HS1

      HS1

    • 🗣️ Plenaries HS1

      HS1

      Convener: Tilman Plehn (Institut für Theoretische Physik, Universität Heidelberg; Interdisciplinary Center for Scientific Computing (IWR), Universität Heidelberg)
      • 2
        From Classifiers to Detectors: Learning Closer to the Data
        Speaker: Jan Kieseler (KIT)
      • 3
        Towards an Artificial Muse for new Ideas in Science
        Speaker: Mario Krenn (University of Tübingen)
    • 15:30
      ☕ Coffee Break Foyer

      Foyer

    • 🔀 Explainability & Theory 1.404

      1.404

      Convener: Gert Aarts (Universitaet Bielefeld)
      • 4
        The Latent Information Geometry of Jet Classification

        Latent representations are an important theme in modern machine learning. Networks trained with a notion of locality encode task-specific similarity as closeness in the latent space. We analyze this latent information for a variational autoencoder with a classifier head using tools from differential geometry, specifically information geometry. Here, we explore the learned latent space using concepts such as geodesic distances and curvature and discuss their significance for the induced decoder and classifier geometries. To link the structure of the latent space to the physics features of the data, we employ the metric tensor induced by the likelihood of the classifier. Furthermore, we transfer the concept of nonmetricity scalars to information geometry and find that they constitute new, coordinate-invariant measures of class separation. We then apply our new methodology to LHC data to understand the physics behind binary quark-gluon classification and three-fold fat jet tagging. Here, information geometry tells us which features a decision is dominantly based on, increasing the interpretability of the network output.

        Speaker: Rebecca Maria Kuntz (Zentrum für Astronomie der Universität Heidelberg, Astronomisches Rechen-Institut)
      • 5
        From Information Geometry to Jet Substructure: A Triality of Cumulant Tensors, Energy Correlators, and Hypergraphs

        Pairwise Fisher graphs capture local covariance information, but they cannot distinguish an irreducible multi-observable radiation pattern from a collection of ordinary pairwise correlations. We show that the higher Fisher tensors supply this missing structure. In a finite basis of binned EECs, ECFs, or EFPs, and in the natural exponential-family coordinates generated by that basis, the same local tensor has three equivalent interpretations: it is a coefficient in the local Kullback–Leibler expansion, a connected cumulant of the chosen correlator observables, and a signed weight on a hyperedge linking those observables. This gives an exact Fisher–correlator–hypergraph triality in the local exponential-family embedding.

        The triality provides a direct construction of physics-informed hypergraphs from measured or simulated correlator data. Extending the quadratic Fisher matrix to the first non-trivial higher tensor identifies connected multi-observable radiation patterns and supplies hyperedge weights for higher-order Laplacians and message passing. It also gives a criterion for compressing observable bases beyond pairwise information. We develop these constructions and explain why the exact cumulant interpretation is specific to natural exponential-family coordinates.

        We illustrate the framework in four applications. In a minimal local-KL study, including the cubic Fisher tensor reduces the KL truncation error by about 30× near the reference point and isolates the dominant triplet structure. In a W → q<span style="text-decoration: overline;">q</span> versus t → bq<span style="text-decoration: overline;">q</span> substructure benchmark, the hypergraph selector improves compressed-basis classification. The Fisher hypergraph retains more third-order local response at twelve observables, with <span style="text-decoration: overline;">R</span>(3) = 0.937 compared with 0.871 for the pairwise graph. A low-capacity learning benchmark then shows that the same Fisher hyperedges serve as an interpretable inductive bias for message passing over correlator observables.

        Speaker: Aritra Bal (Karlsruhe Institute of Technology (KIT))
      • 6
        Generative Criticality in Large Language Model Temperature Scaling

        We propose a statistical-field framework for text generated by large language models (LLMs), treating token embeddings as continuous spin variables on a one-dimensional chain. Defining a susceptibility from the connected two-point correlator and an order parameter from the ensemble-averaged embedding field, we vary the softmax temperature $T$ and observe a sharp susceptibility peak near a characteristic $T_c$ with power-law-like scaling, a concurrent rapid change in the order parameter, and a collapse onto a single semantic direction below $T_c$. The intrinsic dimension estimated by the two nearest neighbor (TwoNN) method independently corroborates these findings, reaching a minimum near $T_c$. Results are robust across model scales (Qwen3: 0.6B--32B) and prompt categories. While the phenomenology closely resembles a continuous phase transition, the non-equilibrium nature of autoregressive generation warrants further investigation. Our framework provides quantitative tools for probing the collective statistical structure of LLM outputs and suggests connections between decoding strategies and critical phenomena.

        Speaker: Xingyu Guo (South China Normal Univeristy)
      • 7
        Steering in Theory Space: Representation Engineering of a Lagrangian Transformer

        Foundation models embed not only their training examples but the space those examples were drawn from. A transformer trained to write Lagrangians symbolically (such as BART-L) thus embeds the space of theories itself. In this work, we investigate the navigation of this learned theory space using activation steering, adapted from interpretability work on LLMs. Using only small contrastive datasets, we build steering vectors that take us from one theory domain to another. This hence reframes theory construction as a navigation problem, where reaching a theory of interest becomes a matter of finding and following directions, rather than an exhaustive scan of a landscape. We demonstrate navigations such as from anomalous to anomaly-free field content, from a base model to its supersymmetric and gauge-extended counterparts, and from charge-violating to charge-conserving Lagrangians. All of this requires no retraining or finetuning at all. The method also serves its original purpose of interpretability. By testing which operators can be recovered as vectors, we can ask whether the learned theory space is sufficient. We find directions corresponding to several Lorentz and gauge representation changes, including one that maps scalars onto fermions in the manner of a supersymmetry generator, as well as failure modes we associate with the limited size of our model and dataset. These results indicate that the model has, to some degree, embedded a navigable space of theories, and point toward a novel approach to theory search which is directed rather than exhaustive.

        Speaker: Yong Sheng Koay (Uppsala University)
      • 8
        StableWein: A Computational Tool for Bounded-From-Below Constraints in the Z2 x Z2-Symmetric Three-Higgs-Doublet Model

        We consider a three-Higgs-doublet extension of the electroweak Standard Model invariant under a $Z_2 \times Z_2$ symmetry. Since necessary and sufficient bounded-from-below (BFB) conditions are not known for this model, we propose an approach based on increasingly stringent necessary conditions. This procedure is implemented in the Mathematica package StableWein, which allows the user to control the accuracy of the BFB analysis. Our results indicate that StableWein can identify BFB potentials with very high precision.
        The package also includes a neural-network option trained to classify $Z_2 \times Z_2$-symmetric potentials. This machine-learning approach does not rely on either sufficient or necessary analytic BFB conditions, yet it achieves high performance, with an accuracy close to (99.9%) and fast inference. Therefore, the proposed method may be adapted to classify BFB parameter points in models for which analytic bounded-from-below conditions are unavailable.

        Speaker: Darius Jurčiukonis (Vilnius University (LT))
      • 9
        Validation, Calibration, and Explainability of Machine Learning Models - From Particle Physics to Eye-Disease Prognosis

        With the increasing volume and complexity of data, machine learning (ML) models are becoming indispensable tools for identifying patterns within structured datasets. However, as these models grow in complexity, it becomes challenging to determine whether they are truly learning meaningful relationships or capturing unintended artifacts. This lack of interpretability leads to mistrust, particularly in high-stakes domains. In clinical settings, prognostic models must be fully explainable to clinicians and patients to ensure confidence in medical decisions. For example, age-related macular degeneration (AMD) is the leading cause of legal blindness in developed countries, and current existing methods of predicting disease progression (survival modelling) can often lack interpretable explanations of progression and can misclassify progressing cases as non-progressing in unbalanced datasets. Similarly, in particle physics, ML-based classification models must be demonstrably aligned with known physical principles, such as those described by the Standard Model. Here, we explore how a Graph Neural Network (GNN) can be adapted from the identification of hadronic tau decays developed for ATLAS to a survival model to predict the risk of disease progression in AMD patients, ensuring robust and interpretable ML applications across disciplines.

        Speaker: Robert McNulty (University of Liverpool (GB))
      • 10
        Explainable Fuzzy-AI Reasoning for Skin Radiation Risk Classification in Medical Physics Applications

        Artificial intelligence methods are increasingly used in physics to support classification, prediction, uncertainty handling, and interpretable decision-making. In medical physics, radiation risk assessment often depends on multiple interacting parameters, including absorbed dose, exposure duration, source activity, irradiated skin area, radiation energy, source-to-skin distance, shielding conditions, and contamination geometry. Conventional numerical approaches provide useful dose estimates but may be less effective when expert reasoning, uncertainty, and transparent risk classification are required.

        This study proposes an explainable fuzzy-AI decision-support framework for skin radiation risk classification in medical physics applications. The model integrates relevant dosimetric and exposure-related parameters within a hierarchical fuzzy reasoning structure. Input variables are represented using linguistically interpretable membership functions, and rule-based inference is applied to classify skin radiation risk into clinically meaningful categories. The framework is designed to support transparent decision-making by explicitly showing how different combinations of dose, activity, exposure time, shielding, and geometry contribute to the final risk level.

        The proposed approach is particularly relevant for scenarios involving occupational exposure, external contamination, radiopharmaceutical handling, and radiation protection training. By combining fuzzy logic with AI-assisted decision support, the framework provides an interpretable alternative to purely black-box classification methods. It may also contribute to educational and practical tools for radiation protection, where explainability, uncertainty awareness, and rapid risk stratification are essential.

        The study demonstrates how explainable AI concepts can be adapted to medical physics and radiation safety problems, providing a bridge between computational intelligence and physics-based risk assessment

        Speaker: Seyedeh Fatemeh Mirekhtiary (Near East University)
    • 🔀 Inference & Uncertainty HS1

      HS1

      Convener: Jonas Spinner (Durham University)
      • 11
        Joint spectral and population inference of Galactic pulsars and extragalactic sources in the Fermi-LAT sky with simulation-based inference

        Since its mission started more than 17 years ago, the Fermi Large Area Telescope (LAT) has significantly advanced our view of the GeV gamma-ray sky, yet several key questions remain - such as the nature of the isotropic diffuse background, the properties of the Galactic pulsar population, and the origin of the GeV excess towards the Galactic Centre. Addressing these challenges requires sophisticated astrophysical modelling and robust statistical methods capable of handling high-dimensional parameter spaces.
        Building on our previous single-population, single-energy-bin analysis of Fermi-LAT data using simulation-based inference (SBI), we present a multi-bin spectral extension of the gamma-ray emission simulator. The updated framework simultaneously models two source populations, Galactic pulsars and extragalactic point sources, with distinct spatial distributions and intrinsic spectral variability informed by the empirical spread observed in detected members of each respective class. This richer forward model enables joint inference of population-level quantities on a per-class basis, including luminosity functions and source-count distributions. We present preliminary results that investigate the extent to which SBI can disentangle the two populations and recover their respective spectral and spatial properties from synthetic observations.

        Speaker: Christopher Eckner (Instituto de Astrofísica de Canarias (IAC))
      • 12
        Simulation-Based Inference of Unresolved Neutrino Source Populations in the Galactic Plane

        The high-energy neutrino emission observed from the Galactic plane by IceCube may contain contributions from both truly diffuse emission produced by cosmic-ray interactions with interstellar gas and a population of individually unresolved hadronic accelerators. Traditional likelihood-based inference is impractical here, since it would require explicit marginalization over the unknown number, positions, and luminosities of the sources, as well as over the events each realization produces, rendering the likelihood computationally intractable.

        We therefore implement a forward model of the Galactic neutrino sky and perform simulation-based inference (SBI) within Falcon, a dynamic SBI framework that constrains population-level parameters directly from energy-binned neutrino sky maps. A convolutional neural network compresses the sparse maps into a learned representation of their spatial and spectral structure. This representation conditions a normalizing flow that approximates the joint posterior of the expected number of sources and the diffuse-emission normalization. Falcon adaptively generates simulations and retrains the neural posterior estimator, concentrating computational effort on the regions of parameter space relevant to the target observation.

        Applied to a synthetic Galactic neutrino sky, the method recovers both injected parameters within their 68% credible intervals, and the inferred posterior exhibits the expected anticorrelation between the unresolved-source contribution and the diffuse normalization. The analysis further shows that SBI can exploit information beyond the total event count: the spatial granularity of the source population, together with a source spectrum harder than that of the diffuse component, provides complementary handles for separating the two contributions.

        This proof of concept establishes a basis for applying SBI to increasingly realistic neutrino observations. Higher-statistics data with improved angular resolution, in particular from KM3NeT/ARCA and future neutrino telescopes, will carry substantially more spatial information and should tighten constraints on the abundance and Galactic distribution of unresolved hadronic accelerators.

        Speaker: Youyou Li (GRAPPA, University of Amsterdam)
      • 13
        Simulation-Based inference for massive black hole binary from mock LISA data

        The Laser Interferometer Space Antenna (LISA) will observe gravitational waves produced by several massive black hole binary (MBHB) mergers per year. While the likelihood can be written in closed form under idealised stationary, Gaussian noise assumptions, including realistic effects — instrumental glitches, gaps in the data, and non-stationary noise — make it intractable or computationally prohibitive. This motivates the development of Simulation-based inference (SBI) pipeline, which requires only forward simulations of the data and can thus in principle include arbitrary complex physics.
        We present proof-of-concept results using truncated marginal neural ratio estimation (TMNRE), a sequential SBI method that iteratively truncates the prior to the region supporting non-negligible posterior mass. One of the challenges is the enormous shrinkage of prior volume under the posterior, a factor ~$10^{-24}$, which makes ordinary MCMC and sequential SBI difficult.
        We demonstrate inference of the parameters of a single MBHB from frequency-domain LISA data with a customised TMNRE algorithm, accommodating the periodicity of the angular parameters and preserving multimodal support during truncation, and validate the method against the MCMC posterior of the tractable case (obtained by artificially reducing the prior volume around the fiducial parameters). For a mock observation with chirp mass $10^{5.25} M_{\odot}$ and SNR 2200, we correctly recover the MCMC posterior over the 11-dimensional parameter space, in about 30 hours of training on a NVIDIA A100 GPU.

        Speaker: Gianmarco Puleo (Scuola Internazionale Superiore di Studi Avanzati (SISSA))
      • 14
        Multimodal Generative Modelling of the Galaxy Population with pop-cosmos

        Panchromatic surveys such as COSMOS have expanded our understanding of how galaxies evolve through cosmic time immensely. Surveys such as the Vera C. Rubin Observatory's LSST are promising an unprecedented view of this evolutionary paradigm provided the massive data challenge they pose is answered. We have been developing pop-cosmos, a comprehensive galaxy population model housing a 16-parameter SPS model parametrized by a flexible diffusion model. The 16D distribution over galaxy properties including their redshift is forward-calibrated on 26-band photometry from COSMOS with a deep infrared selection. In a recent paper we used a galaxy catalogue drawn from the trained pop-cosmos model to investigate the stellar mass assembly and star formation histories of galaxies up to z=4. I will briefly summarize key results such as the cosmic star formation rate density inferred from our model, and will present a look into the quenching mechanisms of galaxy populations. Our investigation finds a shift of the cosmic dawn towards earlier lookback times, and uncovers correlations between star formation, AGN activity and quenching difficult to capture without a population-level analysis. In this talk I will mainly focus on our work expanding the forward process of pop-cosmos to jointly train on the morphology of galaxies and their photometry from profile fitting. The morphology of galaxies encodes important information about their evolutionary stage, and I will talk about how the inclusion of this dimension in the forward modelling impacts the population model. I will describe how we connect the morphology information using the compressed latent representation learned by an encoder-decoder network, like a convolutional variational autoencoder, effectively keeping the computational cost increase due to this new mode to a minimum. To conclude, I will present how this approach unlocks a new avenue for population-level causal inference in astrophysics, using the causal relation between quenching and morphological transformation as the primary example

        Speaker: Sinan Deger (University of Cambridge)
      • 15
        Fast Inferential Diffusion Models for Neural Simulation-based Inference

        Neural simulation-based inference (NSBI) has emerged as a powerful methodology for unbinned likelihood analyses in high-energy physics (HEP), with recent ATLAS [1] results at the Large Hadron Collider (LHC) demonstrating substantial sensitivity gains over binned methods. The dominant NSBI variant in HEP, Neural Likelihood-Ratio Estimation (NLRE), owes its popularity to the simplicity of the likelihood-ratio trick [2]. Despite this success, the implementation of NLRE methods suffers from several structural problems within the context of the profile-likelihood machinery common in HEP analyses. The quality of likelihood ratio estimators (LREs) are highly dependent on the support match, or mismatch, of the classes used in the training of the classifier [3]. When faced with composite likelihoods, often multiple LREs (e.g. one per signal) are used to build the likelihood model of the data, each carrying their own support-mismatch risk, leading to additive calibration/miss-specification errors on the composite likelihood. Continuous parameter scans require parameterised LREs, which consequently requires the base class (y = 0) dataset to be very large, since assigning randomised parameter values to reference events dilutes the effective sample size at each scan point. In contrast, Neural Likelihood Estimation (NLE) returns absolute densities that slot directly into the profile-likelihood workflow, often avoiding these structural problems. Unfortunately, standard NLE routes, such as normalising flows or diffusion models that yield log-densities via trajectory integration, are computationally expensive to train and/or evaluate. The latter is particularly problematic when dealing with the high data volumes of experiments at the LHC, and is a key prohibiting factor in the use of diffusion models for NSBI driven statistical inference tasks.

        To address these problems we introduce a hybrid NLRE-NLE methodology that can learn a log-density of the data x conditioned on a set of parameters of interest (POI) θ (log(p (x | θ))), from a single classifier. Referred to as a Density Ratio Diffusion Probabilistic Model (DRDPM), the DRDPM method replaces the binary classifier of a NLRE method with a multi-class classifier over the discrete timesteps (t) of a diffusion process that interpolates between the data distribution and a standard Gaussian anchor. Applying Bayes’ rule to the classifier’s outputs at the two boundary classes (t = 0 clean data and t = T Gaussian anchor), the conditional log-likelihood of the data can be recovered in a closed form as the difference of the two boundary logits. In summary, the resulting hybrid NLRE–NLE method trains by classification alone, inheriting the classifier-training simplicity of NLRE, and returns a per-event likelihood estimate at inference-time (evaluation) that is O(10 − 100) times faster than current trajectory based Ordinary(Stochastic) Differential Equation (O(S)DE) solver-based solutions, since only a single neural function evaluation call is required per event. We demonstrate this methodology on a representative High Energy Physics problem of proton-proton collisions at the LHC, and benchmark the fast inferential capabilities using a three-body decay toy dataset within the NEEDLE project [4] (a inter-experiment collaboration across the ATLAS, CMS, and ALICE Collaborations).

        [1] The ATLAS Collaboration. Measurement of off-shell higgs boson production in the h → ZZ → 4ℓ decay channel using a neural simulation-based inference technique in 13 TeV pp collisions with the ATLAS detector. Reports on Progress in Physics, 88(5):057803, may 2025.
        [2] Kyle Cranmer, Juan Pavez, and Gilles Louppe. Approximating likelihood ratios with calibrated discriminative classifiers, 2016.
        [3] Benjamin Rhodes, Kai Xu, and Michael U. Gutmann. Telescoping density-ratio estimation. In Advances in Neural Information Processing Systems, volume 33, pages 4905–4916. Curran Associates, Inc., 2020.
        [4] NEEDLE Collaboration. NEEDLE: Neural-based Diffusion Likelihood Estimations. https:://needle-sbi.github.io/, 2026. Orchestration framework and toolkit for the deployment of Neural Simulation-based Inference (NSBI) methods in High Energy Physics. Accessed: July 31, 2026.

        Speaker: Stephen Jiggins (DESY)
      • 16
        Two Phase Rapid Simulation Based Inference with Differentiable Simulators

        I will introduce introduce Rapid Simulation Based Inference (RSBI), a diffusion-based variational approach to likelihood-free Bayesian inference that achieves high sampling efficiency under large prior-to-posterior volumes with multi-modal posterior structure.
        RSBI builds on advances in Schrödinger Bridge diffusion sampling to handle multi-modal posteriors, with the likelihood initially supplied by a surrogate distance measure via a differentiable simulator, followed by neural ratio estimator trained on a constrained corpus of simulation/parameter pairs. The key innovation is that an appropriate surrogate likelihood gives a highly efficient proposal distribution for multi-modal posteriors, replacing sequential rounds of other methods with a one-shot proposal. Subsequent NRE refinement targets the posterior under the simulator and original model prior.
        As a variational method, RSBI does not suffer from leakage and implicit target drift commonly observed in standard sequential posterior estimation methods. To improve mode coverage we optionally utilize the well-tempered meta-dynamics framework, which also encourages exploration of the prior volume. We observe strong performance on standard SBI benchmarks, particularly when the prior volume is scaled up to 2500 times the original, achieving a performance degradation of only $\sim5\%$ on Two Moons under a highly constrained simulation budget. Additionally, we evaluate RSBI's performance for gravitational wave ring-down posterior estimation, in a real-world physics benchmark inspired by black hole spectroscopy.

        Speaker: Andre Scaffidi (SISSA)
      • 17
        Probabilistic prediction of kilonova light curves from observed binary neutron star coalescences

        Kilonovae are the optical counterparts to gravitational wave signals from binary neutron stars. Combining gravitational wave and kilonova observations provide insight into the equation of state of neutron stars. Kilonova detection is challenging and only a few kilonova candidates have been detected to date, with AT2017gfo being the only one associated with a gravitational wave event GW170817. This highlights the need for rapid kilonova predictions to inform follow-up observations for future gravitational wave events. In this work, we develop a probabilistic framework based on normalising flows to predict kilonova spectra and light curves directly from gravitational wave posterior samples of binary neutron star mergers. These outputs can then be used to inform optical and near-infrared follow-up kilonova observation strategies.

        Speaker: Ik Siong Heng (University of Glasgow)
    • 🔀 Patterns & Anomalies 3.404

      3.404

      Convener: Ramon Winterhalder (University of Milan)
      • 18
        Unsupervised Machine Learning for Model-Independent Searches at the LHC

        Since the discovery of the Higgs boson, no experimental evidence for physics beyond the Standard Model has been observed at the LHC. While most searches rely on specific signal hypotheses, anomaly detection aims at identifying potential new-physics signatures without assuming a particular BSM model.

        In recent years, unsupervised machine learning has become an active area of research in high-energy physics. These techniques learn the properties of Standard Model events directly from data and search for deviations that could indicate the presence of new phenomena.

        This contribution reviews recent developments in anomaly detection for fully hadronic final states within the ATLAS Collaboration. Starting from the first fully unsupervised search [1] based on a Variational Recurrent Neural Network trained directly on collision data, we discuss more recent approaches based on Transformer and Graph Neural Network architectures, including EGAT. Their performance is illustrated using the LHC Olympics dataset benchmark [2] and through their first applications to searches for heavy diboson resonances in ATLAS using proton-proton collisions at √s = 13 TeV.

        References
        [1] Phys. Rev. D 108 (2023) 052009.
        [2] The LHC Olympics 2020: A Community Challenge for Anomaly Detection in High Energy Physics, Rep. Prog. Phys. 84 (2021) 124201.

        Speakers: Antonio D'Avanzo (University Federico II and INFN, Naples (IT)), Elvira Rossi (University Federico II and INFN, Naples (IT)), Francesco Cirotto (University Federico II and INFN, Naples (IT)), Francesco Conventi (Università degli studi di Napoli "Parthenope" and INFN Sezione di Napoli (IT)), Graziella Russo (University of California,Santa Cruz (US))
      • 19
        Searching Everything, Everywhere, All at Once

        The LHC search programme remains largely model-driven, probing one final state and one model hypothesis at a time, complemented by occasional model-agnostic searches such as those in dijet spectra. We aim to take the next step: probing many final states at once, in an automatic and robust way, using limited resources efficiently to cover large regions of phase space. This combines reusing our knowledge of established models to boost sensitivity with model-agnostic anomaly-detection techniques.

        We estimate the total number of searches required to cover everything within our reach, and examine how the look-elsewhere effect limits sensitivity in large-scale automated searches, along with strategies to manage it. Finally, we present a proof-of-concept using Gaussian Process Regression as a robust, sensitive, data-driven background-estimation technique across a broad range of smoothly falling spectra.

        Speaker: Theresa Reisch (University of Geneva)
      • 20
        Anomaly detection for multi-jet resonances

        The search for physics beyond the standard model is one of the prime focuses in high-energy physics. Conventional searches at the LHC, though comprehensive, have not yet shown signs for new physics. Machine learning based anomaly detection has emerged offering a model-agnostic way to enhance the sensitivity of generic searches as compared to those targeting specific signal model. CATHODE (Classifying Anomalies THrough Outer Density Estimation), one of these methods, is a two-step method that combines a data driven background density estimation with a classifier flagging potential signal.
        These machine learning methods have shown promising improvements although to date, most studies have mainly focused on dijet resonances. We extend CATHODE to signals with multiple decay modes and varying jet multiplicities, demonstrating its first application to multi-jet resonances and improving the robustness of weakly supervised anomaly detection beyond the dijet regime.

        Speaker: Chitrakshee Yede (Universität Hamburg)
      • 21
        Anomaly detection with normalizing flows for model-agnostic new physics searches

        Discovering new particles from beyond the Standard Model remains one of the main goals of present-day particle physics. Traditional searches for new physics at the Large Hadron Collider rely on specific theoretical scenarios and simulation-based background estimates, limiting their reach and introducing modeling uncertainties. We present an anomaly detection method that uses normalizing flows to identify anomalous events without assuming a particular signal model, with the goal of estimating backgrounds directly from data rather than simulation. By training a flow on data and splitting the resulting latent representation into two independent parts, we construct two decorrelated anomaly scores that allow the background in the signal region to be estimated using the well-established ABCD method. We test this approach on a benchmark new physics scenario and show that it successfully decorrelates the anomaly scores and correctly estimates the S/√B in the signal region.

        Speaker: Rafał Masełek (Jozef Stefan Institute)
      • 22
        HAXAD: Anomaly Detection in the Higgs Peak

        The Higgs boson, with its universal coupling to mass, provides a broadly applicable portal to sectors beyond the Standard Model and is therefore a natural anchor for anomaly detection (AD) at collider experiments. The Higgs And X Anomaly Detection (HAXAD) strategy offers a principled approach to searching for anomalies occurring in association with a Higgs boson. In the kinematic region of the Higgs peak, HAXAD employs machine-learning-based feature embedding, data- and simulation-driven background estimation, and weakly supervised classification.

        We present the HAXAD workflow together with an extensive simulated dataset, including a large benchmark suite of signal models, which we plan to release as an open-data challenge in the future.

        We show that HAXAD achieves strong signal sensitivity, and when benchmarked against a representative multi-category cut-based search, matches or exceeds the best individual cut-based limits for a wide variety of signal models. Together, these results establish HAXAD as a viable and compelling AD-based search strategy with novel discovery potential at colliders.

        Speaker: Dennis Noll
      • 23
        Towards a Statistical Interpretation of the Normalized Autoencoder

        Unsupervised anomaly detection with autoencoders is a promising data-driven and model-agnostic approach for new physics searches at the LHC. However, current anomaly scores assigned by neural networks suffer from a lack of statistical interpretability. The normalized autoencoder (NAE) combines a standard bottleneck architecture with a well-defined probabilistic description. We show that the NAE ties its anomaly score to a learned likelihood via an energy-based training objective, and introduce a Bayesian version of the NAE (BNAE) that additionally provides machine-learned uncertainty estimates. We validate both on a toy model and demonstrate competitive and symmetric anomaly-tagging performance on top-versus-QCD jet tagging.

        Speaker: Jonathan Ostertag-Henning (ITP Heidelberg)
      • 24
        Unsupervised Anomaly Detection for LHC Collision Data Using a Transformer-Based Autoencoder.

        An unsupervised machine learning framework is developed for searching for anomalous events in proton-proton collisions at √s=13 TeV using ATLAS Open Data. The analysis utilizes an object-based Masked Transformer autoencoder trained exclusively on Standard Model Run 2 data to learn complex correlations among reconstructed physics objects and identify events that deviate from expected Standard Model behaviour without relying on signal labels or predefined signal hypotheses. The resulting anomaly scores provide clear separation between Standard Model background and MC generated Beyond-the-Standard-Model benchmark processes.
        To further investigate the most anomalous events in the RUN 2 data, a multi-stage post-processing pipeline is employed to identify anomalous event clusters in the filtered anomalous event space, evaluate their statistical significance, and perform detailed physics characterization. This enables the separation of genuine anomalous event structures from statistical fluctuations and provides physically interpretable anomaly candidates. The proposed framework combines deep learning with unsupervised statistical analysis to provide a transparent, extensible, and data-driven workflow for anomaly discovery, offering a complementary approach to conventional targeted searches for new physics at the LHC.

        Speaker: ASRITH KRISHNA RADHAKRISHNAN (University of Bologna (IT))
      • 25
        Kitchen Sink Anomaly Detection

        Recent years have seen rapid progress in resonant anomaly detection for collider searches, but existing studies often rely on a limited set of signal benchmarks and face a trade-off between sensitive but model-dependent high-level observables and fully agnostic but less performant low-level representations. We address both limitations by introducing new simulated signal benchmarks, publicly released in a format compatible with the LHCO R&D benchmark, and by studying a broad high-level, yet highly agnostic, observable set combining Energy Flow Polynomials with subjettiness variables.
        We evaluate this combined “kitchen sink” representation against several baseline observable sets in both an idealized anomaly-detection setting and the CWoLa hunting task. Across a broad range of signal types, the combined observable set achieves the best overall sensitivity.

        Speaker: Lukas Lang (RWTH Aachen University)
    • 🔀 Simulations & Generative Models HS2

      HS2

      Convener: Andreas Ipp (TU Wien)
      • 26
        Generative Machine Learning for Fast Calorimeter Simulation in ATLAS

        One of the costliest elements of the ATLAS detector simulation is the precise modelling of electromagnetic and hadronic showers. With the aim of lowering CPU usage in Run 3 of the LHC, the collaboration introduced AtlFast3, a fast simulation tool that combines classical histogram-based parameterisations with calorimeter models built on GANs. Once Run 3 was underway, a new effort was launched to refine the voxelisation scheme employed for model training, a scheme that bins energy deposits into small volumetric bins. This reworked voxelisation has produced a more efficient shower description together with a visible boost in physics performance. Despite these advances, GANs still display the well-documented limitations they are known for in terms of stability and accuracy. Guided by the recent CaloChallenge results, ATLAS is consequently looking into newer generative methods such as diffusion models, transformers, and continuous normalizing flows. They are now being evaluated as possible substitutes for parts of the parameterisation used in AtlFast3 today. In summary, this contribution surveys the current machine-learning-based calorimeter simulation within ATLAS, showcases results from the modern generative models being studied, and discusses the key challenges along the way toward a more ML-driven fast simulation for Run 4 and the future.

        Speaker: Florian Ernst (Heidelberg University (DE), CERN)
      • 27
        AllShowers: One model for all calorimeter showers

        To reduce the high computational demand of detector simulations in high-energy physics, various generative surrogate models have been proposed. Classically, one generative model per incident particle type is trained, requiring separate trainings and model weights. This results in increased training and human effort, as well as a larger memory footprint during inference, since multiple sets of weights must be loaded. To address this, we put forward AllShowers, a single generative surrogate capable of generating showers from 12 different particle types that enter the calorimeter system at a wide range of angles and incident energies, without retraining or fine-tuning. This is accomplished by conditioning the model on the incident particle type, angle, and energy. We demonstrate high-fidelity generation of electrons, photons, and charged and neutral hadrons in the highly granular electromagnetic and hadronic calorimeters of the International Large Detector (ILD). In addition to unifying the generation, AllShowers surpasses the fidelity of previous state-of-the-art single-particle-type models for hadronic showers.

        AllShowers is a point-cloud-based flow matching model with a transformer architecture. Key architectural improvements include: a custom attention mask to reduce computational costs and provide a helpful inductive bias; a shower- and layer-wise optimal transport mapping for shorter, more stable flow trajectories; and allowing the model to encode all relevant detector properties in an embedding vector per layer. With AllShowers, we take a significant step towards a general generative model for calorimeter simulations.

        Speaker: Thorsten Buss (RWTH Aachen)
      • 28
        Cross-Geometry Shower Simulation via Pre-training and Transfer

        Deep generative surrogates can cut the computing cost of detailed Geant4 calorimeter simulation, but they remain tied to the detector they were trained on: each new geometry typically demands a large in-domain dataset. We present two studies that reduce this dependence for point cloud generative models.

        In the first, a CaloClouds-based diffusion model pre-trained on ILD photon showers transfers to electron showers in the cylindrical CaloChallenge geometry without re-voxelisation. Fine-tuning on only 100 target showers improves the geometric mean of the Wasserstein distances to Geant4 by about 50% over training from scratch, and bias-only adaptation stays competitive while updating 17% of the parameters.

        The second study scales to multi-geometry pre-training with AllShowers, a transformer flow-matching model. It compares two pre-training pools: a synthetic family of $10^4$ box calorimeters and a set of realistic detectors. On the held-out FCC-ee ALLEGRO calorimeter, with $10^3$ target showers for fine-tuning, the two priors reduce the aggregated sliced Wasserstein distance to Geant4 by factors of 5.2 (synthetic) and 8.0 (realistic) relative to training from scratch. The purely synthetic prior matches the realistic one: geometric diversity alone is enough to pre-train a transferable shower model. The measured pre-training cost amortises within a few fine-tunings, making pre-trained surrogates a practical route to fast simulation for future detectors.

        Speaker: Lorenzo Valente (University of Hamburg)
      • 29
        Experts or Generalists in Calorimeter Simulation

        Physics employs careful factorising of effects to make powerful general predictions. In a Mixture of Experts network, well chosen factorisation of knowledge can improve both memory footprint and generalisation.

        This work explores the relative merits of Expert networks, in comparison to equivalent generalist models. In the setting of fast generative calorimeter simulation, both architectures are explored and the strengths illustrated.

        Speaker: Henry Day-Hall (DESY)
      • 30
        Parnassus: Fast and Accessible Realistic Full Event Simulation

        Detector simulation and event reconstruction constitute the primary computational bottleneck in modern particle physics. To bypass these CPU-intensive processing chains, we present Parnassus (Particle-flow Neural Assisted Simulations), a state-of-the-art generative AI framework for end-to-end fast simulation. Utilizing conditional flow matching, Parnassus maps truth-level stable particles directly to reconstructed particle-flow objects, accurately capturing event-, jet-, and particle-level features within a fully automated, GPU-optimized Python workflow.
        We demonstrate the versatility and generalization power of Parnassus across multiple distinct experimental environments, including CMS, ATLAS, and ALEPH (LEP). Parnassus serves both as a high-throughput GPU alternative to official experiment fast simulations and as an accessible tool for phenomenologists to test the models in a realistic environment. It proves that modern generative approaches generalize seamlessly across diverse geometries and collider eras, offering a scalable path forward for high-energy physics.

        Speaker: Dmitrii Kobylianskii (Weizmann Institute of Science)
      • 31
        Speeding up the MC Background Simulation at Belle II

        By striving for ever-higher luminosities, the Belle II detector is set to observe rare decay signals. However, these high luminosities correspond to an increased demand for MC-simulated events. Therefore, an efficient algorithm to produce and process simulated events in large quantities is vital for the successful interpretation of Belle II’s data. While generating events in the MC simulation chain is quick and easy to compute, detector simulation and particle reconstruction are slow tasks. A fine-tuned version of the Particle Transformer (ParT) is introduced into the chain to classify passing events during analysis pre-selection and to discard those that fail the pre-filter before the detector response simulation. With considerable speedups in the case of static ParT training, this method has been shown to reduce overall time consumption, thereby successfully saving valuable computing resources. This work explores a parallelised approach to the ParT training process in order to match the parallel workflow of data production. It involves an on-the-fly training pipeline across multiple machines, retrieving updated ParT models, and redistributing them back for the next iteration.

        Speaker: Oliver Schumann (LMU Munich)
      • 32
        To surrogate or to differentiate? A Delphes3 optimization study

        Fast parametric detector simulations transform generator-level particles into reconstructed physics objects through a chain of modules controlled by smearing functions. A smearing function contains closed-form resolution and efficiency formulae with numeric coefficients which are traditionally tuned by hand against full simulation or data. We study gradient-based optimization of Delphes3, a multipurpose fast detector simulator for phenomenology studies, using a novel differentiable workflow that can optimize the full set of simulator parameters. We compare this approach to a surrogate network which combines simulation and reconstruction steps. The two workflows provide complementary views on the accuracy and interpretability of fast simulators for the LHC and future experiments.

        Speaker: Luigi Favaro (UCLouvain - CP3)
    • 🧠 Working groups: FAIR-ness & Sustainability (WG3) 2.404

      2.404

      • 33
        FAIR-ness & Sustainability (WG3)
        Speakers: Caterina Doglioni (University of Manchester), Caterina Doglioni, Sascha Diefenbacher (ITP Heidelberg)
    • 17:45
      📸 Conference photo Foyer

      Foyer

    • 18:00
      🍻 Reception + 🖼️ Poster session + 🌟 Highlight vote Foyer

      Foyer

    • 🗣️ Plenaries HS1

      HS1

      Convener: Gabrijela Zaharijaš (University of Nova Gorica)
      • 34
        Theory ex Machina: Machine Learning the LHC Simulation Chain
        Speaker: Ramon Winterhalder (University of Milan)
      • 35
        The Model Says Yes. Now What? And Other Adventures in Scientific AI

        What can we actually conclude when a model gives us the right answer? Did it learn the right correlations? Does it generalize beyond the setting in which we tested it? And what does it mean to validate increasingly complex "scientific AI" models and workflows? I will explore these questions through examples and cautionary tales from language modeling, human modeling, particle physics and astrophysics — from unexpected correlations and hidden systematics to surprising generalization, foundation models and scientific agents producing convincing, but not always plausible, results.

        Speaker: Lucie Flek (University of Bonn)
    • 10:30
      ☕ Coffee Foyer

      Foyer

    • 🔀 Foundation Models 1.404

      1.404

      Convener: Luigi Favaro (UCLouvain - CP3)
      • 36
        A Recipe for Scale: Hyperparameter-Optimal Scaling Laws End-to-End, from Analytic Toy Problems to Jet Tagging

        Recent progress in machine learning has come as much from scale as from architecture, and high-energy physics is starting to follow: work on transformer scaling for flavour tagging in ATLAS [https://inspirehep.net/literature/3114069] has shown that tagger performance follows predictable power laws over orders of magnitude in compute on a multi-billion-jet dataset. This matters particularly in a field built on domain-specific inductive bias: physics-motivated features, symmetry-aware architectures, auxiliary objectives, whose benefits are almost always established at a modest scale. Which of them still pay off as compute grows is therefore an open question, and answering it requires scaling laws that reflect genuine method differences rather than mistuned training dynamics.
        We present a recipe that scales every tunable axis of the system optimally. We first validate the full life of a scaling law: onset, power-law regime, and irreducible floor, on domain-agnostic toy problems where the floor is known analytically, and finally apply the complete procedure to multi-task transformers for jet flavour tagging.
        Tuned this way, design choices can be compared fairly: we quantify how auxiliary objectives and input representations affect both the scaling of the loss with compute and the optimal scaling of the hyperparameters themselves. We further show that the onset of the power-law regime is itself set by scale: below a threshold in model size and training tokens the loss carries no information about high-compute scaling, though fits in that region will still return a biased exponent, showing the need for large open datasets of high-quality full simulation.

        Speaker: Matthias Vigl (TUM)
      • 37
        Neural scaling laws for jet generation

        Scaling laws have become a central topic in modern machine learning, providing a quantitative understanding of how model performance improves with increasing model size, training data, and compute. They also offer insights into whether learning is approaching the information limits of a given dataset. In this contribution, we present the first study of neural scaling laws for generative jet modeling. Beyond the conventional next-token prediction validation loss, we also investigate the scaling behavior of the sliced Wasserstein distance computed for five high-level jet observables, providing a physics-motivated measure of generative performance. As a function of model size, we observe that both metrics exhibit the expected logarithmic scaling behavior. In contrast, scaling with dataset size and training compute is significantly weaker for both the validation loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss several possible explanations for this observation, including the intrinsic stochasticity of QCD jet formation and the fundamental differences between supervised prediction and generative modeling.

        Speaker: Anna Hallin (anna.hallin@uni-hamburg.de)
      • 38
        Scaling laws for amplitude surrogates

        Fast and precise evaluations of scattering amplitudes even in the case of precision calculations is essential for event generation tools at the HL-LHC. We explore the scaling behavior of the achievable precision of neural networks in this regression problem for multiple architectures, including a Lorentz symmetry aware multilayer perceptron and a fully Lorentz equivariant transformer using Lorentz Local Canonicalization (LLoCa). This study addresses in particular the scaling behavior of uncertainty estimations using state of the art methods.

        Speaker: Joaquin Iturriza Ramirez (LPNHE - Sorbonne Université)
      • 39
        Multitask, Every Region, All at Once : Fine-Tunable Region-Conditioned Simulation-Based Inference for Collider Analysis

        Collider simulations simultaneously provide three sources of information: whether an event is signal or background, the physics parameters that generated each signal event, and the event kinematics from which signal regions are defined. Conventional pipelines use these sources in separate stages—training classifiers for signal discrimination, performing parameter inference for fixed analysis regions, and optimizing signal regions for specified signal hypotheses. Because all three are driven by the same simulation and the same underlying statistics, treating them as disjoint tasks leaves shared structure unexploited. We present a proof of concept in which a single region-aware model is trained jointly on all three sources. The same trained model then supports signal detection, parameter inference, and region screening as different projections of a shared conditional distribution, while providing a pretrained backbone that can be fine-tuned to a chosen region with less data than training from scratch. We validate the idea on a Gaussian bump-hunt, where all three tasks and fine-tuning behave largely as intended, and stress-test it on a non-resonant mono-jet dark-matter search that exposes its limitations. These results motivate unified learning from collider simulations as a precursor to foundation models for collider analysis.

        Speaker: Yong Sheng Koay (Uppsala University)
      • 40
        One Generator, Any Process: LLM-Conditioning for the LHC

        Neural network training for LHC event generation should, ideally, benefit from common high-level patterns in different processes. We propose novel conditioning schemes for continuous parameters, process labels, and Feynman diagrams. We employ pre-trained LLMs as multi-modal foundation models to provide descriptive embeddings for an autoregressive transformer. With such high-level physics-inductive bias the generative networks converge faster, provide better result, and generalize to unseen processes.

        Speaker: Thanush Sivagnanalingam (University Heidelberg)
      • 41
        An LLM based AI assistant to support the operation of GSI/FAIR accelerators

        Large-scale accelerator facilities such as GSI/FAIR rely heavily on distributed expertise and fragmented documentation for daily operations, creating persistent challenges in knowledge retention, troubleshooting speed, and workforce continuity. To address these operational bottlenecks, we present an AI assistant designed to deliver context aware, expert level guidance to shift operators. Grounded in large language models (LLM) run locally on the GSI high performance computing infrastructure, the system integrates retrieval augmented generation (RAG) and structured prompt engineering to synthesize domain knowledge from internal wikis, technical manuals, historical logbooks, and online shift data. In addition to the prof of principle status, we present future development plans including incorporation of more knowledge sources, exploration of different LLM models and low-rank adaption and integration with the control system to access accelerator data in real-time.

        Speaker: Philipp Niedermayer (GSI GmbH)
    • 🔀 Inference & Uncertainty HS1

      HS1

      Convener: James Alvey (University of Cambridge)
      • 42
        How to Trust Learned Loop Amplitudes

        Higher-order theory predictions are crucial for the precision LHC program, but the time-consuming amplitude evaluation challenges the corresponding Monte-Carlo simulations. Machine-learned amplitude surrogates can resolve this problem, if we can guarantee their precision over the entire phase space. First, we show that our surrogates provide a calibrated learned uncertainty, even for non-Gaussian systematics; second, we describe how less accurate phase space regions can be identified; third, we demonstrate how the precision in these regions can be improved reliably.

        Speaker: Henning Bahl
      • 43
        Local Conformal Predictions for Calibrated Surrogates

        Neural network surrogates for LHC scattering amplitudes require trustworthy uncertainty estimates, a challenging task given the non-Gaussian systematics. We target it using conformal prediction, a distribution-free post-processing to complement trained surrogates with calibrated uncertainties. We find that standard conformal predictions struggle to provide locally calibrated uncertainties. This leads us to introduce FALCON, a novel conformal prediction method that learns locally calibrated confidence intervals. Our simple examples illustrate the power of distribution-free uncertainty quantification for ultra-fast event generation at the LHC.

        Speaker: Mr Suprio Dubey (Institute for Theoretical Physics and Mannheim Institute for Intelligent Systems in Medicine, Heidelberg University)
      • 44
        Learning to validate generative models

        Generative models are becoming an integral component of scientific workflows, yet their reliable deployment requires a rigorous understanding of their limitations through statistically robust validation. In particular, high-dimensional settings demand powerful goodness-of-fit tests that provide well-defined statistical interpretations and enable hypothesis testing. In this contribution, we propose the use of the New Physics Learning Machine (NPLM), a learning-based goodness-of-fit test inspired by the Neyman–Pearson lemma, to validate generative models trained on high-dimensional scientific data. We demonstrate its performance on two benchmark datasets: normalizing flows trained on Gaussian mixture models of increasing dimensionality, and a publicly available generative model for high-energy physics jet observables. These results demonstrate that learning-based goodness-of-fit tests provide a powerful and statistically rigorous framework for validating generative models in high-dimensional scientific applications.

        Speaker: Humberto Reyes-Gonzalez (RWTH Aachen University)
      • 45
        Forecasting Generative Amplification

        Generative networks are perfect tools to enhance the speed and precision of LHC simulations. It is important to understand their statistical precision, especially when generating events beyond the size of the training dataset. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is possible in specific regions of phase space, but not yet across the entire distribution.

        Speaker: Sascha Diefenbacher (ITP Heidelberg)
      • 46
        Shapes are not enough: CONSERVAttack and its use for finding vulnerabilities and uncertainties in machine learning applications

        In High Energy Physics, as in many other fields of science, the application of machine learning techniques has been crucial in advancing our understanding of fundamental phenomena. Increasingly, deep learning models are applied to analyze both simulated and experimental data. In most experiments, a rigorous regime of testing for physically motivated systematic uncertainties is in place. The numerical evaluation of these tests for differences between the data on the one side and simulations on the other side quantifies the effect of potential sources of mismodelling on the machine learning output. In addition, thorough comparisons of marginal distributions and (linear) feature correlations between data and simulation in "control regions" are applied. However, the guidance by physical motivation, and the need to constrain comparisons to specific regions, does not guarantee that all possible sources of deviations have been accounted for.

        We therefore propose a new adversarial attack -- the CONSERVAttack -- designed to exploit the remaining space of hypothetical deviations between simulation and data after the above mentioned tests. The resulting adversarial perturbations are consistent within the uncertainty bounds --- evading standard validation checks --- while successfully fooling the underlying model. We further propose strategies to mitigate such vulnerabilities and argue that robustness to adversarial effects must be considered when interpreting results from deep learning in particle physics.

        Speaker: Timo Saala (University of Bonn (DE))
      • 47
        Variational Positive-Unlabeled Learning with Observational Bias: Uncertainty-Aware Discovery in Incomplete Binary Data

        Many scientific datasets are fundamentally incomplete: only a biased subset of true positives is ever observed, while the remainder stay unlabeled. This Positive-Unlabeled (PU) setting arises whenever detection efficiency is imperfect and covariate-dependent—a structure shared across many fields in fundamental science. In hadron collider experiments, hardware triggers record only a biased fraction of true signal events, with efficiency depending on transverse momentum and detector geometry. In time-domain astrophysics, spectroscopic confirmation of transients is strongly magnitude- and color-biased, leaving the vast majority of candidates unlabeled. Similarly, in gravitational-wave astronomy, confirmed mergers represent only the loud, well-localized tail of the true population, with detectability driven by
        extrinsic parameters independent of intrinsic source properties.

        We introduce a variational framework for transductive PU classification under observational bias, where labeled positives are drawn non-uniformly from the true positive population according to an unknown,
        covariate-dependent propensity score. The model jointly infers two latent quantities per unlabeled point: the probability of being a hidden positive (classifier score) and the probability of having been observed
        given positivity (propensity score). Both are parameterized as neural network outputs and trained within a unified ELBO that decomposes the marginal likelihood of the observed labeling pattern under the PU
        generative model, yielding calibrated credible intervals over both quantities rather than point estimates. When relational graph structure correlates with ground truth or the sampling mechanism, encoder and/or
        decoder components are replaced by GNNs, allowing uncertainty to propagate through the topology.

        The method is validated on synthetic benchmarks with known propensity functions across varying bias severities, graph topologies, and covariate regimes, demonstrating joint recovery of well-calibrated
        posterior probabilities for unlabeled points alongside accurate propensity score estimates. We then apply the model to mammal–virus association networks, where the observed host–pathogen matrix is severely incomplete due to uneven surveillance across species, viruses, and geographical regions. Our model recovers putative missing mammalian hosts across widespread viral families, providing bias-corrected
        association probabilities with principled uncertainty estimates for ecological risk assessment.

        Across ecology and fundamental physics alike, the underlying challenge is identical. By exploiting the separation between covariates driving physical class membership and those driving detection propensity,
        this framework provides a unified and probabilistic solution for recovering the full positive population from a biased subset of known positive data.

        Speaker: Gabriele Pignalberi (Sapienza Università di Roma and INFN)
      • 48
        Machine learning method for enforcing variable independence in background estimation with LHC data: ABCDisCoTEC

        Estimating the size and shape of backgrounds is a key part in the vast majority of HEP analyses. The ABCD method provides a reliable way to extract this information directly from the experimental data, thus avoiding reliance on simulation. It relies on the assumption that two, statistically independent variables can be constructed, in order to compute the expected background contribution in the signal region of the analysis. This task is however often unfeasible in practice, since even small linear or non-linear correlation can spoil the validity of the ABCD method. A novel Machine Learning approach, known as ABCDisCoTEC, is presented here replacing the two variables with decorrelated neural network scores. The engineering of optimal features for the ABCD method is enforced during the training of the neural networks by minimising a differentiable dedicated loss function. In addition, to deal with such complex minimisation problem involving multiple constraints on different loss terms, a more robust technique, based on the Modified Differential Method for Multipliers, has been employed showing an improvement in the stability and robustness of the method.

        Speaker: Cesare Cazzaniga (DESY)
      • 49
        Know What You Don't Flow

        Calibrated learned uncertainties are a key requirement also for generative neural networks in LHC physics. For a toy model with an explicit likelihood we show how a heteroscedastic and a Bayesian normalizing flow learn the systematic and statistical uncertainties on the underlying phase space density. Without an explicit likelihood we train the heteroscedastic loss on a classifier-reweighted approximate generative network. We illustrate our comprehensive approach for top pair events and show how a conditional heteroscedastic flow propagates calibrated uncertainties to all phase space directions.

        Speaker: Lorenz Vogel (Institute for Theoretical Physics (ITP), Heidelberg University, Germany)
    • 🔀 Patterns & Anomalies 3.404

      3.404

      Convener: Sebastian Dittmeier (PI)
      • 50
        4D particle tracking in the NA62 GigaTracker with transformer-based architectures

        Accurate particle tracking using the GigaTracker (GTK) silicon pixel detector represents a mission-critical stage in the data processing pipeline of the NA62 experiment at CERN, which is dedicated to the precision measurement of the ultra-rare decay $K^{+} \rightarrow \pi^{+} \nu \bar{\nu}$. Operating in a high-intensity environment with a beam rate of up to 750 MHz, the GTK provides the momentum and direction measurement as well as timing information for the incoming beam particles. Traditional tracking approaches, relying on local combinatorial algorithms, suffer from intrinsic limitations due to high pile-up conditions and the resulting combinatorial background, significantly increasing the rate of fake tracks. In this work, we present the first transformer-based reconstruction algorithm developed specifically for the NA62 GTK, designed to exploit the remarkable single hit time resolution of the detector of $\mathcal{O}$(100 ps). Formulating the tracking challenge within an edge classification framework, the architecture employs a transformer encoder to generate rich, global embeddings of each detector hit's features. These representations are subsequently used to compute connectivity scores between admissible hits across consecutive stations. The models have been trained and extensively validated on high-fidelity Monte Carlo simulation samples, demonstrating excellent generalisation capabilities and robustness across varying beam intensities and data-taking periods. The results show a sharp reduction in the number of false-positive tracks and a substantial increase in purity while maintaining a tracking efficiency comparable to or exceeding the standard algorithm. The architecture is currently used as the default reconstruction method in the NA62 C++ software framework. The improved reconstruction directly translates into enhanced performance for the downstream $K-\pi$ matching task, which is central to background rejection in the experiment.

        Speaker: Leonardo Plini (INFN-LNF)
      • 51
        Machine Learning-Based Suppression of Secondary Hit Background in Low-Resistive RPCs

        Resistive Plate Chambers (RPCs) are widely used as tracking detectors in high-energy physics experiments due to their simplicity, robustness, and excellent timing performance. However, low-resistive bakelite RPC prototypes often exhibit secondary hit components that degrade time and position resolution and introduce background, complicating track reconstruction. In this work, we present a data-driven approach to identify and suppress such background clusters using supervised machine-learning techniques. A set of fifteen compact, physically motivated cluster-level observables–capturing statistical and shape-related
        properties of time and ADC distributions–is constructed from controlled laboratory measurements with a single-gap RPC operated under a three-scintillator coincidence trigger.
        Three classification models–Deep Neural Network (DNN), one-dimensional Convolutional Neural Network (1D-CNN), and XGBoost–are trained and evaluated under identical conditions.
        All models achieve efficient signal–background separation, with XGBoost providing the most stable performance. Feature-importance analysis indicates that cluster size and temporal-shape parameters are the most discriminating observables. The proposed method is computationally
        efficient and relies only on compact cluster descriptors, making it well suited for integration into real-time or near real-time reconstruction pipelines in high-rate experiments.

        Speaker: Dr Zubayer Ahammed (Variable Energy Cyclotron Centre, India)
      • 52
        PID Reconstruction in Gaseous Detectors for Future e+e- Experiments

        PID is essential in high energy physics experiments, particularly for future e+e- colliders. The tera-z samples at FCC or CEPC offers vast flavor physics opportunities, necessitating robust hadron identification over a broad momentum range. Moreover, there is growing recognition of the performance gains that PID provides for jet flavor tagging. A key breakthrough in PID is cluster counting (dN/dx), which measures primary ionizations along a particle’s trajectory in gaseous detectors rather than relying on conventional dE/dx measurements, offering intrinsically superior resolution.

        Technically, dN/dx can be implemented in both large-volume time projection chambers (TPCs) and drift chambers (DCHs), where cluster signals are projected to 2D positional readouts in the former and to 1D timing waveforms in the latter. Reconstruction is one of the major challenges of the dN/dx technique, as one must determine the number of primary cluster signals from an environment in which the signals are highly overlapped and contaminated by secondary ionization and noise.

        Supervised learning is a statistical data-driven method capable of extracting complex information from large datasets. For TPCs, we have developed a point‑cloud algorithm using a GNN based on the Point Transformer architecture. Trained on extensive simulated samples, the model enhances pion–kaon separation by 10–20% compared to conventional truncated mean methods. For DCHs, we have developed a model based on LSTM and DGCNN that achieves a remarkable 10% improvement in pion-kaon separation on simulated samples related to traditional methods. Furthermore, for test-beam data samples collected at CERN--where label scarcity and data/MC discrepancies pose challenges--we have developed a semi-supervised domain adaptation model that outperforms traditional methods and maintains consistent performance across varying track lengths.

        Three papers have been published related to this presentation:

        Speaker: Guang Zhao (Institute of High Energy Physics (CAS))
      • 53
        GBDT based unified PID scheme for the CBM experiment

        The Compressed Baryonic Matter (CBM) experiment is a fixed target heavy-ion collision experiment currently being built at the Facility for Antiproton and Ion Research (FAIR), Darmstadt, Germany.
        The CBM experiment is designed to probe the QCD phase diagram at moderate temperatures and high net-baryon density via different physical observables, including flow, fluctuations, and correlations of the final state hadrons.
        Reconstruction and identification of these hadrons is achieved by a combination of a silicon tracking system (STS) and particle identification detectors, specifically the transition radiation detector (TRD) and time of flight detector (TOF).

        It is imperative to preserve a high level of purity in the identification of hadrons in order to investigate rare physics observables.
        In this regard, a unified framework based on machine learning for the identification of hadrons by combining responses from different sub-detectors is developed.
        In the first iteration, gradient-boosted decision trees (xGBOOST) are used as base models.
        This contribution focuses on the implementation of the models and the recent results achieved through their application.
        Furthermore, a comparison to the previously used conventional cut-based method is presented.

        Speaker: Pavish Subramani (University of Wuppertal)
      • 54
        Generative and geometric learning for comprehensive collider event reconstruction

        In particle collider experiments, event reconstruction is the task of inferring the kinematics of the short-lived, 'interesting' particles of the hard scatter from the stable final states recorded by detectors. We decompose event reconstruction into two tasks: assigning measured jets and leptons to parent particles, and regressing unmeasured neutrino kinematics. We present a novel geometric learning framework that represents collider events as hypergraphs, achieving state-of-the-art performance in the assignment task. The hypergraph model is interfaced with a generative diffusion model to infer inherently multimodal neutrino kinematics distributions. We showcase the approach in several proton-proton collision processes, demonstrating new avenues for analysis in Higgs, electroweak boson and top-quark physics.

        Speaker: Ethan Simpson (The University of Manchester (GB))
      • 55
        Encoding locality and geometry in a Transformer encoder for muon triggers

        Muon reconstruction in high-energy physics experiments relies on identifying detector hits originating from muons and fitting trajectories through them to extract track parameters. This process currently consists of separating signal hits from background before the track reconstruction can be performed. In this work, we present a deep learning solution for discriminating between background and signal hits which could be used for a muon trigger.

        Specifically, we propose a compact Transformer encoder architecture to act as a binary classifier that filters out irrelevant background. We investigate the added benefit of a distance-based attention mask for inserting a locality prior, and of the novel "location encoding" that enables the model to learn the detector geometry. The proposed architecture is evaluated on a toy sample simulating muon drift tube chambers, for which we report high background rejection and signal efficiency rates.

        Speaker: Nadezhda Dobreva
      • 56
        Can multi-modality improve IACTs event reconstruction in heterogeneous observing conditions?

        Deep learning methods have emerged as a powerful alternative to reconstruct physical properties directly from images of atmospheric events acquired by Imaging Atmospheric Cherenkov Telescopes (IACTs). In the context of the Cherenkov Telescope Array Observatory (CTAO) first Large-Sized Telescope (LST-1), the specialized architecture gamma-PhysNet has demonstrated strong performances on simulated and real data in constrained conditions. However, image acquisition often involves heterogeneous observation conditions not explicitly incorporated into the learning process, limiting potential performance gains.

        In order to take these additional variables into account, we tested and compared two multi-modal architectures against the vanilla version of gamma-PhysNet: (1) a Conditional Batch Normalization (CBN) model using Night Sky Background (NSB) levels and telescope pointing directions as conditional inputs, and (2) an intermediate information fusion architecture with NSB and pointing directions as additional inputs.

        Using simulated data with zenith angles from 6° to 30° and added NSB noise, we found that neither the CBN nor the fusion architecture significantly improved energy or angular resolution compared to gamma-PhysNet. Additionally, these approaches introduced unnecessary architectural and training complexity.
        Our study suggests that in our case, a simpler architecture with enough training data is a preferable solution.

        Speaker: Thomas Vuillaume (LAPP, CNRS, USMB)
      • 57
        Recent Results from Anomaly Detection Searches on CMS

        In the absence of direct evidence for new physics in targeted searches, model-independent strategies are becoming increasingly important. In this talk, we present recent results of model-agnostic searches that are facilitated by advanced machine learning techniques, opening a new avenue for unbiased detection of potential new physics signals.

        Speaker: Jovin Drews (Universität Hamburg)
    • 🔀 Simulations & Generative Models HS2

      HS2

      Convener: Marco Letizia (Università di Genova)
      • 58
        Generative diffusion models for non-abelian gauge theories on the lattice

        Diffusion models are a widely used method in generative AI to produce images and videos. I will discuss the application to lattice field theory and the connection with methods known from theoretical physics, such as stochastic quantisation. Applications to non-abelian gauge theories using gauge-equivariant convolutional neural networks are presented.

        Speaker: Gert Aarts (Universitaet Bielefeld)
      • 59
        Diffusion Models for SU(2) Lattice Gauge Theory in Two Dimensions

        We apply score-based diffusion models to two-dimensional SU(2) lattice pure gauge theory with the Wilson action, extending recent work on U(1) gauge theories. The SU(2) manifold structure is handled through a quaternion parameterization. The model is trained on 10,000 configurations generated via Hybrid Monte Carlo at a fixed coupling $\beta_0= 2.0$ on an $8\times 8$ lattice, augmented to 20,000 samples via random gauge transformations. Through physics-conditioned sampling exploiting the linear $\beta$-dependence of the score function, we generate configurations at different values of the coupling without retraining; through the fully convolutional U-Net architecture with periodic boundary conditions, we generate configurations on lattices of different spatial extents. We validate our approach by comparing the average plaquette and Wilson action density against exact analytical predictions. At the training lattice size ($8\times 8$), the model reproduces the exact plaquette with biases $|\Delta| \leq 0.001$ for $\beta \in [1.5, 2.5]$ and $|\Delta| < 0.06$ across $\beta \in [1, 4]$.~For lattices sharing the training extent $L=8$ in at least one direction, biases remain below $\sim 0.003$ for $\beta \in [1.5, 2.5]$, with larger deviations at higher couplings. This work demonstrates that diffusion models are a promising tool for non-Abelian gauge field generation and motivates further investigation toward higher-dimensional theories.

        Speaker: Bao-Dong Sun (Ruhr-Universität Bochum)
      • 60
        Improved Lattice Simulations from Machine-Learned Gauge Actions

        Numerical simulations of quantum field theories are indispensable for precision studies of the Standard Model, yet extracting continuum physics from discretized spacetime remains challenging due to lattice artifacts. Renormalization-group improved fixed-point (FP) actions provide an elegant solution by suppressing these artifacts, but their complexity has long prevented their practical use. Advances in machine learning, in particular lattice gauge-equivariant convolutional neural networks (L-CNNs) [1], enable accurate parameterizations of FP actions.

        In this talk, we present recent results from our collaboration [2], demonstrating that Monte Carlo simulations using machine-learned FP actions for four-dimensional SU(3) gauge theory drastically reduce discretization effects in gradient-flow observables, allowing continuum physics to be extracted reliably even at coarse lattice spacings. Furthermore, we establish that the gradient flow of this action is tree-level artifact-free to all orders in the lattice spacing, proving its classically perfect nature and validating the quality of the L-CNN parameterization without introducing secondary artifacts.

        Finally, we report on ongoing work exploring LLM-guided evolutionary optimization of custom GPU kernels for gauge-equivariant neural networks using the OpenEvolve framework. We discuss the challenges and insights of targeting memory consumption and execution efficiency, demonstrating how automated code evolution can help scale physics-informed AI architectures to large-scale scientific applications.

        [1] M. Favoni, A. Ipp, D. I. Müller, D. Schuh, Phys. Rev. Lett. 128 (2022), 032003, https://doi.org/10.1103/PhysRevLett.128.032003 , https://arxiv.org/abs/2012.12901
        [2] K. Holland, A. Ipp, D. I. Müller, U. Wenger, Phys. Rev. Lett. 136, 031901 (2026), https://doi.org/10.1103/k41k-2pnc , https://arxiv.org/abs/2504.15870

        Speaker: Andreas Ipp (TU Wien)
      • 61
        Generative AI for Extreme QCD Matter: from lattice path intagrals to fast heavy-ion event generation

        Generative AI is rapidly becoming more than a tool for producing realistic images: in physics, it can be understood as a learnable machinery for transporting, sampling, and interrogating high-dimensional probability measures. In this talk I will discuss how this viewpoint opens new routes for exploring QCD matter under extreme conditions, where both first-principles theory and event-by-event simulations face severe computational bottlenecks. I will first introduce a unifying physical picture in which normalizing flows, diffusion models, flow matching, and stochastic path samplers are interpreted as finite-time nonequilibrium transports between probability distributions. I will then connect this framework to recent applications in lattice field theory, including diffusion models as stochastic quantization, the reconstruction of higher-order connected correlations, topology-aware sampling, and learning effective real distributions behind complex Langevin dynamics for complex actions. The second part of the talk will focus on heavy-ion collisions, where generative models can act as ultra-fast event-level surrogates of microscopic transport and hybrid simulations. In particular, I will highlight work developed in the joint BMBF-KISS program for HIC simulation sector, including point-cloud diffusion models such as HEIDi for generating complete UrQMD-like final-state hadron events rather than preselected observables. I will close by discussing how such generative models may interface with Bayesian inference, neural unfolding, and inverse problems, potentially enabling a new generation of simulation-accelerated, uncertainty-aware studies of the QCD phase structure.

        Speaker: Prof. Kai Zhou (CUHK-Shenzhen)
      • 62
        Normalising Flow-Assisted Neural Quantum States for Ground-State Estimation

        Neural quantum states provide expressive variational representations of quantum many-body wavefunctions. However, their practical performance depends on how well they can sample the relevant configurations from an exponentially large Hilbert space. Conventional Markov chain Monte Carlo methods can mix slowly between separated high-probability regions, particularly in frustrated and strongly correlated systems.

        We introduce a hybrid variational framework in which a continuous normalising flow learns an auxiliary sampling distribution over a discretised effective subspace, while an independent variational ansatz learns the wavefunction amplitudes. This separates the task of finding the relevant support from the task of estimating the amplitudes within it. The normalising flow can therefore explore non-local regions of configuration space without relying on a sequence of local updates.

        We apply the method to the square-lattice J1–J2 Heisenberg model and compare it with conventional Metropolis sampling using matched variational ansätze and optimisation settings. Our results show that flow-assisted sampling remains competitive across the system sizes studied and improves variational ground-state estimation in several regimes. This suggests that generative-model-assisted sampling can provide a useful practical approach to quantum many-body simulation.

        Speaker: Timur Sypchenko (IPPP)
      • 63
        Reconstruction of Gravitational Form Factors using Generative Machine Learning

        Extracting hadronic form factors from sparse and noisy lattice QCD data typically relies on parametric ansätze, introducing model-dependence. We present a generative framework based on denoising diffusion models for parameterisation-independent reconstruction of these quantities. The generative prior is built from a large ensemble of synthetic curves drawn from distinct functional classes rooted in different theoretical approaches to hadron structure. Applied to the proton gravitational form factors, the framework yields model-independent reconstructions consistent with lattice QCD, remaining robust even with only one or two conditioning points. The densely sampled posterior enables a direct extraction of the chiral low-energy constants c8 and c9​, yielding the nucleon D-term.

        Speaker: Iuliia Panteleeva
      • 64
        Sampling with physics-informed kernels

        The physics-informed renormalisation group describes general reparametrisation of given data distributions. It provides an analytic flow in terms of the distribution itself and the differential reparametrisation, the physics-informed kernel. From the perspective of quantum field theory, this flow can be understood as a general non-linear coarse graining procedure. From the perspective of Machine learning architectures, this flow provides a general layerwise propagation with a general activation function. Hence, it encompasses generative architectures such as normalising flows, PINNs or diffusion models as well as traditional sampling approaches such as sampling with, HMC, (complex) Langevin equations and on Lefschetz thimbles as special cases, endowed with an additional key property, the weight preservation of the map.

        I will provide a short introduction to the approach, concentrating on the key property of weight-preservation of the map with the physics-informed kernel. It can be either used as a standard sampling algorithm or as an additional generative structure in generative NNs. The talk closes with a brief discussion of its implementation for sampling real distributions as well as complex distributions with a sign problem.

        More details and applications will be provided in the talks of Friederike Ihssen and Renzo Kapust.

        Speaker: Jan Pawlowski (Heidelberg university)
      • 65
        Physics-informed kernels: Utilising stochastic processes for complex problems

        Computing observables with respect to highly oscillatory or complex-valued distributions is an important task in physics, in particular in quantum field theory. For complex distributions, however, traditional sampling algorithms either lose their straightforward applicability or are rendered inefficient by the oscillations, which cause a signal-to-noise problem that scales exponentially with the volume - the sign problem.

        Physics-informed kernels [1,2] are a generative method that has been used to tame such sign problems by computing a map that connects a real-valued proper probability distribution to the targeted complex-valued distribution. In principle, stochastic processes can be used to compute these maps. However, the application of stochastic processes to complex-valued theories comes with its own difficulties related to runaway trajectories or possible convergence to wrong solutions.

        In this talk, we will discuss how the physics-informed kernel framework itself can be used to evade these obstructions, allowing transport maps for complex-valued distributions to be computed safely using stochastic processes.

        [1] F. Ihssen, R. Kapust, J. M. Pawlowski , arXiv:2510.26678

        [2] F. Ihssen, R. Kapust, J. M. Pawlowski , arXiv:2603.03159

        Speaker: Renzo Kapust (University of Heidelberg)
    • 🧠 Working groups: AI-assisted Co-Design of Future Ground- and Space-based Detectors (WG2) 2.404

      2.404

      • 66
        AI-assisted Co-Design of Future Ground- and Space-based Detectors (WG2)
    • 12:30
      🍽️ Lunch Foyer

      Foyer

    • 🗣️ Plenaries HS1

      HS1

      Convener: Prof. Kai Zhou (CUHK-Shenzhen)
      • 67
        Architectural Challenges in Enterprise Agentic AI

        At barra, we implement agentic AI projects within complex, large-scale enterprise environments. The enterprise landscape introduces significant organisational complexity as well as rigid requirements for stability, scalability, and economic viability. Navigating these requirements demands a departure from experimental setups. This talk provides an overview of our operational learnings, specifically focusing on the hybrid architectural patterns necessary to reconcile the non-deterministic flexibility of LLMs with the reliability of classical software systems. We will discuss methodologies for choosing optimal agent topologies, such as balancing monolithic versus mesh architectures based on organizational constraints. Furthermore, we address the critical task of integrating deep organizational knowledge into the agent's context. Finally, we explore strategies for observability and output evaluation to handle unpredictable agent behaviors. Our experience offers a perspective on how to mature agentic systems from prototypes into robust, production-ready enterprise tools.

        Speaker: Dario Hügel (Barra Labs)
      • 68
        Foundation Models and Scale: Deep Networks unlock deep mysteries
        Speaker: Nicole Hartman (TUM)
    • 15:30
      ☕ Coffee Foyer

      Foyer

    • 🗣️ Plenaries HS1

      HS1

      Convener: Caterina Doglioni (University of Manchester)
      • 69
        Particle Accelerators, Beam Instabilities and Data-driven Modelling
        Speaker: Adrian Oeftiger (University of Oxford, John Adams Institute)
      • 70
        Machine Learning for the SKA
        Speaker: Caroline Heneka
    • 🗣️ Plenaries HS1

      HS1

      Convener: Johan Messchendorp (GSI GmbH/Ruhr Universität Bochum)
      • 71
        The Mechanics of Learning: Inverse Problems

        ML techniques are routinely used to solve Inverse Problems in particle physics. Focussing on the case of PDF fitting as a testbed, we develop a description of the training dynamics during gradient descent. Our analysis shows that the ML solution is one way of regulating an otherwise ill-defined problem and allows us to address the role of hyperparameters in the training process. A critical comparison with other regulators will yield a better understanding of systematic errors (misspecification) in the different fitting procedures.

        Speaker: Luigi DelDebbio (University of Edinburgh)
      • 72
        AI for Gravitational-Wave Data Analysis
        Speaker: Maximilian Dax (ELLIS Institute Tübingen)
    • 10:30
      ☕ Coffee Foyer

      Foyer

    • 🗣️ Plenaries HS1

      HS1

      Convener: Martin Erdmann
    • 12:30
      🍽️ Lunch Foyer

      Foyer

    • 🔀 Foundation Models 1.404

      1.404

      Convener: Tobias Golling (University of Geneva)
      • 75
        Building foundation models with low-level collision data in particle physics

        The integration of foundation models in particle physics is gaining pace rapidly and has expanded the search for new physics. This talk presents foundation models trained on low-level data from the first fully simulated dataset using Open Data Detector (ColliderML), to distinguish between Standard Model and Beyond Standard Model processes. We compare new physics discovery using only low level tracking information with the inclusion of more sources, such as different scales (tracks, particle flow objects) and different detector regions (calorimetry). The effect of these additions in both the supervised and unsupervised case is presented. The talk will further discuss the performance of these approaches in anomaly detection, highlighting the potential of pre-trained foundation models on low-level data for various downstream tasks.

        Speaker: Shreya Saha (Adelaide University)
      • 76
        JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics

        Jet tagging at the Large Hadron Collider (LHC) increasingly relies on deep learning models trained on massive simulated datasets, leading to high computational costs and limited robustness to detector mismodeling. We introduce JetParticle-JEPA (JP-JEPA), a self-supervised Joint-Embedding Predictive Architecture that learns physically meaningful jet representations directly from continuous particle clouds without tokenization or reconstruction of raw inputs. Built on a Particle Transformer backbone, JP-JEPA predicts latent representations of masked particles while preserving fine-grained kinematic correlations. On the JetClass benchmark, JP-JEPA achieves performance comparable to fully supervised state-of-the-art methods on the full dataset, surpasses supervised baselines in low-label regimes, and significantly outperforms existing SSL approaches. On Top Quark and Quark-Gluon Tagging benchmarks, it remains on par with supervised methods. The learned representations also exhibit strong robustness to missing detector information and improved uncertainty behavior, highlighting JP-JEPA as a promising foundation-model framework for robust and data-efficient jet physics at the LHC.

        Speaker: Guillaume Letellier (GREYC - University of Caen Normandy, France)
      • 77
        Foundation Models for Anomaly Detection in HEP

        Foundation models have recently emerged as powerful tools for learning general-purpose representations that can outperform task-specific models while being generalizable to multiple downstream tasks. In this project, we explore their application to particle physics by training a foundation model on simulated jet data from hadronic collisions. The goal is to build a model that can learn physically meaningful representations of collision events and to use the model as a general tool for different tasks such as anomaly detection, jet reconstruction, and jet tagging. To this end, a foundation model based on the Perceiver architecture was developed and trained on the CoLLiDe-2V dataset, which provides simulations of hadronic background processes in Large Hadron Collider (LHC) detectors. By utilizing self-attention and cross-attention mechanisms, the model processes all jets and their constituents within an event simultaneously to capture inter-jet correlations. Crucially, it maps these tokens onto a fixed-size latent space, maintaining a linear computational complexity relative to the number of input objects, which is essential for scaling to High-Luminosity LHC (HL-LHC) pile-up conditions. The model was pretrained without labeled data via self-supervised learning, using multiple parallel decoding heads tasked with masked-object reconstruction, event-level property prediction, and consistency constraints. After the model was trained on QCD, W and Z into hadronic events, it was tested for anomaly detection using tt̅into hadronic as the signal events. To evaluate the learned representation, a Variational Autoencoder and a Logistic Regression were applied both to the Perceiver's latent space and directly to the raw data. Probing the latent space yielded a competitive performance compared to their raw-data baselines. Transfer learning experiments further validated the representation, showing that a frozen pretrained Perceiver converges faster and achieves comparative performance to training an anomaly detection model from scratch. Finally, an explainability study using Linear Probes and Gaussian Mixture Models (GMM) confirmed that the latent space does not rely on trivial correlations; instead, it geometrically separates anomalies from the background and explicitly encodes physically meaningful quantities, including jet pT, jet mass, and complex inter-object kinematic relationships.
        These results demonstrate that a Perceiver-based foundation model can successfully learn robust, transferable, and physically interpretable representations of collision events without supervision, outperforming conventional task-specific baselines. If extended to the full Standard Model, this approach could provide a model-agnostic way to search for new physics over a wider range than conventional methods.

        Speaker: Desheng Yang (Sapienza Università di Roma and INFN sezione di Roma)
      • 78
        Jet representations transfer across radii: toward foundation models for high-energy physics

        The high-energy physics pipeline already resembles a multi-modal foundation model: a general-purpose representation built once and reused across downstream tasks - but it lacks scale and differentiability. Already, we are seeing the gains from large-scale pre-training in individual pieces of the ATLAS pipeline, such as jet flavour identification. Flavour taggers, however, are normally trained in isolation - a separate model from scratch for each task such as small-R and large-R classification - leaving the benefits of scale untapped. In this work, using ATLAS simulation, we show that the representations learned by a small-R tagger, pre-trained on billions of jets, transfer to the large-R H→bb tagging task, demonstrating that the tagger learns general jet features that generalise across radii. This goes beyond H→bb: the same transfer that carries across radii motivates a unification of taggers - a step toward the foundation-model view of the high-energy physics pipeline.

        Speaker: Una Alberti (University of Bern)
      • 79
        Toward Domain-Invariant Tokenization for Jet Foundation Models

        Tokenization plays a central role in modern foundation models, especially those relying on next-token prediction, and has been extensively studied in domains such as language and vision. In high-energy physics, however, the adaptation of tokenization to continuous jet data from collider experiments remains at an early stage, with few studies on tokenizer performance, comparable evaluation metrics, or generalization across datasets. This is particularly pressing given that jet data are produced by different generators and in various data-taking conditions, raising the question of whether a tokenizer trained on one domain transfers to another.

        When benchmarking standard k-means, product-quantized, and neural (autoencoder-based) tokenizers on four jet datasets (RODEM, JetClass, JetSet, Aspen) using codebook utilization, perplexity, and reconstruction error as comparable evaluation metrics, we encounter domain shift leading to performance degradation. We address this issue by proposing an invariant tokenization strategy, built on a jet-radius-normalized feature space rather than detector-specific raw kinematics. This approach can generalize across jets from different generators, experimental setups and data-tracking conditions.

        Speaker: Giovanni Ottaviano (Sorbonne University)
      • 80
        SPADE: Split-and-Delay Embeddings for Autoregressive High-Granularity Calorimeter Simulation

        Autoregressive transformers are increasingly applied outside the language domain that motivated them, but the tokenization step transfers poorly. Scientific data are already numerical, often partly discrete, and typically live in high-dimensional spaces where each token carries several features. The standard workaround, compressing feature vectors into a single token via a learned codebook, introduces reconstruction loss and a vocabulary that grows multiplicatively with resolution, inflating the embedding and unembedding layers until training becomes prohibitive.
        We introduce SPADE (SPlit And Delay Embeddings), which embeds each feature of a token independently and staggers the resulting streams along the sequence with progressively increasing delays. The vocabulary then scales additively rather than multiplicatively, while intra-token correlations are recovered by the ordinary causal self-attention mechanism: each feature is predicted at its own sequence position, conditioned on the features already emitted for the same object. No auxiliary decoder or quantization stage is required.
        We demonstrate SPADE on point-cloud calorimeter shower generation in the highly granular ILD electromagnetic calorimeter. SPADE is competitive with the state-of-the-art flow-matching model AllShowers on photon showers and substantially outperforms its VQ-VAE-based predecessor OmniJet-$\alpha_\text{C}$. Against a joint-vocabulary baseline at the finest granularity studied, SPADE uses 74× fewer parameters and converges 6.9× faster in GPU hours, while better reproducing observables sensitive to energy–position correlations.
        Because the mechanism assumes only that tokens carry multiple features, discrete or continuous, it offers a route to LLM-style pretraining on high-dimensional sensor data across fundamental physics.

        Speaker: Henning Rose (Uni Hamburg)
    • 🔀 Inference & Uncertainty HS1

      HS1

      Convener: Theo Heimel (UCLouvain)
      • 81
        NS-UNO: Neutron-Star EoS Inference from an Unconstrained Number of Observations

        Future multimessenger observations of neutron stars (NS) are expected to substantially increase both the number and precision of astrophysical constraints on the equation of state (EoS) of dense matter. This motivates inference frameworks capable of accommodating a variable, non fixed number of observations while preserving the posterior information associated with each measurement.
        In this work, we introduce NS-UNO, a Neural Posterior Estimation framework for NS EoS inference designed to accommodate an Unconstrained Number of Observations (UNO). NS-UNO combines a hierarchical DeepSets model with a conditional normalising flow, enabling a single trained model to perform inference from mass radius observation sets of varying size, with each observation represented by a set of posterior samples.
        We demonstrate accurate and well calibrated posterior reconstructions using a model trained jointly on piecewise polytropic and nonparametric Gaussian process EoS ensembles. The reconstruction improves as observations probe a broader range of NS masses, while remaining robust to variations in the number and precision of the observations. The model also generalises to EoSs outside the families used during training. Finally, we qualitatively demonstrate the framework on current multimessenger constraints from NICER and GW170817. NS-UNO provides a flexible and scalable approach to NS EoS inference, naturally suited to the increasingly diverse observational datasets expected from next generation multimessenger astronomy.

        Speaker: Valéria Carvalho (Universidade de Coimbra and Nicolaus copernicus astronomical center)
      • 82
        Generative inference of the neutron-star equation of state

        Neutron-star observations encode the equation of state of cold, dense matter, but reconstructing it is an ill-posed inverse problem, and standard analyses assume fixed functional forms that restrict and unevenly weight the admissible equations of state. We reconstruct the equation of state with a denoising diffusion model trained on a large synthetic ensemble of sound-speed profiles spanning hadronic, hybrid and quark-matter behaviour. Chiral effective field theory anchors generation at low density through inpainting; perturbative QCD and multimessenger data: NICER radii, the tidal deformability of GW170817 and heavy-pulsar masses, enter afterwards through importance reweighting, keeping the learned prior and the observational likelihood cleanly separated. The posterior favours a soft equation of state and indicates near-conformal matter in the cores of the heaviest stars, while current data remain nearly uninformative about a strong first-order phase transition.

        Speaker: Iuliia Panteleeva
      • 83
        Simulation-Based Inference in the Search for the Neutron's Permanent Electric Dipole Moment

        Precision tests of fundamental symmetries frequently rely on multi-stage experimental setups where distinct physical mechanisms shape a common observable via an intractable likelihood. Conditional invertible neural networks, trained on high-fidelity forward simulations, learn to reconstruct the full multidimensional posteriors over the parameter space directly from detector-level observables. Applied to ultracold neutron (UCN) storage experiments, which underlie searches for the CP-violating neutron permanent electric dipole moment (nEDM), the trained network accounts for a complex instrument response to disentangle competing capture and decay loss channels. As the next generation of nEDM experiments addresses the core challenge of limited statistics, generative inference offers a path to resolve new systematics via conditional summaries of particle-level simulations.

        Speaker: Husain Mustansir Manasawala (Universität Heidelberg)
      • 84
        Neural Simulation-Based Inference for the keV-Scale Sterile Neutrino Search with TRISTAN at KATRIN

        Following the completion of its neutrino mass measurement program at the end of 2025, the KATRIN experiment aims to probe keV-scale sterile neutrinos by analyzing the full tritium beta decay spectrum with a novel detector system, TRISTAN. Leveraging KATRIN’s high source activity, this search is sensitive to mixing amplitudes at the parts-per-million level. However, extracting a potential sterile neutrino signature is challenging, as it relies on detailed modeling of the the observed tritium spectrum and requires computationally intensive Monte Carlo simulations. To address this challenge, we implement neural simulation-based inference using normalizing flows to approximate the underlying probability density of the physics simulation. We demonstrate that continuous normalizing flows trained via conditional flow-matching enable highly efficient modeling of experimental spectra. This approach opens up the possibility of a fast surrogate model for rapid sampling and generates a continuous, unbinned representation of the KATRIN beamline response, accelerating and enabling the analysis pipeline.

        Speaker: Luca Fallböhmer (Max-Planck-Institute for Nuclear Physics)
      • 85
        Physics-Guided Diffusion Models for Observable Reconstruction from Sparse and Noisy Data

        Extracting continuous physical quantities from sparse and noisy data is a ubiquitous inverse problem across theoretical physics. In this talk, I will present a physics-guided generative framework based on denoising diffusion probabilistic models that reformulates function reconstruction as a conditional generation (inpainting) problem. Rather than relying on a prescribed functional ansatz, the method learns a generative prior from a diverse ensemble of theoretically motivated functional forms, enabling parametrization-independent reconstructions while preserving known physical constraints. The diffusion process naturally produces posterior ensembles, providing robust uncertainty quantification even in the regime of extremely limited data. I will discuss the methodology, its generality, and its potential as a new tool for inverse problems in theoretical physics. To illustrate its versatility, I will briefly highlight applications to the reconstruction of hadronic gravitational form factors.

        Speaker: Herzallah Alharazin (Ruhr Universität Bochum)
      • 86
        Simulation-Based Inference for Antenna Gain Calibration in 21 cm Cosmology

        The 21 cm signal from neutral hydrogen is a key probe of the Epoch of Reionization (EoR), marking the universe’s transition from a cold, neutral state to a predominantly hot, ionized one, driven by the formation of the first stars and galaxies. Extracting this faint 21 cm signal from radio interferometric data requires precise gain calibration. However, traditional calibration methods are computationally expensive and time-intensive. With next-generation radio telescopes like the Square Kilometer Array (SKA) set to host hundreds of antennas, more efficient calibration techniques are urgently needed.

        To address this challenge, we present a sequential simulation-based inference (SBI) approach for direction-independent gain calibration, designed to automate and accelerate the process while improving scalability and accuracy. Once a forward model is established to generate simulations - transformations of the true sky image due to antenna gain variations - neural posterior estimation (NPE) with embedding networks is employed to infer the correct gain values for multiple antennas from the joint parameter-data distribution. To efficiently estimate the large number of gain parameters within
        a feasible time frame, we leverage GPU-accelerated parallelisation. The Bayesian framework enables robust uncertainty estimation, which traditional methods often overlook, while also facilitating faster and more reliable analysis on real SKA data.

        Future work could extend this approach to direction-dependent gains and other systematic effects or involve validation on existing radio data. By integrating these techniques into the analysis pipeline, we can fully exploit SKA’s unprecedented sensitivity, significantly improving our ability to extract fundamental cosmological insights from large-scale observations.

        Speaker: Joy Sanghavi (UvA)
    • 🔀 Real-Time Data Processing 3.404

      3.404

      Convener: Wouter Verkerke
      • 87
        Real-time data processing model at LHCb and ML/AI: evolution from an extension to a core functionality

        The LHCb experiment has deployed machine learning and artificial intelligence models in its real-time data processing from the start of Run 1 data taking, and by Run 3 such models are an integral part of the trigger.

        This talk will describe the usage of machine learning and AI models and algorithms within the LHCb real-time analysis data processing paradigm as part of both reconstruction and online selection, as well as developments in using such algorithms for online monitoring and software QA.

        Moreover, for the LHC Run 5, the LHCb collaboration is proposing to build a second upgrade of its detector, targeting the creation of an ultimate flavour factory machine in the forward region at the LHC. This talk will also sketch the unprecedented challenges this proposal will pose to the real-time reconstruction and selection of physics-quality signals and the ways in which machine-learning and AI models are anticipated to play a central role.

        Speaker: Miroslav Saur (Lanzhou University)
      • 88
        Metric Learning for Graph Neural Network-Based Track Reconstruction at the ATLAS Event Filter

        The High-Luminosity LHC (HL-LHC) will deliver proton-proton collisions at unprecedented pile-up conditions, making fast and efficient track reconstruction a key computational challenge. Graph Neural Networks (GNNs) offer a promising approach to this problem.

        A critical first step in GNN-based charged-particle tracking is graph construction, which can be carried out with the Metric Learning method. In this approach, a Multi-Layer Perceptron embeds hits into a latent space, where successive hits from the same particle are clustered and unrelated hits are separated; edges are then formed between hits that lie close together in this embedding.

        This contribution presents the Metric Learning stage of a GNN-based tracking pipeline developed for the ATLAS Event Filter, the software-based trigger responsible for the final online event selection at the HL-LHC. Particular emphasis is placed on strategies for hardware-constrained deployment. Regional training, in which models are trained on geometrically partitioned subsets of the ATLAS Inner Tracker to reduce their size and complexity, is investigated. In addition, quantization and compression techniques, originally developed for FPGA implementation within the ATLAS Event Filter, are evaluated. The method’s performance is presented, together with an assessment of these optimization strategies for real-time operation. The optimization strategies discussed are of general relevance to trigger and data acquisition systems operating under similar latency and resource constraints, with the potential for implementation on GPUs as well as FPGAs.

        Speaker: Giulia Fazzino
      • 89
        CNN-Based Trigger for QGP-Sensitive Event Selection in High-Rate Heavy-Ion Experiments

        High-rate heavy-ion experiments require fast and robust event-selection strategies capable of identifying rare physics signatures in real time under severe throughput and storage constraints. We present a convolutional-neural-network-based trigger concept for the selection of events associated with quark–gluon plasma (QGP) formation. The method represents each collision event as a compact multidimensional histogram of reconstructed particle content, encoding particle species, momentum magnitude, polar angle, and azimuthal angle. This representation is processed by a lightweight 3D CNN architecture designed for fast inference in online or quasi-online analysis environments.

        The classifier is trained and validated using events generated with the Parton–Hadron–String Dynamics (PHSD) transport approach, where microscopic QGP-related event labels are available. To assess model dependence, the same event representation and network architecture are further tested with UrQMD-based simulations, including cross-model validation between PHSD and UrQMD. This allows us to investigate whether the network learns robust QGP-sensitive final-state structures rather than generator-specific correlations. In addition, SHAP-based interpretability analysis is used to identify the particle species and phase-space regions most relevant for the network decision, with strange hadrons and antibaryon-related features providing particularly important contributions.

        For realistic deployment, the method is evaluated along the transition from idealized generator-level information to reconstructed events in the CBM/FLES analysis chain. Using ANN4FLES for lightweight C++ inference, the classification accuracy for Au+Au collisions at 30 AGeV decreases from 95.1% at PHSD generator level to 83.7% after full reconstruction, indicating that substantial QGP-sensitive information survives detector acceptance, tracking, and topology reconstruction effects. These results demonstrate the potential of model-robust AI triggers for fast event enrichment in future high-rate heavy-ion experiments.

        Speaker: Olga Soloveva (GSI Helmholtz Zentrum fuer Schwerionenforschung GmbH)
      • 90
        Deep Learning for detector alignment at LHCb

        The LHCb experiment at CERN's Large Hadron Collider (LHC) relies on precise alignment of its tracking detectors comprising the Vertex Locator (VELO), Upstream Tracker (UT), and Scintillating Fibre (SciFi) Tracker to achieve the momentum, mass, and vertex resolutions demanded by the LHCb physics programme. During data-taking, detector alignment at LHCb is performed on CPUs using an iterative χ² minimisation procedure based on Kalman-filter track fits. This algorithm typically requires large event samples and substantial computational resources.

        Within the ErUM-Data project DEEP, we investigate novel Deep Learning (DL) approaches for detector alignment at LHCb, with the long-term goal of deployment on GPUs, FPGAs, and System on Chip (SoC) devices. Such an approach could enable fast, energy-efficient detector alignment, supporting the increasing computing demands of LHCb Run 4 and future upgrades.

        Detector alignment is a challenging inverse problem with inherent ambiguities. Translational and rotational misalignments can produce similar residual signatures. Moreover, the residual measured in a detector module depends not only on its own misalignment but also on the misalignments of all detector modules traversed by the track and on uncertainties in the track reconstruction. Learning these correlations directly from reconstructed tracks therefore represents a non-trivial machine learning task that requires carefully designed model architectures and representative training data.

        We develop and evaluate DL architectures that infer VELO module misalignments directly from track-level information. Initial results on simulated data demonstrate the feasibility of a physics-guided, PointNet-based approach, achieving root mean square errors of approximately 5 μm for translations and 70 μrad for rotations using fewer than 300 events, even when all detector modules are simultaneously and randomly misaligned. These results demonstrate the potential of deep learning-based detector alignment and motivate further development towards deployment on heterogeneous computing platforms for real-time data processing and analyses.

        Speaker: Manjunath Omana Kuttan (Albert-Ludwigs-Universität Freiburg)
      • 91
        Symbolic Regression as a Model Compression Tool for Neural Networks used in Track Seeding

        Machine learning techniques have demonstrated great potential for efficiently addressing the increasing combinatorial complexity of track finding at high-luminosity collider experiments. To achieve high computational efficiency, neural network models often require compression, a necessity that becomes increasingly important for the large models used in GNN-based tracking pipelines. This is particularly true for FPGA implementations targeting low-latency hardware trigger applications, where on-chip compute and memory resources are limited. While common approaches include quantisation and pruning, this contribution investigates symbolic regression (SR) as a model compression tool. SR is a regression algorithm that searches the space of algebraic operators to find analytic functions that best fit an underlying dataset. The aim is to replace the neural network models involved in GNN tracking with a set of functions that require less computational resources.
        In this contribution the open data ColliderML dataset is used to study this approach on networks used for track seeding within regions of the pixel system of the OpenDataDetector. The focus is on Metric Learning which is a machine learning based approach to construct a graph from hit data. Results of applying an SR based approximation scheme to a Metric Learning multilayer perceptron used for regional track seeding are presented. Furthermore, the computational resource usage on FPGA of the found analytic function sets is compared to that of other compression schemes.

        Speaker: Urs Moritz Fischer
      • 92
        Machine Learning Models for Anomaly Detection at HADES

        Maintaining data quality in real-time systems remains a fundamental requirement in nuclear and particle physics experiments. At the HADES collaboration, the Quality Assessment (QA) system generates plots that visualize the real-time performance of detector subsystems. Currently this requires that operators manually inspect these visualizations to identify potential anomalies.

        This contribution presents an approach to automate this task using unsupervised machine learning. To this end, we first established a supervised Convolutional Neural Network (CNN) baseline for comparison, and then propose a fully automated anomaly detection algorithm. We modeled key detector plots by injecting realistic anomalies observed during beamtime. We then tested unsupervised models, choosing to combine a Variational Autoencoder (VAE) with HDBSCAN density-based clustering (VAE-HDBSCAN), where the VAE's CNN-encoded latent space serves as the feature representation for clustering. This approach eliminates the need for prelabeled data or constant human oversight.

        As we demonstrate, VAE-HDBSCAN measurably outperforms non-latent clustering, achieving higher precision in detecting subtle per-bin shifts while maintaining a low false-positive rate. This unsupervised method achieved efficiencies comparable to supervised models on our controlled dataset, demonstrating the feasibility of training and deploying unsupervised machine learning during data-taking periods.

        To facilitate adoption by operators, we are integrating this framework into Jefferson Lab's HYDRA QA system, a web-based data quality platform for online monitoring, labeling, and model training. JLab-HYDRA's accessible front end makes automated anomaly detection practical for operators without a machine learning background.

        Speaker: Óscar Marcos Pérez Cytron (Ruhr-Universität Bochum, GSI Helmholtzzentrum für Schwerionenforschung)
    • 🔀 Simulations & Generative Models HS2

      HS2

      Convener: Andre Scaffidi (SISSA)
      • 93
        Developing Physics-based Gaussian Process Method for Sample-Efficient Inference in Particle Accelerators

        Inferring physical parameters — such as magnetic field and alignment errors -- from sparse, noisy sensor data is essential for maintaining designed performance in machines like the GSI heavy-ion synchrotron SIS18 and the FAIR fragment separator SFRS. Rather than fitting a surrogate to simulation output, we construct a Gaussian Process kernel from an ensemble of forward simulations of the machine lattice, retaining physical fidelity while enabling GP-style uncertainty quantification.
        Compared to LOCO (Linear Optics from Closed Orbits), the accelerator-physics standard, we propose this as a probabilistic alternative offering: (1) parameter estimates with quantified uncertainty rather than point fits, (2) principled incorporation of measurement noise, (3) uncertainty-quantified interpolation between sensors, and (4) an active-learning strategy for improved sample efficiency.
        Demonstrated on a simulated SIS18, the method substantially reduces orbit uncertainty around the ring using a batch-selection strategy for measurement locations. We identify the SFRS, with sparse instrumentation and complex optics, as a promising target application, and discuss open challenges in scaling the approach to more complex beamlines.

        Speaker: Victoria Isensee (TU Darmstadt / GSI)
      • 94
        RHINE: R-process Heating Implementation in hydrodynamic simulations with NEural networks

        Neutron-rich outflows in neutron-star mergers (NSMs) or other explosive events can be subject to substantial heating through the release of rest-mass energy in the course of the rapid neutron-capture (r-) process. This r-process heating can potentially have a significant impact on the dynamics determining the velocity distribution of the ejecta, but due to the complexity of detailed nuclear networks required to describe the r-process self-consistently, hydrodynamic models of NSMs often neglect r-process heating or include it using crude parametrizations. In this talk, I will present a conceptually new method, RHINE, for emulating the r-process and concomitant energy release in hydrodynamic simulations via machine-learning algorithms. I will talk about the effect of r-process heating on the velocity boost, nucleosynthesis yields, and kilonova light curves.

        Speaker: Zewei Xiong (GSI, Darmstadt)
      • 95
        Machine-learning surrogate modeling of redshift observables in matter-perturbed black hole spacetimes

        We investigate whether matter-induced deviations from a vacuum black hole spacetime can be constrained through redshift observations of orbiting stars. Building on our previous study of redshift signatures in fluid-perturbed black hole backgrounds, we consider the problem of efficiently connecting model parameters to observable redshift curves in a setting relevant for data analysis. Since repeated evaluations of the full relativistic model can be computationally expensive in inference workflows, we propose a machine-learning surrogate for the redshift signal. The surrogate is trained on synthetic datasets generated from the underlying general-relativistic model and is designed to reproduce the dependence of the redshift curve on the spacetime's physical parameters and the orbit's parameters. This framework is intended as a first step toward comparisons with real observational data, allowing for fast exploration of parameter space and future parameter-estimation studies. We discuss the construction of the training set, the calibration of the redshift observable, and the accuracy requirements needed for the surrogate to remain physically reliable. Our results aim to provide a practical bridge between theoretical models of non-vacuum black hole environments and data-driven tests based on spectroscopic observations.

        Speaker: Ariadna Uxue Palomino Ylla (Nagoya University)
      • 96
        Field-Level Diffusion Surrogates for Reionization Astrophysics and Cosmology

        The Epoch of Reionization marks the period when the first galaxies ionized the neutral hydrogen in the intergalactic medium. Upcoming 21 cm experiment such as SKA, together with line-intensity mapping and galaxy surveys, will observe overlapping cosmic volumes. Much of their constraining power will come from cross-correlating different tracers, since each observable is affected by different foregrounds and systematics. However, cross-correlation analyses require maps, while many existing reionization emulators predict only summary statistics such as power spectra.

        We train a conditional Diffusion Transformer as a fast generative surrogate for the 21cmFAST forward model within its training prior. The model generates four physically coupled three-dimensional fields: 21 cm brightness temperature, neutral fraction, dark matter density, and line-of-sight velocity. These fields provide the map-level ingredients needed for 21 cm statistics, reionization morphology, redshift-space effects, and cross-correlations with other tracers. The model is conditioned on astrophysical and cosmological parameters, as well as redshift over ($5 \leq z \leq 25$). By learning the joint distribution of these fields, the surrogate captures not only the individual field statistics, but also the physical correlations between them.

        We evaluate the generated fields using physical statistics that are not directly optimized during training. The resulting surrogate turns expensive reionization simulations into fast, field-level generative forward models. This enables efficient parameter sweeps, population studies, and map-level multi-tracer analyses, and provides a route toward simulation-based inference with physically coupled reionization observables.

        Speaker: Divesh Jain (Heidelberg University)
      • 97
        An invertible simulator for efficient reconstruction of dark matter density from line intensity maps

        Observations of the large-scale structure, such as intensity maps of the 21cm line of neutral hydrogen, raise the question whether a precise, robust, and efficient mapping can be established from these observables to the underlying dark matter density field. We present a direct diffusion framework based on Conditional Flow Matching that enables accurate reconstruction of dark matter density from both simulated 21cm brightness temperature maps and realistic mock observations for the Square Kilometre Array, across a broad range of ionization fractions and astrophysical parameters. Furthermore, we demonstrate that our invertible method can also function as a simulator, generating both simulated and mock 21cm observations from dark matter density fields along with the corresponding ionization masks. Our approach provides a robust and efficient bridge between large-scale structure observations and the underlying dark matter distribution.

        Speaker: Thomas Blankenburg
      • 98
        Neural Operator Architectures for the 21cm Reionization Field: Fourier and Beyond

        Fourier Neural Operators (FNOs) and related architectures have emerged as promising surrogates for physical systems, with notable success in fluid dynamics, where they achieve high accuracy at a small fraction of the cost of conventional solvers. Reionization cosmology is a demanding test of that promise: the neutral fraction field in 21cm simulations is near-binary, dominated by sharp ionization fronts separating fully ionized bubbles from neutral gas -  a structure that a smooth global basis represents inefficiently, since a step edge requires all frequencies and any truncation introduces ringing. We train a family of neural operator variants that map the matter density field to the neutral fraction field, using several thousand 21cmFAST lightcones conditioned on eleven astrophysical and cosmological parameters and spanning the full reionization history. All variants share a U-shaped backbone with a distinct local/global architecture: windowed local transforms and a global whole-field bottleneck, which we instantiate with Fourier, wavelet, and Walsh–Hadamard bases, implicit-network-generated spectral kernels, and convolutional paths. This design lets the basis choice be varied independently of the rest of the architecture, enabling controlled comparisons on identical data, splits, and resolution. Validation uses transverse and cylindrical power spectra, cross-correlation coherence, ionized bubble-size distributions, and the global ionization history. We report where each configuration is reliable and where it fails; the most competitive variants match or exceed a standard U-shaped FNO with one to two orders of magnitude fewer parameters, at speeds suitable for use within an inference loop.

        Speaker: Haris Kapetanovic
      • 99
        A compact flow-matching surrogate for gravitational waveforms

        Fast waveform surrogates are essential for gravitational-wave inference, but neural surrogates typically lack uncertainty estimates and offer little guidance on how accuracy scales with training resources. We present an autoregressive flow-matching surrogate for binary-black-hole waveforms that operates on amplitude-phase representations and samples each successive segment with a conditional flow. A compact, few-million-parameter model attains surrogate-grade mismatches, with phase coherence maintained over the full inspiral-merger-ringdown, and stochastic sampling provides uncertainty estimates at no additional training cost. We further measure an empirical scaling law relating accuracy to training-set size, which successfully predicted the performance of our largest training run in advance, offering a quantitative recipe for surrogate development budgets. Across all scales, the dominant residual error is a global phase/time reference offset rather than accumulated instability, identifying anchoring as the central challenge for autoregressive waveform generation. We discuss implications for waveform foundation models and extensions to higher modes and eccentric systems.

        Speaker: Waleed Esmail
      • 100
        Machine Learning Approaches to Black Hole Dynamics: From Perturbation Theory to Numerical Relativity

        Machine learning techniques are increasingly being explored as complementary tools for solving complex problems in gravitational physics. In this talk, I will present recent work on the application of neural-network-based methods to black hole dynamics across different regimes. On the one hand, perturbative models provide an ideal framework for investigating the spectral properties of black hole spacetimes, including quasi-normal modes, scattering phenomena, and the response to generic perturbations. On the other hand, physics-informed neural networks offer a promising approach for solving the nonlinear Einstein equations by directly incorporating the governing equations into the optimization process.

        I will discuss recent progress in applying these techniques to problems ranging from black hole perturbation theory to proof-of-concept numerical relativity simulations of binary black hole head-on collisions. Particular emphasis will be placed on the opportunities and challenges of machine learning methods for solving hyperbolic systems of equations relevant to gravitational-wave physics.

        Speaker: Xisco Jimenez Forteza (Universitat de les Illes Balears)
    • 🧠 Working groups: Large Language Models 2.404

      2.404

      • 101
        Large Language Models
    • 15:30
      Coffee Foyer

      Foyer

    • 🔀 Explainability & Theory 1.404

      1.404

      Convener: Jan Pawlowski
      • 102
        Gravitational Duals from Equations of State: Large Hierarchies and False Vacua

        We investigate the reconstruction of holographic duals for strongly coupled quantum field theories in regimes characterized by large hierarchies and the presence of false vacua. Within the gauge/gravity duality, these features translate into non-trivial thermodynamic behaviour and exotic renormalization group flows, including skipping flows between non-adjacent fixed points. Building on previous work based on Physics-Informed Neural Networks (PINNs), we extend the holographic inverse problem of reconstructing the bulk scalar potential from boundary thermodynamic data into this new regime. This setting presents a variety of conceptual and numerical challenges, such as near-degenerate states, large hierarchies of energy scales, and regions of the potential that are not directly probed by the input data. We develop a set of methodological advances that overcome these obstacles, thereby improving the established PINNs-based methodology and extending it to new physical regimes of interest that were previously out of reach from the model. Applying the developed framework, we demonstrate accurate reconstruction of scalar potentials deep into the false vacuum regime, achieving robust agreement with the physical features of the underlying thermodynamics despite significant numerical stiffness. Our results extend the bridge between holography and machine learning, and suggest that data-driven approaches can provide new insights into the structure of strongly coupled systems.

        Speaker: Pau Solé Vilaró (ICC - Universitat de Barcelona)
      • 103
        Physics-informed Renormalisation Group Flows

        The renormalisation group considers general coarse-graining transformations in lattice systems and more broadly speaking field theories or data spaces and is based on an analytically known optimal transport equation.
        This framework has a natural correspondence with modern generative architectures in terms of pooling and convolutions, but differs in terms of the non-linearity with corresponds to a general field/variable transformation, the physics-informed kernel (PIK) [1,2]. The PIK is a can be chosen in an informed manner to suit the generative task:
        We discuss several examples which are encompassed by this framework, ranging from Langevin evolutions to sampling of complex probability distributions with a sign problem [3].

        [1] FI, Pawlowski, Annals Phys. 481 (2025) 170177
        [2] FI, R. Kapust, J. M. Pawlowski , arXiv:2510.26678
        [3] FI, R. Kapust, J. M. Pawlowski , arXiv:2603.03159

        Speaker: Friederike Ihssen (ITP Heidelberg)
      • 104
        Community Detection on Complex Networks through Transformer-based Message Passing and Denoising Diffusion

        Partitioning a graph into communities is an NP-hard optimization problem with a natural statistical-physics formulation: the optimal partition is the ground state of an antiferromagnetic Potts model, with modularity playing the role of negative energy. Clustering on graph-structured data is a recurring task in fundamental physics, where detector and event data are naturally represented as graphs, as in charged particle tracking and particle-flow reconstruction at colliders and source clustering in astroparticle physics. We frame community detection as energy minimization and solve it with a transformer-based message passing network, where the message and update functions are transformer encoder layers acting on node features. Training follows a denoising-diffusion setting: Gaussian noise is injected into the node community assignments and the network learns to recover the clean configuration, driving the system from local minima towards the minimum-energy state. The loss function couples a physics-derived continuous modularity with a supervised cross-entropy term. We train and validate on LFR benchmarks across different mixing parameters 𝞵, and investigate generalization to real-world networks as an out-of-distribution transfer problem.

        Speaker: Edoardo Murano (Sapienza Università di Roma and INFN)
      • 105
        A graph-bandwidth law for context windows in autoregressive samplers

        An autoregressive neural sampler writes a lattice configuration site by site, each site conditioned on the sites already written, and it returns independent samples with an exact likelihood. Each conditional sees only a context window of fixed reach, and nothing in the architecture says how far that reach has to extend.

        Set the reach one unit short of what the target needs and the interior bonds come out if anything too ordered, while the susceptibility can be statistically indistinguishable from the truth. But the susceptibility only integrates the correlation function, and its shape is wrong by an order of magnitude more.

        The requirement is combinatorial. In a local model the sites already written act on the rest only through the frontier between them, and the generation order never pushes that frontier further back than its own bandwidth. We prove that a contiguous window represents the exact conditionals of every such target precisely when its reach is at least that bandwidth, the largest gap the order leaves between two coupled sites. Across more than a hundred temperature-conditioned samplers, spanning lattice sizes, boundary conditions, orders, depths and targets beyond Ising, the verdict follows the bandwidth. One unit of bandwidth is enough to flip it, and in a separate pair orders that tie on cut width still split.

        The required context can therefore be computed from the interaction graph of any local target before a single training run, and on the torus a well-chosen order drops it from the number of sites to its square root. The thermodynamic observables can pass either way, so a check without ground truth has to read the shape of the correlation function.

        Speaker: Tae-Geun Kim (Fudan U. / RIKEN)
      • 106
        QUIVER: QUantum-Informed Views for Enhanced Representations in Large Machine Learning Models

        Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example. We introduce <span style="font-variant: small-caps;">Quiver</span> (QUantum-Informed Views for Enhanced Representations), a paradigm that enriches classical data-driven features with a quantum Fisher view: a geometrically motivated, basis-independent summary of higher-order correlations captured by a variational quantum circuit (VQC) trained to perform the same task. Unlike classical feature augmentation, the quantum Fisher information matrix encodes the intrinsic geometry of the learned quantum state manifold. This feature map, motivated by quantum information theory, is ordinarily non-trivial to model classically. However, it can reveal statistical structure that additional classical data or model capacity finds difficult to learn, making it a complementary modality. We demonstrate that <span style="font-variant: small-caps;">Quiver</span> improves standard performance metrics on two benchmark datasets from very different fields: the <span style="font-variant: small-caps;">JetClass</span> dataset for predicting jet flavor at the Large Hadron Collider (LHC), and the QM9 dataset for predicting molecule properties. The core contribution, however, is domain-agnostic: the quantum Fisher view can be fused into a broad class of model architectures via targeted modifications to the base architecture, to incorporate information about the quantum geometry of the problem. These results demonstrate that quantum-geometric features, extracted from simulated variational circuits, can deliver measurable value for standard machine learning tasks, well before the advent of fault-tolerant quantum hardware.

        Speaker: Aritra Bal (Karlsruhe Institute of Technology (KIT))
      • 107
        ARQuAdia-Net: Learning Quantum Annealing Dynamics for 3-SAT

        In recent years, approaches inspired by fundamental physics have played an important role in the development of modern machine learning algorithms, from energy based models to diffusion processes. Motivated by this perspective, we propose a hybrid quantum-classical learning framework inspired by adiabatic quantum dynamics and quantum annealing for solving constraint satisfaction problems.In this work, we focus on 3-SAT instances represented as graphs. The training objective combines energy minimization of the 3-SAT Hamiltonian with a residual term derived from McLachlan’s variational principle, encouraging the learned dynamics to approximate an adiabatic path within the accessible variational manifold.Our approach offers a framework for studying how ideas from quantum annealing and variational quantum dynamics can be used to design learning algorithms for NP-complete constraint satisfaction problem. Beyond computational applications, this perspective may also be relevant for fundamental physics, where constraint satisfaction, frustrated energy landscapes and adiabatic state preparation naturally arise. Potential applications include the variational preparation of ground states of many-body quantum systems, the study of lattice gauge theories and constrained Hilbert spaces and the characterization of frustrated spin models and glassy energy landscapes.

        Speaker: Lorenzo Colantonio (Sapienza Università di Roma and INFN)
      • 108
        tn4ml: Tensor Network Training and Customization for Machine Learning

        Tensor Networks have emerged as a prominent alternative to neural networks for addressing machine learning challenges in foundational sciences, offering a bridge between quantum information techniques and classical learning algorithms. This transfer of methods from quantum information to machine learning has the potential to enhance model interpretability and efficiency, paving the way for applications to real-life problems. This work introduces tn4ml, a library designed to seamlessly integrate Tensor Networks into optimization pipelines for machine learning tasks. Inspired by existing machine learning frameworks, the library offers a user-friendly structure with modules for data embedding, objective function definition, model training using diverse optimization strategies and evaluation. We demonstrate its versatility through examples such as supervised learning on tabular data and unsupervised learning on the MNIST dataset. Additionally, we analyze how customizing the parts of the machine learning pipeline for Tensor Networks influences different performance metrics for specific tasks. Beyond these examples, tn4ml has been used in high-energy physics applications, including anomaly detection in the latent space of LHC collision events, and quantized Tensor Network models for low-latency jet tagging.

        Speaker: Ema Puljak (CERN)
      • 109
        Geometry of hidden feature spaces in Quantum Neural Networks

        Classical neural networks have rich feature learning capabilities where each functional composition has the ability to transform the hidden representations into non-isometric geometries. In quantum neural networks, however, depth or state reachability alone does not guarantee this feature-learning capability. We study this question in the pure-state setting by viewing encoded data as embedded in the manifold of pure states and analysing infinitesimal unitary actions through Lie-algebra directions. We introduce Classical-to-Lie-algebra (CLA) maps and the criterion of almost Complete Local Selectivity (aCLS), which combines directional completeness with data-dependent local selectivity. Within this framework, we show that data-independent trainable unitaries are complete but non-selective, i.e. learnable rigid reorientations, whereas pure data encodings are selective but non-tunable, i.e. fixed deformations. Hence, geometric flexibility requires a non-trivial joint dependence on data and trainable weights. We further show that accessing high-dimensional deformations of many-qubit state manifolds requires parametrised entangling directions; fixed entanglers such as CNOT alone do not provide adaptive geometric control. Numerical examples validate that aCLS-satisfying data re-uploading models outperform non-tunable schemes while requiring only a quarter of the gate operations.

        Speaker: Vishal Ngairangbam (Karlsruhe Institute of Technology)
    • 🔀 Hardware & Design HS2

      HS2

      Convener: Bjoern Malte Schaefer
      • 110
        SHiP's Muon Shield Optimization with Machine Learning

        The SHiP experiment is a proposed fixed-target experiment at the CERN SPS designed to search for feebly interacting particles beyond the Standard Model. A critical challenge for SHiP is suppressing the massive flux of beam-dump muons entering the detector acceptance. The Active Muon Shield, a system of magnets designed for muon deflection, must therefore be optimized to minimize this background while satisfying strict geometric and budgetary constraints.
        Conventional design optimization loops, however, are severely constrained by the computational cost of full-scale detector simulations. In this work, we bypass this bottleneck using a three-stage, GPU-accelerated framework: 1) a Normalizing Flow to rapidly sample input muon distributions, replacing Pythia8; 2) an Operator Learning model replacing finite-element magnetic field simulations; and 3) a GPU-native heuristic replacing Geant4 propagation by sampling from pre-computed interaction histograms. By accelerating simulation runtimes from months to hours without compromising fidelity, this pipeline enables the first complete optimization of the SHiP Muon Shield encompassing full muon statistics and all realistic constraints, successfully resolving the shield design challenge.

        Speaker: Luis Felipe Cattelan (University of Zurich)
      • 111
        Large language models for physics instrument design

        Instrument design requires searching a large space of discrete and continuous choices under hard resource constraints, and the effort of setting up such optimizations increasingly limits how far they can scale. We investigate the use of large language models (LLMs) for physics instrument design and compare their performance with reinforcement learning (RL). Using only prompting, the models receive the design constraints and summaries of previously evaluated configurations, then propose complete detector layouts that are assessed with the same simulators and reward functions used for RL optimization. We evaluate four open-weight and proprietary LLMs on two detector-design benchmarks: longitudinal segmentation of a sampling calorimeter and placement and granularity optimization of tracking stations in a magnetic spectrometer. Although RL achieves the strongest final designs, the LLMs consistently generate valid, resource-aware and physically meaningful configurations without task-specific training. We also investigate a hybrid approach in which LLM proposals are refined using a trust-region optimizer, which essentially closes the performance gap with RL on the spectrometer benchmark. These results suggest a role for LLMs as meta-planners in automated instrument-design workflows, generating and organizing design hypotheses while dedicated reward-driven optimizers carry out detailed refinement.

        Speaker: Sara Zoccheddu
      • 112
        Differentiable simulation and optimization of superconducting quantum sensors

        QSOpt (Quantum Sensing Optimization) is an end-to-end differentiable simulation and machine-learning optimization framework for open quantum networks composed of superconducting qubits, bosonic modes and input-output channels with user-defined interactions.
        The advancements in quantum technologies have sparked interest in employing quantum systems as sensors, with superconducting quantum circuits emerging as
        a flexible tool for metrology and quantum sensing, with application in fundamental physics like light dark-matter and axions detection. A promising architecture employs networks of superconducting qubits coupled to microwave bosonic modes. However, accurately modeling and optimizing these systems under realistic noise conditions requires numerical methods beyond tractable analytical approaches.
        QSOpt integrates the QuTiP quantum simulation library with a JAX backend, enabling gradient-based optimization of parametrized quantum circuits for state preparation and readout, together with sweeps over hardware parameters such as couplings and dispersive shifts. This enables systematic exploration of sensing strategies and maximization of sensing performance in multi-qubit architectures under realistic conditions, while facilitating the discovery of previously unexplored
        sensing protocols.

        Speaker: Nathan Campioni (Sapienza Università di Roma and INFN)
      • 113
        Transformer-Based Analysis for Next-Generation Compact Cherenkov Telescopes

        The past decade has witnessed a rapid expansion of machine learning applications in very-high-energy (VHE) gamma-ray astronomy. While many efforts have focused on Convolutional Neural Networks (CNNs), these approaches remain constrained by existing camera geometries and by the limited ability of Monte Carlo simulations to fully capture real telescope performance. In this work, we employ a simulated telescope optimization framework to investigate how advanced machine learning techniques can fundamentally reshape instrumental design requirements.

        We present, to our knowledge, the first transformer-based analysis applied to data from Cherenkov Telescopes. Our analysis achieves an order-of-magnitude improvement in performance compared to a typical standard analysis pipeline. Our results indicate that, within this framework, a single 5-meter-diameter Imaging Atmospheric Cherenkov Telescope (IACT) operating in monoscopic mode can reach an energy threshold below 150GeV, three times lower than the one achieved with standard methods. These findings point toward a viable pathway for the development of cost-effective, potentially autonomous IACTs, capable of substantially increasing the temporal coverage and duty cycle of future gamma-ray observatories.

        Speaker: Elli Jobst (Technical University of Munich, Max Planck Institute for Physics)
      • 114
        Geometry-resolved Characterization of a Dual-detector MA-XRF Scanner

        Macro-X-ray fluorescence (MA-XRF) scanners with two detectors are used, with their signals normally summed to improve the signal-to-noise ratio and the difference between them neglected. We show that this difference, plus one extra measurement, scanning the same painting again with the canvas tilted forward, is enough to characterise the instrument, with no dedicated calibration measurements fully. Tilting the canvas changes the paths that fluorescence photons take to each detector, but leaves the detectors themselves unchanged. Comparing the two scans, therefore, splits the observed differences into two parts: detector properties, which do not depend on tilt, and geometric effects, which change with tilt in a predictable, energy-dependent way. From two routine scans of one painting, we obtain per-element efficiency ratios with uncertainties, the effective viewing angle of each detector, and a flat-field correction map, and we measure how sensitive the element maps are to canvas positioning. The result is an energy-weighted fusion of the two detectors that outperforms simple summing, and a practical estimate of the error caused by imperfect mounting.

        Speakers: Ms Aleksandra Stojanović (School of Electrical Engineering, University of Belgrade), Mr Dimitrije Pešić (School of Electrical Engineering, University of Belgrade), Mr Janko Vukobratović (School of Electrical Engineering, University of Belgrade)
      • 115
        A Low-Latency Neural Network Trigger for SiPM-Based RICH Detectors

        SiPMs (Silicon Photomultipliers) have recently been studied as candidates for building photon cameras in RICH (Ring Imaging Cherenkov) detectors. SiPM-based cameras improve detection efficiency (up to $60\%$), spatial resolution (mm), timing ($< 100$\, ps), scalability, and magnetic field immunity. Nevertheless, for single-photon detection, SiPM thermal noise ($\sim 10^2$\,kHz/mm$^2$) and photon-background sources (scintillation or scattering) pose severe challenges for free-streaming readout systems. In this work, we study the viability of an ML (machine learning) trigger implemented in the CBM (Compressed Baryonic Matter) RICH front-end electronics (edge computing). The ML algorithm aims to filter out fake events caused by the SiPM thermal noise while keeping Cherenkov ring events.

        Speaker: Jesus Pena-Rodriguez (JLU Giessen)
    • 🔀 Inference & Uncertainty HS1

      HS1

      Convener: Humberto Reyes-Gonzalez (RWTH Aachen University)
      • 116
        Recovering end-to-end SBI performance through probabilistic reconstruction

        Simulation-based inference (SBI) has become a key tool for hypothesis testing in high-energy particle physics, enabling likelihood ratio estimation directly from simulated data without requiring tractable likelihoods. In practice, SBI is often applied to high-level variables reconstructed from detector measurements as point estimates, discarding hypothesis-dependent information contained in the reconstruction uncertainty. A direct end-to-end SBI approach on raw detector observables avoids this loss but is often impractical due to high dimensionality, and requires full recalibration whenever the alternative hypothesis changes.
        We present a new approach based on probabilistic reconstruction of the latent variables of interest. Rather than assigning a point estimate, we model the full posterior distribution of the latent variables conditioned on the detector observables. Combined with an SBI classifier in latent space, the resulting likelihood-ratio estimator is equivalent to end-to-end SBI while remaining computationally tractable. The generative component is calibrated only once under the null hypothesis and requires no further recalibration.
        We demonstrate the framework on two problems: a supersymmetric top-quark pair production search using transformer-based flow matching for posterior reconstruction of invisible-particle kinematics, and an image reconstruction benchmark using convolutional flow matching. Across both benchmarks, increasing the number of posterior samples systematically closes the performance gap to end-to-end SBI while consistently outperforming conventional point-estimate reconstruction methods.

        Speaker: Vitus Past (Technische Universität München)
      • 117
        Pairton: Iterative Reconstruction of Short-Lived Particles

        In high-energy collider physics, the reconstruction of complicated event topologies, such as fully hadronic $t\bar{t}$ decays, is a difficult challenge for many analyses due to the large combinatorial backgrounds. While previous machine learning methods have advanced this task, they struggle to learn the exact joint distribution, and instead learn various independent marginal approximations.

        To address this shortcoming, we present Pairton, an iterative framework for reconstructing short-lived particles in high-energy collision events. By formulating particle reconstruction as a masked prediction process over graph structures, Pairton learns conditional distributions consistent with a factorized decomposition of decay products. We demonstrate state-of-the-art performance on fully hadronic $t\bar{t}$ decays. Pairton provides a general, flexible paradigm for particle reconstruction that can be readily extended to other topologies, bridging ideas from modern generative modeling and high-energy physics.

        Speaker: Andreas Hermansen (Universite de Geneve)
      • 118
        Visible Energy Reconstruction in KM3NeT/ARCA using Graph Neural Networks

        Accurate energy reconstruction is a key challenge in neutrino telescopes, where the total neutrino energy is not fully contained in the detector and relying only on the reconstructed muon energy neglects the contribution from hadronic showers. We present a Graph Neural Network (GNN)-based approach for visible energy reconstruction in KM3NeT/ARCA using the DynEdge architecture within the GraphNeT framework, with visible energy as the reconstruction target.
        This first implementation on ARCA with 21 detection units is trained on simulated νμ charged-current events using raw hit-level detector information. The results demonstrate stable performance across different simulation versions and generalization to all neutrino flavors despite training on a single flavor sample. We further investigate training strategies including both neutrino interactions and atmospheric muon backgrounds, revealing a domain adaptation challenge: while the inclusion of atmospheric muons improves their reconstruction, it introduces a degradation in neutrino performance. Understanding and mitigating such effects is essential for robust machine learning (ML) applications in realistic experimental environments. Finally, we discuss initial steps towards interpreting GNN predictions and understanding the detector information driving the reconstruction, an important step towards building trust in ML-based methods for neutrino physics.

        Speaker: Dr Evangelia Drakopoulou (NCSR "Demokritos")
      • 119
        Estimating Neutrino Oscillation Parameters using Structured hierarchical transformers

        Neutrino oscillations encode fundamental information about neutrino masses and mixing parameters, offering a unique window into physics beyond the Standard Model. Estimating these parameters from oscillation probability maps is, however, computationally challenging due to the maps’ high dimensionality and nonlinear dependence on the underlying physics. Traditional inference methods, such as likelihood-based or Monte Carlo sampling approaches, require extensive simulations to explore the parameter space, creating major bottlenecks for large-scale analyses. In this work, we introduce a data-driven framework that reformulates atmospheric neutrino oscillation parameter inference as a supervised regression task over structured oscillation maps. We propose a hierarchical transformer architecture that explicitly models the two-dimensional structure of these maps, capturing angular dependencies at fixed energies and global correlations across the energy spectrum. To improve physical consistency, the model is trained using a surrogate simulation constraint that enforces agreement between the predicted parameters and the reconstructed oscillation patterns. Furthermore, we introduce a neural network-based uncertainty quantification mechanism that produces distribution-free prediction intervals with formal coverage guarantees. Experiments on simulated oscillation maps under Earth-matter conditions demonstrate that the proposed method is comparable to a Markov Chain Monte Carlo baseline in estimation accuracy, with substantial improvements in computational cost (around 240$\times$ fewer FLOPs and 33$\times$ faster in average processing time). Moreover, the conformally calibrated prediction intervals remain narrow while achieving the target nominal coverage of 90\%, confirming both the reliability and efficiency of our approach.

        Speaker: Antonin Vacheret (CNRS - LPC Caen)
      • 120
        Deep Learning for Real-Time Proton CT Reconstruction

        Proton computed tomography (pCT) promises a direct, low-dose measurement of the relative stopping power (RSP) map that governs proton-therapy dose calculations. In a tracking-based scanner every proton is recorded individually, and reconstructing the RSP image requires estimating the path each proton actually took through the patient. This is intrinsically hard: multiple Coulomb scattering makes the path stochastic — two protons entering identically exit along measurably different tracks — so the true per-proton trajectory can only be inferred from the noisy exit measurement. Classical approaches address this by building and repeatedly applying a large system matrix, a step that is memory-hungry and ill-suited to the streaming, high-rate data of a real scanner.

        We explore a machine-learning route to this inference problem: rather than storing a system matrix, a neural network estimates each proton's path on the fly, directly from its exit signature, and feeds the result straight into iterative RSP reconstruction. Framing the per-proton path as a learned inference task naturally raises the question of resolution limits and uncertainty, which we take as central rather than incidental.

        In this talk we motivate the problem from the underlying physics and present our progress toward reconstruction quality and throughput competitive with classical methods — a step toward truly on-the-fly proton CT during treatment.

        Speaker: Gábor Bíró (HUN-REN Wigner RCP)
      • 121
        Kolmogorov-Arnold Networks for Higgs Boson Classification

        Gradient-boosted decision trees such as XGBoost remain the default for tabular high-energy-physics (HEP) classification. We investigate Kolmogorov-Arnold Networks (KAN), which replace fixed activations with learnable B-spline functions on network edges, as an alternative for Higgs boson signal/background separation. We implement KAN from scratch in PyTorch (Cox-de Boor B-splines, adaptive knot placement, coarse-to-fine grid extension, and L1 spline regularisation) and evaluate it on a stratified 2-million-event subset of the FAIR Universe HiggsML dataset (H to tau tau signal). All metrics are reported as bootstrap mean plus/minus sigma from a single deterministic, fully reproducible run.

        A faithful deep KAN (AUC 0.8834, AMS 3.66) and its grid-extended variant (AUC 0.8839, AMS 3.74) both outperform a vanilla XGBoost baseline (0.8802 / 3.31) on discrimination and significance. Under momentum energy-scale systematics (gamma in [0.8, 1.2]), adversarial decorrelation makes KAN roughly 3x more robust (max AUC drop 0.026 vs 0.082). Model-agnostic permutation importance, grouped-feature importance, and a per-class-weighted calibration analysis further show KAN is competitive, well-calibrated, and interpretable for HEP.

        Speaker: Mr Said Abolhassan Razavi
      • 122
        Hybrid GNN-Bayesian Models for Uncertainty-Aware Inference of Rare Events on Spatially Embedded Graphs

        We present a hybrid graph-based Bayesian framework for uncertainty-aware inference of rare events on spatially embedded networks.
        Starting from a Multivariate Conditional Autoregressive Model (MCAR) in a hierarchical Bayesian architecture defined on the dual graph of a street network, events counts are described by Poisson likelihoods with structured graph effects, unstructured heterogeneity and node-level covariates. The approach was validated on georeferenced accidents data from Rome, where it enabled segment-level latent-risk estimation, crashes severity correlation modeling and calibrated posterior uncertainty.
        We propose a physics-oriented extension combining this Bayesian graphical model with GNN-based amortized variational inference. A graph neural encoder learns local and non-local representations and parametrizes variational posteriors or Bayesian hyperparameters, without doing the optimization of the parameters for each observation. Meanwhile, the probabilistic layer preserves interpretable likelihoods, structured priors and uncertainty quantification. This hybrid formulation is scalable beyond MCMC and naturally suited to sparse, noisy, high-dimensional graphs. Potential applications include charged-particle tracking, vertexing, calorimeter clustering, pileup mitigation, jet anomaly detection, continuous gravitational-wave candidate clustering, astroparticle source association and multi-messenger event correlation.

        Speaker: Federica Di Bartolomeo (Sapienza Università di Roma)
    • 🔀 Patterns & Anomalies 3.404

      3.404

      Convener: Dennis Noll
      • 123
        Anomaly-detection of gravitational waves through the 4th observing run of the LIGO detectors

        We present an anomaly detection algorithm based on a deep convolutional autoencoder. The algorithm, initially developed with a focus on the Einstein Telescope, was adjusted to the higher noise levels of the LIGO detectors, implementing coherence, trained and validated using O3 data, and tested on publicly available O4 data. We achieve an excellent recovery rate for short, loud signals and compare our results with those of other unmodeled pipelines. We will also discuss the plans for possible deployment during future observing runs.

        Speakers: Gianluca Inguglia (MBI Vienna), Huw Haigh (MBI Vienna), Ulyana Dupletsa (MBI Vienna)
      • 124
        Anomaly detection for generic-transient gravitational waves

        All gravitational wave detections thus far have been from binary compact object coalescences. These detections rely on analyses that exploit the well-modelled nature of the sources. Generic-transient (burst) gravitational waves, on the other hand, do not have well-modelled gravitational wave signals. Therefore, searches for burst gravitational waves must efficiently scan a broad parameter space while minimising the impact of spurious transient noise. We investigate the application of a transformer-based anomaly detection scheme, TranAD, to burst gravitational wave searches. We characterise the performance of TranAD and demonstrate that this approach is a promising for discovering new types of gravitational wave signals.

        Speaker: Prof. Elena Cuoco (Unversity of Bologna)
      • 125
        DeepExtractor: Glitch Mitigation and Template Free Searches with Model-Agnostic Reconstruction using Deep Learning

        Gravitational-wave detectors such as LIGO, Virgo, and KAGRA, detect faint signals from distant astrophysical events. However, their high sensitivity also exposes them to noise artifacts — including unmodelled transients known as "glitches" — that can mimic or mask genuine signals. Last year we introduced DeepExtractor, a deep learning framework that reconstructs arbitrary signals or glitches with power above the detector noise floor, regardless of its source or morphology, by learning to model and subtract the underlying detector noise rather than the signal itself.

        In this talk we present extensions to the framework. We extend the method to the case where a signal and a glitch overlap in time and frequency, by jointly modelling both components alongside the background noise so that each can be reconstructed via subtraction, and show that this separation step improves downstream parameter estimation on the recovered signal by reducing the glitch bias. Finally, we discuss DeepExtractor's suitability for online, low-latency, template-free searches, where its morphology-agnostic reconstruction offers a potential alternative to matched-filtering and unmodelled search pipelines.

        Speaker: Tom Dooney (Nikhef / Utrecht University)
      • 126
        Anomaly Detection for Faint Galactic Substructure with EagleEye

        Detecting faint stellar substructure in the Milky Way halo and its surroundings favours methods that remain sensitive to weak signals without restrictive assumptions about either the signal morphology or the Galactic background. We present recent applications of EagleEye, a model-independent anomaly detection framework that compares multidimensional data distributions to identify localized over- and underdensities relative to an empirical reference sample. We highlight two applications: the detection of faint dwarf galaxies around the Milky Way and the search for stellar wakes induced by its most massive satellites. As stellar wakes are expected to be exceptionally faint and difficult to model, yet offer a novel probe of dark matter substructure, they are a particularly exciting target for data-driven anomaly detection methods.

        Speaker: Sven Põder (SISSA)
      • 127
        Exploiting the Latent Space of Convolutional Autoencoders to identify low-energy recoil signals in a dual-phase LAr TPC

        In this contribution, we propose a data-driven approach based on unsupervised learning, specifically using deep Convolutional AutoEncoders (CAE), to identify signals from low-energy nuclear recoils in the dual-phase Liquid Argon Time Projection Chamber (LAr TPC) of the Recoil Directionality (ReD) experiment. This is a challenging task for recoil energies $\sim$1 keV, since scintillation remains below detection thresholds and only a few ionization electrons are produced, resulting in a faint electroluminescence pulse (S2).
        The dataset for training, validation, and testing includes $\sim$7,000 gamma-ray events from a $^{252}$Cf source, each corresponding to a 10,000-sample waveform averaged over 22 Silicon photomultiplier channels.
        We trained four distinct CAE models, which heavily compress each input waveform into a latent space of dimensions 1 to 4, respectively. The aim is to directly study the features of such reduced representations, while comparing different compression levels. Indeed, after the training, a region encoding waveforms with negligible signal clearly emerges in every latent space, while larger pulses are encoded further away.
        To calibrate and exploit this mapping as a signal identification method, $\sim$9,900 noise-only waveforms (from runs with the source shielded) are processed, defining the boundaries of pure background regions in the latent spaces; events falling outside are then tagged as signal candidates. The performance is evaluated across the four models on a "golden" neutron dataset of $\sim$600 waveforms, demonstrating that all the events selected in the standard analyses are correctly labelled when the CAE model has at least a 2-dimensional latent space.
        Finally, we present a preliminary evaluation of the signal acceptance for our method, defined as the fraction of correctly tagged signals as a function of their integral, employing a dataset of semi-synthetic waveforms built by overlaying real pedestals with pure log-normal pulses, resembling S2 signals. The results are comparable to conventional strategies, offering a robust alternative for low-energy signal identification.

        Speaker: Gioacchino Alex Anastasi (Università di Catania & INFN Catania)
    • 🧠 Working groups: Simulation-Based Inference (WG6) 2.404

      2.404

      • 128
        Simulation-Based Inference (WG6)
    • 17:30
      🖼️ Poster session + 🍻 Drinks Foyer

      Foyer

    • 🗣️ Plenaries HS1

      HS1

      Convener: Verena Kain
      • 129
        Light in the dark: How machine learning changed observational astronomy
        Speaker: Emille Ishida (Clermont Auvergne)
      • 130
        On the cusp of on-demand material discovery
        Speaker: Johann Brehmer (CuspAI)
    • 10:30
      ☕ Coffee Foyer

      Foyer

    • 🔀 Data & Ethics 1.404

      1.404

      Convener: Kilian Schwarz
      • 131
        TREASURE: Tokenized Event Representations for Collider Foundation Models

        The Tokenized Representations for Energy-frontier AI Searches via Understanding and Reasoning (TREASURE) project aims to make high-energy physics collider data usable by modern AI methods through standardized tokenized representations. Collider experiments produce rich datasets in different, experiment-specific formats, which limits cross-experiment analysis and the reuse of legacy data. TREASURE addresses this challenge by developing tokenized representations of collider data at multiple levels, including events, jets, particles, tracks, clusters, and hits, and by using these representations to build foundation models that can learn across experiments, detectors, and eras of data. This will enable multi-experiment training, improve sensitivity to rare signals, and help extract new information from current and legacy high-energy physics datasets.

        We present a first TREASURE pipeline for LHC open data, which converts experiment-native formats into a common event schema and then into tokenized sequences suitable for transformer-based models. We show initial checks of the tokenized representation, including reconstruction-quality studies and event-level classification benchmarks based on Higgs physics signatures. The goal is to build reusable tokenized collider datasets and shared benchmarks, as a first step toward cross-experiment foundation models for current, legacy, and future high-energy physics experiments.

        Speaker: Merve Nazlim Agaras (Brookhaven National Laboratory (BNL))
      • 132
        A summer of BOA (Bytewise Online Autoregressive Constrictor) for scientific data compression

        The petabyte-scale data generated by High Energy Physics (HEP) experiments presents a significant storage challenge. We present the Bytewise Online Autoregressive (BOA) Constrictor, a new pseudo-streaming lossless neural compressor built upon the Mamba state space model. BOA achieves competitive compression ratios across diverse structured HEP datasets, matching or exceeding LZMA, ZSTD and ZLIB at maximum compression, among other tested algorithms. BOA also demonstrates robust cross-file and cross-condition generalisation on CMS Open Data (NanoAOD format), where it obtains comparable or improved effective compression ratios (within 5%) with respect to the next-best traditional algorithm. Ablation studies show that transitioning to half-precision (FP16) weights reduces the model footprint without degrading predictive accuracy. The model has also been tested in other kinds of scientific data. BOA is supported by a deterministic reference C++ implementation which ensures bit-exact reproducibility across different CUDA architectures. In this proof-of-principle implementation, BOA delivers a decompression throughput that is not yet competitive with optimised algorithms such as ZSTD or LZMA, but still provides a first step towards data compression improvements for next-generation scientific data. This contribution summarises the work done by several summer students on inference acceleration, benchmarking, and environmental impact break-even point between equivalent carbon of inference/training vs embodied carbon for storage.

        Speaker: Caterina Doglioni (University of Manchester)
      • 133
        AI-for-ET: the artificial intelligence division of the Einstein Telescope

        The Einstein Telescope (ET) Collaboration, the European third-generation flagship experiment for gravitational wave searchs, has recently launched a new cross-board division, AI-for-ET, with the aim to explore challenges and opportunities for ET related to the adoption and implementation of AI workflows into the activities of the collaboration. In this talk we will povide an overview of the initial activities and the plans of AI-for-ET with the aim to angage in a frutiful exchnge with other communities.

        Speakers: Gianluca Inguglia (MBI Vienna), Dr Tomislav Andric (GSSI)
      • 134
        COMCHA: Building a Spanish Network on Advanced Computing for Fundamental and Applied Physics

        Artificial intelligence, heterogeneous computing, and scientific software are rapidly transforming the way research is conducted. To address these challenges in a coordinated manner, the Spanish Network on Advanced Computing for Fundamental and Applied Physics (COMCHA) was established approximately a decade ago. Since then, the network has evolved from a collaborative initiative of LHC groups into a well-established national ecosystem that brings together thirteen research groups and two computing infrastructures, spanning computing, high-energy physics, astroparticle physics, nuclear physics, theory, and applied physics, and representing a significant fraction of the Spanish community working on advanced computing for physics.

        COMCHA aims to foster collaboration across institutions, accelerate the adoption of emerging computing technologies, and promote best practices in scientific software, artificial intelligence, and sustainable computing. The network coordinates researcher training, develops strategic roadmaps and policy recommendations, promotes access to national computing infrastructures, and facilitates participation in major international initiatives. Through thematic working groups, schools, workshops, industrial engagement, and outreach activities, COMCHA seeks to build a collaborative and inclusive community capable of addressing the future computing challenges of fundamental and applied physics while maximizing scientific impact.

        In this talk, we present some of the activities and strategies being developed within COMCHA, which naturally dovetail with the European coalition for AI in fundamental physics and may be of interest to other software and computing communities.

        Speaker: Arantza Oyanguren (IFIC - CSIC/UV)
    • 🔀 Inference & Uncertainty HS1

      HS1

      Convener: Skyler Degenkolb (Universität Heidelberg)
      • 135
        Unfolding without Iterations, Adversaries, or Surrogates

        Correcting measurements for detector effects is a pressing inverse problem in LHC physics. Current methods solve this problem by relying on iterative refinement, minimax optimization, or a surrogate forward mapping. In this talk, I present Adversary-free Unfolding SanS Iteration or Emulation (AUSSIE), which dispenses with these mechanisms while remaining asymptotically correct. AUSSIE unfolds by reweighting a reference simulator in similar fashion to OmniFold. However, its new kernel-based loss function yields one-shot solutions with minimal bias toward the reference distribution. I show results for AUSSIE applied to a range of unfolding tasks, from low-dimensional examples to full-phase-space jet substructure.

        Speaker: Ayodele Ore (ITP, Heidelberg University)
      • 136
        Data driven hadronization with HOMER

        Due to the non-perturbative nature of hadronization, its simulation relies on a fragmentation function of a fixed parametric form. We present HOMER, a data-driven alternative based on neural networks that extracts the Lund string fragmentation function directly from data. HOMER addresses the information gap between the latent and observable phase spaces through an iterative reweighting procedure. We explore how far this approach can be pushed by increasing the complexity of the string configurations, which widens the information gap, and by departing from the assumption of the reference simulation.

        Speaker: Susie Kim (ITP, Heidelberg University)
      • 137
        The PAU Survey: Self-supervised image denoising for flux-preserving photometry

        Deep photometric surveys increasingly rely on low signal-to-noise imaging to detect and characterise faint galaxies, where noise affects source detection, flux measurements, and redshift estimation. Self-supervised denoising methods are attractive because they do not require ground-truth images, but their application to quantitative photometry demands control over systematic flux biases.
        We evaluate and improve Noise2Void (N2V) for real astronomical images, using observations from the Physics of the Accelerating Universe Survey (PAUS) and PAUS-like simulations. Standard N2V improves visual quality but systematically underestimates galaxy fluxes, particularly for faint sources. To mitigate this, we modify the N2V architecture with a bounded-output layer and fine-tune the model using low-rank adaptation (LoRA) with two flux-aware loss terms that preserve the flux scale and improve repeated-exposure consistency.
        Standard N2V introduces strong flux biases for faint galaxies. Our bounded-output model restores the flux-ratio peak close to the original scale, while the flux-aware LoRA fine-tuning reduces the repeated-exposure flux-difference scatter by 47% relative to the original PAUS uncertainty, keeping the median flux ratio near unity.
        Our results show that photometric diagnostics, rather than image-quality metrics alone, are essential for evaluating denoising methods in astronomy. Flux-aware modifications to N2V can effectively reduce biases and improve measurement consistency, demonstrating that self-supervised denoising can be adapted to low-SNR survey images while preserving the flux information needed for quantitative analysis.

        Speaker: Martin Boerstad Eriksen (IFAE-PIC)
      • 138
        Constraining the Cosmic Baryon Density from Fast Radio Bursts with a Non-Parametric Reconstruction of the Hubble Parameter

        Fast radio bursts (FRBs) provide a promising tool to investigate the baryonic content of the Universe through their dispersion measures. In this work, we constrain the baryon density parameter $\Omega_b h^2$ using a sample of 130 localized FRBs combined with a non-parametric reconstruction of the Hubble parameter $H(z)$ obtained from cosmic chronometer data through the ReFANN neural-network framework. Within a Bayesian analysis, we simultaneously infer the baryon density and the host-galaxy contribution to the FRB dispersion measure. For the real FRB sample, we obtain $\Omega_b h^2 = 0.02236 \pm 0.00090$, in excellent agreement with Planck and Big Bang Nucleosynthesis results. We also explore the potential of future observations through a mock catalog of 2000 FRBs, showing that upcoming surveys may achieve sub-percent precision on the baryon density. Our results demonstrate that FRBs constitute a robust and complementary late-time cosmological probe.

        Speaker: Dr Klecio Lima (Universidade Federal de Campina Grande)
      • 139
        Neural-field reconstruction of Generalized Parton Distributions

        Generalized parton distributions (GPDs) encode the three-dimensional partonic structure of the nucleon and are inferred from exclusive scattering experiments such as deeply virtual Compton scattering. The reconstruction is an underdetermined inverse problem: the measurements — Compton form factors — constrain only one region of the underlying distribution, leaving an unmeasured complementary domain that the reconstruction must still populate. We present a neural-field framework that parametrizes the underlying double distribution F(β,α), a parametrization that links the measured DGLAP region (|x|>ξ) to the unmeasured ERBL region (|x|<ξ), with a Fourier-feature MLP and propagates it to the observable via the Radon transform. Closure tests against the Goloskokov–Kroll model recover the measured-region GPD across a range of skewness ξ; the unmeasured region is reproduced when an appropriate prior is supplied, and is provably unconstrained otherwise. This last point is the central methodological observation: the data admits a non-trivial shadow distribution — a function reproducing the observable everywhere it is measured but differing arbitrarily where it is not. Quantifying the uncertainty in the unmeasured region therefore requires sampling this null space, not propagating data noise alone. We discuss the inverse-problem framing, the prior-dependence of the extrapolation, and next steps: MC-dropout and stochastic weight averaging for null-space sampling, sensitivity analysis via the input–output Jacobian to identify the most informative kinematic regions, which provides opportunities to understand the impact of data provided by future experiments and facilities.

        Speaker: Marija Čuić (Irfu, CEA, Université Paris-Saclay/Aidas)
      • 140
        Towards the new NNPDF4.1 PDF set

        Parton distribution functions (PDFs) are a key ingredient for precision predictions at particle colliders, describing the dynamics of quarks and gluons inside the proton. Their determination is a challenging inverse problem, where PDFs are extracted from data via a convolution with theoretical predictions. The NNPDF methodology uses a neural network as a flexible parameterisation and a Monte Carlo replica method for a faithful representation of uncertainties. In the new NNPDF release, we introduce a novel hyperoptimisation to determine the set of hyperparameters used in the fit. The main novelty is the ability to determine the uncertainty associated with the hyperparameter choice, along with an optimal set of parameters.

        Speaker: Eva Groenendijk (University of Milan and INFN Milan)
    • 🔀 Patterns & Anomalies 3.404

      3.404

      Convener: Nicole Hartman (TUM)
      • 141
        Multi-task, Multi-modal Transformers for Jet Flavour Tagging in ATLAS

        Identifying the flavour of hadronic jets is a cornerstone of the ATLAS physics programme, underpinning measurements of the Higgs boson and top quark as well as searches for di-Higgs production and physics beyond the Standard Model. We present the latest generation of ATLAS flavour-tagging algorithms, which have evolved from hybrid multi-stage approaches into unified transformer architectures trained end-to-end on low-level detector information.

        The current-generation tagger, GN2, recently published in Nature Communications, improves the rejection of $c$-jets (light-flavour jets) by a factor of 3.5 (1.8) with respect to its predecessor at a 70% $b$-jet efficiency. Its successor, GN3, pushes this paradigm further: a multi-modal transformer combining charged-particle tracks, neutral particle-flow objects, and soft-lepton information via a learned object-association scheme, trained with physics-informed auxiliary objectives such as track-origin classification and jet-$p_\mathrm{T}$ regression. GN3 extends the classification space to strange-quark, up/down-quark, and gluon jets, and delivers a further 5.2% (6.1%) absolute gain in $b$-jet ($c$-jet) efficiency at fixed background rejection. The multi-task design also enables qualitatively new capabilities: an auxiliary charge-prediction objective yields the first transformer-based ATLAS tagger to separate $b$- from $\bar{b}$-initiated and $c$- from $\bar{c}$-initiated jets.

        We discuss the architecture, training workflow, the newest performance results and give an outlook on new developments.

        Speaker: Alexander Froch (Université de Genève)
      • 142
        Lorentz Local Canonicalization: How to Make Any Network Lorentz-Equivariant

        Lorentz-equivariant neural networks are becoming the leading architectures for high-energy physics. Current implementations rely on specialized layers, limiting architectural choices. We introduce Lorentz Local Canonicalization (LLoCa), a general framework that renders any backbone network exactly Lorentz-equivariant. Using equivariantly predicted local reference frames, we construct LLoCa-transformers and graph networks. We adapt a recent approach for geometric message passing to the non-compact Lorentz group, allowing propagation of space-time tensorial features. Data augmentation emerges from LLoCa as a special choice of reference frame. Our models achieve competitive and state-of-the-art accuracy on relevant particle physics tasks, while being 4× faster and using 10× fewer FLOPs.

        Speaker: Sebastian Pitz (LPNHE Paris)
      • 143
        Virtues and Vices of Equivariant Transformers

        We study for the first time the benefit of Lorentz-equivariant transformers for large-size jet tagging and flavor tagging. To control their computing demands, we optimize all implementations for inference cost metrics. In our scaling studies, we find that Lorentz-equivariant networks outperform standard transformers, provided geometric features are relevant. This holds true in an idealized world as well as for limited resources. The conditional gain from Lorentz equivariance provides interesting input to the development of foundation models for LHC data.

        Speaker: Jonas Spinner (IPPP, Durham University)
      • 144
        Gauge equivariant transformer for lattice gauge theories

        Gauge equivariant convolutional neural networks have shown that enforcing exact lattice gauge symmetry in a neural network significantly improves regression accuracy on gauge invariant observables. However, their convolutional kernels stay fixed after training and cannot adapt to configuration-dependent features. We introduce GELT (Gauge Equivariant Lattice Transformer), a gauge equivariant attention-based encoder network for lattice gauge theories that exploits the configuration-dependent attention mechanism to focus on the most physically relevant input features.
        GELT employs gauge invariant attention scores together with a gauge equivariant matrix bilinear value path. Keys and values are parallel transported from neighboring lattice sites within a L1-ball along shortest lattice paths. We inject a geometrical prior in the attention weights with rotary positional encoding (RoPE).
        We test GELT on various regression targets involving physical observables and compare its performance against existing gauge-equivariant architectures. We also investigate the interpretability of the attention filters, studying whether attention localizes in topologically rich regions and whether the attention range correlates with physical correlation length.

        Speaker: Fracesco Passante (Sapienza Università di Roma and INFN)
      • 145
        Towards a comprehensive high-z quasar selection: multimodal self-supervised learning on Euclid data

        Within the first billion years of cosmic time, galaxies and supermassive black holes (SMBHs) had already formed. Especially, rapidly growing billion solar mass black holes - shining as quasars - discovered at z > 7 challenge standard models of SMBH growth. To constrain their formation, expanding the quasar redshift frontier is critical.
        The deep, wide-area photometry of the Euclid mission, is revolutionizing this frontier and is expected to discover hundreds of new quasars. Over the past two years, about 50 new quasars have been confirmed, more than doubling the number of z > 7 quasars. This success was facilitated by state-of-the-art supervised machine learning methods, efficiently selecting candidates from billions of sources.
        However, they are naturally biased toward the training dataset, potentially running the risk of missing quasars with properties different from known populations.
        In this work, we explore self-supervised machine learning methods for a less biased census of the quasar population. Specifically, we build a multimodal probabilistic autoencoder (PAE) for anomaly detection and apply it to Euclid data. We combine image cutouts and catalog data via cross-attention in our architecture. The PAE training consists of two stages: first, latent representations are learned using a standard autoencoder; second, the latent space is mapped to a Gaussian using normalizing flow. We define an anomaly score based on the source distribution likelihood.
        In current tests, known quasars are well separated from contaminants and form a compact region in the representation space, well consistent with high anomaly scores.
        With this study we are laying the foundation for a less biased, more comprehensive quasar selection methodology, which will be applied to the full Euclid DR1 dataset.

        Speaker: Tatsuyuki Sekine (University of Hamburg)
      • 146
        SURFing to the Fundamental Limit of Jet Tagging

        Beyond the practical goal of improving search and measurement sensitivity through better jet tagging algorithms, there is a deeper question: what are their upper performance limits? Generative surrogate models with learned likelihood functions offer a new approach to this problem, provided the surrogate correctly captures the underlying data distribution. In this work, we introduce the SUrrogate ReFerence (SURF) method, a new approach to validating generative models. This framework enables exact Neyman–Pearson tests by training the target model on samples from another tractable surrogate, which is itself trained on real data. We argue that the EPiC-FM generative model is a valid surrogate reference for JetClass jets and apply SURF to show that modern jet taggers may already be operating close to the true statistical limit. By contrast, we find that autoregressive GPT models unphysically exaggerate top vs. QCD separation power encoded in the surrogate reference, implying that they are giving a misleading picture of the fundamental limit.

        Speaker: Ranit Das (Heidelberg University)
      • 147
        Multiclass classification without labels (multiCWoLa)

        In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.
        In this talk, we also present applications to particle physics and astrophysics.

        Speakers: Johann Ioannou-Nikolaides (Niels Bohr Institute), Raphaël Bonnet-Guerrini (Computer Science Dep. University of Milan)
      • 148
        A Machine-Learning Analysis of the Higgs Boson in the H → ZZ∗ →4ℓ Channel

        The H → ZZ → 4ℓ channel remains one of the cleanest probes of Higgs boson properties at the LHC, owing to its fully reconstructable final state and well-understood background composition. This work addresses the problem of identifying an optimal machine-learning classifier for extracting the H → ZZ → 4ℓ signal from the ATLAS Open Data 2025 release (√s = 13 TeV, 36.6 fb⁻¹), comparing seven algorithms, XGBoost, LightGBM, Random Forest, a multilayer perceptron, Logistic Regression, QDA, and Gaussian Naive Bayes, under a single, controlled experimental protocol. The specialty of this analysis is twofold. First, every classifier is trained exclusively on Monte Carlo simulation and then applied, without retraining or recalibration, to the real ATLAS dataset, so that the reported significance reflects genuine generalization from simulation to data rather than performance on held-out simulated events alone. Second, background estimation itself is treated as a robustness test: each classifier's signal significance on real data is computed twice, once under a Monte Carlo background prediction and once under a data-driven sideband extrapolation, and a classifier is judged reliable only if the two estimates agree within uncertainties, since the two methods rest on largely independent assumptions. Four-lepton events are reconstructed from the ATLAS Open Data samples and reduced to a set of 31 physics-motivated kinematic and angular features, subsequently pruned to 20 via separation-power ranking and correlation filtering. All seven classifiers are trained and cross-validated on Monte Carlo simulation under a 5-fold stratified scheme, then applied at a fixed classifier-score threshold to the full real-data sample, with the signal region and sideband control regions defined directly in the four-lepton invariant mass spectrum. Signal significance is computed independently under both background estimators for each classifier, yielding a direct test of consistency rather than a single, isolated figure of merit. Under this framework, LightGBM emerges as the best-performing and most consistent classifier overall, achieving a signal significance of Z = 4.66 ± 1.22σ (p = 1.55 × 10⁻⁶) under the Monte Carlo background estimate and Z = 5.08 ± 1.50σ (p = 1.90 × 10⁻⁷) under the sideband estimate. XGBoost follows closely, with Z = 4.24 ± 1.10σ (p = 1.14 × 10⁻⁵) and Z = 4.70 ± 1.35σ (p = 1.31 × 10⁻⁶) under the same two methods, respectively. Both classifiers reach evidence-level significance under both background estimators, with their results agreeing within uncertainties across two independent methods, confirming that the models generalize from simulation to real collision data and that the resulting evidence is not an artifact of a particular background model. This consistency establishes gradient-boosted decision trees as a robust, reproducible means of recovering Higgs boson evidence from public ATLAS Open Data using standard machine-learning methods alone.

        Speaker: Mane Papoyan (American University of Armenia (AUA))
    • 🔀 Simulations & Generative Models HS2

      HS2

      Convener: Andrew Pilkington (University of Manchester)
      • 149
        MadNIS at NLO

        We combine fast amplitude surrogates with neural importance sampling to accelerate NLO calculations. For virtual corrections, a learned ratio to the Born matrix element with calibrated uncertainties guarantees reliable precision across phase space. For real emission, we stick to the standard FKS subtraction and train sector-conditioned surrogates of the regularized integrands away from divergences. MadNIS then uses multi-channel mappings and FKS sectors as conditions. We validate our approach for electron-positron scattering to three and four jets and find significant speed-ups and variance reduction in the integration.

        Speaker: Giovanni De Crescenzo
      • 150
        Neural Control Variates for LO and NLO

        We employ neural control variates to minimize the range of event weights and avoid negative weights for phase-space integration and event generation. A signed control variate, built from two normalizing flows, fulfills both tasks. Combined with neural importance sampling, it significantly reduces the computational cost of LO and NLO predictions. For the NLO case, our conditional neural control variate can be viewed as a trainable subtraction term, complementing the established physics subtraction schemes for enhanced sampling performance.

        Speaker: Sophia Vent (Heidelberg University)
      • 151
        Billion-Event Generation for Complete Partonic Many-Jet Processes: An End-to-End GPU Workflow with Normalizing Flows

        Producing very large unweighted event samples for high-multiplicity processes is limited by expensive matrix-element evaluations and low unweighting efficiencies. We present the first end-to-end GPU-resident event-generation workflow controlled from Python that integrates normalizing-flow proposals with the parton-level event generator \Pepper. Helicity-conditioned coupling flows are trained using online updates supplemented by sample replay and deployed across all subprocesses of complete proton--proton collision processes with many final-state jets. In this workflow, Python and \Pepper exchange flow-generated phase-space points and the corresponding target-density evaluations directly in device memory. Python performs flow sampling, proposal-density evaluation, and unweighting, while \Pepper evaluates the matrix elements, PDFs, and phase-space factors defining the target density and writes the accepted events in standard formats. We compare subprocess-specific flows, with one flow per partonic subprocess, to grouped conditional flows that share parameters among subprocesses with related parton content. The workflow is benchmarked for $pp \to e^+e^- + 4j$, $pp \to e^+e^- + 5j$, $pp \to t \bar t + 4j$, $pp \to 4j$, and $pp \to 5j$ production. On four H100 GPUs, we generate (10^9) unweighted events for each benchmark process. Including the cost of flow training, the workflow achieves end-to-end speedups of up to two orders of magnitude over \Pepper-direct generation and turns a multi-week task into a sub-day computation. It thereby makes billion-event production more practical and offers a pathway to alleviating the Monte Carlo statistics bottleneck in high-multiplicity collider physics.

        Speaker: Daohan Wang (Marietta Blau Institute for Particle Physics, Austrian Academy of Sciences)
      • 152
        Generative Amplification with Surrogate Monte Carlo

        Amplitude surrogates speed up one of the most computationally expensive steps in the LHC simulation chain. We adapt the generative amplification framework to amplitude surrogates evaluated by reweighting and apply it to uncertainty-aware surrogates. We show that amplitude surrogates, like generative event generators, can exhibit generative amplification when trained on a limited set of expensive amplitude evaluations. In particular, we study gluon-associated $Z$ production and find that amplification emerges in the sparsely populated high-$p_T^Z$ tails, where it matters most.

        Speaker: Rebecca Revelli (Heidelberg University)
      • 153
        FASTColor – Full-color Amplitude Surrogate Toolkit for QCD

        High-multiplicity events remain a bottleneck for LHC simulations due to their computational cost. We present a ML-surrogate approach to accelerate matrix element reweighting from leading-color (LC) to full-color (FC) accuracy, building on recent advancements in LC event generation. Comparing a variety of modern network architectures for rep-
        resentative QCD processes, we achieve speed-up of around a factor two over the current LC-to-FC baseline. We also show how transformers learn and exploit underlying symmetries, to improve generalization. Given the gained trust in trained networks and developments in learned uncertainties, the LC-to-FC approach will eventually benefit further
        from not needing a final classic unweighting step.

        Speaker: Javier Mariño Villadamigo (University of Heidelberg)
      • 154
        Estimating event-by-event multiplicity by a Machine Learning Method for Hadronization Studies

        Hadronization is a non-perturbative process, which theoretical description can not be deduced from first principles. Modeling hadron formation requires several assumptions and various phenomenological approaches. Utilizing state-of-the-art Deep Learning algorithms, it is eventually possible to train neural networks to learn non-linear and non-perturbative features of the physical processes. In this presentetion, the prediction results of three trained ResNet networks are presented, by investigating charged particle multiplicities at event-by-event level [1]. The widely used Lund string fragmentation model is applied as a training-baseline at $\sqrt{s}=7$ TeV proton-proton collisions. We found that neural-networks with $≳O(10^3)$ parameters can predict the event-by-event charged hadron multiplicity values up to $N_{ch}≲90$. We also present results on how ML methods can preserve the KNO-scaling [2,3] of PYTHIA8 and HIJING++ generated events.

        [1] Int.J.Mod.Phys.A 40 (2025) 21, 2542011
        [2] J.Phys.Conf.Ser. 3206 (2026) 1, 012129
        [3] PoS ICHEP2022 (2022) 1188

        Speaker: Gergely Gábor Barnaföldi (HUN-REN Wigner RCP)
    • 🧠 Working groups: Foundation Models & Discovery (WG1) 2.404

      2.404

      • 155
        Foundation Models & Discovery (WG1)
    • 12:30
      🍽️ Lunch Foyer

      Foyer

    • 🔀 Agentic AI HS2

      HS2

      Convener: Anna Hallin (Uni Hamburg)
      • 156
        Development of AI agents in LHCb

        The LHCb collaboration at CERN spans from individual physics analyses to production computing. That breadth makes it a natural place for the agentic AI. The software stack grew over decades around human experts. The physics is split into internal working groups, which standard LLMs were not trained to distinguish. This HEP-specific workflow cannot be integrated with commercial tools without custom pipelines.
        The agentic AI landscape also changes quickly. In an isolated environment with large amounts of internal knowledge, we cannot adopt new model or agent architectures on intuition alone: we need a scientific methodology to track impact, avoid regressions, and manage risks. Within LHCb we are therefore building dedicated tooling to give agents structured access to our data and services, and benchmarking them with metrics tailored to our workflows so that we can quantify how each change affects real tasks.
        We are developing the LHCb Brain, an agentic chatbot whose router can dispatch requests across many internal services, much like branching pathways. At the same time, we are exploring IDE analysis assistance and fully encapsulated agents running inside the LHCb stack. What we have learned is that usefulness depends less on generic agent capabilities and more on overcoming fragmentation: connecting legacy services, domain conventions, and workflows that were never meant to be automated. This talk will give an overview of the current developments happening in LHCb.

        Speaker: Sergio Arguedas Cuendis (CONARE)
      • 157
        MadAgents

        We present MadAgents, an effective and communicative set of agents for working with MadGraph. Agentic installation, learning-by-doing training, user support, and autonomous simulation campaigns provide easy access to state-of-the-art simulations and accelerate LHC research. We show how MadAgents interact with inexperienced and advanced users, support a range of simulation tasks, and analyze the results. We also discuss the updated MadAgents implementation, with significantly improved answer reliability and the ability to improve through use.

        Speaker: Daniel Schiller (Institute for Theoretical Physics, Heidelberg University)
      • 158
        Agentic Re-Casting using Agentic Re-Simulations

        Analysis re-casting at the LHC is highly standardized and nevertheless requires resources, time, and expert physics input. Building on the newly developed MadAgents.v3 framework, we demonstrate how a global SMEFT analysis can be updated through an agentic workflow with a physicist in the loop. Although demonstrated within the SFitter framework, the underlying technical aspects of the agentic recasting we present can be readily extended to other analysis frameworks.

        Speaker: Nikita Schmal (ITP, Heidelberg University)
      • 159
        AgentRivet: an automated system for producing Rivet routines from journal publications

        Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measurements. Rivet is a C++ toolkit that allow new theoretical models to be compared to the measurements, thus aiding the development and tuning of Monte Carlo event generators as well as searches for physics beyond the Standard Model. However, analysis coverage is known to be incomplete, with only 39% of measurements having documented and publicly available Rivet routines. In this article, we design and implement an automated workflow based on Large Language Models with the goal of providing the missing routines. This multi-step workflow, referred to as AgentRivet, extracts the physics analysis information from published papers and writes the missing Rivet routines, with intermediate code- and physics- reviews as part of an autonomous quality control. We report the results obtained using commercial Large Language Models, provided by OpenAI, Anthropic, and Google, for two recent measurements from the ATLAS and CMS experiments. We find that AgentRivet produces competent Rivet routines with few syntax errors. The physics fidelity of the routines is reasonable and follows the explanations given in the relevant publications. Nevertheless, physics-implementation issues do arise and are investigated using the artefacts produced by AgentRivet. The majority of physics implementation issues arise from subtle-but-ambiguous definitions in the given publication, although some models struggle to implement complex observables even when clear definitions are given.

        Speaker: Andrew Pilkington (University of Manchester)
      • 160
        Agentic Deciphering of Complex Computation Graphs

        Natural science increasingly relies on complex models embedded in large computational frameworks. Motivated by the DEMOS (DEmocratizing MOdelS) project's goal of making research models interoperable and encoding them into framework-independent formats, we address a necessary prerequisite: reconstructing an analysis-specific model and translating it between frameworks without changing its numerical predictions.

        As a case study in hadron physics, we reimplement the $B^+ \to D^{*\pm}D^{\mp}K^+$ decay amplitude -- originally analyzed with the TF-PWA framework by the LHCb collaboration -- into CascadeDecays.jl, serving as a standard for the DEMOS project. Due to the complexity of the amplitude analysis framework through deeply nested, hardware optimized function calls, manual code tracing is extremely slow and error-prone. For the same reason, agentic forward tracing of the execution flow from input parameters to the final amplitude fails due to the combinatorial explosion of execution paths.

        To overcome this, we introduce a method that recursively applies a standardized analysis protocol to each function layer: starting with the single top-level amplitude function, the AI agent identifies all underlying dependencies and sequentially progresses down the call tree to the input parameters and kinematics.

        This reconstruction enabled the identification of convention mismatch factors, resulting in a framework alignment within numerical precision. The validated implementation establishes the basis for serializing this decay model, while the AI-driven backward-tracing approach provides a general method for deciphering codebases that collapse nested model definitions into single numerical outputs.

        Speaker: Alexander Kazatsky (Ruhr University Bochum)
      • 161
        Path-based RAG on particle physics literature

        Across many fields of fundamental physics, researchers must reason over large and growing corpora of specialized publications and increasingly turn to language models for help — yet these models rarely connect their answers to the primary sources they draw on and often lack the information to give a reliable answer. Building on the PathRAG graph-retrieval framework, we adapt it to fundamental-physics literature and demonstrate it on the full corpus of LHCb collaboration papers (about 850), grounding the model in the underlying physics. Documents are turned into an entity–relationship graph, and a query is answered by walking relational paths between the entities it mentions, so the model receives a concise chain of evidence rather than an unstructured collection of similar passages — following, for example, how a measured quantity traces to the decay channel that determines it and on to the particles in its final state, or how two decay channels are linked through the symmetry they share. We anchor the graph in the Particle Data Group database, mapping the many spellings a particle takes across papers to one canonical node, and we specialize entity extraction to high-energy physics so that detectors and datasets do not dominate the graph. The system has ingested the complete corpus using a small 31b Gemma model and retrieves in seconds. Because the model can be run locally, we can process analyses that cannot be submitted to public models. More importantly, the graph retains source provenance that graph-based retrieval pipelines typically lack: each entity and relation records the exact document passages behind it, the context is grouped by source paper, and every model claim carries a citation that resolves to a verbatim source passage, checked by an optional LLM-as-judge pass. The result is a literature-aware assistant whose answers are verifiable and reliable by design.

        Speaker: Aleksei Mikhasenko (Universität Bonn)
      • 162
        Head-to-head comparison of agentic AI systems applied to gravitational wave studies

        We present a first head-to-head comparison of different agentic AI systems applied to the study of gravitational waves. Agentic systems, supervised either via a human-in-the-loop or in fully autonomous mode, are tasked with performing matched-filter analyses in the context of the Einstein Telescope mock data challenge and with fully writing a scientific paper in the style of PRD. While all systems successfully complete the pipeline, we observe very different patterns depending on the chosen backend LLM, with different computational costs and varying quality and correctness of the produced manuscripts.

        Speaker: Gianluca Inguglia (MBI Vienna)
      • 163
        Plausible but Wrong: A Case Study on Agentic Failures in Astrophysical Workflows

        Agentic AI systems are increasingly deployed in scientific workflows, yet existing evaluations primarily measure task completion and provide limited insight into scientific reliability. We argue that such metrics are insufficient for assessing autonomous systems in research settings, where plausible but incorrect results may be more dangerous than overt failures.

        We present a structured evaluation framework for scientific AI agents and apply it to CMBAgent across eighteen astrophysical tasks spanning tool-grounded computation, Bayesian inference, and multi-step research workflows. The framework combines execution success, numerical fidelity, parameter recovery, physical plausibility, and failure transparency to assess both performance and scientific validity.

        Our analysis reveals a consistent pattern across workflow paradigms. In the One-Shot setting, domain-specific context retrieval drives a ~6x performance improvement (Final Score 0.85 vs. ~0 without context); the primary failure mode without context is silent wrong computation — syntactically valid code producing plausible but numerically incorrect results. In the Deep Research setting, under-constrained inference tasks fail systematically: parameter degeneracies go undetected, yielding physically inconsistent posteriors reported as valid results. Compositional workflows exhibit inter-trial inconsistency and systematic bias without self-diagnosis. Across all tasks, failure transparency is the weakest dimension: the agent never proactively flags known pathologies in its own outputs.

        These findings highlight a critical challenge for physics and astronomy: the dominant failure mode of agentic AI is not crashing or refusing a task, but confidently producing incorrect scientific conclusions without warning. We argue that reliability, physical consistency, and error detection should be treated as first-class evaluation objectives.

        Speaker: Shivam Rawat (University of Bonnn)
    • 🔀 Explainability & Theory 1.404

      1.404

      Convener: Anja Butter (LPNHE, Sorbonne Université, Université Paris Cité, CNRS/IN2P3, Paris, France, Institut für Theoretische Physik, Universität Heidelberg, Germany)
      • 164
        Discovering Euler-Lagrange equations from trajectory data

        Neural ODEs enable the automatic discovery of a system's ordinary differential equations (ODEs) using only trajectory measurements. We combine this powerful and widely used machine learning method with the principle of stationary action from theoretical physics to learn only ODEs that are admissible as fundamental physical laws. To this end, we develop Helmholtz metrics, a machine learning architecture that quantifies violations of the Helmholtz conditions arising in the inverse problem of the calculus of variations. These conditions determine whether a Lagrangian exists that yields a given ODE as an Euler-Lagrange equation. We then use this measure as a regularization term for Neural ODEs to obtain Lagrangian Neural ODEs, which recover Euler-Lagrange equations directly from positional data. When no Lagrangian exists for a given system, the models reflect this incomaptibility through their convergence behaviour, indicating inconsistencies in the system's physical description. For systems that admit Lagrangian description, ODEs learned in this way are not only more theoretically grounded, but also yield more physically consistent and robust predictions, particularly in sparse and noisy data regimes.

        Speaker: Luca Wolf
      • 165
        Learning Relativistic Geodesics and Chaotic Dynamics via Stabilized Lagrangian Neural Networks

        Lagrangian Neural Networks (LNNs) can learn arbitrary Lagrangians from trajectory data, but their unusual optimization objective leads to significant training instabilities that limit their application to complex systems. We propose several improvements that address these fundamental challenges, namely, a Hessian regularization scheme that penalizes unphysical signatures in the Lagrangian’s second derivatives with respect to velocities, preventing the network from learning unstable dynamics, activation functions that are better suited to the problem of learning Lagrangians, and a physics-aware coordinate scaling that improves stability. We systematically evaluate these techniques alongside previously proposed methods for improving stability. Our improved architecture successfully trains on systems of unprecedented complexity, including triple pendulums, and achieved 96.6% lower validation loss value and 90.68% better stability than baseline LNNs in double pendulum systems. With the improved framework, we show that our LNNs can learn Lagrangians representing geodesic motion in both non-relativistic and general relativistic settings. To deal with the relativistic setting, we extended our regularization to penalize violations of Lorentzian signatures, which allowed us to predict a geodesic Lagrangian under AdS4 spacetime metric directly from trajectory data, which to our knowledge has not been done in the literature before. This opens new possibilities for automated discovery of geometric structures in physics, including extraction of spacetime metric tensor components from geodesic trajectories. While our approach inherits some limitations of the original LNN framework, particularly the requirement for invertible Hessians, it significantly expands the practical applicability of LNNs for scientific discovery tasks.

        Speaker: Abdullah Umut Hamzaogullari (Bogazici University)
      • 166
        Recovering Irreducible Three-Body Interactions with Lagrangian Neural Networks

        Lagrangian Neural Networks (LNNs) learn mechanical systems directly from trajectory data by parameterizing a scalar Lagrangian and deriving the dynamics through the Euler-Lagrange equations, so that the learned model is strongly biased toward motion that respects the variational structure of mechanics. The interactions studied in this way have so far been pairwise-additive, leaving open whether LNNs can recover an irreducible many-body interaction from motion alone. We investigate this question using the Axilrod-Teller triple-dipole interaction, the leading three-body dispersion contribution, whose dependence on the collective geometry of three particles cannot be represented by any sum of pairwise interactions. Our strategy is to prescribe the kinetic energy exactly and learn only the potential as an unknown function of symmetry-preserving geometric features comprising interparticle distances and angles. Together with a signal-adaptive sampling strategy that concentrates training data where the three-body signal is strongest while preserving coverage of the broader configuration space, we recover the interaction from acceleration observations alone. The learned potential accurately reproduces the radial scaling, angular dependence, and characteristic sign reversal of the Axilrod-Teller interaction, and generates stable trajectories. Embedding this interaction within the standard pairwise Lennard-Jones potential further lets us quantify its observability. The three-body contribution is recovered once its dynamical signature rises above the residual modeling error, and we bracket this transition between the two couplings, recovering it cleanly at amplified coupling while finding it drops below the error floor at the coupling characteristic of a real noble gas. These results demonstrate that LNNs are capable of recovering irreducible three-body interactions directly from trajectory data when their dynamical signature is sufficiently strong, extending their demonstrated capabilities beyond previously studied pairwise-additive systems.

        Speaker: Sena Kalabalik
      • 167
        Solving Schwarzschild Geodesic Motion with Lagrangian Neural Networks

        Lagrangian Neural Networks (LNNs) learn dynamics from trajectory data by parameterizing a scalar Lagrangian and deriving accelerations through the Euler–Lagrange equation, providing a variational inductive bias that improves physical consistency over purely data-driven models. LNNs have recently been extended to general relativistic settings, leaving Schwarzschild geodesics as an open challenge. However, training on complex relativistic systems is hindered by the diversity of the dynamics, as well as numerical instabilities and the difficulty of adequately sampling physically relevant, diverse orbital regimes. We introduce two coordinated improvements: a decoupled position–velocity normalization scheme that independently rescales each coordinate and its time derivatives, reweighting acceleration targets to balance contributions across all channels; and an energy-constrained log-uniformly sampled dataset generation strategy informed by the innermost stable circular orbit (ISCO) structure. With these improvements, our model accurately reproduces perihelion precession and relativistic time dilation, conserves energy and angular momentum to within a few percent even after completing multiple orbits around the central mass and predicts orbital trajectories across eight orbital scenarios in close, medium, and far-field regimes. These results establish that LNNs, augmented with physically motivated normalization and sampling, can reliably learn geodesic dynamics in curved spacetime. Our work also represents the first successful training of LNNs on Schwarzschild geodesics and offers a promising direction for data-driven modeling of complex relativistic systems.

        Speaker: Mr Sukru Caglar (Bogazici University)
      • 168
        Lagrangian Neural Networks with Simultaneously Learned Constraints

        Lagrangian Neural Networks (LNNs) learn dynamical laws from trajectory data by representing a system’s Lagrangian with a neural network. In their standard formulation, however, LNNs are typically applied using carefully chosen generalized coordinates, such as angular coordinates for a pendulum, which already encode the system’s constraints and true degrees of freedom. Existing approaches permit the use of redundant Cartesian coordinates, but require the governing holonomic constraints to be supplied explicitly. We introduce, to our knowledge, the first framework that instead learns the Lagrangian and previously unknown holonomic constraints simultaneously from trajectory data. The proposed architecture consists of two jointly trained neural networks: a Lagrangian network, which learns the dynamics, and a constraint network, which discovers the constraint manifold. Derivatives obtained from the two networks are combined through the constrained Euler–Lagrange equations, coupling the learning of dynamics directly to the learning of geometry. Constraint discovery is enabled by additional loss terms enforcing geometric and kinematic consistency with the observed trajectories, and via representing the constraints as signed distance functions provides a normalized description of the same physical constraint manifold. Using pendulum systems observed only through Cartesian trajectories as representative examples, the framework recovers the hidden holonomic structure while learning dynamics that generate accurate, constraint-preserving motion. By removing the need for either predefined generalized coordinates or explicitly supplied constraint equations, this work extends LNNs to genuinely black-box settings in which only trajectory observations are available, providing a foundation for discovering both dynamical laws and unknown constraints in more complex systems, including future applications to astrophysical data.

        Speaker: Abdullah Umut Hamzaogullari (Bogazici University)
      • 169
        Black Hole Spectroscopy via Physics-Informed Neural Networks: The Case of Massive Scalar Fields in Schwarzschild Spacetime

        The detection of gravitational waves has turned the study of quasinormal modes (QNMs) into a fundamental tool for testing General Relativity. However, traditional numerical methods often face stability challenges when dealing with massive fields or high-order overtones. In this work, we present a robust computational framework based on Physics-Informed Neural Networks (PINNs) to solve the eigenvalue problem associated with the Regge-Wheeler equation. Our approach incorporates a tailored ansatz to enforce radiation boundary conditions and utilizes a hybrid optimization strategy, combining the Adam algorithm for global exploration with the L-BFGS method for high-precision refinement. We demonstrate that this architecture achieves unprecedented residual accuracy, in the order of $10^{-7}$, effectively capturing the complex frequency spectrum for both massless and massive scalar fields ($\mu = 0.1$). Our results show excellent agreement with traditional spectral methods and highlight the PINN’s stability in resolving overtones up to $n=3$. This framework establishes a flexible and efficient path for exploring modified gravity theories in future research.

        Speaker: Prof. M. B. Cruz (State University of Paraiba)
      • 170
        Neural Boltzmann Equations

        The dynamics of particles in the early universe are described by Boltzmann equations, which involve high-dimensional phase space integrals. Classical approaches use quadrature integration and evolve the system on a fixed momentum grid, which scales poorly to high-dimensional integrals and parameter scans, severely limiting the complexity of processes that can be studied. We introduce Neural Boltzmann Equations (NBE), which use three coupled concepts to overcome these limitations. First, particle properties are encoded in physics-inspired neural distribution functions, with parameters that can be predicted using neural networks, enabling efficient parameter scans. Second, phase space integrals are evaluated with Monte Carlo methods, using established importance sampling tools from collider physics that benefit from the efficiency of neural distribution functions. Third, we use the natural gradient method to train the networks and evolve the system. After demonstrating the individual benefits of NBEs, we use our framework to perform a precision calculation of the number of effective neutrino degrees in the early universe.

        Speaker: Jonas Spinner (Durham University)
    • 🔀 Foundation Models HS1

      HS1

      Convener: Luca Fallböhmer (Max-Planck-Institute for Nuclear Physics)
      • 171
        A Foundation Model for 21cm cosmology: Cross-Simulator Robustness via Self-Supervised Summaries

        Simulation-based inference (SBI) has become the default for intractable-likelihood problems across fundamental physics, but is limited by model misspecification: a density estimator trained on one simulator becomes inaccurate or miscalibrated when applied to data drawn from a different forward model or with unseen instrumental systematics. We study this in 21cm cosmology, where the brightness-temperature field of the Epoch of Reionization is the key target for the Square Kilometre Array. Competing simulators disagree and the true instrument response is unknown in advance. We treat SKATR, a Vision Transformer pretrained with a self-supervised Joint Embedding Predictive Architecture (JEPA) as a foundation model for the field: it is trained once, label-free, on a large set of cheap semi-numerical lightcones, then frozen and used purely as a summarizer, with a lightweight conditional-flow-matching head performing the downstream parameter inference. Despite never seeing the target simulator, its parameters, or any noise, the frozen encoder transfers under a joint domain shift, to a more physically accurate hydrodynamical radiative-transfer simulator and to realistic interferometer noise, matching or exceeding a supervised network trained directly on the noisy target while remaining the only pipeline whose posteriors stay calibrated. In a controlled ablation that holds architecture fixed and varies only the pretraining objective, we find that the self-supervised objective, not data scale, is what buys this robustness. These results position self-supervised pretraining as a general recipe for trustworthy, physics- and systematics-agnostic inference.

        Speaker: Yannic Pietschke (Institute for Theoretical Physics, Heidelberg)
      • 172
        Sparse autoencoders for a neutrino foundation model: From physical concepts to interpretable uncertainty

        Pretrained foundation models are emerging in neutrino telescopes, but what their internal representations actually contain, and whether the tasks built on top of them put it to use, is largely unknown. We study PolarBERT, a transformer pretrained on IceCube data and fine-tuned for direction reconstruction. Sparse autoencoders uncover a validated atlas of physics across the frozen backbone, including event quality, auxiliary activity, brightness, and detector depth at the layers where these concepts are most clearly represented. Causal interventions show the direction head barely draws on this atlas, relying instead on a single axis with no simple physical meaning.
        We then show that a different task can make better use of the same representation. A second head, trained on the same frozen summary to predict the model’s angular reconstruction error, turns out to depend causally on precisely the quality and brightness concepts the direction head ignores.
        The result is a interpretable and trustworthy per-event uncertainty estimator that reaches $3.2^\circ$ median resolution at 20% selection efficiency, against about $20^\circ$ for the best detector observable, and its reliance on that physics is established by intervention rather than assumed.

        Speakers: Johann Ioannou-Nikolaides (Niels Bohr Institute), Raphaël Bonnet-Guerrini (Computer Science Dep. University of Milan)
      • 173
        The Low-Level Inverse Jet-Quenching Problem: Learning Vacuum Constituents from Medium-Modified Jets

        Jet quenching modifies the energy and substructure of jets in heavy-ion collisions, but the corresponding inverse problem remains largely unexplored at the constituent level: given a medium-modified jet, infer the vacuum jet associated with the same parton shower. We formulate this low-level inverse jet-quenching task using paired HYBRID simulations, in which each quenched shower is obtained by modifying a known vacuum shower and therefore provides direct jet-by-jet supervision. Central to this construction is, to our knowledge, the first constituent-level treatment of HYBRID medium response. Positive and negative wake contributions are handled event-wise through constituent subtraction, conceptually paralleling JEWEL’s 4MomSub approach while retaining a physical constituent list for downstream analysis. We subsequently add and subtract a realistic underlying event according to the Apples-to-Apples prescription, ensuring that residual background contamination enters the learning problem consistently. Our deterministic baseline combines the pretrained PET v2 body of OmniLearned with a cross-attention decoder that predicts variable constituent multiplicity, identified-particle labels, and an on-shell four-vector representation. We compare fine-tuning of pretrained representations with training the same architecture from scratch and evaluate constituent-level reconstruction together with the closure of jet kinematics and substructure observables. The inferred vacuum jet enables jet-by-jet comparisons of kinematics and substructure in settings where a paired vacuum shower is unavailable, including standard JEWEL samples and, subject to domain and uncertainty validation, experimental data. This provides a controlled baseline for extending the inverse map to probabilistic conditional transport.

        Speaker: João A. Gonçalves (U. Bonn, B-it, Lamarr Institute)
    • 🔀 Real-Time Data Processing 3.404

      3.404

      Convener: Thorsten Buss (RWTH Aachen)
      • 174
        EPIGRAPHY: Advancing Edge AI for Particle Physics through Community Benchmarks and Data Challenges

        Machine learning is playing an increasingly important role in particle physics, from offline data analysis to real-time event selection at the Large Hadron Collider. The COST Action EPIGRAPHY (Edge deeP learnIng foR pArticle PHYsics) brings together researchers across Europe to advance efficient deep learning methods for resource-constrained, low-latency computing environments. This contribution will provide an overview of the activities of the EPIGRAPHY network, with a particular focus on the development of a community-wide machine learning benchmark and data challenge. The initiative defines a suite of standardised online and offline challenges covering a broad range of physics tasks, accompanied by a common dataset (Collide-2V), evaluation metrics, baseline models, and submission infrastructure. The online challenges additionally incorporate realistic FPGA latency and resource constraints, enabling fair comparison of algorithms for edge deployment.

        Speaker: Sebastian Dittmeier (PI)
      • 175
        Economical Jet Taggers -- Equivariant, Slim and Quantized

        Modern machine learning is transforming jet tagging at the LHC, but the leading transformer architectures are large, not particularly fast, and training-intensive. We present a slim version of the L-GATr tagger, reduce the number of parameters of jet-tagging transformers, and quantize them. We compare different quantization methods for standard and Lorentz-equivariant transformers and estimate their gains in resource efficiency. We find an order-of-magnitude reduction in energy cost for an moderate performance decrease, down to 1000-parameter taggers. This might be a step towards trigger-level jet tagging with small and quantized versions of the leading equivariant transformer architectures.

        Speaker: Antoine Petitjean (ITP, Heidelberg University)
      • 176
        Enabling portable and optimized Machine Learning Inference through code generation

        SOFIE, or the System for Optimized Fast Inference code Emit, is a tool developed by the ML4EP team at CERN that translates trained machine learning models into self-contained, low-latency C++ code. The generated code is portable, hardware-agnostic, and highly optimized while depending only on BLAS libraries.

        Code generated by SOFIE achieves portability across different hardware architectures by leveraging the abstract buffer definitions provided by the alpaka[1] library for heterogeneous programming. In addition to generating portable C++ code, SOFIE incorporates several optimization methods that improve inference efficiency, including aggressive kernel fusion, efficient memory allocation and management, and support for inference on quantized models.

        Beyond model inference, SOFIE serves as a component in several projects. These include Yukti, a header-only interface that provides a unified zero-copy API for machine learning inference runtimes. Furthermore, SOFIE is used in RooFit as a neural surrogate for likelihood evaluation, with automatic differentiation provided by CLAD, a source-to-source automatic differentiation tool for C++. Embedding BOA Constrictor[2], SOFIE provides a hardware-agnostic pipeline for highly optimized compression and decompression while remaining independent of the underlying hardware architecture.

        [1] Matthes, A., Widera, R., Zenker, E., Worpitz, B., Huebl, A., & Bussmann, M. (2017, June 30). Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library. Retrieved from http://arxiv.org/abs/1706.10086
        [2] Gupta, A., Doglioni, C., & Elliott, T. J. (2026). BOA constrictor: a Mamba-based lossless compressor for scientific data. Machine Learning: Science and Technology, 7(3), 035014. doi:10.1088/2632-2153/ae64a9

        Speaker: Sanjiban Sengupta (CERN, University of Manchester)
      • 177
        From slow to real-time control in accelerator-based research infrastructures

        The KIT accelerator team has contributed significantly to the state-of-the-art of accelerator control especially in controlling electron beams near real-time in storage rings and synchrotrons. Adapted to the latency of components like magnets, slow control optimizes beam transfer and injection into a storage ring [1]. Fast control by online reinforcement learning with AI-on-hardware acceleration reached unprecedented latencies of less than 3 µs without prior training via low-level radiofrequency (RF) manipulation [2]. This enables the control of non-equilibrium processes and dynamics [3]. On a larger scale and on several levels of complexity does the electrical grid also affect the performance of accelerator-based research infrastructures leading to real-time digital twins [4]. The presentation will give an overview of AI deployment for accelerator beam control and accompanying benefits in energy efficiency and sustainability. Our work also prepares for the EC/EU supported project TwinRISE starting in 2027 [5], an initiative of the ARTIFACT network [6].

        References
        [1] Bayesian optimization of the beam injection process into a storage ring; Xu, C.; Boltz, T.; Mochihashi, A.; Santamaria Garcia, A.; Schuh, M.; Müller, A.-S.; 2023. Phys. Review Accel. Beams 26, 034601. doi: https://doi.org/10.1103/PhysRevAccelBeams.26.034601 ; Top-10 Downloaded Paper in 2023 of Phys. Rev. Accel. Beams.
        [2] Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware; Scomparin, L.; Caselle, M.; Santamaria Garcia, A.; Xu, C.; Blomley, E.; Dritschler, T.; Mochihashi, A.; Schuh, M.; Steinmann, J. L.; Bründermann, E.; Kopmann, A.; Becker, J.; Müller, A.-S.; Weber, M.; 2026. Machine Learning: Sci. and Technol. 7, 025056. doi: https://doi.org/10.1088/2632-2153/ae5b20
        [3] Preliminary results on the reinforcement learning-based control of the microbunching instability; Scomparin, L.; Santamaria Garcia, A.; Kopmann, A.; Mueller, A.-S.; Xu, C.; Blomley, E.; Bründermann, E.; Steinmann, J. L.; Becker, J.; Schuh, M.; Schuh, M.; Caselle, M.; Dritschler, T.; Mochibashi, A.; Weber, M.; 2024. Proceedings of 15th Int. Particle Accel. Conf. 1808–1811. doi: https://doi.org/10.18429/JACoW-IPAC2024-TUPS61
        [4] Development of Real-Time Digital Twins for Particle Accelerators, Mohammad Zadeh, M.; Gethmann, J.; Müller, A.-S.; Carne, G. De, 2024. IEEE Energy Conversion Congress and Expo Europe (ECCE 2024), 5p. doi: https://doi.org/10.1109/ECCEEurope62508.2024.10752039
        [5] https://cordis.europa.eu/project/id/101287548
        [6] https://artifact-network.org/initiatives/twinrise/

        Speaker: Erik Bründermann (Karlsruhe Institute of Technology (KIT))
      • 178
        Real-time Gravitational Wave Parameter Estimation

        In this talk I will present a specialised GPU-native nested sampling kernel targeting rapid parameter estimation for gravitational wave inference problems. Building upon a Slice-within-Gibbs (SwiG) structure for rapid mixing, we investigated how far we can push baseline stochastic sampling techniques on modern GPU hardware. I will show that for typical long-duration binary neutron star signals observed by the LIGO and Virgo detectors, we can achieve well calibrated posterior inference on the full uncompressed data of a three detector network in a median of twelve minutes on a single GPU. Utilising heterodyning to compress the data reduces the median wall time to 89 seconds – less than the length of the segment itself – and enables inference with precessing spin, tidal waveforms on GW170817 in around two minutes. This pushes stochastic sampling techniques using full physical waveform calculations, launched from an uninformed prior state, towards real-time gravitational wave parameter estimation.

        Speaker: James Alvey (University of Cambridge)
      • 179
        Pre-localization of Massive Black Hole Binaries in the Millihertz Band

        The space-borne gravitational-wave (GW) detectors will open a new mass and redshift regime, allowing us to observe massive black hole binaries (MBHBs) throughout the Universe. A subset of these systems is expected to produce electromagnetic (EM) counterparts, offering a unique opportunity to follow the continuous evolution of massive black holes through joint GW and EM observations. Realizing this potential, however, requires low-latency, high-throughput data-analysis pipelines that can extract reliable source parameters and sky localizations from space-borne data streams fast enough to trigger EM follow-up. In this work we develop a fast, normalising flow-based inference pipeline designed for early-warning analysis of MBHB signals in a TianQin-like configuration. Our method combines a learned embedding of the detector time series with a neural spline flow (NSF) to perform amortized Bayesian inference, producing posterior samples for the main source parameters in roughly one minute per event. For a representative MBHB whose merger occurs 15 minutes after the end of the analyzed GW observation, the pipeline achieves pre-merger sky localizations of order about 20 square degree , recovers the same number of sky modes as a reference parallel-tempered Markov chain Monte Carlo (PTMCMC) analysis, and yields parameter uncertainties of comparable scale, while still operating within a practically useful pre-merger warning window. These results demonstrate that NSF-based inference can deliver accurate, near-real-time parameter estimation for space-borne MBHB GW signals, and that the resulting early-warning localizations are sufficiently precise to make rapid EM follow-up.

        Speaker: Dr Xue-Ting Zhang (Max Planck Institute for Gravitational Physics (Albert Einstein Institute))
      • 180
        Is Data Compression an Anomaly Detector? Applications to High-Energy Physics Datasets

        The ATLAS experiment at the CERN Large Hadron Collider (LHC) records and processes vast amounts of data from proton-proton collisions. With the High-Luminosity LHC (HL-LHC), the expected increase in data volume by more than an order of magnitude will place unprecedented demands on storage, data throughput, and analysis. In this contribution, we will start with the comparison of two novel ML-based compression methods (using Transformers and Mamba networks) in terms of throughput and compression rate.

        We then move on to investigate the interplay between data compression and anomaly detection. Traditional lossless compression algorithms exploit statistical redundancies in the data and become less efficient when encountering previously unseen patterns. This observation motivates the hypothesis that the inability to compress an event efficiently may indicate that it deviates from the distribution on which the compressor was optimized.

        We investigate whether ML-based compression can be used as an anomaly detection technique for high-energy physics. We employ the Byte-Oriented Autoregressive (BOA) compression model using Mamba networks, which learns the probability distribution of Standard Model background events and assigns likelihood-based compression scores to unseen data. Events that compress poorly are interpreted as candidates for anomalous behaviour. The approach is evaluated using LHC open data.

        Speaker: Prof. Caterina Doglioni
      • 181
        Accelerating Cosmological Inference: A Re-usable Library of JAX-based Machine Learning Emulators for Next-Generation Surveys

        The era of precision cosmology has revealed persistent tensions between independent measurements of fundamental parameters within the concordance model, most notably the Hubble constant ($H_0$), the clustering amplitude ($\sigma_8$), and spatial curvature ($\Omega_K$). The study of these discrepancies between independent datasets, which are theoretically predicted to agree, is known as tension quantification. Upcoming astronomical surveys, such as the Euclid mission and the Vera C. Rubin Observatory, will deliver data of unprecedented precision and volume, providing new insights into these tensions. However, conducting joint-constraint analyses with existing legacy datasets will require marginalising over tens of instrumental nuisance parameters, rendering traditional sampling methods computationally prohibitive.

        We approach this problem by producing a comprehensive, re-usable library of machine learning emulators trained on a massive grid of nested sampling chains across 8 cosmological models and 12 astronomical surveys. Trained using DiRAC GPUs (DP456), this framework is released as part of the pip-installable package $\texttt{unimpeded}$ [2511.04661, 2511.05470] (\url{https://github.com/handley-lab/unimpeded}). Serving as a machine-learning-enhanced analogue to the Planck Legacy Archive (PLA), it enables rapid parameter estimation, cosmological model comparison and tension quantification.

        The emulators are implemented with normalising flows using the JAX-native density estimation package $\texttt{margarine}$ [2205.12841]. By learning the marginal posterior directly from these pre-existing nested sampling chains, the normalising flows act as ultra-fast ``nuisance-free likelihoods'' and informative true priors. This eliminates the need to re-sample legacy instrumental systematics (e.g. calibration, foregrounds) or rely on standard Gaussian approximations. The combination of rigorous sampling and density estimation reproduces the exact posterior distributions of a full nuisance-marginalised run, but many orders of magnitude faster. Furthermore, hyperparameter tuning for these normalising flows has been systematically explored with different combinations of network architecture, learning scheduling and activation functions for optimal performance.

        We believe this work represents a significant step forward in cosmological data analysis. By collapsing multi-survey joint analyses down to just the cosmological subspace, it provides a versatile, efficient and equitable platform to address current observational tensions and advance our understanding of the Universe.

        Speaker: Dily Duan Yi Ong (University of Cambridge)
    • 🧠 Working groups: Building Bridges - Community, Connections and Funding (WG5) 2.404

      2.404

      • 182
        Building Bridges - Community, Connections and Funding (WG5)
    • 15:30
      🖼️ Poster session + ☕ Coffee Foyer

      Foyer

    • 🧠 EuCAIF Fellow Meeting HS1

      HS1

    • 🗣️ Plenaries HS1

      HS1

      Convener: Caterina Doglioni (University of Manchester)
    • 🌟 Highlight talks HS1

      HS1

      Convener: Maurizio Pierini (CERN)
      • 186
        Highlight Talk: SPADE: Split-and-Delay Embeddings for Autoregressive High-Granularity Calorimeter Simulation

        Autoregressive transformers are increasingly applied outside the language domain that motivated them, but the tokenization step transfers poorly. Scientific data are already numerical, often partly discrete, and typically live in high-dimensional spaces where each token carries several features. The standard workaround, compressing feature vectors into a single token via a learned codebook, introduces reconstruction loss and a vocabulary that grows multiplicatively with resolution, inflating the embedding and unembedding layers until training becomes prohibitive.
        We introduce SPADE (SPlit And Delay Embeddings), which embeds each feature of a token independently and staggers the resulting streams along the sequence with progressively increasing delays. The vocabulary then scales additively rather than multiplicatively, while intra-token correlations are recovered by the ordinary causal self-attention mechanism: each feature is predicted at its own sequence position, conditioned on the features already emitted for the same object. No auxiliary decoder or quantization stage is required.
        We demonstrate SPADE on point-cloud calorimeter shower generation in the highly granular ILD electromagnetic calorimeter. SPADE is competitive with the state-of-the-art flow-matching model AllShowers on photon showers and substantially outperforms its VQ-VAE-based predecessor OmniJet-
        . Against a joint-vocabulary baseline at the finest granularity studied, SPADE uses 74× fewer parameters and converges 6.9× faster in GPU hours, while better reproducing observables sensitive to energy–position correlations.
        Because the mechanism assumes only that tokens carry multiple features, discrete or continuous, it offers a route to LLM-style pretraining on high-dimensional sensor data across fundamental physics.

        Speaker: Henning Rose (Uni Hamburg)
      • 187
        Highlight Talk: Low-cost Adaptation of a Geometry Foundation Model for Bayesian Shape Inference

        Inferring a three-dimensional shape from sparse, indirect measurements is a severely ill-posed inverse problem, which arises throughout the physical and biological sciences. Doing so successfully with principled Bayesian uncertainties requires a tractable parametrisation of geometry, and a strong prior encoding domain knowledge of plausible shapes.

        In this talk we show how a pretrained geometry foundation model, in our case one built for engineering optimisation, can be adapted to supply both, and embedded in an end-to-end differentiable simulator. We do so using only O(100) example geometries and without retraining or fine-tuning the foundation model. We instead fit a shape prior in the model's latent space using a covariance-aware variant of probabilistic PCA that maximises the likelihood of generating the decoded geometries rather than their latent codes, while remaining fully differentiable. The resulting low-dimensional, well-conditioned parameter space opens the inverse problem to a range of inference strategies.

        As an example application, we consider asteroid lightcurve inversion: using gradient-based MAP estimation we infer shapes for 3,236 asteroids from Gaia DR3 photometry without the convexity assumption standard to that the field, while truncated sequential neural posterior estimation yields the first Bayesian uncertainties on non-convex asteroid geometries. Requiring only a foundation model with a vector latent space and a handful of domain example geometries, the outlined recipe should transfer cheaply to other geometric inverse problems, and to Bayesian inference with latent-variable generative models more broadly.

        Speaker: Thomas Gessey-Jones (PhysicsX)
      • 188
        Highlight Talk: Steering in Theory Space: Representation Engineering of a Lagrangian Transformer

        Foundation models embed not only their training examples but the space those examples were drawn from. A transformer trained to write Lagrangians symbolically (such as BART-L) thus embeds the space of theories itself. In this work, we investigate the navigation of this learned theory space using activation steering, adapted from interpretability work on LLMs. Using only small contrastive datasets, we build steering vectors that take us from one theory domain to another. This hence reframes theory construction as a navigation problem, where reaching a theory of interest becomes a matter of finding and following directions, rather than an exhaustive scan of a landscape. We demonstrate navigations such as from anomalous to anomaly-free field content, from a base model to its supersymmetric and gauge-extended counterparts, and from charge-violating to charge-conserving Lagrangians. All of this requires no retraining or finetuning at all. The method also serves its original purpose of interpretability. By testing which operators can be recovered as vectors, we can ask whether the learned theory space is sufficient. We find directions corresponding to several Lorentz and gauge representation changes, including one that maps scalars onto fermions in the manner of a supersymmetry generator, as well as failure modes we associate with the limited size of our model and dataset. These results indicate that the model has, to some degree, embedded a navigable space of theories, and point toward a novel approach to theory search which is directed rather than exhaustive.

        Speaker: Yong Sheng Koay (Uppsala University)
    • 10:30
      ☕ Coffee Foyer

      Foyer

    • 🌟 Highlight talks HS1

      HS1

      Convener: Jan M. Pawlowski (Heidelberg University)
      • 189
        Highlight Talk: A Low-Latency Neural Network Trigger for SiPM-Based RICH Detectors

        SiPMs (Silicon Photomultipliers) have recently been studied as candidates for building photon cameras in RICH (Ring Imaging Cherenkov) detectors. SiPM-based cameras improve detection efficiency (up to
        ), spatial resolution (mm), timing (
        \, ps), scalability, and magnetic field immunity. Nevertheless, for single-photon detection, SiPM thermal noise (
        \,kHz/mm
        ) and photon-background sources (scintillation or scattering) pose severe challenges for free-streaming readout systems. In this work, we study the viability of an ML (machine learning) trigger implemented in the CBM (Compressed Baryonic Matter) RICH front-end electronics (edge computing). The ML algorithm aims to filter out fake events caused by the SiPM thermal noise while keeping Cherenkov ring events.

        Speaker: Jesus Pena-Rodriguez (JLU Giessen)
      • 190
        Highlight Talk: Anomaly-detection of gravitational waves through the 4th observing run of the LIGO detectors

        We present an anomaly detection algorithm based on a deep convolutional autoencoder. The algorithm, initially developed with a focus on the Einstein Telescope, was adjusted to the higher noise levels of the LIGO detectors, implementing coherence, trained and validated using O3 data, and tested on publicly available O4 data. We achieve an excellent recovery rate for short, loud signals and compare our results with those of other unmodeled pipelines. We will also discuss the plans for possible deployment during future observing runs.

        Speaker: Gianluca Inguglia (MBI Vienna)
    • 191
      🙏 Closing words HS1

      HS1