24–28 Aug 2026
Kirchhoff Institute for Physics (KIP)
Europe/Berlin timezone

A Machine-Learning Analysis of the Higgs Boson in the H → ZZ∗ →4ℓ Channel

27 Aug 2026, 12:20
8m
3.404

3.404

Patterns & Anomalies 🔀 Patterns & Anomalies

Speaker

Mane Papoyan (American University of Armenia (AUA))

Description

The H → ZZ → 4ℓ channel remains one of the cleanest probes of Higgs boson properties at the LHC, owing to its fully reconstructable final state and well-understood background composition. This work addresses the problem of identifying an optimal machine-learning classifier for extracting the H → ZZ → 4ℓ signal from the ATLAS Open Data 2025 release (√s = 13 TeV, 36.6 fb⁻¹), comparing seven algorithms, XGBoost, LightGBM, Random Forest, a multilayer perceptron, Logistic Regression, QDA, and Gaussian Naive Bayes, under a single, controlled experimental protocol. The specialty of this analysis is twofold. First, every classifier is trained exclusively on Monte Carlo simulation and then applied, without retraining or recalibration, to the real ATLAS dataset, so that the reported significance reflects genuine generalization from simulation to data rather than performance on held-out simulated events alone. Second, background estimation itself is treated as a robustness test: each classifier's signal significance on real data is computed twice, once under a Monte Carlo background prediction and once under a data-driven sideband extrapolation, and a classifier is judged reliable only if the two estimates agree within uncertainties, since the two methods rest on largely independent assumptions. Four-lepton events are reconstructed from the ATLAS Open Data samples and reduced to a set of 31 physics-motivated kinematic and angular features, subsequently pruned to 20 via separation-power ranking and correlation filtering. All seven classifiers are trained and cross-validated on Monte Carlo simulation under a 5-fold stratified scheme, then applied at a fixed classifier-score threshold to the full real-data sample, with the signal region and sideband control regions defined directly in the four-lepton invariant mass spectrum. Signal significance is computed independently under both background estimators for each classifier, yielding a direct test of consistency rather than a single, isolated figure of merit. Under this framework, LightGBM emerges as the best-performing and most consistent classifier overall, achieving a signal significance of Z = 4.66 ± 1.22σ (p = 1.55 × 10⁻⁶) under the Monte Carlo background estimate and Z = 5.08 ± 1.50σ (p = 1.90 × 10⁻⁷) under the sideband estimate. XGBoost follows closely, with Z = 4.24 ± 1.10σ (p = 1.14 × 10⁻⁵) and Z = 4.70 ± 1.35σ (p = 1.31 × 10⁻⁶) under the same two methods, respectively. Both classifiers reach evidence-level significance under both background estimators, with their results agreeing within uncertainties across two independent methods, confirming that the models generalize from simulation to real collision data and that the resulting evidence is not an artifact of a particular background model. This consistency establishes gradient-boosted decision trees as a robust, reproducible means of recovering Higgs boson evidence from public ATLAS Open Data using standard machine-learning methods alone.

Authors

Mane Papoyan (American University of Armenia (AUA)) Dr Houry Keoshkerian (American University of Armenia (AUA))

Presentation materials