Speakers
Description
Pretrained foundation models are emerging in neutrino telescopes, but what their internal representations actually contain, and whether the tasks built on top of them put it to use, is largely unknown. We study PolarBERT, a transformer pretrained on IceCube data and fine-tuned for direction reconstruction. Sparse autoencoders uncover a validated atlas of physics across the frozen backbone, including event quality, auxiliary activity, brightness, and detector depth at the layers where these concepts are most clearly represented. Causal interventions show the direction head barely draws on this atlas, relying instead on a single axis with no simple physical meaning.
We then show that a different task can make better use of the same representation. A second head, trained on the same frozen summary to predict the model’s angular reconstruction error, turns out to depend causally on precisely the quality and brightness concepts the direction head ignores.
The result is a interpretable and trustworthy per-event uncertainty estimator that reaches $3.2^\circ$ median resolution at 20% selection efficiency, against about $20^\circ$ for the best detector observable, and its reliance on that physics is established by intervention rather than assumed.