Speaker
Description
Latent representations are an important theme in modern machine learning. Networks trained with a notion of locality encode task-specific similarity as closeness in the latent space. We analyze this latent information for a variational autoencoder with a classifier head using tools from differential geometry, specifically information geometry. Here, we explore the learned latent space using concepts such as geodesic distances and curvature and discuss their significance for the induced decoder and classifier geometries. To link the structure of the latent space to the physics features of the data, we employ the metric tensor induced by the likelihood of the classifier. Furthermore, we transfer the concept of nonmetricity scalars to information geometry and find that they constitute new, coordinate-invariant measures of class separation. We then apply our new methodology to LHC data to understand the physics behind binary quark-gluon classification and three-fold fat jet tagging. Here, information geometry tells us which features a decision is dominantly based on, increasing the interpretability of the network output.