Speaker
Description
Many scientific datasets are fundamentally incomplete: only a biased subset of true positives is ever observed, while the remainder stay unlabeled. This Positive-Unlabeled (PU) setting arises whenever detection efficiency is imperfect and covariate-dependent—a structure shared across many fields in fundamental science. In hadron collider experiments, hardware triggers record only a biased fraction of true signal events, with efficiency depending on transverse momentum and detector geometry. In time-domain astrophysics, spectroscopic confirmation of transients is strongly magnitude- and color-biased, leaving the vast majority of candidates unlabeled. Similarly, in gravitational-wave astronomy, confirmed mergers represent only the loud, well-localized tail of the true population, with detectability driven by
extrinsic parameters independent of intrinsic source properties.
We introduce a variational framework for transductive PU classification under observational bias, where labeled positives are drawn non-uniformly from the true positive population according to an unknown,
covariate-dependent propensity score. The model jointly infers two latent quantities per unlabeled point: the probability of being a hidden positive (classifier score) and the probability of having been observed
given positivity (propensity score). Both are parameterized as neural network outputs and trained within a unified ELBO that decomposes the marginal likelihood of the observed labeling pattern under the PU
generative model, yielding calibrated credible intervals over both quantities rather than point estimates. When relational graph structure correlates with ground truth or the sampling mechanism, encoder and/or
decoder components are replaced by GNNs, allowing uncertainty to propagate through the topology.
The method is validated on synthetic benchmarks with known propensity functions across varying bias severities, graph topologies, and covariate regimes, demonstrating joint recovery of well-calibrated
posterior probabilities for unlabeled points alongside accurate propensity score estimates. We then apply the model to mammal–virus association networks, where the observed host–pathogen matrix is severely incomplete due to uneven surveillance across species, viruses, and geographical regions. Our model recovers putative missing mammalian hosts across widespread viral families, providing bias-corrected
association probabilities with principled uncertainty estimates for ecological risk assessment.
Across ecology and fundamental physics alike, the underlying challenge is identical. By exploiting the separation between covariates driving physical class membership and those driving detection propensity,
this framework provides a unified and probabilistic solution for recovering the full positive population from a biased subset of known positive data.