arXiv:2608.19885v2 Announce Type: replace-cross
Abstract: After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single quantity — estimable as the difference of two discriminators’ held-out cross-entropies, D_CJS = CE(Z|X) – CE(Z|X,Y) — is exactly the part of a dataset difference that covariate shift cannot explain. We propose the Conditional Jensen-Shannon Discrepancy (CJSD): with a task indicator Z, the chain rule I(Z;X,Y) = I(Z;X) + I(Z;Y|X) splits total task discrepancy exactly into a covariate axis and a functional axis, both estimable from two ordinary classifiers, with no task-specific predictors, generative models, or bootstrap surrogates. We prove a covariate-null property (the functional axis is exactly zero under pure covariate shift, however severe), a drift-mass law (D_CJS/ln2 equals the mass of the disagreement region for deterministic labels), a one-sided misspecification-control inequality (each direction of estimation error is bounded, unconditionally, by the excess risk of a single discriminator), and a fixed-measure metrization via an identifiability lemma. Empirically, on a ten-measure battery over 202 dataset pairs (synthetic, Electricity, Covertype), only the two conditional-information estimators — CJSD and a kNN plug-in for the same estimand — separate concept from covariate shift with AUC 1.0; the case for CJSD is the estimator: under controlled dimensionality scaling the kNN plug-in fails from d=64 while the discriminator route holds to d=256 with a swappable classifier, and it alone yields paired confidence intervals and sequential extensions from the same learned object. The same estimator audits the conditional fidelity of synthetic-data generators that marginal and joint QA metrics pass, detects annotation-guideline changes invisible to input-space monitors, and supports null-calibrated fairness audits.
THE AI TODAY 