Uncertainty under Composition: Joint Confidence Regions and an Expiring Calibration Certificate for Composed AI-Trustworthiness Scores
Preprint (working draft v0.1) — deposited to date-stamp priority. Composite trustworthiness scores for machine-learning systems are increasingly used to compare and select models. The rank-graduation family developed at the University of Pavia measures accuracy (RGA), robustness (RGR) and explainability (RGE) on a common rank-based footing, and recent work by the same group integrates them into a single compliance score. Inference exists for the individual metrics: a significance test and variance for RGA, and two-model comparison tests within one metric at a time. The composed score, however, is reported as a bare point value. This article supplies the missing layer. Because all component metrics are estimated on the same test sample, their sampling errors are correlated, and any uncertainty statement for the composition must carry the cross-metric covariance. We construct joint confidence regions for the component vector, and delta-method and paired-bootstrap intervals for the composed score, cross-checked by a leave-one-out jackknife, and we extend the existing per-metric two-model test to the composition. A worked example on the Statlog German credit data, run against a pinned commit of the group's own open-source package, finds cross-metric correlations up to +0.60, a cross-covariance share of up to 11% of the composed-score variance, and a two-model composed-score difference whose confidence interval comfortably includes zero even though the point values differ in the third decimal. Because the components sit near their upper bound, the symmetric delta interval misses the percentile bounds by about 0.002; a logit-scale delta interval and BCa both close that gap, and all three are reported. We then bind the resulting uncertainty budget into a signed, dated, machine-readable calibration certificate with an explicit recalibrate-by date, and instantiate its propagation across a two-link delegation chain: with links evaluated on a shared sample the cross-link correlation is measured at +0.72, and treating the links as independent understates the chain interval by 23.8% of its width. The certificate, not the point score, is the object that has to travel. Prior work is conceded by name. The rank-graduation metrics, their per-metric inference and the integrated compliance score are the work of the Pavia group (Giudici, Raffinetti, Babaei and colleagues); the substrate package used here is authored by Vasily Kolesnikov and is MIT-licensed, cloned unmodified at the pinned commit. This article claims only the joint/composed uncertainty layer and the certificate object. Reproducibility. Every number, the figure and both example certificates regenerate from one self-bootstrapping script in the attached reproduction package (seed 20260723, B = 2000).