Comment by reexpressionist
> "This means that a human does not necessarily need to be in the loop for agentic decisions anymore"That's only true in a practical sense if you can actually rely on the probabilities estimated by the model. That's a non-trivial problem for multiple reasons, among them: 1. What is the particular quantity you seek to estimate (marginal, approximately conditional, etc.)? 2. What is the reference class for that quantity? 3. What method are you going to use to estimate that quantity? 4. What is the error in your method to estimate that quantity (e.g., as via accounting for the effective sample size)?
A further practical challenge is that the output logits of neural networks are in effect a highly lossy compression of the epistemic (reducible) uncertainty. Even if your estimates are well-calibrated (for some definition of well-calibrated) on in-distribution data using the output logits, those estimates can be grossly uncalibrated in the presence of covariate shifts, and the logits themselves are not reliable signals of such shifts, nor of being out-of-distribution. Informally, the output logits themselves do not encode a good sense of what they do[n't] know.
Additionally, ideally the probability estimates are interpretable in the sense that there is some instance-wise connection to the training/calibration data. If the estimates are being used for decision-making, you need to be able to post-hoc audit the estimates to be able to modify the data for future decision-making, if needed.
Growing evidence in ML/NLP/Stats from the last few years is that with neural networks, as a starting point for constructing reliable estimates of the predictive uncertainty, we need to control for metric-learner signals over the support/training set (e.g., the L^2 distance to the nearest training instance and depth-matches into training). Once you have that, then you can choose your desired quantity of interest (e.g., class- and prediction-conditional accuracy at least some given value). Concretely, here's a tutorial (along with Apache-2.0 code) that steps through a simple, illustrative example: https://reexpressai.github.io/reexpress_sdm/tutorials/gettin...
More context is in the link, but at a high-level from an engineering perspective, just as dense vector matching is used by RAG for information retrieval, we can also use dense vector matching in this way to estimate the predictive uncertainty, getting around the limitations of the output logits.