Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Mark Steyvers

Bayesian Inference for Correlated Human Experts and Classifiers

Jun 05, 2025

Markelle Kelly, Alex Boyd, Sam Showalter, Mark Steyvers, Padhraic Smyth

Abstract:Applications of machine learning often involve making predictions based on both model outputs and the opinions of human experts. In this context, we investigate the problem of querying experts for class label predictions, using as few human queries as possible, and leveraging the class probability estimates of pre-trained classifiers. We develop a general Bayesian framework for this problem, modeling expert correlation via a joint latent representation, enabling simulation-based inference about the utility of additional expert queries, as well as inference of posterior distributions over unobserved expert labels. We apply our approach to two real-world medical classification problems, as well as to CIFAR-10H and ImageNet-16H, demonstrating substantial reductions relative to baselines in the cost of querying human experts while maintaining high prediction accuracy.

* accepted to ICML 2025

Via

Access Paper or Ask Questions

Metacognition and Uncertainty Communication in Humans and Large Language Models

Apr 18, 2025

Mark Steyvers, Megan A. K. Peters

Abstract:Metacognition, the capacity to monitor and evaluate one's own knowledge and performance, is foundational to human decision-making, learning, and communication. As large language models (LLMs) become increasingly embedded in high-stakes decision contexts, it is critical to assess whether, how, and to what extent they exhibit metacognitive abilities. Here, we provide an overview of current knowledge of LLMs' metacognitive capacities, how they might be studied, and how they relate to our knowledge of metacognition in humans. We show that while humans and LLMs can sometimes appear quite aligned in their metacognitive capacities and behaviors, it is clear many differences remain. Attending to these differences is crucial not only for enhancing human-AI collaboration, but also for promoting the development of more capable and trustworthy artificial systems. Finally, we discuss how endowing future LLMs with more sensitive and more calibrated metacognition may also help them develop new capacities such as more efficient learning, self-direction, and curiosity.

Via

Access Paper or Ask Questions

Hybrid Forecasting of Geopolitical Events

Dec 14, 2024

Daniel M. Benjamin, Fred Morstatter, Ali E. Abbas, Andres Abeliuk, Pavel Atanasov, Stephen Bennett, Andreas Beger, Saurabh Birari, David V. Budescu, Michele Catasta(+19 more)

Figure 1 for Hybrid Forecasting of Geopolitical Events

Figure 2 for Hybrid Forecasting of Geopolitical Events

Figure 3 for Hybrid Forecasting of Geopolitical Events

Figure 4 for Hybrid Forecasting of Geopolitical Events

Abstract:Sound decision-making relies on accurate prediction for tangible outcomes ranging from military conflict to disease outbreaks. To improve crowdsourced forecasting accuracy, we developed SAGE, a hybrid forecasting system that combines human and machine generated forecasts. The system provides a platform where users can interact with machine models and thus anchor their judgments on an objective benchmark. The system also aggregates human and machine forecasts weighting both for propinquity and based on assessed skill while adjusting for overconfidence. We present results from the Hybrid Forecasting Competition (HFC) - larger than comparable forecasting tournaments - including 1085 users forecasting 398 real-world forecasting problems over eight months. Our main result is that the hybrid system generated more accurate forecasts compared to a human-only baseline which had no machine generated predictions. We found that skilled forecasters who had access to machine-generated forecasts outperformed those who only viewed historical data. We also demonstrated the inclusion of machine-generated forecasts in our aggregation algorithms improved performance, both in terms of accuracy and scalability. This suggests that hybrid forecasting systems, which potentially require fewer human resources, can be a viable approach for maintaining a competitive level of accuracy over a larger number of forecasting questions.

* AI Magazine, Volume 44, Issue 1, Pages 112-128, Spring 2023
* 20 pages, 6 figures, 4 tables

Via

Access Paper or Ask Questions

Perceptions of Linguistic Uncertainty by Language Models and Humans

Jul 22, 2024

Catarina G Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, Padhraic Smyth

Abstract:Uncertainty expressions such as ``probably'' or ``highly unlikely'' are pervasive in human language. While prior work has established that there is population-level agreement in terms of how humans interpret these expressions, there has been little inquiry into the abilities of language models to interpret such expressions. In this paper, we investigate how language models map linguistic expressions of uncertainty to numerical responses. Our approach assesses whether language models can employ theory of mind in this setting: understanding the uncertainty of another agent about a particular statement, independently of the model's own certainty about that statement. We evaluate both humans and 10 popular language models on a task created to assess these abilities. Unexpectedly, we find that 8 out of 10 models are able to map uncertainty expressions to probabilistic responses in a human-like manner. However, we observe systematically different behavior depending on whether a statement is actually true or false. This sensitivity indicates that language models are substantially more susceptible to bias based on their prior knowledge (as compared to humans). These findings raise important questions and have broad implications for human-AI alignment and AI-AI communication.

* In submission

Via

Access Paper or Ask Questions

The Calibration Gap between Model and Human Confidence in Large Language Models

Jan 24, 2024

Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas Mayer, Padhraic Smyth

Abstract:For large language models (LLMs) to be trusted by humans they need to be well-calibrated in the sense that they can accurately assess and communicate how likely it is that their predictions are correct. Recent work has focused on the quality of internal LLM confidence assessments, but the question remains of how well LLMs can communicate this internal model confidence to human users. This paper explores the disparity between external human confidence in an LLM's responses and the internal confidence of the model. Through experiments involving multiple-choice questions, we systematically examine human users' ability to discern the reliability of LLM outputs. Our study focuses on two key areas: (1) assessing users' perception of true LLM confidence and (2) investigating the impact of tailored explanations on this perception. The research highlights that default explanations from LLMs often lead to user overestimation of both the model's confidence and its' accuracy. By modifying the explanations to more accurately reflect the LLM's internal confidence, we observe a significant shift in user perception, aligning it more closely with the model's actual confidence levels. This adjustment in explanatory approach demonstrates potential for enhancing user trust and accuracy in assessing LLM outputs. The findings underscore the importance of transparent communication of confidence levels in LLMs, particularly in high-stakes applications where understanding the reliability of AI-generated information is essential.

* 27 pages, 10 figures

Via

Access Paper or Ask Questions

Bayesian Online Learning for Consensus Prediction

Dec 12, 2023

Sam Showalter, Alex Boyd, Padhraic Smyth, Mark Steyvers

Figure 1 for Bayesian Online Learning for Consensus Prediction

Figure 2 for Bayesian Online Learning for Consensus Prediction

Figure 3 for Bayesian Online Learning for Consensus Prediction

Figure 4 for Bayesian Online Learning for Consensus Prediction

Abstract:Given a pre-trained classifier and multiple human experts, we investigate the task of online classification where model predictions are provided for free but querying humans incurs a cost. In this practical but under-explored setting, oracle ground truth is not available. Instead, the prediction target is defined as the consensus vote of all experts. Given that querying full consensus can be costly, we propose a general framework for online Bayesian consensus estimation, leveraging properties of the multivariate hypergeometric distribution. Based on this framework, we propose a family of methods that dynamically estimate expert consensus from partial feedback by producing a posterior over expert and model beliefs. Analyzing this posterior induces an interpretable trade-off between querying cost and classification performance. We demonstrate the efficacy of our framework against a variety of baselines on CIFAR-10H and ImageNet-16H, two large-scale crowdsourced datasets.

Via

Access Paper or Ask Questions

Capturing Humans' Mental Models of AI: An Item Response Theory Approach

May 15, 2023

Markelle Kelly, Aakriti Kumar, Padhraic Smyth, Mark Steyvers

Abstract:Improving our understanding of how humans perceive AI teammates is an important foundation for our general understanding of human-AI teams. Extending relevant work from cognitive science, we propose a framework based on item response theory for modeling these perceptions. We apply this framework to real-world experiments, in which each participant works alongside another person or an AI agent in a question-answering setting, repeatedly assessing their teammate's performance. Using this experimental data, we demonstrate the use of our framework for testing research questions about people's perceptions of both AI agents and other people. We contrast mental models of AI teammates with those of human teammates as we characterize the dimensionality of these mental models, their development over time, and the influence of the participants' own self-perception. Our results indicate that people expect AI agents' performance to be significantly better on average than the performance of other humans, with less variation across different types of problems. We conclude with a discussion of the implications of these findings for human-AI interaction.

* FAccT 2023

Via

Access Paper or Ask Questions

Combining Human Predictions with Model Probabilities via Confusion Matrices and Calibration

Oct 01, 2021

Gavin Kerrigan, Padhraic Smyth, Mark Steyvers

Figure 1 for Combining Human Predictions with Model Probabilities via Confusion Matrices and Calibration

Figure 2 for Combining Human Predictions with Model Probabilities via Confusion Matrices and Calibration

Figure 3 for Combining Human Predictions with Model Probabilities via Confusion Matrices and Calibration

Figure 4 for Combining Human Predictions with Model Probabilities via Confusion Matrices and Calibration

Abstract:An increasingly common use case for machine learning models is augmenting the abilities of human decision makers. For classification tasks where neither the human or model are perfectly accurate, a key step in obtaining high performance is combining their individual predictions in a manner that leverages their relative strengths. In this work, we develop a set of algorithms that combine the probabilistic output of a model with the class-level output of a human. We show theoretically that the accuracy of our combination model is driven not only by the individual human and model accuracies, but also by the model's confidence. Empirical results on image classification with CIFAR-10 and a subset of ImageNet demonstrate that such human-model combinations consistently have higher accuracies than the model or human alone, and that the parameters of the combination method can be estimated effectively with as few as ten labeled datapoints.

* NeurIPS 2021

Via

Access Paper or Ask Questions

Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

Oct 19, 2020

Disi Ji, Padhraic Smyth, Mark Steyvers

Figure 1 for Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

Figure 2 for Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

Figure 3 for Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

Figure 4 for Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

Abstract:We investigate the problem of reliably assessing group fairness when labeled examples are few but unlabeled examples are plentiful. We propose a general Bayesian framework that can augment labeled data with unlabeled data to produce more accurate and lower-variance estimates compared to methods based on labeled data alone. Our approach estimates calibrated scores for unlabeled examples in each group using a hierarchical latent variable model conditioned on labeled examples. This in turn allows for inference of posterior distributions with associated notions of uncertainty for a variety of group fairness metrics. We demonstrate that our approach leads to significant and consistent reductions in estimation error across multiple well-known fairness datasets, sensitive attributes, and predictive models. The results show the benefits of using both unlabeled data and Bayesian inference in terms of assessing whether a prediction model is fair or not.

* 27 pages

Via

Access Paper or Ask Questions

Active Bayesian Assessment for Black-Box Classifiers

Feb 16, 2020

Disi Ji, Robert L. Logan IV, Padhraic Smyth, Mark Steyvers

Figure 1 for Active Bayesian Assessment for Black-Box Classifiers

Figure 2 for Active Bayesian Assessment for Black-Box Classifiers

Figure 3 for Active Bayesian Assessment for Black-Box Classifiers

Figure 4 for Active Bayesian Assessment for Black-Box Classifiers

Abstract:Recent advances in machine learning have led to increased deployment of black-box classifiers across a wide variety of applications. In many such situations there is a crucial need to assess the performance of these pre-trained models, for instance to ensure sufficient predictive accuracy, or that class probabilities are well-calibrated. Furthermore, since labeled data may be scarce or costly to collect, it is desirable for such assessment be performed in an efficient manner. In this paper, we introduce a Bayesian approach for model assessment that satisfies these desiderata. We develop inference strategies to quantify uncertainty for common assessment metrics (accuracy, misclassification cost, expected calibration error), and propose a framework for active assessment using this uncertainty to guide efficient selection of instances for labeling. We illustrate the benefits of our approach in experiments assessing the performance of modern neural classifiers (e.g., ResNet and BERT) on several standard image and text classification datasets.

Via

Access Paper or Ask Questions