Zuchen Gao, Lingbo Tong, Zhiyong Zhang
2026.2.17Journal of Behavioral Data Science
tlooto Summary
This survey presents a structured review of methods for detecting and evaluating bias in LLMs, and introduces the conceptual foundations, including representational versus allocational harms and taxonomies of bias.
Abstract
Large Language Models (LLMs) are increasingly deployed in sensitive real-world contexts, yet concerns remain about their biases and the harms they can cause. Existing surveys mostly discuss sources of bias and mitigation techniques, but give less systematic attention to how bias in LLMs should be detected, measured, and reported. This survey addresses that gap. We present a structured review of methods for detecting and evaluating bias in LLMs. We first introduce the conceptual foundations, including representational versus allocational harms and taxonomies of bias. We then discuss how to design evaluations in practice: specifying measurement targets, choosing datasets and metrics, and reasoning about validity and reliability. Building on this, we review intrinsic methods that probe representations and likelihoods, and extrinsic methods that assess bias in classification, question answering, open-ended generation, and dialogue. We further highlight recent advances in counterfactual and certification-based evaluation, which aim to provide stronger guarantees on fairness metrics. Beyond English-centric settings, we survey cross-lingual and application-specific evaluations, intersectional bias analysis, and meta-level issues such as evaluator reliability, metric robustness, reproducibility, and governance. The review concludes by synthesizing best practices and offering a practitioner-oriented checklist, providing both a conceptual map and a practical toolkit for evaluating bias in LLMs.
Citation format
GAO, Zuchen; TONG, Lingbo; ZHANG, Zhiyong. Detecting and evaluating bias in large language models: Concepts, methods, and challenges. Journal of Behavioral Data Science, 2026.