arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00222v1 [cs.CL] 31 Aug 2026

LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts

Daniela Occhipinti Affiliation: Fondazione Bruno Kessler, Via Sommarive 18, Povo, Trento, Italy Email: docchipinti@fbk.eu    Andrea Piergentili Affiliation: Almawave Labs, Via Di Casal Boccone, 188/190, Rome Email: a.piergentili@almawavelabs.it    Marco Guerini Affiliation: Fondazione Bruno Kessler, Via Sommarive 18, Povo, Trento, Italy Email: guerini@fbk.eu
Abstract

Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator’s demographic profile to align its judgments with the corresponding group’s. We test whether this alignment emerges distributionally, comparing the predicted label distributions of 23 open-weight LLMs on three subjective tasks against those of real annotator groups, under three conditions: no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education. Three findings emerge. First, a judge prompted with no demographics is not perspective-neutral: models best reproduce the judgments of White, college-educated annotators. Second, demographic conditioning is asymmetric: it moves the judge toward majority groups and away from minority groups, most strongly on offensiveness, where intersectional profiles amplify the harm. Third, by comparing base and instruct models we identify instruction-tuning as a possible source of the asymmetry. Demographic conditioning should therefore be used with caution to estimate group judgments: conditioning moves predictions away from the reference distributions of the minority groups the method is often invoked to serve.

††footnotetext: ∗ These authors contributed equally.

Warning: This work contains unobfuscated examples that some readers may find offensive

1 Introduction

LLMs are increasingly employed as automated judges: they assess text in place of human annotators to evaluate NLP systems and produce annotated data at scale (Li et al., 2025; Gu et al., 2026; Amin et al., 2026; Xu et al., 2026). Since human judgments vary partly with annotators’ sociodemographic background (Pei and Jurgens, 2023; Diaz et al., 2018), sociodemographic prompting has emerged as a way to simulate groups of annotators: conditioning the judge on a demographic profile, a form of persona prompting (Chen et al., 2024b; Tseng et al., 2024), should steer it to answer as an annotator from that group would. Yet such simulation is only as valid as its agreement with human judgments, which varies sharply across tasks (Bavaresco et al., 2025). This is especially delicate for subjective judgments, where no single ground truth exists: annotators disagree, and this disagreement should be treated as signal rather than noise (Pavlick and Kwiatkowski, 2019; Plank, 2022). This is the concern of pluralistic alignment: a model that reflects average human preferences may still fail to reflect the judgments of specific groups (Sorensen et al., 2024; Feng et al., 2024). Thus, for subjective tasks, what matters is not only how accurate an LLM judge is, but whose judgments it reproduces.

Recent analyses on sociodemographic prompting report that effects are mixed, prompt-sensitive, and sometimes negative (Beck et al., 2024; Sun et al., 2025; Gupta et al., 2024), and demographic labels predict a rater’s judgments worse than the rater’s own past annotations or values (Orlikowski et al., 2025; Sorensen et al., 2025). These analyses typically measure how well a judge simulates labels and assess conditioning by a single (averaged) effect per model. What remains untested is whether conditioning moves the model’s label distribution toward that of the target annotator group, rather than toward a stereotype.

Refer to caption
Figure 1: The evaluation pipeline. A model judges a subjective property on a five-point scale, either unconditioned or with a profile prepended (Your Profile). Its log-probabilities over the five options are renormalized into a predicted distribution pp and scored against the matching group’s label distribution qq.

We address this gap distributionally, adopting a perspectivist view of sociodemographic prompting (Cabitza et al., 2023; Basile et al., 2021; Frenda et al., 2024): the object of evaluation is the distribution of human judgments for each group rather than an aggregated gold label (Figure 1). Following Santurkar et al. (2023), we ask whose judgments LLMs reflect through three research questions. RQ1: with no demographic information, do LLM judges tend to a default demographic profile? RQ2: does conditioning on (intersectional) demographic profiles move the judge toward those groups’ judgments? RQ3: is demographic conditioning’s effect equitable across groups?

We evaluate 23 open-weight LLMs as judges on three subjective tasks (politeness, intimacy, and offensiveness), comparing unconditioned judges against judges conditioned on single-attribute and intersectional profiles, each scored against the matching human group’s label distribution. Unconditioned judges tend to align with White, college-educated annotators. Averaged across groups, demographic conditioning apparently has little effect on instruction-tuned judges. However, this hides opposite effects across groups, moving the judge closer to some and further from others, especially those disadvantaged by known LLM biases (Navigli et al., 2023; Sap et al., 2022; Santy et al., 2023), here grouped under the term minority groups. Larger models are not immune, and the asymmetry is specific to instruction-tuned judges.

Our contributions are: (i) a framework for evaluating demographic conditioning of LLM judges against annotator groups’ label distributions, rather than against their mean or majority label; (ii) a per-group decomposition of that effect, revealing biases in both unconditioned and conditioned models; (iii) a matched base-vs-instruct comparison identifying instruction-tuning as a source of bias in demographic conditioning.

2 Related Work

LLM-as-a-Judge.

LLMs are increasingly used to evaluate text, from scoring open-ended generations to ranking responses (Li et al., 2025; Gu et al., 2026; Zheng et al., 2023; Wang et al., 2024c). Judge–human agreement varies sharply across tasks (Bavaresco et al., 2025), including under persona conditioning (Dong et al., 2024), and verdicts are sensitive to surface factors, such as response order (Wang et al., 2024a) and human-like response biases (Basile et al., 2021; Tjuatja et al., 2024; Chen et al., 2024a). Crucially, prior work validates a judge against a single target, such as a mean score or majority label (Gordon et al., 2022; Mostafazadeh Davani et al., 2022), which does not exist for subjective tasks: annotator groups disagree systematically, so a judge matching the aggregate may diverge from every group. We therefore evaluate alignment with each group’s label distribution, assessing the judge as an estimator of group judgments rather than a rater of quality.

Persona-Based Generation.

LLMs can be prompted to role-play specific people or groups, and some models are fine-tuned to better capture individual preferences or personas (Horton et al., 2023; Shao et al., 2023; Occhipinti et al., 2024). Assigning a persona (a role, identity, or demographic profile) is widely used to steer model behavior, assuming the model adopts the assigned perspective faithfully (Argyle et al., 2022; Chen et al., 2024b; Tseng et al., 2024). This assumption does not always hold: rather than reproducing a group’s genuine perspective, a persona may instead activate the stereotypes the model associates with that group, degrading performance for some identities (Gupta et al., 2024; Agnew et al., 2024; Dong et al., 2024) or increasing toxic output (Deshpande et al., 2023). Conditioning a judge on demographics could therefore move its predictions toward a group’s real judgments or move them away. We investigate the effect of conditioning on alignment with real label distributions.

Demographic Simulation with LLMs.

Prior work uses LLMs to simulate human populations (Argyle et al., 2022; Mou et al., 2026), showing that unconditioned models can align more closely with some groups than others (Santurkar et al., 2023; Durmus et al., 2023). Such bias is also visible on broader demographic imbalances in NLP datasets and models (Santy et al., 2023; Alipour et al., 2025). Demographic conditioning has therefore been studied as a way to recover group-specific judgments, but its effectiveness is fragile and sensitive to the evaluation setup (Simmons and Savinov, 2024; Adilazuarda et al., 2024; Gao et al., 2025), motivating calls for principled evaluation (Luz de Araujo et al., 2025; Neumann et al., 2026). Recent work has compared demographic prompting with richer ways of modeling annotators, including fine-tuning on past judgments (Orlikowski et al., 2025) and conditioning on elicited values (Sorensen et al., 2025). These studies reduce in-group variability to a single aggregate label. The distributional alternative has been developed on opinion surveys (Meister et al., 2025; Lutz et al., 2025): on annotation, where the reference is disagreement about the same text, judges are still scored against a group’s mean (Sun et al., 2025; Schäfer et al., 2025) or against individual annotations (Hu and Collier, 2024). We instead ask how well models reproduce each demographic group’s label distribution, rather than its mean, and whether assigning a demographic profile moves the judge closer to or further from the judgments of the group named in the prompt.

3 Methodology

This section describes the data used (§3.1), the evaluation metrics (§3.2), the generation configurations and analysis (§3.3), and models (§3.4).

3.1 Data

Our study requires parallel subjective judgments: multiple ratings of the same text, each linked to the demographic profile of the annotator who produced it. We therefore leverage the DeMo dataset (Orlikowski et al., 2025) which combines five annotator-level corpora and maps self-reported demographics onto four dimensions: gender, age, race, and education. Since each annotation is aligned with the annotator’s profile, we can construct the gold label distribution of judgments for any text and demographic group.

We retain three of the five tasks in DeMo, all rated on a five-point ordinal scale:11 1 We exclude safety (DICES-350; Aroyo et al., 2023), which uses a three-way categorical scale and generational cohorts rather than age brackets, and sentiment (Diaz et al., 2018), which recruited only annotators over 5050.

  • •

    Intimacy: rating how intimate a Twitter post is, from not intimate at all to very intimate (MINT; Pei et al., 2023);

  • •

    Offensiveness: rating how offensive a Reddit comment is, from not offensive at all to very offensive (POPQUORN; Pei and Jurgens, 2023);

  • •

    Politeness: rating how polite a workplace email is, from not polite at all to very polite (POPQUORN).

Table 1 shows an example of an offensiveness entry and the corresponding judgment distribution of two annotator groups (men and women).

Text: “thats fucking hilarious, but also sad at the same time because I know it won’t change anything. Here’s to hoping these laws get struck down, and old pseudo-religious trying to win political brownie points men stop trying to tell women what they can and can’t do with their bodies”
A B C D E ss
Humans
Men (n=4n{=}4) 00 22 22 00 00
Women (n=4n{=}4) 22 00 11 00 11
Llama 70B
Unconditioned .00.00 .07.07 .40.40 .52.52 .01.01 —
Man profile .12.12 .46.46 .28.28 .13.13 .01.01 0.910.91
Woman profile .84.84 .13.13 .02.02 .01.01 .00.00 0.670.67
Table 1: Offensiveness example. A–E are 5-point ratings from not offensive at all to very offensive; ss measures distance between rating distributions (see 3.2). Human rows report annotator counts by gender (n=4n{=}4 each). Llama 70B rows are predicted probabilities, with ss computed for man/woman profiles against the corresponding human distribution.

For our analysis, we remove categories too sparsely represented to support a stable per-group label distribution. The complete filtering procedure is described in Appendix A. Statistics about the resulting dataset are reported in Table 2, whereas the demographic categories retained after filtering are reported in Table 3.

Task Genre Texts Judgments Annotators
Intimacy tweets 1,992 11,355 237
Offensiveness comments 1,500 12,489 251
Politeness emails 3,718 23,896 483
Total 7,210 43,402 971
Table 2: Statistics of the 3 tasks retained from the DeMo dataset, after filtering (see § 3.1). The average number of judgements per entry is 6.02.
Dimension Values
Gender (2) Man, Woman
Age (8)
18–24, 25–29, 30–34, 35–39,
40–44, 45–49, 50–59, 60–69
Race (3)
Asian, Black, White
Education (3)
High school or below,
College degree, Graduate degree
Table 3: Demographic dimensions and values used in our experiments. Their cross-product (2×8×3×32\times 8\times 3\times 3) defines the space of demographics we consider.

3.2 Evaluation Metrics

Distributions as targets.

We evaluate a judge against the full distribution of human labels, at the level of demographic groups. Let xix_{i} be a text and gg a demographic profile, i.e. a combination of values for one or two of the four dimensions. The human reference q(i,g)q^{(i,g)} is the distribution over the K=5K{=}5 scale points of the labels that annotators matching gg assign to xix_{i}. Each generation configuration yields a model distribution p(i,g)p^{(i,g)} over the same points, from the log-probabilities the model assigns to the five answer tokens. Evaluating a model thus consists in measuring, for each (xi,g)(x_{i},g) pair, how close p(i,g)p^{(i,g)} is to q(i,g)q^{(i,g)} (§3.3).

Table 1 shows why we compare distributions rather than aggregated labels. Mapping the ordinal points A–E to equally spaced values in [0,1][0,1] (i.e., {0,0.25,0.5,0.75,1}\{0,0.25,0.5,0.75,1\}), we can summarize a distribution by its mean rating, its expected value under this mapping. The two groups in the example have the same mean rating (0.3750.375), so an averaged label would make the judge equally close to both. However, the distributions differ: men concentrate on the middle of the scale, women split between its two ends. Unlike the score ss defined below, the mean rating treats the scale as interval. We use it only to report the direction and magnitude of conditioning shifts.

Comparing ordinal distributions.

Because our labels are ordinal, the discrepancy between two distributions should grow with the distance over which probability mass is misplaced. Divergence measures such as KL treat labels as nominal, penalizing misplaced mass equally regardless of where in the scale it lands. We therefore use the Earth Mover’s Distance (EMD) (Rubner et al., 2000), the probability mass that must be moved to turn one distribution into the other, weighted by the ordinal distance moved:22 2 EMD is used in NLP for subjective tasks, from document similarity (Kusner et al., 2015) to aligning model and human response distributions on ordinal survey scales (Santurkar et al., 2023; Tjuatja et al., 2024).

EMD⁡(p,q)=∑k=1K−1|∑j≤kpj−∑j≤kqj|.\mathrm{EMD}(p,q)=\sum_{k=1}^{K-1}\Bigl|\,\textstyle\sum_{j\leq k}p_{j}-\sum_{j\leq k}q_{j}\,\Bigr|.

Following Santurkar et al. (2023), we report it as a bounded similarity score, inverted so that higher is better:

s⁡(p,q)=1−EMD⁡(p,q)K−1.s(p,q)=1-\frac{\mathrm{EMD}(p,q)}{K-1}.

Because the maximum EMD on a KK-point scale is K−1K-1, reached when the two distributions place all their mass on opposite extremes, ss ranges from 00 (maximal divergence) to 11 (perfect match) and is comparable across tasks and configurations. To isolate the effect of demographic conditioning, we score both the unconditioned and the conditioned prediction (see §3.3) for the same text against the same group reference:

Δ​s=s⁡(p(i,g),q(i,g))−s⁡(p∅(i),q(i,g)).\Delta s=s\bigl(p^{(i,g)},q^{(i,g)}\bigr)-s\bigl(p^{(i)}_{\emptyset},q^{(i,g)}\bigr).

where p∅(i)p^{(i)}_{\emptyset} is the model’s prediction for xix_{i} without a profile. While ss measures how closely a judge matches a group, Δ​s\Delta s isolates the effect of the profile alone: Δ​s>0\Delta s>0 means conditioning moved the judge toward the group, Δ​s<0\Delta s<0 that it moved the judge away from the group it was instructed to represent.

Aggregation and uncertainty.

Both ss and Δ​s\Delta s can be affected by the shape of the prediction and reference distributions: a more spread-out prediction tends to score better against a more spread-out reference, even without better group alignment. This can distort comparisons in two ways: (i) on the human labels side, groups can have different numbers of annotators per item, so some reference distributions are more sparse than others; (ii) on the model side, instruction-tuning is known to sharpen answer-token distributions (Santurkar et al., 2023; Durmus et al., 2023; Sorensen et al., 2024), so any comparison of base and instruct model variants risks being an artifact of sharpness. We therefore validate our results against mode accuracy, which is unaffected by distribution sharpness. We average scores within each cell (a model, configuration, task, and demographic group) and then macro-average cell means with equal weight, preventing groups with more observations from dominating the average. Confidence intervals are 95%95\% percentile bootstraps over cells (2,000 resamples). We validate all group comparisons against two independent spread-insensitive controls: density-matched references and mode accuracy. Mode accuracy also mitigates the concern that first-token probabilities may diverge from the model’s generated answer (Wang et al., 2024b). We verify token coverage but do not compare against generated text (Appendix F).

3.3 Experimental Design

We elicit a judgment from a model in a single forward pass and read off the full distribution it places on the five scale points, rather than sampling a discrete answer. We follow the prompt format of Orlikowski et al. (2025): each prompt presents the task question, the text to be judged, and the five labeled options (A)–(E) in order, and ends with an assistant-turn prefix constraining the model’s next token to be one of the option letters (Figure 1).33 3 The order is fixed, but the models show no choice-position bias (Zheng et al., 2024): pooled over the 23 models, option mass tracks the human labels (r=0.83r=0.83) rather than letter order, and the modal option (A) (0.275CLOSE(0.275 against 0.2120.212 for (B)) is also the modal human label (0.323)(0.323). All results in §4.3 use this single template, since conditioned and unconditioned predictions must be scored under the same wording to be paired. Because prompt format and wording could play a relevant role in the distributions (Beck et al., 2024), we also run all 2323 models under a structurally different template, an interview-style profile in which the annotator states their own demographics (Lutz et al., 2025) (see Appendix B).

Following prior work that reads judgment distributions directly from answer-token probabilities (Santurkar et al., 2023; Durmus et al., 2023), we take the predicted distribution pp to be the softmax over the log-probabilities of the five option tokens A–E.44 4 We retrieve the top-20 next-token log-probabilities and read each option’s first token, assigning −∞-\infty to any option absent from that set. The configurations below differ only in (i) the demographic profile, if any, prepended to the prompt, and (ii) the human reference against which each prediction is scored (§3.2).

Unconditioned Baseline.

In this configuration the prompt carries no demographic information (Figure 1, without the profile line), thus the model judges without being assigned a perspective. For each text, the baseline prediction is compared with the human label distribution of every demographic group that annotated it. This unconditioned setting (i) provides the reference point for score difference Δ​s\Delta s (§3.2), isolating the effect of adding a demographic profile to the prompt, and (ii) reveals possible models’ default demographic alignment (i.e., the perspective it tends to adopt in the absence of demographic conditioning).

Demographic Conditioning.

We prompt the model to judge the text from the perspective of a specific demographic group, i.e. a profile (Figure 1), and score its prediction against that group’s human label distribution. A profile assigns a value to one or two of the four dimensions of §3.1: gender, age, race, and education. We consider the four single-attribute configurations and the six two-attribute combinations.55 5 We stop at two attributes, since narrower profiles leave too few annotators per cell to estimate a stable reference. The former isolate each dimension’s contribution, the latter test whether intersectional profiles help or hurt. Across the ten configurations this yields 268K (xi,g)(x_{i},g) pairs per model. For each, we compute ss twice against the same reference q(i,g)q^{(i,g)} (once for the conditioned prediction p(i,g)p^{(i,g)}, once for the unconditioned p∅(i)p^{(i)}_{\emptyset}) and take their difference as Δ​s\Delta s.

Figure 2: Alignment of the unconditioned baseline prediction with each demographic group’s human label distribution, one panel per dimension and one row per model. Hollow markers are base models, filled markers instruct models. Color identifies the group.

3.4 Models

We evaluate models from five families, ranging from 4B to 70B parameters: Gemma 3 4B, 12B, and 27B (Team et al., 2025); Llama 3.1 8B and 70B (Grattafiori et al., 2024); Ministral 3 8B and 14B (Liu et al., 2026); OLMo 3 7B and 32B (Olmo et al., 2026); and Qwen 3 8B, 14B, and 32B (Yang et al., 2025). We evaluate models in both base and instruction-tuned form,66 6 Instruction-tuned models receive the prompt in their chat template, with the final assistant turn left open at the answer prefix. Base models receive the same messages as plain concatenation, so each is prompted in its native format. except Qwen 3 32B, which is available only instruction-tuned, for 23 models in total. Variation in size within each family lets us study scale while keeping the model family fixed. The 11 matched base–instruct pairs isolate the effect of instruction-tuning.

4 Results and Discussion

We organize the results around the three research questions of §1.

4.1 RQ1: Default Profiles of LLM-Judges

In line with Schäfer et al. (2025), who found judges furthest from Black annotators and unaffected by gender on two proprietary models, we find a shared default profile in open-weight models’ label distributions.

All models align best with White and college-educated annotators, equally well with men and women, and differ from one another in how they relate to age.

Figure 2 reports the alignment of the unconditioned predictions with each demographic group’s label distributions, one panel per dimension. Race shows the largest gaps: all base models and the majority of instruct models match White annotators more closely than Black and Asian ones, with an average White–Black gap of 0.0890.089 for base models and 0.0470.047 for instruct models.

For base judges, this corresponds to moving more than one third of the probability mass by one response step away from Black annotators’ judgments.77 7 The ordering is robust to annotator-count differences for instruction-tuned models, and holds for most base models under matching, even as the gap sizes shrink (Appendix F.1). Such group-level differences are visible only under a distributional comparison, while scoring against a single aggregate label would collapse (§3.2). Education is the most consistent dimension: pooled across tasks, every model is closest to college-educated annotators, ahead of annotators with a high-school education. Gender, by contrast, shows no default: the markers largely overlap, with mean absolute gaps below 0.010.01. Age is the only dimension on which models disagree, and the split tracks model type: base models are generally closest to the oldest group, whereas most instruct models are closest to the youngest. This is a first indication that instruction-tuning changes whose perspective a judge adopts even before any profile is assigned. We investigate the impact of instruction-tuning further in §4.2. These defaults are largely stable across tasks: race and gender patterns hold on all three, while age (for all models) and education (for instruct models only) vary more by task (see Appendix C). Appendix F verifies that these orderings are not artifacts of distributional spread.

An unconditioned judge is not a neutral annotator.

LLMs are often used as substitutes for human annotators, but an unconditioned model should not be treated as an average annotator. Without any demographic profile, every instruction-tuned judge aligns more closely with White than Black or Asian annotators, and with college-educated than less-educated annotators. Thus, using an unconditioned judge as a generic annotator introduces a systematic demographic perspective. This confirms, on open-weight judges and distributional scoring, the conclusion of Schäfer et al. (2025) that LLMs do not represent all social groups equally. The analyses that follow ask what happens when a profile is assigned: whether conditioning corrects it or compounds it.

4.2 RQ2: The Effect of Demographic Conditioning

Demographic conditioning helps base models, but not instruct models (on average).

Figure 3 compares the conditioned and unconditioned predictions of each judge against the same group references. Assigning a profile improves 10 of 11 base models, while among the 12 instruct models it helps five, leaves two unchanged, and hurts the rest: on average, conditioning appears to have little to no effect on instruction-tuned models.

The base model gain is not an artifact of spread-out predictions.

Base models distribute their probability more evenly across the five options than instruct models, which tend to concentrate it on one or two.88 8 Measured as entropy normalized by its maximum log⁡K\log K, so that 00 is a one-point distribution and 11 a uniform one: 0.9190.919 on average for base models, 0.2740.274 for the instruct. This asymmetry admits an alternative explanation for the base model gain: a spread-out prediction necessarily overlaps part of any reference distribution, so conditioning could improve base scores simply by reshaping an almost-flat prediction, without moving the judge toward the group’s actual judgments. We rule out this explanation using mode accuracy, which ignores probability spread and gives credit only when the model’s most likely option matches the group’s majority label. Under this metric, the mean conditioning gain for base models increases from +0.014+0.014 to +0.032+0.032 and remains positive for 1010 of 1111 models (see Appendix F.2). Instruct models remain near zero under both metrics.

Figure 3: Alignment with annotator groups, unconditioned (triangle) vs. conditioned (circle), for the matched base models (hollow) and the instruction-tuned judges (filled). The horizontal distance between a row’s triangle and circle represents the effect of conditioning for that model (Δ​s\Delta s).

Neither scale nor instruction-tuning makes judges more steerable.

Larger models align better with annotators: within the instruction-tuned pool, r=0.85r=0.85 between log parameter count and ss, and instruct models start 0.0770.077 ahead of the base ones in the unconditioned setting. Yet neither gain carries over to conditioning: log parameter count and Δ​s\Delta s are essentially uncorrelated across the 23 judges (r=0.016r=0.016), and the largest gains fall on the base Qwen models regardless of scale (Qwen 8B +0.041+0.041, Qwen 14B +0.044+0.044).

Better judges, however, are not more steerable judges. A judge can be wrong about a text in two ways: it can misjudge the text itself, which puts it off for every group at once, or it can miss the way one group departs from the others. Instruction-tuning only improves the first, and every group’s score rises accordingly.

The instruct models’ null effect is not uniform across tasks.

For base models, conditioning helps on all three tasks. For instruct models, the Δ​s\Delta s is below 0.010.01 on politeness and intimacy but turns negative on offensiveness (−0.011-0.011 on average, down to −0.080-0.080 for Ministral 14B). Appendix D reports the per-model effects: on the one task where demographic perspective shows the greatest variability, assigning a profile makes several instruct judges worse at representing the group they are instructed to represent.

However, the almost flat average conceals relevant per-group behavior, both positive and negative (Simpson, 1951; Blyth, 1972). The next section decomposes it, asking which groups conditioning helps and which it hurts.

Figure 4: Conditioning effect (Δ​s\Delta s) by demographic group for the 11 models released in both configurations. Bars are hollow for base, filled for instruct, colored by sign. Whiskers are 95%95\% cluster-bootstrap intervals over judges. Left: single attributes. Right: gender×\timesrace. Age is omitted, as no bracket differs from zero for instruct models.
Figure 5: As Figure 4, one column per task. Top row: single attributes. Bottom row: gender×\timesrace. Diamonds mark effects below 0.0020.002, too small to draw.

4.3 RQ3: Conditioning Helps Majority Groups and Hurts Minority Groups

Instruction-tuning makes demographic conditioning uneven across groups.

Figure 4 shows the conditioning effect separately for each demographic group. For base models, conditioning is positive for every group: assigning a profile moves the judge toward that group’s judgments regardless of which group it is. With instruction-tuned models, the effect splits: conditioning remains positive for White, Man, College degree, and High school, but becomes negative for Asian, Woman, Graduate degree, and Black. The largest harm falls on Black annotators (−0.033-0.033): instructing the judge to answer as a Black annotator makes it less aligned with Black annotators’ judgments than giving it no profile at all.

Intersectional profiles do not combine additively: the joint effect is attenuated toward the gender marginal, landing above the sum of the two single-attribute effects. Man ×\times White benefits (+0.009+0.009), while against Black, the near-zero Man effect dilutes the harm (Man ×\times Black −0.019-0.019) while Woman deepens it (Woman ×\times Black −0.036-0.036).

The harm to minority profiles is concentrated on offensiveness.

Figure 5 splits the same breakdown by task. Unlike the defaults of §4.1 and the pooled effects of §4.2, which apply similarly to all three tasks, the harm to minority groups is more evident in offensiveness. On offensiveness, conditioning an instruct judge reduces alignment with Black annotators by −0.074-0.074, with Asian annotators and with women by −0.035-0.035. Every gender ×\times race profile containing a minority attribute is negative, with Woman ×\times Black at −0.087-0.087, the largest effect in the study. The pattern is weaker on intimacy, where Black is the only group with a reliable negative effect (−0.018-0.018), and absent on politeness, where no minority profile is harmed. Base models show none of this task specificity: their conditioning effect is positive for nearly every group on all three tasks. Thus, the asymmetry is not a general property of all LLMs under demographic conditioning. It is rather introduced by instruction-tuning and expressed on the task where demographic perspective is most contested.

Prompting moves the judge toward a stereotype, not toward the group.

Figure 6: One point per judge and profile, one panel per task. The horizontal axis shows the change in predicted rating; the vertical axis shows the change in group alignment (Δ​s\Delta s). The yellow point marks Ministral 14B with Woman×\timesBlack. Annotations report Pearson rr, overall and by model pool.
Group Human Δ\Delta base Δ\Delta instruct
Offensiveness
White −0.018-0.018 −0.070-0.070 −0.026-0.026
Black +0.052+0.052 +0.005+0.005 +0.080+0.080
Asian −0.003-0.003 −0.031-0.031 +0.010+0.010
Man +0.006+0.006 −0.044-0.044 −0.004-0.004
Woman +0.003+0.003 −0.026-0.026 +0.039+0.039
Politeness
White −0.001-0.001 +0.054+0.054 +0.051+0.051
Black +0.028+0.028 +0.022+0.022 +0.003+0.003
Asian −0.027-0.027 +0.047+0.047 +0.029+0.029
Man +0.005+0.005 +0.027+0.027 +0.020+0.020
Woman +0.001+0.001 +0.029+0.029 +0.009+0.009
Intimacy
White +0.005+0.005 −0.044-0.044 −0.017-0.017
Black +0.048+0.048 −0.025-0.025 +0.030+0.030
Asian −0.025-0.025 −0.036-0.036 +0.001+0.001
Man −0.018-0.018 −0.047-0.047 −0.005-0.005
Woman +0.011+0.011 −0.039-0.039 +0.009+0.009
Table 4: Group deviation from the pooled human mean and the rating shift induced by conditioning for all models (Δ\Delta), by task and single-attribute group (mean rating on the [0,1] scale described in §3.2). Faithful conditioning would give Δ≈\Delta~\approx Human.

Demographic conditioning is effective when it moves the judge’s rating by the amount the group actually differs from the mean of all annotators. The two shifts are reported in Table 4. Two patterns emerge. On offensiveness, Black annotators rate +0.052+0.052 above the mean, yet the Black profile moves instruct judges by +0.080+0.080, about 50%50\% more than the real difference: conditioning amplifies the gap rather than reproducing it. On politeness, Asian is the only group rating below the mean (−0.027-0.027), yet its profile shifts judges upward (+0.029+0.029 instruct, +0.047+0.047 base), opposite to the group it names. As the table shows, the shift is largely insensitive to the group’s actual deviation: on offensiveness it is upward for instruct judges regardless of the group, and only base judges, which overestimate offensiveness unconditioned, move back toward the human ratings. In general, the average shift is upward, whether the group differs from the mean or not: the profile triggers the expectation that the named group perceives more offensiveness, not that group’s actual judgments. This confirms the stereotype activation reported for persona-prompted generation (Gupta et al., 2024; Deshpande et al., 2023), and quantifies it, as against real annotator distributions, the shift can be measured against the group’s actual deviation. Intersectional profiles show the same pattern as the corresponding single-attribute ones. We report them in Appendix E.

Gender shows that judges shift even when there is no group-related difference.

As shown in Table 4, gender is the one dimension with no default (§4.1) and no group deviation to reproduce: women rate at the pooled mean (e.g., +0.003+0.003 in offensiveness). A faithful judge conditioned on Woman would therefore not move. Instead the profile shifts instruction-tuned judges upward (e.g., +0.039+0.039 in offensiveness). Here, on the contrary, the distributions unduly move because of the models’ stereotyped expectations of the group.

Conditioning shifts ratings, not perspectives.

Figure 6 shows the correlation between shifts in the label distributions and Δ​s\Delta s. On offensiveness, the more conditioning raises a judge’s rating, the more alignment it loses (r=−0.93r=-0.93, p<.001p<.001). The same trend holds when comparing judges given the same profile (within-profile r=−0.92r=-0.92).

Base models tend to move downward, toward the human ratings, and therefore improve; instruct models tend to move upward, away from them, and therefore worsen. Intimacy shows the same pattern. Politeness reverses the direction. Judges initially underrate politeness, predicting 0.380.38 on average against a human mean of 0.580.58. Conditioning tends to push their ratings upward, so larger shifts improve alignment (r=+0.75r=+0.75). Yet this does not mean the judges recover each group’s perspective: only 5 of the 11 profiles shift the judge in the same direction as the group’s actual deviation from the overall human mean. The gain therefore comes mainly from correcting the judges’ general underestimation of politeness, not from reproducing group-specific judgments.

Across all three tasks, the mechanism is the same: demographic conditioning pushes all predictions in a direction that is not reliably tied to the group being represented. It helps when that push happens to move the judge toward the human ratings, and hurts when it moves the judge away.

5 Conclusions

Across 23 open-weight LLM judges, we asked whether conditioning a judge on an annotator’s demographic profile brings its judgments closer to that group’s. An unconditioned judge is not perspective-neutral: every model aligns more closely with White than with Black or Asian annotators, and with college-educated annotators than with either other band. Conditioning an instruction-tuned judge appears to have no effect on average. However, that average hides gains for the groups the judge already leans toward, and losses for minority groups. The harm concentrates on offensiveness, is sharpened rather than repaired by intersectional profiling, and is not recovered by model size. Comparing each instruction-tuned judges with their base counterparts locates the asymmetry in instruction-tuning: conditioning positively impacts base models for every group, whereas with instruction-tuned models the profile does not activate the group’s perspective but an expectation about it, exaggerating differences where they exist and introducing them where they do not. Thus the practical question is not simply whether demographic conditioning helps, but whom it helps: its benefits are smallest for the very groups such interventions are meant to represent.

Limitations

Our findings rest on three English tasks from one annotator-level corpus (Orlikowski et al., 2025), so we cannot claim the pattern transfers to other languages or annotation schemes. Those corpora are public, and any leakage into pretraining would favor the alignment we measure. Per-group distributions thin as profiles narrow, so we condition on at most two attributes, and minority groups carry fewer parallel annotations (Difallah et al., 2018; Prabhakaran et al., 2021).

Ethics Statement

No new annotation was collected to conduct the analysis described in this work and no annotator is identifiable. We use coarse demographic categories only to define the group reference distributions, not as proxies for individual perspectives. Our results caution against that stronger reading, since conditioning helped the groups judges already favour while harming the rest. Thus, demographic prompting should not be treated as a substitute for recruiting annotators from the groups of interest.

References

  • Adilazuarda et al. (2024) M. F. Adilazuarda, S. Mukherjee, P. Lavania, S. S. Singh, A. F. Aji, J. O’Neill, A. Modi, and M. Choudhury Towards measuring and modeling “culture” in LLMs: a survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 15763–15784. External Links: Link, Document Cited by: §2.
  • Agnew et al. (2024) W. Agnew, A. S. Bergman, J. Chien, M. Díaz, S. El-Sayed, J. Pittman, S. Mohamed, and K. R. McKee The illusion of artificial inclusion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA. External Links: ISBN 9798400703300, Link, Document Cited by: §2.
  • Alipour et al. (2025) S. Alipour, I. Sen, M. Samory, and T. Mitra Robustness and confounders in the demographic alignment of LLMs with human perceptions of offensiveness. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 22025–22047. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
  • Amin et al. (2026) H. Amin, H. Y. Tian, X. Duan, C. Ho, R. Khanna, and M. Yin From fallback to frontline: when can LLMs be superior annotators of human perspectives?. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 42798–42830. External Links: Link, ISBN 979-8-89176-395-1 Cited by: §1.
  • Argyle et al. (2022) L. P. Argyle, E. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate Out of one, many: using language models to simulate human samples. Political Analysis 31, pp. 337 – 351. External Links: Link Cited by: §2, §2.
  • Aroyo et al. (2023) L. Aroyo, A. Taylor, M. Díaz, C. Homan, A. Parrish, G. Serapio-García, V. Prabhakaran, and D. Wang DICES dataset: diversity in conversational ai evaluation for safety. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, pp. 53330–53342. External Links: Link Cited by: footnote 1.
  • Basile et al. (2021) V. Basile, M. Fell, T. Fornaciari, D. Hovy, S. Paun, B. Plank, M. Poesio, and A. Uma We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, K. Church, M. Liberman, and V. Kordoni (Eds.), Online, pp. 15–21. External Links: Link, Document Cited by: §1, §2.
  • Bavaresco et al. (2025) A. Bavaresco, R. Bernardi, L. Bertolazzi, D. Elliott, R. Fernández, A. Gatt, E. Ghaleb, M. Giulianelli, M. Hanna, A. Koller, A. Martins, P. Mondorf, V. Neplenbroek, S. Pezzelle, B. Plank, D. Schlangen, A. Suglia, A. K. Surikuchi, E. Takmaz, and A. Testoni LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 238–255. External Links: Link, Document, ISBN 979-8-89176-252-7 Cited by: §1, §2.
  • Beck et al. (2024) T. Beck, H. Schuff, A. Lauscher, and I. Gurevych Sensitivity, performance, robustness: deconstructing the effect of sociodemographic prompting. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Y. Graham and M. Purver (Eds.), St. Julian’s, Malta, pp. 2589–2615. External Links: Link, Document Cited by: §1, §3.3.
  • Blyth (1972) C. R. Blyth On Simpson’s paradox and the sure-thing principle. Journal of the American Statistical Association 67 (338), pp. 364–366. External Links: Document Cited by: §4.2.
  • Cabitza et al. (2023) F. Cabitza, A. Campagner, and V. Basile Toward a perspectivist turn in ground truthing for predictive computing. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’23/IAAI’23/EAAI’23. External Links: ISBN 978-1-57735-880-0, Link, Document Cited by: §1.
  • Chen et al. (2024a) G. H. Chen, S. Chen, Z. Liu, F. Jiang, and B. Wang Humans or LLMs as the judge? a study on judgement bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 8301–8327. External Links: Link, Document Cited by: §2.
  • Chen et al. (2024b) J. Chen, X. Wang, R. Xu, S. Yuan, Y. Zhang, W. Shi, J. Xie, S. Li, R. Yang, T. Zhu, et al. From persona to personalization: a survey on role-playing language agents. Transactions on Machine Learning Research. Cited by: §1, §2.
  • Deshpande et al. (2023) A. Deshpande, V. Murahari, T. Rajpurohit, A. Kalyan, and K. Narasimhan Toxicity in chatgpt: analyzing persona-assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 1236–1270. External Links: Link, Document Cited by: §2, §4.3.
  • Diaz et al. (2018) M. Diaz, I. Johnson, A. Lazar, A. M. Piper, and D. Gergle Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI ’18, New York, NY, USA, pp. 1–14. External Links: ISBN 9781450356206, Link, Document Cited by: §1, footnote 1.
  • Difallah et al. (2018) D. Difallah, E. Filatova, and P. Ipeirotis Demographics and dynamics of mechanical turk workers. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, WSDM ’18, New York, NY, USA, pp. 135–143. External Links: ISBN 9781450355810, Link, Document Cited by: Limitations.
  • Dong et al. (2024) Y. R. Dong, T. Hu, and N. Collier Can LLM be a personalized judge?. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 10126–10141. External Links: Link, Document Cited by: §2, §2.
  • Durmus et al. (2023) E. Durmus, K. Nyugen, T. I. Liao, N. Schiefer, A. Askell, A. Bakhtin, C. Chen, Z. Hatfield-Dodds, D. Hernandez, N. Joseph, L. Lovitt, S. McCandlish, O. Sikder, A. Tamkin, J. Thamkul, J. Kaplan, J. Clark, and D. Ganguli Towards measuring the representation of subjective global opinions in language models. External Links: 2306.16388 Cited by: §2, §3.2, §3.3.
  • Feng et al. (2024) S. Feng, T. Sorensen, Y. Liu, J. Fisher, C. Y. Park, Y. Choi, and Y. Tsvetkov Modular pluralism: pluralistic alignment via multi-LLM collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 4151–4171. External Links: Link, Document Cited by: §1.
  • Frenda et al. (2024) S. Frenda, G. Abercrombie, V. Basile, A. Pedrani, R. Panizzon, A. T. Cignarella, C. Marco, and D. Bernardi Perspectivist approaches to natural language processing: a survey: perspectivist approaches to natural language processing…. Lang. Resour. Eval. 59 (2), pp. 1719–1746. External Links: ISSN 1574-020X, Link, Document Cited by: §1.
  • Gao et al. (2025) Y. Gao, D. Lee, G. Burtch, and S. Fazelpour Take caution in using llms as human surrogates. Proceedings of the National Academy of Sciences 122 (24), pp. e2501660122. External Links: Document, Link, https://www.pnas.org/doi/pdf/10.1073/pnas.2501660122 Cited by: §2.
  • Gordon et al. (2022) M. L. Gordon, M. S. Lam, J. S. Park, K. Patel, J. Hancock, T. Hashimoto, and M. S. Bernstein Jury learning: integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY, USA. External Links: ISBN 9781450391573, Link, Document Cited by: §2.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §3.4.
  • Gu et al. (2026) J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Z. Lin, B. Zhang, L. Ni, W. Gao, Y. Wang, and J. Guo A survey on llm-as-a-judge. The Innovation 7 (6), pp. 101253. External Links: ISSN 2666-6758, Document, Link Cited by: §1, §2.
  • Gupta et al. (2024) S. Gupta, V. Shrivastava, A. Deshpande, A. Kalyan, P. Clark, A. Sabharwal, and T. Khot Bias runs deep: implicit reasoning biases in persona-assigned llms. In International Conference on Learning Representations, Vol. 2024, pp. 21849–21874. Cited by: §1, §2, §4.3.
  • Horton et al. (2023) J. J. Horton, A. Filippas, and B. S. Manning Large language models as simulated economic agents: what can we learn from homo silicus?. Technical report National Bureau of Economic Research. Cited by: §2.
  • Hu and Collier (2024) T. Hu and N. Collier Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 10289–10307. External Links: Link, Document Cited by: §2.
  • Kusner et al. (2015) M. J. Kusner, Y. Sun, N. I. Kolkin, and K. Q. Weinberger From word embeddings to document distances. In Proceedings of the 32nd International Conference on Machine Learning - Volume 37, ICML’15, pp. 957–966. Cited by: footnote 2.
  • Li et al. (2025) D. Li, B. Jiang, L. Huang, A. Beigi, C. Zhao, Z. Tan, A. Bhattacharjee, Y. Jiang, C. Chen, T. Wu, K. Shu, L. Cheng, and H. Liu From generation to judgment: opportunities and challenges of LLM-as-a-judge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2757–2791. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1, §2.
  • Liu et al. (2026) A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, A. Sablayrolles, A. Héliou, A. You, A. Ehrenberg, A. Lo, A. Eliseev, A. Calvi, A. Sooriyarachchi, B. Bout, B. Rozière, B. D. Monicault, C. Lanfranchi, C. Barreau, C. Courtot, D. Grattarola, D. Dabert, D. de las Casas, E. Chane-Sane, F. Ahmed, G. Berrada, G. Ecrepont, G. Guinet, G. Novikov, G. Kunsch, G. Lample, G. Martin, G. Gupta, J. Ludziejewski, J. Rute, J. Studnia, J. Amar, J. Delas, J. S. Roberts, K. Yadav, K. Chandu, K. Jain, L. Aitchison, L. Fainsin, L. Blier, L. Zhao, L. Martin, L. Saulnier, L. Gao, M. Buyl, M. Jennings, M. Pellat, M. Prins, M. Poirée, M. Guillaumin, M. Dinot, M. Futeral, M. Darrin, M. Augustin, M. Chiquier, M. Schimpf, N. Grinsztajn, N. Gupta, N. Raghuraman, O. Bousquet, O. Duchenne, P. Wang, P. von Platen, P. Jacob, P. Wambergue, P. Kurylowicz, P. R. Muddireddy, P. Chagniot, P. Stock, P. Agrawal, Q. Torroba, R. Sauvestre, R. Soletskyi, R. Menneer, S. Vaze, S. Barry, S. Gandhi, S. Waghjale, S. Gandhi, S. Ghosh, S. Mishra, S. Aithal, S. Antoniak, T. L. Scao, T. Cachet, T. S. Sorg, T. Lavril, T. N. Saada, T. Chabal, T. Foubert, T. Robert, T. Wang, T. Lawson, T. Bewley, T. Bewley, T. Edwards, U. Jamil, U. Tomasini, V. Nemychnikova, V. Phung, V. Maladière, V. Richard, W. Bouaziz, W. Li, W. Marshall, X. Li, X. Yang, Y. E. Ouahidi, Y. Wang, Y. Tang, and Z. Ramzi Ministral 3. External Links: 2601.08584, Link Cited by: §3.4.
  • Lutz et al. (2025) M. Lutz, I. Sen, G. Ahnert, E. Rogers, and M. Strohmaier The prompt makes the person(a): a systematic evaluation of sociodemographic persona prompting for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 23212–23237. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: Appendix B, §2, §3.3.
  • Luz de Araujo et al. (2025) P. H. Luz de Araujo, P. Röttger, D. Hovy, and B. Roth Principled personas: defining and measuring the intended effects of persona prompting on task performance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26857–26886. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2.
  • Meister et al. (2025) N. Meister, C. Guestrin, and T. Hashimoto Benchmarking distributional alignment of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 24–49. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2.
  • Mostafazadeh Davani et al. (2022) A. Mostafazadeh Davani, M. Díaz, and V. Prabhakaran Dealing with disagreements: looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics 10, pp. 92–110. External Links: Link, Document Cited by: §2.
  • Mou et al. (2026) X. Mou, X. Ding, Q. He, L. Wang, J. Liang, X. Zhang, L. Sun, J. Lin, J. Zhou, H. Xuanjing, and Z. Wei From individual to society: a survey on social simulation driven by large language model-based agents. ACM Comput. Surv. 58 (11). External Links: ISSN 0360-0300, Link, Document Cited by: §2.
  • Navigli et al. (2023) R. Navigli, S. Conia, and B. Ross Biases in large language models: origins, inventory, and discussion. J. Data and Information Quality 15 (2). External Links: ISSN 1936-1955, Link, Document Cited by: §1.
  • Neumann et al. (2026) T. Neumann, M. De-Arteaga, and S. Fazelpour Should you use llms to simulate opinions? quality checks for early-stage deliberation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 39070–39079. Cited by: §2.
  • Occhipinti et al. (2024) D. Occhipinti, S. S. Tekiroğlu, and M. Guerini PRODIGy: a PROfile-based DIalogue generation dataset. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 3500–3514. External Links: Link, Document Cited by: §2.
  • Olmo et al. (2026) T. Olmo, :, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A. Asai, A. Bhagia, A. Wettig, A. Liu, A. Rangapur, C. Anastasiades, C. Huang, D. Schwenk, H. Trivedi, I. Magnusson, J. Lochner, J. Liu, L. J. V. Miranda, M. Sap, M. Morgan, M. Schmitz, M. Guerquin, M. Wilson, R. Huff, R. L. Bras, R. Xin, R. Shao, S. Skjonsberg, S. Z. Shen, S. S. Li, T. Wilde, V. Pyatkin, W. Merrill, Y. Chang, Y. Gu, Z. Zeng, A. Sabharwal, L. Zettlemoyer, P. W. Koh, A. Farhadi, N. A. Smith, and H. Hajishirzi Olmo 3. External Links: 2512.13961, Link Cited by: §3.4.
  • Orlikowski et al. (2025) M. Orlikowski, J. Pei, P. Röttger, P. Cimiano, D. Jurgens, and D. Hovy Beyond demographics: fine-tuning large language models to predict individuals’ subjective text perceptions. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 2092–2111. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1, §2, §3.1, §3.3, Limitations.
  • Pavlick and Kwiatkowski (2019) E. Pavlick and T. Kwiatkowski Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics 7, pp. 677–694. External Links: Link, Document Cited by: §1.
  • Pei and Jurgens (2023) J. Pei and D. Jurgens When do annotator demographics matter? measuring the influence of annotator demographics with the POPQUORN dataset. In Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII), J. Prange and A. Friedrich (Eds.), Toronto, Canada, pp. 252–265. External Links: Link, Document Cited by: §1, 2nd item.
  • Pei et al. (2023) J. Pei, V. Silva, M. Bos, Y. Liu, L. Neves, D. Jurgens, and F. Barbieri SemEval-2023 task 9: multilingual tweet intimacy analysis. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), A. Kr. Ojha, A. S. Doğruöz, G. Da San Martino, H. Tayyar Madabushi, R. Kumar, and E. Sartori (Eds.), Toronto, Canada, pp. 2235–2246. External Links: Link, Document Cited by: 1st item.
  • Plank (2022) B. Plank The “problem” of human label variation: on ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 10671–10682. External Links: Link, Document Cited by: §1.
  • Prabhakaran et al. (2021) V. Prabhakaran, A. Mostafazadeh Davani, and M. Diaz On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, C. Bonial and N. Xue (Eds.), Punta Cana, Dominican Republic, pp. 133–138. External Links: Link, Document Cited by: Limitations.
  • Rubner et al. (2000) Y. Rubner, C. Tomasi, and L. J. Guibas The earth mover’s distance as a metric for image retrieval. International Journal of Computer Vision 40 (2), pp. 99–121. External Links: ISSN 1573-1405, Document, Link Cited by: §3.2.
  • Santurkar et al. (2023) S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto Whose opinions do language models reflect?. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. Cited by: §1, §2, §3.2, §3.2, §3.3, footnote 2.
  • Santy et al. (2023) S. Santy, J. Liang, R. Le Bras, K. Reinecke, and M. Sap NLPositionality: characterizing design biases of datasets and models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 9080–9102. External Links: Link, Document Cited by: §1, §2.
  • Sap et al. (2022) M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y. Choi, and N. A. Smith Annotators with attitudes: how annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M. de Marneffe, and I. V. Meza Ruiz (Eds.), Seattle, United States, pp. 5884–5906. External Links: Link, Document Cited by: §1.
  • Schäfer et al. (2025) J. Schäfer, A. Combs, C. Bagdon, J. Li, N. Probol, L. Greschner, S. Papay, Y. Menchaca Resendiz, A. Velutharambath, A. Wuehrl, S. Weber, and R. Klinger Which demographics do LLMs default to during annotation?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 17331–17348. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2, §4.1, §4.1.
  • Shao et al. (2023) Y. Shao, L. Li, J. Dai, and X. Qiu Character-LLM: a trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 13153–13187. External Links: Link Cited by: §2.
  • Simmons and Savinov (2024) G. Simmons and V. Savinov Assessing generalization for subpopulation representative modeling via in-context learning. In Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024), A. Deshpande, E. Hwang, V. Murahari, J. S. Park, D. Yang, A. Sabharwal, K. Narasimhan, and A. Kalyan (Eds.), St. Julians, Malta, pp. 18–35. External Links: Link, Document Cited by: §2.
  • Simpson (1951) E. H. Simpson The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society: Series B (Methodological) 13 (2), pp. 238–241. External Links: Document Cited by: §4.2.
  • Sorensen et al. (2025) T. Sorensen, P. Mishra, R. Patel, M. H. Tessler, M. A. Bakker, G. Evans, I. Gabriel, N. Goodman, and V. Rieser Value profiles for encoding human variation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2047–2095. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1, §2.
  • Sorensen et al. (2024) T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, T. Althoff, and Y. Choi Position: a roadmap to pluralistic alignment. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. Cited by: §1, §3.2.
  • Sun et al. (2025) H. Sun, J. Pei, M. Choi, and D. Jurgens Sociodemographic prompting is not yet an effective approach for simulating subjective judgments with LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 845–854. External Links: Link, Document, ISBN 979-8-89176-190-2 Cited by: §1, §2.
  • Team et al. (2025) G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. György, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. Klimczak-Plucińska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. Põder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot Gemma 3 technical report. External Links: 2503.19786, Link Cited by: §3.4.
  • Tjuatja et al. (2024) L. Tjuatja, V. Chen, T. Wu, A. Talwalkwar, and G. Neubig Do LLMs exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics 12, pp. 1011–1026. External Links: Link, Document Cited by: §2, footnote 2.
  • Tseng et al. (2024) Y. Tseng, Y. Huang, T. Hsiao, W. Chen, C. Huang, Y. Meng, and Y. Chen Two tales of persona in LLMs: a survey of role-playing and personalization. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 16612–16631. External Links: Link, Document Cited by: §1, §2.
  • Wang et al. (2024a) P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, L. Kong, Q. Liu, T. Liu, and Z. Sui Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 9440–9450. External Links: Link, Document Cited by: §2.
  • Wang et al. (2024b) X. Wang, B. Ma, C. Hu, L. Weber-Genzel, P. Röttger, F. Kreuter, D. Hovy, and B. Plank “My answer is C”: first-token probabilities do not match text answers in instruction-tuned language models. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7407–7416. External Links: Link, Document Cited by: §3.2.
  • Wang et al. (2024c) Y. Wang, Z. Yu, W. Yao, Z. Zeng, L. Yang, C. Wang, H. Chen, C. Jiang, R. Xie, J. Wang, et al. Pandalm: an automatic evaluation benchmark for llm instruction tuning optimization. In International Conference on Learning Representations, Vol. 2024, pp. 43573–43593. Cited by: §2.
  • Xu et al. (2026) Y. Xu, V. Derricks, A. Earl, and D. Jurgens Modeling annotator disagreement with demographic-aware experts and synthetic perspectives. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 41260–41294. External Links: Link, ISBN 979-8-89176-390-6 Cited by: §1.
  • Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: §3.4.
  • Zheng et al. (2024) C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: footnote 3.
  • Zheng et al. (2023) L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §2.

Appendix A Filtering of Under-Represented Categories

As described in §3.1, we remove annotations from demographic categories too sparsely represented in the source data to support reliable group-level comparison. Within each dimension we drop values contributed by too few annotators to estimate a stable per-group label distribution. Counting distinct annotators on the three retained tasks, this removes Non-binary from gender, Hispanic/Latino from race, Less than high school from education, and the two oldest age brackets, 7070–7979 and 8080–8989, each populated only by a single annotator.

A cell is dropped if any value in its attribute conjunction is excluded, since the cell’s human reference is defined by the full conjunction. Filtering is applied before any scoring, so unconditioned and conditioned predictions are evaluated against references drawn only from retained categories.

Appendix B Prompt Templates

Every task uses the same two-message template: the user message states the task, presents the item, and lists the five options in ascending order, and the assistant message is a continued prefix, so the next token is the option letter. A conditioned prompt prepends the **Your Profile** line, one clause per active attribute in the fixed order race, age, gender, education.

[user]
Your Profile: In terms of race or ethnicity, you are White. You are 35 to 39 years old.
Question: Consider you read this comment on Reddit, how offensive do you think it is?
Comment:: <text>
(A) not offensive at all
[…]
(E) very offensive
[assistant]
Answer: (
Figure 7: The prompt template (offensiveness, two-attribute profile). The final assistant turn is left open at the answer prefix (§3.3).

The unconditioned prompt is the same with the Your Profile line removed, so both conditions put the same question to the model and differ only in whether a profile precedes it. For base models the same messages are concatenated in plain text, without chat-template.

The other two framings ask how intimate a Twitter **Post** is and how polite an **Email:** from a colleague is, with options from not intimate/polite at all to very intimate/polite. All eleven configurations are released with the code as fully rendered prompts.

Robustness to the template.

To test whether the demographic conditioning effects depend on the wording, we repeat the experiment with an interview template in which the model states the profile itself (Lutz et al., 2025). Question, items and options are unchanged; only the lines before them differ.

[user]
Interviewer: Before we begin, tell me a little about yourself.
You: In terms of race or ethnicity, I am White.
Interviewer: Thank you. Now please answer the following.
Question: […]
Comment: <text>
(A) not offensive at all
[…]
(E) very offensive
[assistant]
Answer: (
Figure 8: The interview-style template: the annotator states the demographic attribute in first person instead of receiving it as an assigned profile.

Its unconditioned arm drops the exchange and opens with Interviewer: Please answer the following., so what separates the two arms is the profile alone. Figure 9 compares the conditioning effects across groups under the two templates.

Figure 9: The conditioning effect per group when the profile is asserted to the model against when the model states it itself. Bars are Δ​s\Delta s, with 95%95\% cluster bootstrap intervals over judges.

The negative effects do not disappear under the alternative template. For instruct models, conditioning becomes more negative across groups: for Black annotators, for example, Δ​s\Delta s changes from −0.027-0.027 to −0.035-0.035. On average, the self-report template lowers Δ​s\Delta s by 0.0200.020, making the effect negative for all sixteen groups. Base models are much less sensitive to the change in wording: no group changes by more than 0.0070.007, and conditioning remains positive for all sixteen groups. The negative effects are therefore not specific to the wording of the main prompt, and are stronger when the model states the identity itself.

Appendix C Per-Group Alignment of the Unconditioned Judge

This appendix reports the per-model, per-task view behind §4.1: unconditioned predictions only, scored and macro-averaged as in §3.2, over the retained groups. The base pool barely moves across tasks (Figure 10): White is the closest race group and College degree the closest education band in every base model on every task, and the oldest bracket leads in all but one model, except on intimacy, where the closest bracket splits between the youngest and 50–59. The instruct pool varies more. White stays ahead of Black in all 1212 models on politeness and intimacy but in 99 on offensiveness, where the mean gap drops to +0.023+0.023; Asian is the closest race group in half the models on politeness. Education reverses: College degree is closest in all 1212 models on politeness, the split is even on intimacy, and High school or below leads in 99 of 1212 on offensiveness. Gender stays nearly flat on every task; age changes closest bracket from task to task.

Figure 10: Group alignment by task, one dot per demographic group.

Appendix D Per-Model Conditioning Effects

This appendix reports the per-model, per-task view behind §4.2. Figure 11 splits the per-model view of Figure 3 by task, with rows in the same order. Two things are worth reading off directly: (i) on every task the hollow pairs shift right more often than the filled ones, and the base mean is positive on all three tasks and exceeds the instruct mean on all three; (ii) the tasks differ in dispersion rather than in direction. Politeness holds the smallest effects in both pools, while the largest sit on offensiveness for instruct models and on intimacy for base models. Offensiveness is the only task whose instruct mean is negative, and since the pooled instruct effect averages the three tasks, it is what places the pooled figure below zero.

Figure 11: Conditioning effects by task. Each row is one model. The horizontal distance between the markers in a row is that model’s conditioning.

Appendix E Rating-Shift Faithfulness for Intersectional Profiles

Table 5 repeats the analysis of 4 for the six gender×\timesrace profiles. Each profile shifts in the direction of its race attribute, with magnitudes attenuated toward the gender component. The reference deviations are computed from the annotators matching both attributes, so the intervals are wider than in the single-attribute figure.

Group Human Δ\Delta base Δ\Delta instruct
Offensiveness
Man ×\times White −0.008-0.008 −0.068-0.068 −0.025-0.025
Woman ×\times White −0.019-0.019 −0.049-0.049 +0.005+0.005
Man ×\times Asian +0.016+0.016 −0.041-0.041 +0.001+0.001
Woman ×\times Asian −0.010-0.010 −0.018-0.018 +0.043+0.043
Man ×\times Black +0.061+0.061 −0.011-0.011 +0.063+0.063
Woman ×\times Black +0.061+0.061 +0.011+0.011 +0.104+0.104
Politeness
Man ×\times White +0.000+0.000 +0.048+0.048 +0.039+0.039
Woman ×\times White +0.003+0.003 +0.046+0.046 +0.032+0.032
Man ×\times Asian −0.027-0.027 +0.041+0.041 +0.026+0.026
Woman ×\times Asian −0.030-0.030 +0.037+0.037 +0.011+0.011
Man ×\times Black +0.050+0.050 +0.023+0.023 +0.005+0.005
Woman ×\times Black +0.016+0.016 +0.020+0.020 −0.011-0.011
Intimacy
Man ×\times White −0.019-0.019 −0.054-0.054 −0.028-0.028
Woman ×\times White +0.009+0.009 −0.044-0.044 −0.018-0.018
Man ×\times Asian −0.045-0.045 −0.047-0.047 −0.018-0.018
Woman ×\times Asian −0.023-0.023 −0.039-0.039 −0.004-0.004
Man ×\times Black +0.037+0.037 −0.038-0.038 +0.011+0.011
Woman ×\times Black +0.060+0.060 −0.030-0.030 +0.018+0.018
Table 5: Group deviation from the pooled human mean and the rating shift induced by conditioning for all models (Δ\Delta), by task and gender×\timesrace profile (mean rating on the [0,1] scale described in §3.2). Faithful conditioning would give Δ≈\Delta\approx Human.

Appendix F Controls for Distributional Spread

Scoring predictions introduces a sensitivity that a point-estimate comparison never faces: both sides of the comparison can be inflated by spread alone. A reference distribution built from many annotators covers several scale points, becoming a wide target that almost any prediction lands close to; a spread-out prediction, in turn, overlaps part of any reference by construction. A judge can therefore score higher on a group without simulating it any better. Because this risk is specific to distributional scoring, we introduce two controls that isolate and remove it: density matching equalizes the human reference, and a mode-accuracy reading equalizes the prediction. Effects that survive both are properties of the judge, not of distribution width. A final check (§F.3) confirms that the models are placing probability mass on tokens relevant to the question rather than elsewhere.

F.1 Density Matching

Groups differ in how many annotators rate each item. A reference built from five annotators spreads over several scale points; one built from a single annotator is a point, which a prediction either matches or misses. To measure how much of the alignment gaps this width difference explains, we replay every comparison with references of equal width: for each item we retain one random annotation per group, rebuild the references from those single labels, and recompute the gaps, averaging over 2020 random draws. We match at one annotation because many Black and Asian cells contain only one. Under matching (Figure 12), the base-pool gaps shrink substantially and a few change sign, indicating that most of the base pool’s raw gaps reflect annotator counts rather than judge behavior. The instruct gaps also shrink but remain positive for race in every model, and reverse for education in only two. Gap magnitudes are therefore partly a property of the data, whereas the instruct pool’s orderings are a property of the judges; this is why §4.1 reports the orderings but does not interpret their sizes. The control thus separates a data artifact from a model effect that a single-target comparison could not have told apart.

F.2 Mode Accuracy

The second control removes the corresponding advantage on the prediction side. We recompute every comparison under mode accuracy: on each item the judge is credited only when its single most probable option is also the most frequent human label (ties included), on the same items and estimator as before. Because this reading compares only the top choice on each side, a spread-out prediction gains nothing from its spread.

Figure 13 compares the group gaps under the two readings, and the pattern mirrors density matching. In the base pool the gaps largely close and a few change sign, confirming that most base-pool gaps come from spread. In the instruct pool the gaps do not close; for race they are mostly larger under mode accuracy than under ss, and every instruct model remains positive under both readings. The instruct gaps therefore do not rest on distribution width.

Mode accuracy also tests whether the base-model conditioning gain of §4.2 is an artifact of near-flat predictions. It is not: under mode accuracy the gain doubles, from +0.014+0.014 to +0.032+0.032, and stays positive in 10 of 11 base models, so conditioning changes which option a base model selects, toward the option the group chose. The instruct effect remains near zero under both readings. The per-group result of §4.3 likewise holds: 14 of the 16 retained groups keep their sign, except for College degree, whose gain appears only under ss. The two controls agree at the model level: both leave every instruct contrast standing and attenuate the base contrasts, with only Gemma 4B, OLMo 7B failing both. We also confirm that the conditioning effects are not driven by reference noise: we resample annotators within cells jointly with judges (2,0002{,}000 replicates). The effects are mostly unchanged (Black −0.033-0.033, [−0.048,−0.020][-0.048,-0.020]; Woman×\timesBlack on offensiveness −0.087-0.087, [−0.132,−0.051][-0.132,-0.051]), and single-annotator cells, which contribute no reference noise, are already covered by the density matching.

F.3 Answer-Token Coverage

Finally, we verify that the predicted distribution reflects a genuine judgment rather than an evasion. Because pp is a softmax over the five option tokens, a model that declined to answer would still yield a distribution over the scale. In practice the option tokens carry at least 96.7%96.7\% of the next-token mass in every base model and 99.7%99.7\% in every instruction-tuned one, and on offensiveness, where refusal is likeliest, adding a demographic profile shifts that share by under 10−410^{-4}. The effects we report are therefore changes in how the models judge, not changes in whether they answer.

Refer to caption
Figure 12: Density matching per model (averaged over 2020 draws in filled bars).
Refer to caption
Figure 13: Group gaps under the two readings.