LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts
Abstract
Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator’s demographic profile to align its judgments with the corresponding group’s. We test whether this alignment emerges distributionally, comparing the predicted label distributions of 23 open-weight LLMs on three subjective tasks against those of real annotator groups, under three conditions: no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education. Three findings emerge. First, a judge prompted with no demographics is not perspective-neutral: models best reproduce the judgments of White, college-educated annotators. Second, demographic conditioning is asymmetric: it moves the judge toward majority groups and away from minority groups, most strongly on offensiveness, where intersectional profiles amplify the harm. Third, by comparing base and instruct models we identify instruction-tuning as a possible source of the asymmetry. Demographic conditioning should therefore be used with caution to estimate group judgments: conditioning moves predictions away from the reference distributions of the minority groups the method is often invoked to serve.
Warning: This work contains unobfuscated examples that some readers may find offensive
1 Introduction
LLMs are increasingly employed as automated judges: they assess text in place of human annotators to evaluate NLP systems and produce annotated data at scale (Li et al., 2025; Gu et al., 2026; Amin et al., 2026; Xu et al., 2026). Since human judgments vary partly with annotators’ sociodemographic background (Pei and Jurgens, 2023; Diaz et al., 2018), sociodemographic prompting has emerged as a way to simulate groups of annotators: conditioning the judge on a demographic profile, a form of persona prompting (Chen et al., 2024b; Tseng et al., 2024), should steer it to answer as an annotator from that group would. Yet such simulation is only as valid as its agreement with human judgments, which varies sharply across tasks (Bavaresco et al., 2025). This is especially delicate for subjective judgments, where no single ground truth exists: annotators disagree, and this disagreement should be treated as signal rather than noise (Pavlick and Kwiatkowski, 2019; Plank, 2022). This is the concern of pluralistic alignment: a model that reflects average human preferences may still fail to reflect the judgments of specific groups (Sorensen et al., 2024; Feng et al., 2024). Thus, for subjective tasks, what matters is not only how accurate an LLM judge is, but whose judgments it reproduces.
Recent analyses on sociodemographic prompting report that effects are mixed, prompt-sensitive, and sometimes negative (Beck et al., 2024; Sun et al., 2025; Gupta et al., 2024), and demographic labels predict a rater’s judgments worse than the rater’s own past annotations or values (Orlikowski et al., 2025; Sorensen et al., 2025). These analyses typically measure how well a judge simulates labels and assess conditioning by a single (averaged) effect per model. What remains untested is whether conditioning moves the model’s label distribution toward that of the target annotator group, rather than toward a stereotype.
We address this gap distributionally, adopting a perspectivist view of sociodemographic prompting (Cabitza et al., 2023; Basile et al., 2021; Frenda et al., 2024): the object of evaluation is the distribution of human judgments for each group rather than an aggregated gold label (Figure 1). Following Santurkar et al. (2023), we ask whose judgments LLMs reflect through three research questions. RQ1: with no demographic information, do LLM judges tend to a default demographic profile? RQ2: does conditioning on (intersectional) demographic profiles move the judge toward those groups’ judgments? RQ3: is demographic conditioning’s effect equitable across groups?
We evaluate 23 open-weight LLMs as judges on three subjective tasks (politeness, intimacy, and offensiveness), comparing unconditioned judges against judges conditioned on single-attribute and intersectional profiles, each scored against the matching human group’s label distribution. Unconditioned judges tend to align with White, college-educated annotators. Averaged across groups, demographic conditioning apparently has little effect on instruction-tuned judges. However, this hides opposite effects across groups, moving the judge closer to some and further from others, especially those disadvantaged by known LLM biases (Navigli et al., 2023; Sap et al., 2022; Santy et al., 2023), here grouped under the term minority groups. Larger models are not immune, and the asymmetry is specific to instruction-tuned judges.
Our contributions are: (i) a framework for evaluating demographic conditioning of LLM judges against annotator groups’ label distributions, rather than against their mean or majority label; (ii) a per-group decomposition of that effect, revealing biases in both unconditioned and conditioned models; (iii) a matched base-vs-instruct comparison identifying instruction-tuning as a source of bias in demographic conditioning.
2 Related Work
LLM-as-a-Judge.
LLMs are increasingly used to evaluate text, from scoring open-ended generations to ranking responses (Li et al., 2025; Gu et al., 2026; Zheng et al., 2023; Wang et al., 2024c). Judge–human agreement varies sharply across tasks (Bavaresco et al., 2025), including under persona conditioning (Dong et al., 2024), and verdicts are sensitive to surface factors, such as response order (Wang et al., 2024a) and human-like response biases (Basile et al., 2021; Tjuatja et al., 2024; Chen et al., 2024a). Crucially, prior work validates a judge against a single target, such as a mean score or majority label (Gordon et al., 2022; Mostafazadeh Davani et al., 2022), which does not exist for subjective tasks: annotator groups disagree systematically, so a judge matching the aggregate may diverge from every group. We therefore evaluate alignment with each group’s label distribution, assessing the judge as an estimator of group judgments rather than a rater of quality.
Persona-Based Generation.
LLMs can be prompted to role-play specific people or groups, and some models are fine-tuned to better capture individual preferences or personas (Horton et al., 2023; Shao et al., 2023; Occhipinti et al., 2024). Assigning a persona (a role, identity, or demographic profile) is widely used to steer model behavior, assuming the model adopts the assigned perspective faithfully (Argyle et al., 2022; Chen et al., 2024b; Tseng et al., 2024). This assumption does not always hold: rather than reproducing a group’s genuine perspective, a persona may instead activate the stereotypes the model associates with that group, degrading performance for some identities (Gupta et al., 2024; Agnew et al., 2024; Dong et al., 2024) or increasing toxic output (Deshpande et al., 2023). Conditioning a judge on demographics could therefore move its predictions toward a group’s real judgments or move them away. We investigate the effect of conditioning on alignment with real label distributions.
Demographic Simulation with LLMs.
Prior work uses LLMs to simulate human populations (Argyle et al., 2022; Mou et al., 2026), showing that unconditioned models can align more closely with some groups than others (Santurkar et al., 2023; Durmus et al., 2023). Such bias is also visible on broader demographic imbalances in NLP datasets and models (Santy et al., 2023; Alipour et al., 2025). Demographic conditioning has therefore been studied as a way to recover group-specific judgments, but its effectiveness is fragile and sensitive to the evaluation setup (Simmons and Savinov, 2024; Adilazuarda et al., 2024; Gao et al., 2025), motivating calls for principled evaluation (Luz de Araujo et al., 2025; Neumann et al., 2026). Recent work has compared demographic prompting with richer ways of modeling annotators, including fine-tuning on past judgments (Orlikowski et al., 2025) and conditioning on elicited values (Sorensen et al., 2025). These studies reduce in-group variability to a single aggregate label. The distributional alternative has been developed on opinion surveys (Meister et al., 2025; Lutz et al., 2025): on annotation, where the reference is disagreement about the same text, judges are still scored against a group’s mean (Sun et al., 2025; Schäfer et al., 2025) or against individual annotations (Hu and Collier, 2024). We instead ask how well models reproduce each demographic group’s label distribution, rather than its mean, and whether assigning a demographic profile moves the judge closer to or further from the judgments of the group named in the prompt.
3 Methodology
This section describes the data used (§3.1), the evaluation metrics (§3.2), the generation configurations and analysis (§3.3), and models (§3.4).
3.1 Data
Our study requires parallel subjective judgments: multiple ratings of the same text, each linked to the demographic profile of the annotator who produced it. We therefore leverage the DeMo dataset (Orlikowski et al., 2025) which combines five annotator-level corpora and maps self-reported demographics onto four dimensions: gender, age, race, and education. Since each annotation is aligned with the annotator’s profile, we can construct the gold label distribution of judgments for any text and demographic group.
We retain three of the five tasks in DeMo, all rated on a five-point ordinal scale:11 1 We exclude safety (DICES-350; Aroyo et al., 2023), which uses a three-way categorical scale and generational cohorts rather than age brackets, and sentiment (Diaz et al., 2018), which recruited only annotators over .
- •
Intimacy: rating how intimate a Twitter post is, from not intimate at all to very intimate (MINT; Pei et al., 2023);
- •
Offensiveness: rating how offensive a Reddit comment is, from not offensive at all to very offensive (POPQUORN; Pei and Jurgens, 2023);
- •
Politeness: rating how polite a workplace email is, from not polite at all to very polite (POPQUORN).
Table 1 shows an example of an offensiveness entry and the corresponding judgment distribution of two annotator groups (men and women).
| Text: “thats fucking hilarious, but also sad at the same time because I know it won’t change anything. Here’s to hoping these laws get struck down, and old pseudo-religious trying to win political brownie points men stop trying to tell women what they can and can’t do with their bodies” | ||||||
|---|---|---|---|---|---|---|
| A | B | C | D | E | ||
| Humans | ||||||
| Men () | ||||||
| Women () | ||||||
| Llama 70B | ||||||
| Unconditioned | — | |||||
| Man profile | ||||||
| Woman profile | ||||||
For our analysis, we remove categories too sparsely represented to support a stable per-group label distribution. The complete filtering procedure is described in Appendix A. Statistics about the resulting dataset are reported in Table 2, whereas the demographic categories retained after filtering are reported in Table 3.
| Task | Genre | Texts | Judgments | Annotators |
|---|---|---|---|---|
| Intimacy | tweets | 1,992 | 11,355 | 237 |
| Offensiveness | comments | 1,500 | 12,489 | 251 |
| Politeness | emails | 3,718 | 23,896 | 483 |
| Total | 7,210 | 43,402 | 971 |
| Dimension | Values | ||
|---|---|---|---|
| Gender (2) | Man, Woman | ||
| Age (8) |
| ||
| Race (3) |
| ||
| Education (3) |
|
3.2 Evaluation Metrics
Distributions as targets.
We evaluate a judge against the full distribution of human labels, at the level of demographic groups. Let be a text and a demographic profile, i.e. a combination of values for one or two of the four dimensions. The human reference is the distribution over the scale points of the labels that annotators matching assign to . Each generation configuration yields a model distribution over the same points, from the log-probabilities the model assigns to the five answer tokens. Evaluating a model thus consists in measuring, for each pair, how close is to (§3.3).
Table 1 shows why we compare distributions rather than aggregated labels. Mapping the ordinal points A–E to equally spaced values in (i.e., ), we can summarize a distribution by its mean rating, its expected value under this mapping. The two groups in the example have the same mean rating (), so an averaged label would make the judge equally close to both. However, the distributions differ: men concentrate on the middle of the scale, women split between its two ends. Unlike the score defined below, the mean rating treats the scale as interval. We use it only to report the direction and magnitude of conditioning shifts.
Comparing ordinal distributions.
Because our labels are ordinal, the discrepancy between two distributions should grow with the distance over which probability mass is misplaced. Divergence measures such as KL treat labels as nominal, penalizing misplaced mass equally regardless of where in the scale it lands. We therefore use the Earth Mover’s Distance (EMD) (Rubner et al., 2000), the probability mass that must be moved to turn one distribution into the other, weighted by the ordinal distance moved:22 2 EMD is used in NLP for subjective tasks, from document similarity (Kusner et al., 2015) to aligning model and human response distributions on ordinal survey scales (Santurkar et al., 2023; Tjuatja et al., 2024).
Following Santurkar et al. (2023), we report it as a bounded similarity score, inverted so that higher is better:
Because the maximum EMD on a -point scale is , reached when the two distributions place all their mass on opposite extremes, ranges from (maximal divergence) to (perfect match) and is comparable across tasks and configurations. To isolate the effect of demographic conditioning, we score both the unconditioned and the conditioned prediction (see §3.3) for the same text against the same group reference:
where is the model’s prediction for without a profile. While measures how closely a judge matches a group, isolates the effect of the profile alone: means conditioning moved the judge toward the group, that it moved the judge away from the group it was instructed to represent.
Aggregation and uncertainty.
Both and can be affected by the shape of the prediction and reference distributions: a more spread-out prediction tends to score better against a more spread-out reference, even without better group alignment. This can distort comparisons in two ways: (i) on the human labels side, groups can have different numbers of annotators per item, so some reference distributions are more sparse than others; (ii) on the model side, instruction-tuning is known to sharpen answer-token distributions (Santurkar et al., 2023; Durmus et al., 2023; Sorensen et al., 2024), so any comparison of base and instruct model variants risks being an artifact of sharpness. We therefore validate our results against mode accuracy, which is unaffected by distribution sharpness. We average scores within each cell (a model, configuration, task, and demographic group) and then macro-average cell means with equal weight, preventing groups with more observations from dominating the average. Confidence intervals are percentile bootstraps over cells (2,000 resamples). We validate all group comparisons against two independent spread-insensitive controls: density-matched references and mode accuracy. Mode accuracy also mitigates the concern that first-token probabilities may diverge from the model’s generated answer (Wang et al., 2024b). We verify token coverage but do not compare against generated text (Appendix F).
3.3 Experimental Design
We elicit a judgment from a model in a single forward pass and read off the full distribution it places on the five scale points, rather than sampling a discrete answer. We follow the prompt format of Orlikowski et al. (2025): each prompt presents the task question, the text to be judged, and the five labeled options (A)–(E) in order, and ends with an assistant-turn prefix constraining the model’s next token to be one of the option letters (Figure 1).33 3 The order is fixed, but the models show no choice-position bias (Zheng et al., 2024): pooled over the 23 models, option mass tracks the human labels () rather than letter order, and the modal option (A) against for (B)) is also the modal human label . All results in §4.3 use this single template, since conditioned and unconditioned predictions must be scored under the same wording to be paired. Because prompt format and wording could play a relevant role in the distributions (Beck et al., 2024), we also run all models under a structurally different template, an interview-style profile in which the annotator states their own demographics (Lutz et al., 2025) (see Appendix B).
Following prior work that reads judgment distributions directly from answer-token probabilities (Santurkar et al., 2023; Durmus et al., 2023), we take the predicted distribution to be the softmax over the log-probabilities of the five option tokens A–E.44 4 We retrieve the top-20 next-token log-probabilities and read each option’s first token, assigning to any option absent from that set. The configurations below differ only in (i) the demographic profile, if any, prepended to the prompt, and (ii) the human reference against which each prediction is scored (§3.2).
Unconditioned Baseline.
In this configuration the prompt carries no demographic information (Figure 1, without the profile line), thus the model judges without being assigned a perspective. For each text, the baseline prediction is compared with the human label distribution of every demographic group that annotated it. This unconditioned setting (i) provides the reference point for score difference (§3.2), isolating the effect of adding a demographic profile to the prompt, and (ii) reveals possible models’ default demographic alignment (i.e., the perspective it tends to adopt in the absence of demographic conditioning).
Demographic Conditioning.
We prompt the model to judge the text from the perspective of a specific demographic group, i.e. a profile (Figure 1), and score its prediction against that group’s human label distribution. A profile assigns a value to one or two of the four dimensions of §3.1: gender, age, race, and education. We consider the four single-attribute configurations and the six two-attribute combinations.55 5 We stop at two attributes, since narrower profiles leave too few annotators per cell to estimate a stable reference. The former isolate each dimension’s contribution, the latter test whether intersectional profiles help or hurt. Across the ten configurations this yields 268K pairs per model. For each, we compute twice against the same reference (once for the conditioned prediction , once for the unconditioned ) and take their difference as .
3.4 Models
We evaluate models from five families, ranging from 4B to 70B parameters: Gemma 3 4B, 12B, and 27B (Team et al., 2025); Llama 3.1 8B and 70B (Grattafiori et al., 2024); Ministral 3 8B and 14B (Liu et al., 2026); OLMo 3 7B and 32B (Olmo et al., 2026); and Qwen 3 8B, 14B, and 32B (Yang et al., 2025). We evaluate models in both base and instruction-tuned form,66 6 Instruction-tuned models receive the prompt in their chat template, with the final assistant turn left open at the answer prefix. Base models receive the same messages as plain concatenation, so each is prompted in its native format. except Qwen 3 32B, which is available only instruction-tuned, for 23 models in total. Variation in size within each family lets us study scale while keeping the model family fixed. The 11 matched base–instruct pairs isolate the effect of instruction-tuning.
4 Results and Discussion
We organize the results around the three research questions of §1.
4.1 RQ1: Default Profiles of LLM-Judges
In line with Schäfer et al. (2025), who found judges furthest from Black annotators and unaffected by gender on two proprietary models, we find a shared default profile in open-weight models’ label distributions.
All models align best with White and college-educated annotators, equally well with men and women, and differ from one another in how they relate to age.
Figure 2 reports the alignment of the unconditioned predictions with each demographic group’s label distributions, one panel per dimension. Race shows the largest gaps: all base models and the majority of instruct models match White annotators more closely than Black and Asian ones, with an average White–Black gap of for base models and for instruct models.
For base judges, this corresponds to moving more than one third of the probability mass by one response step away from Black annotators’ judgments.77 7 The ordering is robust to annotator-count differences for instruction-tuned models, and holds for most base models under matching, even as the gap sizes shrink (Appendix F.1). Such group-level differences are visible only under a distributional comparison, while scoring against a single aggregate label would collapse (§3.2). Education is the most consistent dimension: pooled across tasks, every model is closest to college-educated annotators, ahead of annotators with a high-school education. Gender, by contrast, shows no default: the markers largely overlap, with mean absolute gaps below . Age is the only dimension on which models disagree, and the split tracks model type: base models are generally closest to the oldest group, whereas most instruct models are closest to the youngest. This is a first indication that instruction-tuning changes whose perspective a judge adopts even before any profile is assigned. We investigate the impact of instruction-tuning further in §4.2. These defaults are largely stable across tasks: race and gender patterns hold on all three, while age (for all models) and education (for instruct models only) vary more by task (see Appendix C). Appendix F verifies that these orderings are not artifacts of distributional spread.
An unconditioned judge is not a neutral annotator.
LLMs are often used as substitutes for human annotators, but an unconditioned model should not be treated as an average annotator. Without any demographic profile, every instruction-tuned judge aligns more closely with White than Black or Asian annotators, and with college-educated than less-educated annotators. Thus, using an unconditioned judge as a generic annotator introduces a systematic demographic perspective. This confirms, on open-weight judges and distributional scoring, the conclusion of Schäfer et al. (2025) that LLMs do not represent all social groups equally. The analyses that follow ask what happens when a profile is assigned: whether conditioning corrects it or compounds it.
4.2 RQ2: The Effect of Demographic Conditioning
Demographic conditioning helps base models, but not instruct models (on average).
Figure 3 compares the conditioned and unconditioned predictions of each judge against the same group references. Assigning a profile improves 10 of 11 base models, while among the 12 instruct models it helps five, leaves two unchanged, and hurts the rest: on average, conditioning appears to have little to no effect on instruction-tuned models.
The base model gain is not an artifact of spread-out predictions.
Base models distribute their probability more evenly across the five options than instruct models, which tend to concentrate it on one or two.88 8 Measured as entropy normalized by its maximum , so that is a one-point distribution and a uniform one: on average for base models, for the instruct. This asymmetry admits an alternative explanation for the base model gain: a spread-out prediction necessarily overlaps part of any reference distribution, so conditioning could improve base scores simply by reshaping an almost-flat prediction, without moving the judge toward the group’s actual judgments. We rule out this explanation using mode accuracy, which ignores probability spread and gives credit only when the model’s most likely option matches the group’s majority label. Under this metric, the mean conditioning gain for base models increases from to and remains positive for of models (see Appendix F.2). Instruct models remain near zero under both metrics.
Neither scale nor instruction-tuning makes judges more steerable.
Larger models align better with annotators: within the instruction-tuned pool, between log parameter count and , and instruct models start ahead of the base ones in the unconditioned setting. Yet neither gain carries over to conditioning: log parameter count and are essentially uncorrelated across the 23 judges (), and the largest gains fall on the base Qwen models regardless of scale (Qwen 8B , Qwen 14B ).
Better judges, however, are not more steerable judges. A judge can be wrong about a text in two ways: it can misjudge the text itself, which puts it off for every group at once, or it can miss the way one group departs from the others. Instruction-tuning only improves the first, and every group’s score rises accordingly.
The instruct models’ null effect is not uniform across tasks.
For base models, conditioning helps on all three tasks. For instruct models, the is below on politeness and intimacy but turns negative on offensiveness ( on average, down to for Ministral 14B). Appendix D reports the per-model effects: on the one task where demographic perspective shows the greatest variability, assigning a profile makes several instruct judges worse at representing the group they are instructed to represent.
However, the almost flat average conceals relevant per-group behavior, both positive and negative (Simpson, 1951; Blyth, 1972). The next section decomposes it, asking which groups conditioning helps and which it hurts.
4.3 RQ3: Conditioning Helps Majority Groups and Hurts Minority Groups
Instruction-tuning makes demographic conditioning uneven across groups.
Figure 4 shows the conditioning effect separately for each demographic group. For base models, conditioning is positive for every group: assigning a profile moves the judge toward that group’s judgments regardless of which group it is. With instruction-tuned models, the effect splits: conditioning remains positive for White, Man, College degree, and High school, but becomes negative for Asian, Woman, Graduate degree, and Black. The largest harm falls on Black annotators (): instructing the judge to answer as a Black annotator makes it less aligned with Black annotators’ judgments than giving it no profile at all.
Intersectional profiles do not combine additively: the joint effect is attenuated toward the gender marginal, landing above the sum of the two single-attribute effects. Man White benefits (), while against Black, the near-zero Man effect dilutes the harm (Man Black ) while Woman deepens it (Woman Black ).
The harm to minority profiles is concentrated on offensiveness.
Figure 5 splits the same breakdown by task. Unlike the defaults of §4.1 and the pooled effects of §4.2, which apply similarly to all three tasks, the harm to minority groups is more evident in offensiveness. On offensiveness, conditioning an instruct judge reduces alignment with Black annotators by , with Asian annotators and with women by . Every gender race profile containing a minority attribute is negative, with Woman Black at , the largest effect in the study. The pattern is weaker on intimacy, where Black is the only group with a reliable negative effect (), and absent on politeness, where no minority profile is harmed. Base models show none of this task specificity: their conditioning effect is positive for nearly every group on all three tasks. Thus, the asymmetry is not a general property of all LLMs under demographic conditioning. It is rather introduced by instruction-tuning and expressed on the task where demographic perspective is most contested.
Prompting moves the judge toward a stereotype, not toward the group.
| Group | Human | base | instruct |
|---|---|---|---|
| Offensiveness | |||
| White | |||
| Black | |||
| Asian | |||
| Man | |||
| Woman | |||
| Politeness | |||
| White | |||
| Black | |||
| Asian | |||
| Man | |||
| Woman | |||
| Intimacy | |||
| White | |||
| Black | |||
| Asian | |||
| Man | |||
| Woman | |||
Demographic conditioning is effective when it moves the judge’s rating by the amount the group actually differs from the mean of all annotators. The two shifts are reported in Table 4. Two patterns emerge. On offensiveness, Black annotators rate above the mean, yet the Black profile moves instruct judges by , about more than the real difference: conditioning amplifies the gap rather than reproducing it. On politeness, Asian is the only group rating below the mean (), yet its profile shifts judges upward ( instruct, base), opposite to the group it names. As the table shows, the shift is largely insensitive to the group’s actual deviation: on offensiveness it is upward for instruct judges regardless of the group, and only base judges, which overestimate offensiveness unconditioned, move back toward the human ratings. In general, the average shift is upward, whether the group differs from the mean or not: the profile triggers the expectation that the named group perceives more offensiveness, not that group’s actual judgments. This confirms the stereotype activation reported for persona-prompted generation (Gupta et al., 2024; Deshpande et al., 2023), and quantifies it, as against real annotator distributions, the shift can be measured against the group’s actual deviation. Intersectional profiles show the same pattern as the corresponding single-attribute ones. We report them in Appendix E.
Gender shows that judges shift even when there is no group-related difference.
As shown in Table 4, gender is the one dimension with no default (§4.1) and no group deviation to reproduce: women rate at the pooled mean (e.g., in offensiveness). A faithful judge conditioned on Woman would therefore not move. Instead the profile shifts instruction-tuned judges upward (e.g., in offensiveness). Here, on the contrary, the distributions unduly move because of the models’ stereotyped expectations of the group.
Conditioning shifts ratings, not perspectives.
Figure 6 shows the correlation between shifts in the label distributions and . On offensiveness, the more conditioning raises a judge’s rating, the more alignment it loses (, ). The same trend holds when comparing judges given the same profile (within-profile ).
Base models tend to move downward, toward the human ratings, and therefore improve; instruct models tend to move upward, away from them, and therefore worsen. Intimacy shows the same pattern. Politeness reverses the direction. Judges initially underrate politeness, predicting on average against a human mean of . Conditioning tends to push their ratings upward, so larger shifts improve alignment (). Yet this does not mean the judges recover each group’s perspective: only 5 of the 11 profiles shift the judge in the same direction as the group’s actual deviation from the overall human mean. The gain therefore comes mainly from correcting the judges’ general underestimation of politeness, not from reproducing group-specific judgments.
Across all three tasks, the mechanism is the same: demographic conditioning pushes all predictions in a direction that is not reliably tied to the group being represented. It helps when that push happens to move the judge toward the human ratings, and hurts when it moves the judge away.
5 Conclusions
Across 23 open-weight LLM judges, we asked whether conditioning a judge on an annotator’s demographic profile brings its judgments closer to that group’s. An unconditioned judge is not perspective-neutral: every model aligns more closely with White than with Black or Asian annotators, and with college-educated annotators than with either other band. Conditioning an instruction-tuned judge appears to have no effect on average. However, that average hides gains for the groups the judge already leans toward, and losses for minority groups. The harm concentrates on offensiveness, is sharpened rather than repaired by intersectional profiling, and is not recovered by model size. Comparing each instruction-tuned judges with their base counterparts locates the asymmetry in instruction-tuning: conditioning positively impacts base models for every group, whereas with instruction-tuned models the profile does not activate the group’s perspective but an expectation about it, exaggerating differences where they exist and introducing them where they do not. Thus the practical question is not simply whether demographic conditioning helps, but whom it helps: its benefits are smallest for the very groups such interventions are meant to represent.
Limitations
Our findings rest on three English tasks from one annotator-level corpus (Orlikowski et al., 2025), so we cannot claim the pattern transfers to other languages or annotation schemes. Those corpora are public, and any leakage into pretraining would favor the alignment we measure. Per-group distributions thin as profiles narrow, so we condition on at most two attributes, and minority groups carry fewer parallel annotations (Difallah et al., 2018; Prabhakaran et al., 2021).
Ethics Statement
No new annotation was collected to conduct the analysis described in this work and no annotator is identifiable. We use coarse demographic categories only to define the group reference distributions, not as proxies for individual perspectives. Our results caution against that stronger reading, since conditioning helped the groups judges already favour while harming the rest. Thus, demographic prompting should not be treated as a substitute for recruiting annotators from the groups of interest.
References
- Towards measuring and modeling “culture” in LLMs: a survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 15763–15784. External Links: Link, Document Cited by: §2.
- The illusion of artificial inclusion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA. External Links: ISBN 9798400703300, Link, Document Cited by: §2.
- Robustness and confounders in the demographic alignment of LLMs with human perceptions of offensiveness. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 22025–22047. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.
- From fallback to frontline: when can LLMs be superior annotators of human perspectives?. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 42798–42830. External Links: Link, ISBN 979-8-89176-395-1 Cited by: §1.
- Out of one, many: using language models to simulate human samples. Political Analysis 31, pp. 337 – 351. External Links: Link Cited by: §2, §2.
- DICES dataset: diversity in conversational ai evaluation for safety. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, pp. 53330–53342. External Links: Link Cited by: footnote 1.
- We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, K. Church, M. Liberman, and V. Kordoni (Eds.), Online, pp. 15–21. External Links: Link, Document Cited by: §1, §2.
- LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 238–255. External Links: Link, Document, ISBN 979-8-89176-252-7 Cited by: §1, §2.
- Sensitivity, performance, robustness: deconstructing the effect of sociodemographic prompting. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Y. Graham and M. Purver (Eds.), St. Julian’s, Malta, pp. 2589–2615. External Links: Link, Document Cited by: §1, §3.3.
- On Simpson’s paradox and the sure-thing principle. Journal of the American Statistical Association 67 (338), pp. 364–366. External Links: Document Cited by: §4.2.
- Toward a perspectivist turn in ground truthing for predictive computing. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’23/IAAI’23/EAAI’23. External Links: ISBN 978-1-57735-880-0, Link, Document Cited by: §1.
- Humans or LLMs as the judge? a study on judgement bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 8301–8327. External Links: Link, Document Cited by: §2.
- From persona to personalization: a survey on role-playing language agents. Transactions on Machine Learning Research. Cited by: §1, §2.
- Toxicity in chatgpt: analyzing persona-assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 1236–1270. External Links: Link, Document Cited by: §2, §4.3.
- Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI ’18, New York, NY, USA, pp. 1–14. External Links: ISBN 9781450356206, Link, Document Cited by: §1, footnote 1.
- Demographics and dynamics of mechanical turk workers. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, WSDM ’18, New York, NY, USA, pp. 135–143. External Links: ISBN 9781450355810, Link, Document Cited by: Limitations.
- Can LLM be a personalized judge?. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 10126–10141. External Links: Link, Document Cited by: §2, §2.
- Towards measuring the representation of subjective global opinions in language models. External Links: 2306.16388 Cited by: §2, §3.2, §3.3.
- Modular pluralism: pluralistic alignment via multi-LLM collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 4151–4171. External Links: Link, Document Cited by: §1.
- Perspectivist approaches to natural language processing: a survey: perspectivist approaches to natural language processing…. Lang. Resour. Eval. 59 (2), pp. 1719–1746. External Links: ISSN 1574-020X, Link, Document Cited by: §1.
- Take caution in using llms as human surrogates. Proceedings of the National Academy of Sciences 122 (24), pp. e2501660122. External Links: Document, Link, https://www.pnas.org/doi/pdf/10.1073/pnas.2501660122 Cited by: §2.
- Jury learning: integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY, USA. External Links: ISBN 9781450391573, Link, Document Cited by: §2.
- The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §3.4.
- A survey on llm-as-a-judge. The Innovation 7 (6), pp. 101253. External Links: ISSN 2666-6758, Document, Link Cited by: §1, §2.
- Bias runs deep: implicit reasoning biases in persona-assigned llms. In International Conference on Learning Representations, Vol. 2024, pp. 21849–21874. Cited by: §1, §2, §4.3.
- Large language models as simulated economic agents: what can we learn from homo silicus?. Technical report National Bureau of Economic Research. Cited by: §2.
- Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 10289–10307. External Links: Link, Document Cited by: §2.
- From word embeddings to document distances. In Proceedings of the 32nd International Conference on Machine Learning - Volume 37, ICML’15, pp. 957–966. Cited by: footnote 2.
- From generation to judgment: opportunities and challenges of LLM-as-a-judge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2757–2791. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1, §2.
- Ministral 3. External Links: 2601.08584, Link Cited by: §3.4.
- The prompt makes the person(a): a systematic evaluation of sociodemographic persona prompting for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 23212–23237. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: Appendix B, §2, §3.3.
- Principled personas: defining and measuring the intended effects of persona prompting on task performance. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26857–26886. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2.
- Benchmarking distributional alignment of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 24–49. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2.
- Dealing with disagreements: looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics 10, pp. 92–110. External Links: Link, Document Cited by: §2.
- From individual to society: a survey on social simulation driven by large language model-based agents. ACM Comput. Surv. 58 (11). External Links: ISSN 0360-0300, Link, Document Cited by: §2.
- Biases in large language models: origins, inventory, and discussion. J. Data and Information Quality 15 (2). External Links: ISSN 1936-1955, Link, Document Cited by: §1.
- Should you use llms to simulate opinions? quality checks for early-stage deliberation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 39070–39079. Cited by: §2.
- PRODIGy: a PROfile-based DIalogue generation dataset. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 3500–3514. External Links: Link, Document Cited by: §2.
- Olmo 3. External Links: 2512.13961, Link Cited by: §3.4.
- Beyond demographics: fine-tuning large language models to predict individuals’ subjective text perceptions. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 2092–2111. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1, §2, §3.1, §3.3, Limitations.
- Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics 7, pp. 677–694. External Links: Link, Document Cited by: §1.
- When do annotator demographics matter? measuring the influence of annotator demographics with the POPQUORN dataset. In Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII), J. Prange and A. Friedrich (Eds.), Toronto, Canada, pp. 252–265. External Links: Link, Document Cited by: §1, 2nd item.
- SemEval-2023 task 9: multilingual tweet intimacy analysis. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), A. Kr. Ojha, A. S. Doğruöz, G. Da San Martino, H. Tayyar Madabushi, R. Kumar, and E. Sartori (Eds.), Toronto, Canada, pp. 2235–2246. External Links: Link, Document Cited by: 1st item.
- The “problem” of human label variation: on ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 10671–10682. External Links: Link, Document Cited by: §1.
- On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, C. Bonial and N. Xue (Eds.), Punta Cana, Dominican Republic, pp. 133–138. External Links: Link, Document Cited by: Limitations.
- The earth mover’s distance as a metric for image retrieval. International Journal of Computer Vision 40 (2), pp. 99–121. External Links: ISSN 1573-1405, Document, Link Cited by: §3.2.
- Whose opinions do language models reflect?. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. Cited by: §1, §2, §3.2, §3.2, §3.3, footnote 2.
- NLPositionality: characterizing design biases of datasets and models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 9080–9102. External Links: Link, Document Cited by: §1, §2.
- Annotators with attitudes: how annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M. de Marneffe, and I. V. Meza Ruiz (Eds.), Seattle, United States, pp. 5884–5906. External Links: Link, Document Cited by: §1.
- Which demographics do LLMs default to during annotation?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 17331–17348. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2, §4.1, §4.1.
- Character-LLM: a trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 13153–13187. External Links: Link Cited by: §2.
- Assessing generalization for subpopulation representative modeling via in-context learning. In Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024), A. Deshpande, E. Hwang, V. Murahari, J. S. Park, D. Yang, A. Sabharwal, K. Narasimhan, and A. Kalyan (Eds.), St. Julians, Malta, pp. 18–35. External Links: Link, Document Cited by: §2.
- The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society: Series B (Methodological) 13 (2), pp. 238–241. External Links: Document Cited by: §4.2.
- Value profiles for encoding human variation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2047–2095. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1, §2.
- Position: a roadmap to pluralistic alignment. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. Cited by: §1, §3.2.
- Sociodemographic prompting is not yet an effective approach for simulating subjective judgments with LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 845–854. External Links: Link, Document, ISBN 979-8-89176-190-2 Cited by: §1, §2.
- Gemma 3 technical report. External Links: 2503.19786, Link Cited by: §3.4.
- Do LLMs exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics 12, pp. 1011–1026. External Links: Link, Document Cited by: §2, footnote 2.
- Two tales of persona in LLMs: a survey of role-playing and personalization. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 16612–16631. External Links: Link, Document Cited by: §1, §2.
- Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 9440–9450. External Links: Link, Document Cited by: §2.
- “My answer is C”: first-token probabilities do not match text answers in instruction-tuned language models. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7407–7416. External Links: Link, Document Cited by: §3.2.
- Pandalm: an automatic evaluation benchmark for llm instruction tuning optimization. In International Conference on Learning Representations, Vol. 2024, pp. 43573–43593. Cited by: §2.
- Modeling annotator disagreement with demographic-aware experts and synthetic perspectives. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 41260–41294. External Links: Link, ISBN 979-8-89176-390-6 Cited by: §1.
- Qwen3 technical report. External Links: 2505.09388, Link Cited by: §3.4.
- Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: footnote 3.
- Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §2.
Appendix A Filtering of Under-Represented Categories
As described in §3.1, we remove annotations from demographic categories too sparsely represented in the source data to support reliable group-level comparison. Within each dimension we drop values contributed by too few annotators to estimate a stable per-group label distribution. Counting distinct annotators on the three retained tasks, this removes Non-binary from gender, Hispanic/Latino from race, Less than high school from education, and the two oldest age brackets, – and –, each populated only by a single annotator.
A cell is dropped if any value in its attribute conjunction is excluded, since the cell’s human reference is defined by the full conjunction. Filtering is applied before any scoring, so unconditioned and conditioned predictions are evaluated against references drawn only from retained categories.
Appendix B Prompt Templates
Every task uses the same two-message template: the user message states the task, presents the item, and lists the five options in ascending order, and the assistant message is a continued prefix, so the next token is the option letter. A conditioned prompt prepends the **Your Profile** line, one clause per active attribute in the fixed order race, age, gender, education.
Your Profile: In terms of race or ethnicity, you are White. You are 35 to 39 years old.
Question: Consider you read this comment on Reddit, how offensive do you think it is?
Comment:: <text>
(A) not offensive at all
[…]
(E) very offensive
[assistant]
Answer: (
The unconditioned prompt is the same with the Your Profile line removed, so both conditions put the same question to the model and differ only in whether a profile precedes it. For base models the same messages are concatenated in plain text, without chat-template.
The other two framings ask how intimate a Twitter **Post** is and how polite an **Email:** from a colleague is, with options from not intimate/polite at all to very intimate/polite. All eleven configurations are released with the code as fully rendered prompts.
Robustness to the template.
To test whether the demographic conditioning effects depend on the wording, we repeat the experiment with an interview template in which the model states the profile itself (Lutz et al., 2025). Question, items and options are unchanged; only the lines before them differ.
Interviewer: Before we begin, tell me a little about yourself.
You: In terms of race or ethnicity, I am White.
Interviewer: Thank you. Now please answer the following.
Question: […]
Comment: <text>
(A) not offensive at all
[…]
(E) very offensive
[assistant]
Answer: (
Its unconditioned arm drops the exchange and opens with Interviewer: Please answer the following., so what separates the two arms is the profile alone. Figure 9 compares the conditioning effects across groups under the two templates.
The negative effects do not disappear under the alternative template. For instruct models, conditioning becomes more negative across groups: for Black annotators, for example, changes from to . On average, the self-report template lowers by , making the effect negative for all sixteen groups. Base models are much less sensitive to the change in wording: no group changes by more than , and conditioning remains positive for all sixteen groups. The negative effects are therefore not specific to the wording of the main prompt, and are stronger when the model states the identity itself.
Appendix C Per-Group Alignment of the Unconditioned Judge
This appendix reports the per-model, per-task view behind §4.1: unconditioned predictions only, scored and macro-averaged as in §3.2, over the retained groups. The base pool barely moves across tasks (Figure 10): White is the closest race group and College degree the closest education band in every base model on every task, and the oldest bracket leads in all but one model, except on intimacy, where the closest bracket splits between the youngest and 50–59. The instruct pool varies more. White stays ahead of Black in all models on politeness and intimacy but in on offensiveness, where the mean gap drops to ; Asian is the closest race group in half the models on politeness. Education reverses: College degree is closest in all models on politeness, the split is even on intimacy, and High school or below leads in of on offensiveness. Gender stays nearly flat on every task; age changes closest bracket from task to task.
Appendix D Per-Model Conditioning Effects
This appendix reports the per-model, per-task view behind §4.2. Figure 11 splits the per-model view of Figure 3 by task, with rows in the same order. Two things are worth reading off directly: (i) on every task the hollow pairs shift right more often than the filled ones, and the base mean is positive on all three tasks and exceeds the instruct mean on all three; (ii) the tasks differ in dispersion rather than in direction. Politeness holds the smallest effects in both pools, while the largest sit on offensiveness for instruct models and on intimacy for base models. Offensiveness is the only task whose instruct mean is negative, and since the pooled instruct effect averages the three tasks, it is what places the pooled figure below zero.
Appendix E Rating-Shift Faithfulness for Intersectional Profiles
Table 5 repeats the analysis of 4 for the six genderrace profiles. Each profile shifts in the direction of its race attribute, with magnitudes attenuated toward the gender component. The reference deviations are computed from the annotators matching both attributes, so the intervals are wider than in the single-attribute figure.
| Group | Human | base | instruct |
|---|---|---|---|
| Offensiveness | |||
| Man White | |||
| Woman White | |||
| Man Asian | |||
| Woman Asian | |||
| Man Black | |||
| Woman Black | |||
| Politeness | |||
| Man White | |||
| Woman White | |||
| Man Asian | |||
| Woman Asian | |||
| Man Black | |||
| Woman Black | |||
| Intimacy | |||
| Man White | |||
| Woman White | |||
| Man Asian | |||
| Woman Asian | |||
| Man Black | |||
| Woman Black | |||
Appendix F Controls for Distributional Spread
Scoring predictions introduces a sensitivity that a point-estimate comparison never faces: both sides of the comparison can be inflated by spread alone. A reference distribution built from many annotators covers several scale points, becoming a wide target that almost any prediction lands close to; a spread-out prediction, in turn, overlaps part of any reference by construction. A judge can therefore score higher on a group without simulating it any better. Because this risk is specific to distributional scoring, we introduce two controls that isolate and remove it: density matching equalizes the human reference, and a mode-accuracy reading equalizes the prediction. Effects that survive both are properties of the judge, not of distribution width. A final check (§F.3) confirms that the models are placing probability mass on tokens relevant to the question rather than elsewhere.
F.1 Density Matching
Groups differ in how many annotators rate each item. A reference built from five annotators spreads over several scale points; one built from a single annotator is a point, which a prediction either matches or misses. To measure how much of the alignment gaps this width difference explains, we replay every comparison with references of equal width: for each item we retain one random annotation per group, rebuild the references from those single labels, and recompute the gaps, averaging over random draws. We match at one annotation because many Black and Asian cells contain only one. Under matching (Figure 12), the base-pool gaps shrink substantially and a few change sign, indicating that most of the base pool’s raw gaps reflect annotator counts rather than judge behavior. The instruct gaps also shrink but remain positive for race in every model, and reverse for education in only two. Gap magnitudes are therefore partly a property of the data, whereas the instruct pool’s orderings are a property of the judges; this is why §4.1 reports the orderings but does not interpret their sizes. The control thus separates a data artifact from a model effect that a single-target comparison could not have told apart.
F.2 Mode Accuracy
The second control removes the corresponding advantage on the prediction side. We recompute every comparison under mode accuracy: on each item the judge is credited only when its single most probable option is also the most frequent human label (ties included), on the same items and estimator as before. Because this reading compares only the top choice on each side, a spread-out prediction gains nothing from its spread.
Figure 13 compares the group gaps under the two readings, and the pattern mirrors density matching. In the base pool the gaps largely close and a few change sign, confirming that most base-pool gaps come from spread. In the instruct pool the gaps do not close; for race they are mostly larger under mode accuracy than under , and every instruct model remains positive under both readings. The instruct gaps therefore do not rest on distribution width.
Mode accuracy also tests whether the base-model conditioning gain of §4.2 is an artifact of near-flat predictions. It is not: under mode accuracy the gain doubles, from to , and stays positive in 10 of 11 base models, so conditioning changes which option a base model selects, toward the option the group chose. The instruct effect remains near zero under both readings. The per-group result of §4.3 likewise holds: 14 of the 16 retained groups keep their sign, except for College degree, whose gain appears only under . The two controls agree at the model level: both leave every instruct contrast standing and attenuate the base contrasts, with only Gemma 4B, OLMo 7B failing both. We also confirm that the conditioning effects are not driven by reference noise: we resample annotators within cells jointly with judges ( replicates). The effects are mostly unchanged (Black , ; WomanBlack on offensiveness , ), and single-annotator cells, which contribute no reference noise, are already covered by the density matching.
F.3 Answer-Token Coverage
Finally, we verify that the predicted distribution reflects a genuine judgment rather than an evasion. Because is a softmax over the five option tokens, a model that declined to answer would still yield a distribution over the scale. In practice the option tokens carry at least of the next-token mass in every base model and in every instruction-tuned one, and on offensiveness, where refusal is likeliest, adding a demographic profile shifts that share by under . The effects we report are therefore changes in how the models judge, not changes in whether they answer.