arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2607.04590v1 [cs.AI] 06 Jul 2026

Attention Limited Reward Learning

Wenqian Xing Management Science and Engineering Department, Stanford University, wxing@stanford.edu
Abstract

Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley–Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption misses through a reduced-form model motivated by rational inattention, in which each label is generated by a low-capacity evaluation channel. The model separates two forms of ambiguity that standard reward modeling tends to conflate: a comparison may be difficult because the two candidates are genuinely close in value, or because the relevant distinction is hard to detect under limited attention. We show that limited attention can fundamentally distort what pairwise comparisons reveal. In particular, passive comparison data cannot generally distinguish reward, attention, and default tendencies, and heterogeneous attention can make standard Bradley–Terry reward modeling recover misleading rankings. Our analysis shows that learning is governed not by the raw number of labels, but by the amount of attended information each label carries. A case study on human votes over language-model pairs from Chatbot Arena exhibits the predicted signature, a cyclic component of the comparison data that exceeds sampling noise and that no scalar reward can represent; a second case study on perceptual comparisons shows that response times and gaze carry gap information that the labels do not. This perspective suggests that human feedback should be treated not as direct revealed preference, but as an attention-limited measurement process: a weak preference signal may reflect hidden evaluation difficulty rather than genuine indifference.

1 Introduction

Human feedback has become a central part of how modern AI systems are trained and evaluated. In reinforcement learning from human feedback (RLHF), preference-based policy optimization, and related alignment pipelines, the basic measurement primitive is often simple. A human is shown two candidate responses, rankings, plans, allocations, or trajectories and asked which one is better. A reward model is then trained so that its pairwise differences predict these choices. This template underlies much of contemporary reward learning for language models and agentic systems [34, 7, 36, 31, 23]. Direct preference optimization changes the downstream policy update, but it still relies on the same basic comparison signal linking human choices to reward differences [24].

The standard statistical abstraction is the Bradley–Terry, or conditional-logit, model [3, 18, 20]. For a query consisting of a context xx and two candidates y0,y1y^{0},y^{1}, it assumes

[y1y0x]=σ(r(x,y1)r(x,y0)),σ(t)=11+et.\mathbb{P}[y^{1}\succ y^{0}\mid x]=\sigma\bigl(r(x,y^{1})-r(x,y^{0})\bigr),\qquad\sigma(t)=\frac{1}{1+e^{-t}}. (1)

Under this model, comparison labels are noisy but direct measurements of a single latent reward. If two candidates are chosen at roughly equal rates, the model interprets this as evidence that their rewards are nearly equal. Much of the statistical ranking literature studies estimation and ranking in this correctly specified regime [22, 26].

For alignment, however, the most important comparisons are often not the easiest ones. A response may be fluent but subtly misleading. A plan may appear helpful while creating long-run risks. A model output may satisfy the literal request while violating the user’s underlying intent. In such cases, the relevant distinction is not necessarily small; it may simply be hard to notice. This observation is the premise of work on scalable oversight, which anticipates that as systems become more capable, the outputs whose evaluation matters most are the ones that strain unaided human judgment [1, 17, 13, 2]. Human evaluators operate with limited time, limited context, and limited cognitive bandwidth. They attend to surface features that are cheap to check, such as fluency, length, format, and confidence, and can miss properties that require costly deliberation, a pattern documented empirically in reward models that absorb such surface regularities [29, 10]. This is not just random noise around a fixed reward. It is a structured measurement problem.

We model this problem using an attention-scaled reduced form motivated by rational inattention. A rationally inattentive evaluator allocates scarce information-processing effort before making a choice [28, 19, 4, 9]. Some comparisons are cheap to evaluate because one answer is incoherent, one plan obviously fails, or one candidate plainly violates an instruction. Others require careful reasoning about factual dependencies, safety implications, long-horizon consequences, or subtle forms of misalignment. A Shannon rational-inattention first-order condition, derived in Section 2, motivates a marginal comparison channel of the form

z=ηz+βzΔz,\ell_{z}=\eta_{z}+\beta_{z}\Delta_{z}^{*}, (2)

where Δz\Delta_{z}^{*} is the deliberative reward difference for query zz, βz0\beta_{z}\geq 0 is an attention multiplier, and ηz\eta_{z} is a default, order, prior, or salience term. Thus the observed label is not just a noisy draw from the reward difference. It is the output of an attention-limited channel whose strength varies across queries.

This channel highlights a distinction that is central to alignment. A comparison can be hard because close, in that the two candidates genuinely have similar deliberative value. A comparison can also be hard because hidden, in that the reward-relevant difference is large but the evidence needed to see it is costly or non-salient, so the observed choice probability remains near 50505050. Standard Bradley–Terry modeling treats both cases as small reward gaps. For alignment this conflation is dangerous, because the hard-because-hidden case is where subtle misalignment lives. A weak preference signal need not mean that two outputs are equally good; it may mean that the evaluation protocol failed to elicit the information needed to distinguish them.

Contributions.

This paper develops the theoretical implications of treating comparisons as attention-limited measurements. The argument proceeds in three steps, each answering a question the previous step raises. Figure 1 previews the central failure the paper isolates, i.e., attention-limited comparisons whose every pairwise majority is correct drive the Bradley–Terry fit to the wrong ranking.

RR^{*}AA (22)BB (11)CC (0)\ellAABBCC=ε\ell=\varepsilonβ=ε\beta=\varepsilon=ε\ell=\varepsilonβ=ε/2\beta=\varepsilon/2=5ε\ell=5\varepsilon,  β=5ε\beta=5\varepsilonε+5εε0\varepsilon+5\varepsilon-\varepsilon\neq 0  (Prop. 1)rr^{\dagger}BB (10ε/310\varepsilon/3)AA (8ε/38\varepsilon/3)CC (0)BAB\succ Ae=βeΔe\ell_{e}=\beta_{e}\Delta^{*}_{e}heterogeneous βe\beta_{e}(attention)(BWB)+BW(B^{\top}WB)^{+}B^{\top}W\ell(Prop. 2)
Figure 1: The projection reversal (Example 1). Left to right: deliberative reward RR^{*}, observed log odds \ell produced by the attention channel, and the Bradley–Terry fit rr^{\dagger}. Attention is heterogeneous across pairs, with βBC=5ε\beta_{BC}=5\varepsilon but βAC=ε/2\beta_{AC}=\varepsilon/2. Every observed majority favors the better item, yet the heterogeneity makes the cycle sum nonzero, so no scalar reward represents \ell (Proposition 1); the fit is the projection of \ell onto potential fields (Proposition 2) and ranks BB above AA (red).

The first step asks when attention-filtered comparisons are compatible with reward fitting at all. Section 2 motivates the channel (2) from a Shannon rational-inattention first-order condition and states it as a reduced-form measurement model. Section 3 then gives an exact answer on a finite comparison graph. A scalar Bradley–Terry reward can represent the observed log odds if and only if their circulation around every cycle vanishes (Proposition 1). Homogeneous attention passes this test, because a common multiplier merely rescales reward; heterogeneous attention generically fails it. So the object that reward fitting presumes need not exist, and the question becomes what a Bradley–Terry learner converges to instead.

The second step characterizes that limit and shows the resulting error is one of kind, not degree. Near the low-information regime, the population Bradley–Terry solution is the weighted projection of the observed log-odds field onto the space of potential fields (Proposition 2), an object shaped by the comparison graph and query distribution as much as by the underlying values. A three-item example makes the danger concrete. The projection ranks item BB above item AA even though AA is deliberatively better and every pairwise majority points in the correct direction (Example 1). One might hope this is an artifact of the Bradley–Terry functional form, to be repaired by a richer model class, a prior, or more data. Proposition 3 shows that this is not the case. Reward, attention, and defaults are not separately identified from passive comparison probabilities, so the same data are consistent with many incompatible reward vectors no matter what estimator consumes them.

The third step explains why these failures are inevitable and what they cost, moving from the Bradley–Terry learner to arbitrary procedures by working at the level of the labels themselves. A comparison label is a one-bit message emitted after costly information acquisition, and its entropy is not its information content. A label can be maximally random while carrying arbitrarily little information about the reward gap (Proposition 4). Per-label mutual and Fisher information about reward scale as βz2\beta_{z}^{2}, and when attention is unknown the local information matrix is rank one, so the likelihood sees only the product βΔ\beta\Delta (Proposition 5). This is the local face of the non-identification above. KL and Fano arguments turn the scaling into sample complexity. Distinguishing Δ=δ\Delta=\delta from Δ=δ\Delta=-\delta requires on the order of 1/(β2δ2)1/(\beta^{2}\delta^{2}) labels, and reward recovery over any hypothesis class is governed by total attended information rather than the number of annotations (Theorem 1).

Finally, Section 5 tests the theory empirically in two settings. In human votes over language-model pairs from Chatbot Arena, the observed log-odds field carries cyclic energy beyond sampling noise, rejecting representation by any scalar score, and the one-eighth law of Proposition 7 prices the misfit to within about ten percent. In a perceptual comparison dataset with response times and gaze, the channel’s ingredients are visible directly. Psychometric slopes vary sixfold across evaluators, gaze appears as an additive default, and response time reveals gap-magnitude information that is absent from single labels alone.

Related work.

Reward learning from human feedback and its failure modes. Pairwise comparisons are the standard interface for aligning language models and agentic systems [34, 7, 36, 31, 23]; direct preference optimization removes the explicit reward model but keeps the Bradley–Terry link between choices and reward differences [24]. A growing literature documents failure modes of this pipeline. Reward models degrade under optimization pressure [10], admit reward hacking through gaps between proxy and target [30], absorb surface regularities such as response length [29], and inherit fundamental limitations of human oversight [5]. We contribute an attention-based account of one such limitation. Even with unlimited data and no optimization pressure, the Bradley–Terry estimand itself can be wrong once evaluator attention is heterogeneous.

Human models in reward inference. Any reward-learning method embeds a model of the human data generator [14]. The classical choices are noiseless or Boltzmann-rational behavior [18]; richer models treat systematic human biases as objects to be learned [8, 27] or cast value learning as a cooperative game between human and machine [12]. The attention-scaled channel is motivated by endogenous attention allocation. Costly evaluation makes the distortion query-specific, correlated with evaluation difficulty, and, as Section 3 shows, unidentifiable from passive labels.

Beyond Bradley–Terry. Motivated in part by intransitivity in human preference data, recent work replaces the scalar-reward assumption with general preference models and solves for von Neumann winners or Nash equilibria of a preference game [11, 21, 32]. Our results are complementary. The attention-scaled channel shows how cyclic comparison data can arise from heterogeneous attention, and Proposition 7 prices the information that any scalar reward must discard.

Statistical ranking and rational inattention. The statistical ranking literature studies estimation from pairwise comparisons when Bradley–Terry or a parametric relative is correctly specified [22, 26, 35]; we instead characterize the pseudo-true target under the misspecification that attention induces. The potential–cyclic decomposition of Section 4.4 is the combinatorial Hodge decomposition introduced to statistical ranking by Jiang et al. [15]; here the cyclic component arises from heterogeneous attention, and we quantify its KL cost for scalar reward fitting. A behavioral motivation is Shannon rational inattention. The generalized-logit structure of optimal behavior follows Matějka and McKay [19], building on Sims [28], with revealed-preference and discrete-choice formulations in Caplin and Dean [4] and Fosgerau et al. [9]. We use this theory to motivate, rather than fully identify from first principles, a measurement channel for alignment feedback.

Scalable oversight. The premise that the most alignment-relevant comparisons are the hardest to evaluate motivates scalable-oversight proposals such as recursive reward modeling, debate, and evaluation assistance [1, 17, 13, 2]. Our lower bounds formalize what such interventions must accomplish. Because sign, ranking, and reward recovery are governed by attended information, an oversight protocol adds value insofar as it raises the attention multiplier on the comparisons that matter, and active query selection [25] cannot substitute for attention it does not change.

2 Model

2.1 Pairwise reward learning

Let x𝒳x\in\mathcal{X} denote a context, prompt, market state, or initial state, and let 𝒴(x)\mathcal{Y}(x) be the feasible candidate set. A comparison query is a triple

z=(x,y0,y1)𝒵,z=(x,y^{0},y^{1})\in\mathcal{Z},

and the observed label is c{0,1}c\in\{0,1\}, where c=1c=1 records a choice of y1y^{1} over y0y^{0}. Queries are drawn from a distribution ρ\rho determined by the data-collection policy.

The target is a deliberative reward R(x,y)R^{*}(x,y)\in\mathbb{R}, interpreted as the evaluation the principal wants the human to apply after processing all reward-relevant information under the intended criterion. Write

Δz=R(x,y1)R(x,y0)\Delta_{z}^{*}=R^{*}(x,y^{1})-R^{*}(x,y^{0})

for the deliberative difference in query zz. A standard reward learner chooses a score function rr\in\mathcal{R} by logistic maximum likelihood,

r^nargmaxr1nt=1n{ctlogσ(dr(zt))+(1ct)logσ(dr(zt))},dr(z)=r(x,y1)r(x,y0),\widehat{r}_{n}\in\operatorname*{arg\,max}_{r\in\mathcal{R}}\frac{1}{n}\sum_{t=1}^{n}\left\{c_{t}\log\sigma(d_{r}(z_{t}))+(1-c_{t})\log\sigma(-d_{r}(z_{t}))\right\},\qquad d_{r}(z)=r(x,y^{1})-r(x,y^{0}), (3)

under the maintained hypothesis that [c=1z]=σ(dr(z))\mathbb{P}[c=1\mid z]=\sigma(d_{r}(z)) for some reward aligned with RR^{*}. The paper studies what comparison data reveal when human labels are instead attention-filtered measurements of deliberative reward.

2.2 Rationally inattentive comparison behavior

Fix a query zz and let ωΩz\omega\in\Omega_{z} denote the evidence relevant to the comparison: facts in the context, latent consequences of the two candidates, safety implications, or other payoff-relevant information. For the rational-inattention benchmark in this subsection, assume Ωz\Omega_{z} is finite and the prior μz\mu_{z} has full support. A comparator chooses an attention-decision rule πz(aω)\pi_{z}(a\mid\omega) over actions a{0,1}a\in\{0,1\}, where action a=1a=1 means choosing y1y^{1} and action a=0a=0 means choosing y0y^{0}. The marginal action probability induced by πz\pi_{z} is

π¯z(a)=ωΩzμz(ω)πz(aω).\bar{\pi}_{z}(a)=\sum_{\omega\in\Omega_{z}}\mu_{z}(\omega)\pi_{z}(a\mid\omega).

Let uz(a,ω)u_{z}(a,\omega) be the comparator’s payoff, encoding the evaluation criterion the principal intends. The comparator solves the Shannon rational-inattention problem

maxπz𝔼ωμz,aπz(ω)[uz(a,ω)]κzIz(ω;a),\max_{\pi_{z}}\mathbb{E}_{\omega\sim\mu_{z},\,a\sim\pi_{z}(\cdot\mid\omega)}[u_{z}(a,\omega)]-\kappa_{z}\mathrm{I}_{z}(\omega;a), (4)

where κz>0\kappa_{z}>0 is the marginal cost of information and

Iz(ω;a)=ω,aμz(ω)πz(aω)logπz(aω)π¯z(a).\mathrm{I}_{z}(\omega;a)=\sum_{\omega,a}\mu_{z}(\omega)\pi_{z}(a\mid\omega)\log\frac{\pi_{z}(a\mid\omega)}{\bar{\pi}_{z}(a)}. (5)

The next lemma is the standard generalized-logit form of the Shannon solution. It is included to pin down where the attention multiplier and the endogenous marginal action tendency enter.

Lemma 1 (Shannon rational inattention implies generalized logit).

At any interior optimum of (4),

πz(aω)=π¯z(a)exp(uz(a,ω)/κz)b{0,1}π¯z(b)exp(uz(b,ω)/κz).\pi_{z}(a\mid\omega)=\frac{\bar{\pi}_{z}(a)\exp(u_{z}(a,\omega)/\kappa_{z})}{\sum_{b\in\{0,1\}}\bar{\pi}_{z}(b)\exp(u_{z}(b,\omega)/\kappa_{z})}. (6)

Consequently,

logπz(1ω)πz(0ω)=αz+βz{uz(1,ω)uz(0,ω)},αz=logπ¯z(1)π¯z(0),βz=κz1.\log\frac{\pi_{z}(1\mid\omega)}{\pi_{z}(0\mid\omega)}=\alpha_{z}+\beta_{z}\{u_{z}(1,\omega)-u_{z}(0,\omega)\},\qquad\alpha_{z}=\log\frac{\bar{\pi}_{z}(1)}{\bar{\pi}_{z}(0)},\quad\beta_{z}=\kappa_{z}^{-1}. (7)

Equation (7) is a conditional first-order characterization. The payoff difference is multiplied by the query-specific inverse information cost βz\beta_{z}, while the endogenous marginal action tendency αz\alpha_{z} enters additively. It does not by itself pin down the ex-ante marginal label probability obtained by integrating over evidence states, and when the payoff difference is deterministic and nonzero the optimum is a boundary case to which the interior logit formula does not apply. The rest of the paper therefore takes the following equation as a reduced-form marginal measurement channel, motivated by the conditional rational-inattention first-order condition rather than derived from it. The additive term ηz\eta_{z} in the reduced form should be read as standing in for the endogenous action tendency αz\alpha_{z} together with genuinely exogenous order, format, and salience effects; Remark 1 gives one environment in which the channel holds exactly.

Definition 1 (Attention-scaled comparison channel).

An attention-scaled comparison channel is a map zqz(0,1)z\mapsto q_{z}\in(0,1) satisfying

qz=σ(z),z=ηz+βzΔz,βz0.q_{z}=\sigma(\ell_{z}),\qquad\ell_{z}=\eta_{z}+\beta_{z}\Delta_{z}^{*},\qquad\beta_{z}\geq 0. (8)

The case ηz=0\eta_{z}=0 and βzβ>0\beta_{z}\equiv\beta>0 is called homogeneous attention.

Remark 1 (An exact conditional interpretation).

There is one environment in which the channel holds exactly rather than as a reduced form. Suppose the evidence state ω\omega is realized but unknown to the comparator ex ante, and let Δz=uz(1,ω)uz(0,ω)\Delta_{z}^{*}=u_{z}(1,\omega)-u_{z}(0,\omega) be the payoff difference in the realized state, which is the deliberative difference of Section 2 when the payoff encodes the intended criterion. If the optimum of (4) is interior, Lemma 1 implies that the label distribution conditional on the realized state is exactly σ(αz+βzΔz)\sigma(\alpha_{z}+\beta_{z}\Delta_{z}^{*}), which is Definition 1 with ηz=αz\eta_{z}=\alpha_{z}. In this paper, the default is endogenous, determined by the prior μz\mu_{z} and the information cost κz\kappa_{z}, and this is why priors and salience enter additively rather than being scaled like reward. The reduced form frees ηz\eta_{z} from this benchmark to accommodate order and format effects as well.

All graph-theoretic and information-theoretic results below are statements about the reduced-form channel in Definition 1. If βz=0\beta_{z}=0, the comparison label is independent of the deliberative reward difference after conditioning on the default term. Such a query may still produce random-looking labels, but it carries no information about the sign or magnitude of Δz\Delta_{z}^{*} through the reward channel.

With the channel in hand, the paper’s question can be stated. A learner observes the choice probabilities qzq_{z}, while the object of interest is Δz\Delta_{z}^{*}; what does the former reveal about the latter? Section 3 answers for the standard Bradley–Terry reward learner, at the level of representation and identification. Section 4 then drops the estimator from the picture and answers at the level of information, bounding what any procedure could distinguish with any amount of data.

3 Why Raw Comparisons Are Not Enough

This section studies the standard Bradley–Terry reward learner under the attention-scaled channel. We first ask whether the observed log odds are compatible with any scalar reward at all, and the answer is a cycle criterion that heterogeneous attention generically violates. Practitioners fit Bradley–Terry whether or not the criterion holds, so the next question is what the fit converges to when it fails. It converges to a projection of the log-odds field, and the projection can reverse the deliberative ranking (Section 3.2). The last question is whether a different parameterization of the reward model could avoid these problems. It cannot, because without further restrictions reward, attention, and defaults are observationally equivalent along a continuum of explanations (Section 3.3).

3.1 Cycle obstructions to scalar representability

The first question is whether attention-filtered log odds can be represented by any scalar reward at all. This is a purely algebraic question about the structure of the log-odds field, and it has a complete answer. Fix one context and a finite candidate set V={1,,m}V=\{1,\ldots,m\}. Let G=(V,E)G=(V,E) be the undirected comparison graph. Choose an arbitrary orientation for each edge e=(i,j)Ee=(i,j)\in E and let qe=qijq_{e}=q_{ij} be the probability that ii is chosen over jj in that oriented comparison. Define e=logit(qe)\ell_{e}=\operatorname{logit}(q_{e}), and use the antisymmetric extension ji=ij\ell_{ji}=-\ell_{ij} when an edge is traversed against its chosen orientation. Let BE×VB\in\mathbb{R}^{E\times V} be the signed incidence matrix with (Br)e=rirj(Br)_{e}=r_{i}-r_{j} for e=(i,j)e=(i,j). A Bradley–Terry representation of the channel is exactly a potential representation =Br\ell=Br. The only obstruction is circulation around cycles.

Proposition 1 (Cycle consistency).

Assume GG is connected and qe(0,1)q_{e}\in(0,1) on every compared edge. There exists a score vector rmr\in\mathbb{R}^{m} such that e=(Br)e\ell_{e}=(Br)_{e} on every edge if and only if

k=1Kikik+1=0,iK+1=i1,\sum_{k=1}^{K}\ell_{i_{k}i_{k+1}}=0,\qquad i_{K+1}=i_{1}, (9)

for every directed cycle (i1,i2,,iK,i1)(i_{1},i_{2},\ldots,i_{K},i_{1}) in GG, where reversing an oriented edge changes the sign of its log odds. When such a score exists, it is unique up to an additive constant.

The proposition gives an immediate diagnostic for attention-filtered data. Homogeneous attention is harmless because it merely rescales the potential. Potential defaults are representable but still misaligned, because the fitted reward absorbs the default. Heterogeneous edge-specific attention generally creates nonzero cycle sums.

Corollary 1 (Representability of attention-scaled comparisons).

On a finite connected comparison graph, attention-scaled log odds

ij=ηij+βij(RiRj)\ell_{ij}=\eta_{ij}+\beta_{ij}(R_{i}^{*}-R_{j}^{*})

admit an exact Bradley–Terry reward if and only if their cycle sums vanish. In particular:

  1. 1.

    if ηij=0\eta_{ij}=0 and βij=β>0\beta_{ij}=\beta>0 on all edges, then ri=βRir_{i}=\beta R_{i}^{*} is an exact representation;

  2. 2.

    if βij=β>0\beta_{ij}=\beta>0 and ηij=bibj\eta_{ij}=b_{i}-b_{j} for some potential bmb\in\mathbb{R}^{m}, then ri=βRi+bir_{i}=\beta R_{i}^{*}+b_{i} is an exact representation.

The second case illustrates a subtle failure mode. A default with potential structure, such as a systematic preference for familiar formats or concise answers, is perfectly representable by Bradley–Terry scores. It is nevertheless not deliberative reward; it reverses pairs whenever the default difference dominates the attended reward difference.

3.2 The attention-blind Bradley–Terry target

A failed cycle criterion does not stop anyone from fitting Bradley–Terry. The population likelihood remains strictly concave, so the fit converges regardless, and the question is what it converges to. The next proposition identifies the limit near the low-information regime as the weighted least-squares projection of the human log-odds field onto the space of potential fields. The fitted reward is therefore not a noisy copy of deliberative value but a graph-dependent compression of it, shaped by the comparison graph and the sampling policy as much as by human values.

Let W=diag(ρe)W=\operatorname{diag}(\rho_{e}) collect the edge sampling probabilities, with ρe>0\rho_{e}>0 and eρe=1\sum_{e}\rho_{e}=1. Normalize scores by 𝟏r=0\mathbf{1}^{\top}r=0. For an oriented edge vector \ell, define the population objective

L(r;)=eEρe{σ(e)logσ((Br)e)+(1σ(e))logσ((Br)e)}.L(r;\ell)=\sum_{e\in E}\rho_{e}\left\{\sigma(\ell_{e})\log\sigma((Br)_{e})+\bigl(1-\sigma(\ell_{e})\bigr)\log\sigma(-(Br)_{e})\right\}. (10)
Proposition 2 (Local projection of log odds).

Suppose GG is connected, ρe>0\rho_{e}>0 on every edge, and ε\|\ell\|_{\infty}\leq\varepsilon. Let r()r^{\dagger}(\ell) be the unique maximizer of (10) on the subspace 𝟏r=0\mathbf{1}^{\top}r=0. Then, as ε0\varepsilon\to 0,

r()=(BWB)+BW+O(ε3),r^{\dagger}(\ell)=(B^{\top}WB)^{+}B^{\top}W\ell+O(\varepsilon^{3}), (11)

where ()+(\cdot)^{+} denotes the Moore–Penrose inverse acting on the subspace orthogonal to constants. The remainder is in Euclidean norm and is uniform over sufficiently small \|\ell\|_{\infty}.

The projection in Equation (11) discards all cyclic components of the human log odds. The following example shows that the induced error is not merely cardinal; it can reverse the learned ranking.

Example 1 (A three-item projection reversal).

Consider one context with candidates A,B,CA,B,C and deliberative rewards

RA=2,RB=1,RC=0.R_{A}^{*}=2,\qquad R_{B}^{*}=1,\qquad R_{C}^{*}=0.

There is no default bias. Human log odds for the higher-reward item over the lower-reward item are

AB=ε,AC=ε,BC=5ε,\ell_{AB}=\varepsilon,\qquad\ell_{AC}=\varepsilon,\qquad\ell_{BC}=5\varepsilon, (12)

for small ε>0\varepsilon>0. These log odds arise from the scalar attention channel with βAB=ε\beta_{AB}=\varepsilon, βAC=ε/2\beta_{AC}=\varepsilon/2, and βBC=5ε\beta_{BC}=5\varepsilon. Thus every individual pairwise majority points in the deliberatively correct direction. Sampling the three pairs equally and normalizing rC=0r_{C}=0, Proposition 2, equivalently the direct first-order calculation in the appendix, yields

rA=83ε+O(ε3),rB=103ε+O(ε3),rC=0.r_{A}^{\dagger}=\frac{8}{3}\varepsilon+O(\varepsilon^{3}),\qquad r_{B}^{\dagger}=\frac{10}{3}\varepsilon+O(\varepsilon^{3}),\qquad r_{C}^{\dagger}=0. (13)

For sufficiently small ε\varepsilon, the attention-blind Bradley–Terry target ranks BB above AA although AA is deliberatively best. The reversal is not a knife-edge feature of the expansion; solving the population first-order conditions exactly preserves the same ordering at moderate scales such as ε=0.4\varepsilon=0.4.

3.3 Non-identification

Example 1 might suggest that the problem lies with the Bradley–Terry functional form, and that a richer model class, a Bayesian prior over rewards, or simply more data would recover the deliberative ranking. The deeper obstacle is that comparison probabilities alone cannot disentangle reward, attention, and defaults, so any method that consumes only passive labels inherits the same ambiguity. The next proposition makes this point in finite-dimensional form.

Proposition 3 (Non-identification of reward, attention, and defaults).

Fix comparison probabilities qij(0,1)q_{ij}\in(0,1) on a finite set of oriented pairs, and write ij=logit(qij)\ell_{ij}=\operatorname{logit}(q_{ij}).

  1. 1.

    For any candidate reward vector vmv\in\mathbb{R}^{m} and any nonnegative attention multipliers βij0\beta_{ij}\geq 0, the defaults

    ηij=ijβij(vivj)\eta_{ij}=\ell_{ij}-\beta_{ij}(v_{i}-v_{j}) (14)

    rationalize the data through the attention-scaled channel.

  2. 2.

    Suppose defaults are restricted to zero. If, for every compared pair, either vivjv_{i}\neq v_{j} and ij(vivj)0\ell_{ij}\,(v_{i}-v_{j})\geq 0, or vi=vjv_{i}=v_{j} and ij=0\ell_{ij}=0, then the same data are rationalized by the pair-specific multipliers

    βij=ijvivj\beta_{ij}=\frac{\ell_{ij}}{v_{i}-v_{j}} (15)

    on all compared pairs with vivjv_{i}\neq v_{j}, and by any nonnegative βij\beta_{ij} on zero-gap pairs with ij=0\ell_{ij}=0.

This proposition is the formal reason that a Bayesian prior, a larger neural reward model, or more passive samples cannot by itself recover deliberative reward. Without restrictions that distinguish reward from attention and defaults, the likelihood is flat along observationally equivalent explanations.

Non-identification says the likelihood is flat; it does not yet say why the flatness arises, or how much data it costs even where the parameters are partially informative. Both questions have exact answers once the label is treated as a communication channel and its information content is accounted for directly. That accounting is the subject of the next section.

4 Information-Theoretic Limits of Reward Learning

Section 3 located three failures of the Bradley–Terry reward learner. The log odds need not be representable by any scalar reward, the fit converges to a graph-dependent projection, and the likelihood is flat across observationally equivalent explanations. All three statements concern one particular estimator, so one might still hope that a cleverer use of the same labels escapes them. This section shows that they are instead properties of the labels themselves, by re-deriving each failure at the level of information, where no estimator appears. The weak-signal ambiguity becomes a separation between label entropy and reward information (Section 4.1); non-identification becomes a rank-one information matrix (Section 4.2); the estimation failures become sample-complexity lower bounds that bind every procedure, adaptive or not (Section 4.3); and the cycle obstruction acquires an exact price in KL divergence (Section 4.4). Throughout, a comparison label is viewed as a binary message emitted after information acquisition. Its entropy may be high while the information it carries about deliberative reward is arbitrarily small. Defaults are suppressed when they are not essential, since adding a known default shifts the log odds but does not change the information bottleneck created by β\beta.

4.1 High label entropy is not high reward information

A nearly even split of labels is often treated as the most informative region of a logistic model because the Bernoulli variance is largest near one half. That intuition is incomplete. If the split is even because attention is near zero, the label may be maximally random while revealing almost nothing about reward.

Consider a single pair and let Δ\Delta be an unknown deliberative reward difference with a finite prior distribution Π\Pi. Conditional on Δ\Delta, the observed label satisfies

CΔBernoulli(σ(η+βΔ)),β0.C\mid\Delta\sim\operatorname{Bernoulli}\bigl(\sigma(\eta+\beta\Delta)\bigr),\qquad\beta\geq 0. (16)

Let 𝖧(C)\mathsf{H}(C) denote the entropy of the marginal label and I(Δ;C)\mathrm{I}(\Delta;C) the mutual information between the reward gap and the label, measured in nats.

Proposition 4 (Entropy-information separation).

In the channel (16):

  1. 1.

    if β=0\beta=0, then CC is independent of Δ\Delta, so I(Δ;C)=0\mathrm{I}(\Delta;C)=0;

  2. 2.

    if, in addition, η=0\eta=0, then CC is a fair coin, so 𝖧(C)=log2\mathsf{H}(C)=\log 2 while I(Δ;C)=0\mathrm{I}(\Delta;C)=0;

  3. 3.

    if the prior on Δ\Delta has finite support, then as β0\beta\to 0,

    I(Δ;C)=12σ(η)(1σ(η))β2Var(Δ)+O(β3).\mathrm{I}(\Delta;C)=\frac{1}{2}\sigma(\eta)\bigl(1-\sigma(\eta)\bigr)\beta^{2}\operatorname{Var}(\Delta)+O(\beta^{3}). (17)

Thus a 50505050 outcome has two distinct interpretations. It may mean the reward gap is close to zero, or it may mean the channel is low-transmission because β\beta is small. The entropy of the label alone cannot distinguish these cases.

4.2 Fisher information and local non-identification

Proposition 4 is a statement about a prior over the gap. The same distinction appears in local, estimation-theoretic form through Fisher information, and in that form it also exposes the mechanism behind the non-identification of Proposition 3. In the standard Bradley–Terry model, one label contains the most local information about the reward difference near a 50505050 split. Under the attention-scaled channel, that maximum is multiplied by β2\beta^{2}.

Proposition 5 (Fisher information collapse and rank-one information).

For one comparison with

CBernoulli(q),q=σ(η+βΔ),C\sim\operatorname{Bernoulli}(q),\qquad q=\sigma(\eta+\beta\Delta), (18)

the following hold.

  1. 1.

    If η\eta and β\beta are known, the Fisher information in one label about Δ\Delta is

    Δ(Δ)=β2q(1q)β24.\mathcal{I}_{\Delta}(\Delta)=\beta^{2}q(1-q)\leq\frac{\beta^{2}}{4}. (19)
  2. 2.

    If η\eta is known but (Δ,β)(\Delta,\beta) are both unknown, the Fisher information matrix for (Δ,β)(\Delta,\beta) is

    (Δ,β)=q(1q)(β2βΔβΔΔ2)=q(1q)(βΔ)(βΔ),\mathcal{I}_{(\Delta,\beta)}=q(1-q)\begin{pmatrix}\beta^{2}&\beta\Delta\\ \beta\Delta&\Delta^{2}\end{pmatrix}=q(1-q)\begin{pmatrix}\beta\\ \Delta\end{pmatrix}\begin{pmatrix}\beta&\Delta\end{pmatrix}, (20)

    and therefore has rank at most one.

  3. 3.

    If η\eta, Δ\Delta, and β\beta are all unknown, the Fisher information matrix for (η,Δ,β)(\eta,\Delta,\beta) is

    (η,Δ,β)=q(1q)(1βΔ)(1βΔ),\mathcal{I}_{(\eta,\Delta,\beta)}=q(1-q)\begin{pmatrix}1\\ \beta\\ \Delta\end{pmatrix}\begin{pmatrix}1&\beta&\Delta\end{pmatrix}, (21)

    and again has rank at most one.

When β\beta is constrained to be nonnegative, the usual regular Fisher-information interpretation for parameters involving β\beta applies at interior points β>0\beta>0; at the boundary β=0\beta=0, the displays are local derivative formulas for the logit index.

The first claim gives a sample-complexity penalty, since low attention reduces per-label information about reward quadratically. The rank-one claims give a local version of Proposition 3. The likelihood identifies the log-odds index η+βΔ\eta+\beta\Delta, not the reward gap, the attention multiplier, and the default separately.

4.3 Lower bounds for sign and ranking recovery

Fisher information measures local difficulty; the operational question is global. How many labels does it take to answer the coarsest question the data could settle, namely which of the two candidates is better? The next result shows that ordinal recovery, too, is bought only with attended information. To isolate the effect, consider the clean channel with no default and known attention,

CBernoulli(σ(βΔ)).C\sim\operatorname{Bernoulli}(\sigma(\beta\Delta)). (22)

The task is to decide whether Δ=δ\Delta=\delta or Δ=δ\Delta=-\delta for a fixed δ>0\delta>0.

Proposition 6 (KL lower bound for sign recovery).

Let P+P_{+} be the distribution of one label under Δ=δ\Delta=\delta and PP_{-} the distribution under Δ=δ\Delta=-\delta in (22). If β=0\beta=0, then P+=PP_{+}=P_{-} and no test can have worst-case error probability below 1/21/2. If β>0\beta>0, then

DKL(P+P)=βδtanh(βδ2)(βδ)22.D_{\mathrm{KL}}(P_{+}\|P_{-})=\beta\delta\tanh\left(\frac{\beta\delta}{2}\right)\leq\frac{(\beta\delta)^{2}}{2}. (23)

For any test based on nn independent labels and β>0\beta>0, if the worst-case probability of sign error is at most α<1/2\alpha<1/2, then necessarily

n2(12α)2βδtanh(βδ/2)4(12α)2β2δ2.n\geq\frac{2(1-2\alpha)^{2}}{\beta\delta\tanh(\beta\delta/2)}\geq\frac{4(1-2\alpha)^{2}}{\beta^{2}\delta^{2}}. (24)

Consequently, without a positive lower bound on attention, no uniform finite-sample sign guarantee is possible.

The lower bound depends on the product βδ\beta\delta. Bradley–Terry interprets a small product as a small reward gap. The attention-scaled model shows that it may instead be a large reward gap passing through a low-attention channel.

The same logic extends from one pair to many reward hypotheses. Let Θ\Theta be a finite set of possible reward functions and suppose the learner observes labels C1,,CnC_{1},\ldots,C_{n}, possibly under adaptively chosen queries. Let Ht1H_{t-1} denote the history before label tt, including previous labels, adaptively chosen queries, and any external randomization used by the query rule. Let Pθ,t(ht1)P_{\theta,t}(\cdot\mid h_{t-1}) denote the conditional distribution of label tt under reward hypothesis θ\theta and history ht1h_{t-1}.

Theorem 1 (Fano bound for attention-limited reward recovery).

Assume θ\theta is uniform on a finite class Θ\Theta with |Θ|=M2|\Theta|=M\geq 2. Suppose that for every time tt, every history ht1h_{t-1} with positive probability under the joint mixture distribution induced by the uniform prior over Θ\Theta, and every pair θ,θΘ\theta,\theta^{\prime}\in\Theta,

DKL(Pθ,t(ht1)Pθ,t(ht1))dt.D_{\mathrm{KL}}\bigl(P_{\theta,t}(\cdot\mid h_{t-1})\,\|\,P_{\theta^{\prime},t}(\cdot\mid h_{t-1})\bigr)\leq d_{t}. (25)

Then any estimator based on the observed history satisfies

[θ^θ]1t=1ndt+log2logM.\mathbb{P}[\widehat{\theta}\neq\theta]\geq 1-\frac{\sum_{t=1}^{n}d_{t}+\log 2}{\log M}. (26)

In attention-scaled logistic channels, the constants dtd_{t} are of order βt2\beta_{t}^{2} for small attended reward separations. The bound therefore says that reward recovery is limited by total attended information tdt\sum_{t}d_{t}, not simply by the number of labels. Repeating low-attention comparisons can be much less valuable than collecting fewer high-attention comparisons.

4.4 Potential information and cyclic information on comparison graphs

The bounds so far concern single pairs and finite hypothesis classes. Returning to the comparison graph closes the loop with Section 3. The cycle obstruction, similarly, has an information-theoretic meaning, and the projection of Proposition 2 turns out to be the object that discards it. A scalar reward can represent only potential fields on the comparison graph. The part of the observed log-odds field orthogonal to all potentials is cyclic comparison information, present in the human comparisons but impossible to encode in any scalar reward. The decomposition below is the combinatorial Hodge decomposition of Jiang et al. [15], applied to the weighted log-odds field generated by the attention-scaled channel.

Let the weighted inner product on edge fields be

a,bW=aWb,aW2=aWa.\langle a,b\rangle_{W}=a^{\top}Wb,\qquad\|a\|_{W}^{2}=a^{\top}Wa.

Let 𝒫=im(B)\mathcal{P}=\operatorname{im}(B) be the potential subspace. Define

rH=(BWB)+BW,pot=BrH,cyc=pot.r^{H}=(B^{\top}WB)^{+}B^{\top}W\ell,\qquad\ell^{\mathrm{pot}}=Br^{H},\qquad\ell^{\mathrm{cyc}}=\ell-\ell^{\mathrm{pot}}. (27)
Proposition 7 (Cyclic information loss of scalar rewards).

Suppose GG is connected, W=diag(ρe)W=\operatorname{diag}(\rho_{e}) has positive diagonal entries, and ε\|\ell\|_{\infty}\leq\varepsilon. Then:

  1. 1.

    =pot+cyc\ell=\ell^{\mathrm{pot}}+\ell^{\mathrm{cyc}} is the weighted orthogonal decomposition of the log-odds field into the potential subspace and its orthogonal complement: BWcyc=0B^{\top}W\ell^{\mathrm{cyc}}=0 and

    W2=potW2+cycW2.\|\ell\|_{W}^{2}=\|\ell^{\mathrm{pot}}\|_{W}^{2}+\|\ell^{\mathrm{cyc}}\|_{W}^{2}. (28)
  2. 2.

    The best scalar-reward approximation in weighted KL divergence satisfies

    infr: 1r=0eEρeDKL(Bernoulli(σ(e))Bernoulli(σ((Br)e)))\displaystyle\inf_{r:\,\mathbf{1}^{\top}r=0}\sum_{e\in E}\rho_{e}\,D_{\mathrm{KL}}\Bigl(\operatorname{Bernoulli}(\sigma(\ell_{e}))\,\Big\|\,\operatorname{Bernoulli}(\sigma((Br)_{e}))\Bigr)
    =18cycW2+O(ε3),\displaystyle\hskip 50.00008pt=\frac{1}{8}\|\ell^{\mathrm{cyc}}\|_{W}^{2}+O(\varepsilon^{3}), (29)

    where the remainder is uniform over sufficiently small \|\ell\|_{\infty}.

The coefficient rHr^{H} in (27) is computed by the same operator as the pseudo-true target in (11). To first order, the attention-blind Bradley–Terry limit is the potential component of the human log odds. Proposition 7 therefore prices what Proposition 2 discards. The cycle obstruction is not merely a consistency condition. It quantifies how much comparison information is irreducibly non-reward-like, and the cyclic energy measures the second-order information loss of compressing the human comparison channel to its potential component.

5 Empirical Case Studies

The theory leaves two observable footprints. The geometric footprint is cyclic energy in comparison fields, and the mechanistic footprint is that measurements of the evaluation process, such as response times and gaze, should carry information that labels do not. Two case studies check one footprint each.

5.1 Cyclic comparison information in Arena votes

The theory predicts that the log-odds field of real comparison data need not be a potential field, and Proposition 7 prices what any scalar score then discards. We test this signature on a canonical human preference dataset, 57,47757{,}477 pairwise battles between 6464 language models collected on Chatbot Arena [6].111Publicly available at https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k. Items are models, edges are model pairs, and labels are human votes, so the empirical vote frequencies provide exactly the weight matrix WW of Section 3. We drop ties (31%31\% of battles), keep pairs with at least 100100 decisive votes, and restrict to the largest connected component, which leaves m=32m=32 models, 7474 pairs, 15,00115{,}001 votes, and a cycle space of dimension 7431=4374-31=43. Edge log odds use the Haldane–Anscombe estimate q^e=(we+12)/(ne+1)\hat{q}_{e}=(w_{e}+\tfrac{1}{2})/(n_{e}+1).

Sampling noise alone creates cyclic energy, so significance is judged against a parametric bootstrap null. We fit Bradley–Terry by weighted maximum likelihood, regenerate binomial votes at the observed nen_{e} two thousand times, and recompute the cyclic energy share of each regenerated field. The observed share is 4.2%4.2\%, against a null mean of 2.5%2.5\% and a null 9595th percentile of 3.5%3.5\%, for a bootstrap pp-value of 0.0080.008 (Figure 2); the likelihood-ratio statistic gives the same verdict (70.670.6 on 4343 degrees of freedom, bootstrap p=0.007p=0.007). The hypothesis that some scalar score generates these win rates is rejected at the one percent level.

The misfit is small but priced accurately. The best scalar fit loses 0.00340.0034 bits per comparison in weighted KL, and the prediction 18^cycW2\tfrac{1}{8}\|\hat{\ell}^{\mathrm{cyc}}\|_{W}^{2} of Proposition 7 is 0.00380.0038 bits, a ratio of 0.890.89, even though the largest observed log odds is near 1.81.8 and the proposition’s small-ε\varepsilon hypothesis is far from satisfied. Part of the raw loss reflects sampling noise; the deviance in excess of its degrees of freedom gives a rough noise-corrected estimate of the population-level misfit near 0.00130.0013 bits per comparison. Individual triangle circulations reach at most z3.0z\approx 3.0, which is unremarkable given the number of triangles examined, so the evidence lies in the global excess rather than in any single cycle; we note, without attaching significance, that the top-ranked triangles involve closely related model versions such as claude-2.1, gpt-4-0314, and gpt-4-1106-preview.

The excess cyclic energy survives vote thresholds of 5050 and 200200 per pair (bootstrap p=0.046p=0.046 and 0.0080.008). Counting ties as half wins shrinks every log odds toward zero and weakens the excess, to p=0.06p=0.06 on the same graph and to insignificance when the threshold also admits sparsely compared pairs, consistent with dilution by a large mass of uninformative labels. Two caveats bound the interpretation. The votes aggregate over prompts and annotators, so cyclic energy can also arise from aggregation rather than per-query attention, and the analysis conditions on the pairs the platform chose to sample. What the case study establishes is a statistical rejection of every scalar score for this comparison data, at a misfit priced well by the one-eighth law; the finding is consistent with, though not unique evidence for, the attention mechanism. A self-contained script reproduces all numbers.

Refer to caption
Figure 2: Cyclic comparison information in Chatbot Arena votes (3232 models, 7474 pairs, 15,00115{,}001 decisive votes). Left: observed log odds against the best Bradley–Terry fit, with ±2\pm 2 standard errors per pair. Right: observed cyclic energy share against its sampling-noise null under the fitted Bradley–Terry model (2,0002{,}000 parametric bootstrap replicates); the observed share exceeds the null 9595th percentile (p=0.008p=0.008).

5.2 Response times carry information that labels do not

The second footprint concerns the channel itself, and it needs data in which the deliberative gap and the evaluator’s information acquisition are both measured. The perceptual comparison dataset of Tavares et al. [33] provides both.222Distributed with the aDDM toolbox, publicly available at https://github.com/goptavares/aDDM-Toolbox. Twenty-five participants made 31,85431{,}854 binary comparisons between two rotated bars, choosing the one closer to a target orientation, with angular distances set by the experimenter; the quality gap Δ{3,,3}\Delta\in\{-3,\ldots,3\} is therefore ground truth rather than a rating proxy, and every trial records the choice, the response time, and gaze fixations. Pooled choices follow a logistic in Δ\Delta with slope 1.141.14, and per-participant slopes range from 0.450.45 to 2.702.70, a factor of six (Figure 3, left). The same objective gap passes through channels of very different gain across evaluators. The reduced form absorbs attention, effort, and acuity alike into β\beta, so this is heterogeneous β\beta measured directly, whatever its source; and because preference datasets pool annotators over different pairs, evaluator-level heterogeneity of this size induces the edge-level heterogeneity that generates cyclic energy of the kind found in Section 5.1.

The headline is an information decomposition. Because the design is symmetric, a single label reveals which item is better but almost nothing about by how much; empirically the label carries 0.330.33 bits about the sign of the gap and 0.00010.0001 bits about its magnitude. Response time carries 0.0350.035 bits per trial about the magnitude (permutation-corrected, p<0.002p<0.002), while the label carries essentially none; mean response time falls from 2.22.2 seconds on ties to 1.41.4 seconds at the largest gap (Figure 3, right). Repeated labels do not close this gap, since they reveal only the confounded index βΔ\beta\Delta of Proposition 5, while response time gives a reading from outside the label channel. The quantity an annotation protocol needs in order to separate a near-tie from a large-but-unresolved gap is absent from the label and present in the response time, which is exactly the measurement the theory recommends collecting.

Gaze is associated with an additive shift of the kind the default η\eta describes. At fixed Δ\Delta, one second of relative dwell toward an item is associated with a 0.910.91 increase in its choice log odds (standard error 0.030.03), about as much as a full quality level, consistent with the gaze bias documented by Krajbich et al. [16]; the association is observational, and gaze may follow an emerging choice as well as shape it. Two further caveats apply. The task is perceptual rather than preferential, and response time is endogenous to difficulty, so these are the channel’s ingredients observed in a controlled comparison task rather than in RLHF annotation itself. Annotation platforms typically record decision latency and could release it; the analysis here is what that would enable.

Refer to caption
Figure 3: Choices and response times in a perceptual comparison task [33]; 2525 participants, 31,85431{,}854 trials. Left: psychometric curves per participant (gray) and pooled (blue); slopes vary by a factor of six across evaluators. Right: mean response time against gap magnitude, per participant and pooled with 95%95\% intervals. The label carries 0.330.33 bits about the sign of the gap and 0.00010.0001 bits about its magnitude, while response time carries 0.0350.035 bits about the magnitude.

6 Conclusion

Reward learning from pairwise comparisons changes qualitatively when humans are rationally inattentive. The reduced-form attention-scaled channel motivated by Shannon rational inattention does not simply add noise to reward differences; it scales them by query-dependent attention and adds defaults that do not scale with reward. Raw comparison log odds therefore need not form a potential field, the Bradley–Terry pseudo-true reward can reverse the deliberative ranking, and reward is not separately identified from attention or defaults. The information-theoretic analysis explains why these are not mere modeling inconveniences. A near 50505050 label may reflect true closeness or an attention bottleneck, per-label information about reward scales as β2\beta^{2}, unknown attention makes the local information matrix rank one, and KL and Fano bounds show that sign, ranking, and reward recovery require total attended information rather than more labels. On graphs, the Hodge decomposition prices the cyclic comparison information that no scalar reward can represent. The case studies find that signature in Chatbot Arena votes, with the one-eighth law pricing the misfit accurately, and show in a perceptual comparison task that response times and gaze carry information about the evaluation process that labels do not.

The main implication for alignment practice is that reward learning under bounded attention is measurement under an endogenous information constraint, not curve fitting. Annotation protocols should treat weak pairwise signals as ambiguous between true indifference and failed evaluation, and should be especially suspicious of near-even splits on safety-relevant pairs, where the hard-because-hidden interpretation is most plausible. Our bounds also give scalable-oversight interventions [17, 13, 2] a precise target. Such interventions help insofar as they raise the attention multiplier on the comparisons that matter, since no amount of passive relabeling can substitute for attended information. Natural next steps include oversight interventions that demonstrably shift attention, response-time or deliberation-effort measurement as proxies for information acquisition, partial-identification bounds under explicit attention constraints, and extensions of the graph information decomposition to function approximation over contexts and trajectories.

References

  • [1] D. Amodei, C. Olah, J. Steinhardt, P. F. Christiano, J. Schulman, and D. Mané (2016) Concrete problems in AI safety. External Links: 1606.06565, Document, Link Cited by: §1, §1.
  • [2] S. R. Bowman, J. Hyun, E. Perez, E. Chen, C. Pettit, S. Heiner, K. Lukošiūtė, A. Askell, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Olah, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, J. Clark, J. Kernion, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, L. Lovitt, N. Elhage, N. Schiefer, N. Joseph, N. Mercado, N. DasSarma, R. Larson, S. McCandlish, S. Kundu, S. Johnston, S. Kravec, S. El Showk, S. Fort, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, B. Mann, and J. Kaplan (2022) Measuring progress on scalable oversight for large language models. External Links: 2211.03540, Document, Link Cited by: §1, §1, §6.
  • [3] R. A. Bradley and M. E. Terry (1952) Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika 39 (3/4), pp. 324–345. External Links: Document Cited by: §1.
  • [4] A. Caplin and M. Dean (2015) Revealed preference, rational inattention, and costly information acquisition. American Economic Review 105 (7), pp. 2183–2203. External Links: Document Cited by: §1, §1.
  • [5] S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, T. Wang, S. Marks, C. Segerie, M. Carroll, A. Peng, P. Christoffersen, M. Damani, S. Slocum, U. Anwar, A. Siththaranjan, M. Nadeau, E. J. Michaud, J. Pfau, D. Krasheninnikov, X. Chen, L. Langosco, P. Hase, E. Bıyık, A. Dragan, D. Krueger, D. Sadigh, and D. Hadfield-Menell (2023) Open problems and fundamental limitations of reinforcement learning from human feedback. Transactions on Machine Learning Research. External Links: 2307.15217, Document, Link Cited by: §1.
  • [6] W. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, and I. Stoica (2024-21–27 Jul) Chatbot Arena: an open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 8359–8388. External Links: Link Cited by: §5.1.
  • [7] P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei (2017) Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, Vol. 30. External Links: Link Cited by: §1, §1.
  • [8] O. Evans, A. Stuhlmüller, and N. D. Goodman (2016) Learning the preferences of ignorant, inconsistent agents. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 30. External Links: Document, Link Cited by: §1.
  • [9] M. Fosgerau, E. Melo, A. de Palma, and M. Shum (2020) Discrete choice and rational inattention: a general equivalence result. International Economic Review 61 (4), pp. 1569–1589. External Links: Document Cited by: §1, §1.
  • [10] L. Gao, J. Schulman, and J. Hilton (2023-23–29 Jul) Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, pp. 10835–10866. External Links: Link Cited by: §1, §1.
  • [11] M. Gheshlaghi Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello (2024-02–04 May) A general theoretical paradigm to understand learning from human preferences. In Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, S. Dasgupta, S. Mandt, and Y. Li (Eds.), Proceedings of Machine Learning Research, Vol. 238, pp. 4447–4455. External Links: Link Cited by: §1.
  • [12] D. Hadfield-Menell, S. Russell, P. Abbeel, and A. Dragan (2016) Cooperative inverse reinforcement learning. In Advances in Neural Information Processing Systems, Vol. 29. External Links: Link Cited by: §1.
  • [13] G. Irving, P. F. Christiano, and D. Amodei (2018) AI safety via debate. External Links: 1805.00899, Document, Link Cited by: §1, §1, §6.
  • [14] H. J. Jeon, S. Milli, and A. Dragan (2020) Reward-rational (implicit) choice: a unifying formalism for reward learning. In Advances in Neural Information Processing Systems, Vol. 33. External Links: Link Cited by: §1.
  • [15] X. Jiang, L. Lim, Y. Yao, and Y. Ye (2011) Statistical ranking and combinatorial Hodge theory. Mathematical Programming 127 (1), pp. 203–244. External Links: Document Cited by: §1, §4.4.
  • [16] I. Krajbich, C. Armel, and A. Rangel (2010) Visual fixations and the computation and comparison of value in simple choice. Nature Neuroscience 13 (10), pp. 1292–1298. External Links: Document Cited by: §5.2.
  • [17] J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg (2018) Scalable agent alignment via reward modeling: a research direction. External Links: 1811.07871, Document, Link Cited by: §1, §1, §6.
  • [18] R. D. Luce (1959) Individual choice behavior: a theoretical analysis. Wiley, New York. Cited by: §1, §1.
  • [19] F. Matějka and A. McKay (2015) Rational inattention to discrete choices: a new foundation for the multinomial logit model. American Economic Review 105 (1), pp. 272–298. External Links: Document Cited by: §1, §1.
  • [20] D. McFadden (1974) Conditional logit analysis of qualitative choice behavior. In Frontiers in Econometrics, P. Zarembka (Ed.), pp. 105–142. Cited by: §1.
  • [21] R. Munos, M. Valko, D. Calandriello, M. Gheshlaghi Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, C. Fiegel, A. Michi, M. Selvi, S. Girgin, N. Momchev, O. Bachem, D. J. Mankowitz, D. Precup, and B. Piot (2024-21–27 Jul) Nash learning from human feedback. In Proceedings of the 41st International Conference on Machine Learning, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 36743–36768. External Links: Link Cited by: §1.
  • [22] S. Negahban, S. Oh, and D. Shah (2017) Rank Centrality: ranking from pairwise comparisons. Operations Research 65 (1), pp. 266–287. External Links: Document Cited by: §1, §1.
  • [23] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35, pp. 27730–27744. External Links: Link Cited by: §1, §1.
  • [24] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn (2023) Direct preference optimization: your language model is secretly a reward model. In Advances in Neural Information Processing Systems, Vol. 36. External Links: Link Cited by: §1, §1.
  • [25] D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia (2017) Active preference-based learning of reward functions. In Robotics: Science and Systems XIII, External Links: Document, Link Cited by: §1.
  • [26] N. B. Shah and M. J. Wainwright (2018) Simple, robust and optimal ranking from pairwise comparisons. Journal of Machine Learning Research 18 (199), pp. 1–38. External Links: Link Cited by: §1, §1.
  • [27] R. Shah, N. Gundotra, P. Abbeel, and A. Dragan (2019-09–15 Jun) On the feasibility of learning, rather than assuming, human biases for reward inference. In Proceedings of the 36th International Conference on Machine Learning, K. Chaudhuri and R. Salakhutdinov (Eds.), Proceedings of Machine Learning Research, Vol. 97, pp. 5670–5679. External Links: Link Cited by: §1.
  • [28] C. A. Sims (2003) Implications of rational inattention. Journal of Monetary Economics 50 (3), pp. 665–690. External Links: Document Cited by: §1, §1.
  • [29] P. Singhal, T. Goyal, J. Xu, and G. Durrett (2024) A long way to go: investigating length correlations in RLHF. In Proceedings of the First Conference on Language Modeling, Note: Spotlight; arXiv:2310.03716 External Links: 2310.03716, Document, Link Cited by: §1, §1.
  • [30] J. Skalse, N. Howe, D. Krasheninnikov, and D. Krueger (2022) Defining and characterizing reward gaming. In Advances in Neural Information Processing Systems, Vol. 35, pp. 9460–9471. External Links: 2209.13085, Document, Link Cited by: §1.
  • [31] N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano (2020) Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, Vol. 33, pp. 3008–3021. External Links: Link Cited by: §1, §1.
  • [32] G. Swamy, C. Dann, R. Kidambi, S. Wu, and A. Agarwal (2024-21–27 Jul) A minimaximalist approach to reinforcement learning from human feedback. In Proceedings of the 41st International Conference on Machine Learning, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 47345–47377. External Links: Link Cited by: §1.
  • [33] G. Tavares, P. Perona, and A. Rangel (2017) The attentional drift diffusion model of simple perceptual decision-making. Frontiers in Neuroscience 11, pp. 468. External Links: Document, Link Cited by: Figure 3, §5.2.
  • [34] C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz (2017) A survey of preference-based reinforcement learning methods. Journal of Machine Learning Research 18 (136), pp. 1–46. External Links: Link Cited by: §1, §1.
  • [35] B. Zhu, M. Jordan, and J. Jiao (2023-23–29 Jul) Principled reinforcement learning with human feedback from pairwise or K-wise comparisons. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, pp. 43037–43067. External Links: Link Cited by: §1.
  • [36] D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. F. Christiano, and G. Irving (2019) Fine-tuning language models from human preferences. External Links: 1909.08593, Document, Link Cited by: §1, §1.

Appendix A Proofs

A.1 Proof of Lemma 1

Proof.

Fix zz and suppress the query subscript. Write u(a,ω)=uz(a,ω)u(a,\omega)=u_{z}(a,\omega), μ=μz\mu=\mu_{z}, κ=κz\kappa=\kappa_{z}, and π=πz\pi=\pi_{z}. Because the optimum is interior, π(aω)>0\pi(a\mid\omega)>0 and π¯(a)>0\bar{\pi}(a)>0 for both actions. The mutual information can be written as

I(ω;a)=ω,aμ(ω)π(aω)logπ(aω)aπ¯(a)logπ¯(a).\mathrm{I}(\omega;a)=\sum_{\omega,a}\mu(\omega)\pi(a\mid\omega)\log\pi(a\mid\omega)-\sum_{a}\bar{\pi}(a)\log\bar{\pi}(a).

For each aa and ω\omega,

Iπ(aω)=μ(ω){logπ(aω)+1}μ(ω){logπ¯(a)+1}=μ(ω)logπ(aω)π¯(a).\frac{\partial\mathrm{I}}{\partial\pi(a\mid\omega)}=\mu(\omega)\{\log\pi(a\mid\omega)+1\}-\mu(\omega)\{\log\bar{\pi}(a)+1\}=\mu(\omega)\log\frac{\pi(a\mid\omega)}{\bar{\pi}(a)}.

The Lagrangian for the simplex constraints aπ(aω)=1\sum_{a}\pi(a\mid\omega)=1 is

(π,λ)=ω,aμ(ω)π(aω)u(a,ω)κI(ω;a)+ωλω(aπ(aω)1).\mathcal{L}(\pi,\lambda)=\sum_{\omega,a}\mu(\omega)\pi(a\mid\omega)u(a,\omega)-\kappa\mathrm{I}(\omega;a)+\sum_{\omega}\lambda_{\omega}\left(\sum_{a}\pi(a\mid\omega)-1\right).

The first-order condition for an interior optimum is

μ(ω)u(a,ω)κμ(ω)logπ(aω)π¯(a)+λω=0.\mu(\omega)u(a,\omega)-\kappa\mu(\omega)\log\frac{\pi(a\mid\omega)}{\bar{\pi}(a)}+\lambda_{\omega}=0.

Since μ(ω)>0\mu(\omega)>0, this is equivalent to

π(aω)=π¯(a)exp{u(a,ω)/κ+m(ω)},\pi(a\mid\omega)=\bar{\pi}(a)\exp\{u(a,\omega)/\kappa+m(\omega)\},

where m(ω)=λω/(κμ(ω))m(\omega)=\lambda_{\omega}/(\kappa\mu(\omega)) is independent of aa. Enforcing aπ(aω)=1\sum_{a}\pi(a\mid\omega)=1 determines the normalizing factor and gives Equation (6). Taking the ratio between actions 11 and 0 gives Equation (7) with αz=log{π¯z(1)/π¯z(0)}\alpha_{z}=\log\{\bar{\pi}_{z}(1)/\bar{\pi}_{z}(0)\} and βz=1/κz\beta_{z}=1/\kappa_{z}. ∎

A.2 Proof of Proposition 1

Proof.

If =Br\ell=Br, then the sum of \ell around any directed cycle telescopes:

k=1Kikik+1=k=1K(rikrik+1)=0.\sum_{k=1}^{K}\ell_{i_{k}i_{k+1}}=\sum_{k=1}^{K}(r_{i_{k}}-r_{i_{k+1}})=0.

Conversely, assume all directed cycle sums are zero. Fix a reference vertex v0v_{0}. For any vertex ii, choose a path PiP_{i} from ii to v0v_{0} and define rir_{i} as the signed sum of \ell along that path, with an edge contributing ab\ell_{ab} when traversed in its chosen orientation (a,b)(a,b) and ab-\ell_{ab} when traversed in reverse. If two paths from ii to v0v_{0} produced different sums, traversing one path and then the reverse of the other would give a closed walk with nonzero total circulation. Removing repeated vertices decomposes that closed walk into simple cycles, at least one of which would have nonzero circulation, contradicting the assumption. Hence rir_{i} is well-defined.

For an oriented edge e=(i,j)e=(i,j), compare the path from ii to v0v_{0} with the path that first traverses ee from ii to jj and then follows the chosen path from jj to v0v_{0}. Path independence gives ri=ij+rjr_{i}=\ell_{ij}+r_{j}, so rirj=ijr_{i}-r_{j}=\ell_{ij}. Thus =Br\ell=Br. If rr and rr^{\prime} both represent \ell, then rirj=rirjr_{i}-r_{j}=r^{\prime}_{i}-r^{\prime}_{j} on every edge, so rrr-r^{\prime} is constant on every connected component. Since GG is connected, the difference is a global additive constant. ∎

A.3 Proof of Corollary 1

Proof.

The first sentence follows by substituting ij=ηij+βij(RiRj)\ell_{ij}=\eta_{ij}+\beta_{ij}(R_{i}^{*}-R_{j}^{*}) into Proposition 1. If ηij=0\eta_{ij}=0 and βij=β\beta_{ij}=\beta, then ij=βRiβRj\ell_{ij}=\beta R_{i}^{*}-\beta R_{j}^{*}, so ri=βRir_{i}=\beta R_{i}^{*} represents the field. If ηij=bibj\eta_{ij}=b_{i}-b_{j} and βij=β\beta_{ij}=\beta, then

ij=(βRi+bi)(βRj+bj),\ell_{ij}=(\beta R_{i}^{*}+b_{i})-(\beta R_{j}^{*}+b_{j}),

so ri=βRi+bir_{i}=\beta R_{i}^{*}+b_{i} represents the field. ∎

A.4 Proof of Proposition 2

Proof.

The objective in (10) is continuous and concave in rr. It is coercive on the normalized subspace 𝟏r=0\mathbf{1}^{\top}r=0: if r\|r\|\to\infty with 𝟏r=0\mathbf{1}^{\top}r=0, then, because BB is injective on 𝟏\mathbf{1}^{\perp} in finite dimensions, Br\|Br\|_{\infty}\to\infty. Since σ(e)(0,1)\sigma(\ell_{e})\in(0,1) for every finite e\ell_{e}, the Bernoulli logistic term on any edge whose fitted logit diverges tends to -\infty, while the remaining edge terms are nonpositive. Hence L(r;)L(r;\ell)\to-\infty along diverging normalized sequences, so a maximizer exists. On 𝟏\mathbf{1}^{\perp}, the objective is strictly concave because GG is connected and all ρe\rho_{e} are positive: if BrBrBr\neq Br^{\prime} on some edge, strict concavity of the Bernoulli log-likelihood is strict along that edge, and if Br=BrBr=Br^{\prime} then rrr-r^{\prime} is constant and therefore zero on the normalized subspace. Hence the maximizer is unique.

The first-order condition is

BW{σ()σ(Br)}=0.B^{\top}W\{\sigma(\ell)-\sigma(Br)\}=0. (30)

At =0\ell=0, the unique normalized solution is r=0r=0. The derivative of the left-hand side of (30) with respect to rr at (,r)=(0,0)(\ell,r)=(0,0) is (1/4)BWB-(1/4)B^{\top}WB, which is nonsingular on 𝟏\mathbf{1}^{\perp} because GG is connected and WW has positive diagonal entries. The implicit-function theorem therefore gives a smooth map r()r^{\dagger}(\ell) in a neighborhood of zero, with r(0)=0r^{\dagger}(0)=0 and r()=O()\|r^{\dagger}(\ell)\|=O(\|\ell\|).

For |t||t| small, σ(t)=1/2+t/4+O(t3)\sigma(t)=1/2+t/4+O(t^{3}); the quadratic term is absent because σ(t)1/2\sigma(t)-1/2 is odd. Since ε\|\ell\|_{\infty}\leq\varepsilon and Br()=O(ε)\|Br^{\dagger}(\ell)\|_{\infty}=O(\varepsilon), substituting this expansion into (30) yields

BW{Br()4}=O(ε3).B^{\top}W\left\{\frac{\ell-Br^{\dagger}(\ell)}{4}\right\}=O(\varepsilon^{3}).

Multiplying by 44 and solving on 𝟏\mathbf{1}^{\perp} gives

r()=(BWB)+BW+O(ε3),r^{\dagger}(\ell)=(B^{\top}WB)^{+}B^{\top}W\ell+O(\varepsilon^{3}),

with a uniform remainder on a sufficiently small neighborhood of zero for the fixed finite graph and fixed positive weights. ∎

A.5 Derivation for Example 1

Proof.

Let rA=xr_{A}=x, rB=yr_{B}=y, and rC=0r_{C}=0. With equal edge weights, the Bradley–Terry fitted probabilities are

pAB=σ(xy),pAC=σ(x),pBC=σ(y).p_{AB}=\sigma(x-y),\qquad p_{AC}=\sigma(x),\qquad p_{BC}=\sigma(y).

The true probabilities corresponding to (12) are

qAB=σ(ε),qAC=σ(ε),qBC=σ(5ε).q_{AB}=\sigma(\varepsilon),\qquad q_{AC}=\sigma(\varepsilon),\qquad q_{BC}=\sigma(5\varepsilon).

The first-order conditions for xx and yy are

(qABpAB)+(qACpAC)\displaystyle(q_{AB}-p_{AB})+(q_{AC}-p_{AC}) =0,\displaystyle=0,
(qABpAB)+(qBCpBC)\displaystyle-(q_{AB}-p_{AB})+(q_{BC}-p_{BC}) =0.\displaystyle=0.

Using σ(t)=1/2+t/4+O(t3)\sigma(t)=1/2+t/4+O(t^{3}) and the smoothness of the optimum from Proposition 2, the first-order terms satisfy

2xy=2ε,2yx=4ε.2x-y=2\varepsilon,\qquad 2y-x=4\varepsilon.

Solving gives x=8ε/3x=8\varepsilon/3 and y=10ε/3y=10\varepsilon/3. The omitted terms are O(ε3)O(\varepsilon^{3}) by the same Taylor expansion and implicit-function argument used in Proposition 2. ∎

A.6 Proof of Proposition 3

Proof.

For arbitrary vv and nonnegative multipliers βij\beta_{ij}, define ηij\eta_{ij} by (14). Then

σ{ηij+βij(vivj)}=σ(ij)=qij,\sigma\{\eta_{ij}+\beta_{ij}(v_{i}-v_{j})\}=\sigma(\ell_{ij})=q_{ij},

which proves the first claim. For the second claim, first consider a compared pair with vivjv_{i}\neq v_{j}. The condition ij(vivj)0\ell_{ij}\,(v_{i}-v_{j})\geq 0 makes (15) nonnegative, and it satisfies

σ{βij(vivj)}=σ(ij)=qij.\sigma\{\beta_{ij}(v_{i}-v_{j})\}=\sigma(\ell_{ij})=q_{ij}.

If vi=vjv_{i}=v_{j}, the stated condition gives ij=0\ell_{ij}=0, so setting, for example, βij=0\beta_{ij}=0 yields σ{βij(vivj)}=σ(0)=qij\sigma\{\beta_{ij}(v_{i}-v_{j})\}=\sigma(0)=q_{ij}. Thus all compared pairs are rationalized with zero defaults. ∎

A.7 Proof of Proposition 4

Proof.

If β=0\beta=0, then [C=1Δ]=σ(η)\mathbb{P}[C=1\mid\Delta]=\sigma(\eta) for every value of Δ\Delta, so the conditional distribution of CC is independent of Δ\Delta and I(Δ;C)=0\mathrm{I}(\Delta;C)=0. If also η=0\eta=0, then σ(η)=1/2\sigma(\eta)=1/2, so CC is a fair coin and 𝖧(C)=log2\mathsf{H}(C)=\log 2.

For the expansion, let q0=σ(η)q_{0}=\sigma(\eta) and s=q0(1q0)s=q_{0}(1-q_{0}). Because the prior has finite support, all Taylor remainders below are uniform over the support of Δ\Delta. Write qβ(Δ)=σ(η+βΔ)q_{\beta}(\Delta)=\sigma(\eta+\beta\Delta). Then

qβ(Δ)=q0+sβΔ+O(β2),q¯β:=𝔼[qβ(Δ)]=q0+sβ𝔼[Δ]+O(β2).q_{\beta}(\Delta)=q_{0}+s\beta\Delta+O(\beta^{2}),\qquad\bar{q}_{\beta}:=\mathbb{E}[q_{\beta}(\Delta)]=q_{0}+s\beta\mathbb{E}[\Delta]+O(\beta^{2}).

The mutual information is

I(Δ;C)=𝔼[DKL(Bernoulli(qβ(Δ))Bernoulli(q¯β))].\mathrm{I}(\Delta;C)=\mathbb{E}\left[D_{\mathrm{KL}}\bigl(\operatorname{Bernoulli}(q_{\beta}(\Delta))\,\|\,\operatorname{Bernoulli}(\bar{q}_{\beta})\bigr)\right].

For Bernoulli parameters uu and vv in a compact subinterval of (0,1)(0,1),

DKL(Bernoulli(u)Bernoulli(v))=(uv)22q0(1q0)+O(|uv|3+|vq0||uv|2).D_{\mathrm{KL}}(\operatorname{Bernoulli}(u)\|\operatorname{Bernoulli}(v))=\frac{(u-v)^{2}}{2q_{0}(1-q_{0})}+O(|u-v|^{3}+|v-q_{0}||u-v|^{2}).

Here

qβ(Δ)q¯β=sβ(Δ𝔼Δ)+O(β2),q_{\beta}(\Delta)-\bar{q}_{\beta}=s\beta(\Delta-\mathbb{E}\Delta)+O(\beta^{2}),

so

I(Δ;C)=12s𝔼[s2β2(Δ𝔼Δ)2]+O(β3)=12sβ2Var(Δ)+O(β3),\mathrm{I}(\Delta;C)=\frac{1}{2s}\mathbb{E}\left[s^{2}\beta^{2}(\Delta-\mathbb{E}\Delta)^{2}\right]+O(\beta^{3})=\frac{1}{2}s\beta^{2}\operatorname{Var}(\Delta)+O(\beta^{3}),

which is (17). ∎

A.8 Proof of Proposition 5

Proof.

For a Bernoulli observation with parameter q=σ(s)q=\sigma(s) and scalar index ss, the score with respect to ss is CqC-q and the Fisher information for ss is q(1q)q(1-q). By the chain rule, for any parameter vector θ\theta entering only through s(θ)s(\theta),

θ=q(1q)s(θ)s(θ).\mathcal{I}_{\theta}=q(1-q)\,\nabla s(\theta)\nabla s(\theta)^{\top}.

If s=η+βΔs=\eta+\beta\Delta and η,β\eta,\beta are known, then s/Δ=β\partial s/\partial\Delta=\beta, giving (19). Since q(1q)1/4q(1-q)\leq 1/4, the upper bound follows. If η\eta is known and θ=(Δ,β)\theta=(\Delta,\beta), then s=(β,Δ)\nabla s=(\beta,\Delta)^{\top}, giving (20). If θ=(η,Δ,β)\theta=(\eta,\Delta,\beta), then s=(1,β,Δ)\nabla s=(1,\beta,\Delta)^{\top}, giving (21). Each displayed matrix is an outer product times the positive scalar q(1q)q(1-q), so its rank is at most one. If the model constrains β0\beta\geq 0, the regular Fisher-information interpretation for the parameters involving β\beta is at interior values β>0\beta>0; the same displayed derivatives remain the local logit-index derivatives at the boundary. ∎

A.9 Proof of Proposition 6

Proof.

If β=0\beta=0, then both hypotheses generate Bernoulli(1/2)\operatorname{Bernoulli}(1/2) labels, so P+=PP_{+}=P_{-} and every test has worst-case error probability at least 1/21/2. Now assume β>0\beta>0. Let a=βδa=\beta\delta and q=σ(a)q=\sigma(a). Under Δ=δ\Delta=\delta, one label is Bernoulli(q)\operatorname{Bernoulli}(q); under Δ=δ\Delta=-\delta, one label is Bernoulli(1q)\operatorname{Bernoulli}(1-q). Thus

DKL(P+P)\displaystyle D_{\mathrm{KL}}(P_{+}\|P_{-}) =qlogq1q+(1q)log1qq\displaystyle=q\log\frac{q}{1-q}+(1-q)\log\frac{1-q}{q}
=(2q1)logq1q=atanh(a/2).\displaystyle=(2q-1)\log\frac{q}{1-q}=a\tanh(a/2).

Because tanh(a/2)a/2\tanh(a/2)\leq a/2 for a0a\geq 0, the upper bound in (23) follows.

For nn independent labels, product additivity gives

DKL(P+nPn)=nDKL(P+P).D_{\mathrm{KL}}(P_{+}^{\otimes n}\|P_{-}^{\otimes n})=nD_{\mathrm{KL}}(P_{+}\|P_{-}).

Let s^{+,}\widehat{s}\in\{+,-\} be any test and let α\alpha bound its worst-case error probability. The total variation distance between the two product distributions must satisfy

TV(P+n,Pn)12α,\mathrm{TV}(P_{+}^{\otimes n},P_{-}^{\otimes n})\geq 1-2\alpha,

because the optimal testing error equals (1TV)/2(1-\mathrm{TV})/2 and no test can outperform the optimal test. Pinsker’s inequality gives

TV(P+n,Pn)12nDKL(P+P).\mathrm{TV}(P_{+}^{\otimes n},P_{-}^{\otimes n})\leq\sqrt{\frac{1}{2}nD_{\mathrm{KL}}(P_{+}\|P_{-})}.

Combining the two displays yields

nDKL(P+P)2(12α)2,nD_{\mathrm{KL}}(P_{+}\|P_{-})\geq 2(1-2\alpha)^{2},

which gives the first lower bound in (24). Substituting DKL(P+P)β2δ2/2D_{\mathrm{KL}}(P_{+}\|P_{-})\leq\beta^{2}\delta^{2}/2 gives the second. If β\beta can be arbitrarily close to zero, the necessary sample size can be made arbitrarily large. ∎

A.10 Proof of Theorem 1

Proof.

Let Ht1H_{t-1} denote the history before label tt, including past labels, chosen queries, and any external randomization used by the adaptive query rule. Query choices and external randomization add no information about θ\theta except through past labels because the randomization is independent of θ\theta conditional on the past. Thus the chain rule for mutual information gives

I(θ;Hn)=t=1nI(θ;CtHt1).\mathrm{I}(\theta;H_{n})=\sum_{t=1}^{n}\mathrm{I}(\theta;C_{t}\mid H_{t-1}).

For any fixed history ht1h_{t-1}, the conditional mutual information between a discrete parameter and one observation is bounded by the average pairwise KL divergence among the conditional observation laws, and hence by their maximum. Assumption (25) therefore implies

I(θ;CtHt1=ht1)dt\mathrm{I}(\theta;C_{t}\mid H_{t-1}=h_{t-1})\leq d_{t}

for every history with positive probability under the joint mixture distribution. Taking expectations over histories and summing gives

I(θ;Hn)t=1ndt.\mathrm{I}(\theta;H_{n})\leq\sum_{t=1}^{n}d_{t}.

Fano’s inequality for a uniform parameter on MM hypotheses states that any estimator based on the observed history satisfies

[θ^θ]1I(θ;Hn)+log2logM.\mathbb{P}[\widehat{\theta}\neq\theta]\geq 1-\frac{\mathrm{I}(\theta;H_{n})+\log 2}{\log M}.

Combining the two displays proves (26). ∎

A.11 Proof of Proposition 7

Proof.

The vector rHr^{H} in (27) is the weighted least-squares projection coefficient of \ell onto im(B)\operatorname{im}(B) under the normalization orthogonal to constants. The normal equations are

BW(BrH)=0,B^{\top}W(\ell-Br^{H})=0,

which is BWcyc=0B^{\top}W\ell^{\mathrm{cyc}}=0. Hence pot𝒫\ell^{\mathrm{pot}}\in\mathcal{P} and cyc𝒫W\ell^{\mathrm{cyc}}\in\mathcal{P}^{\perp_{W}}, proving the weighted orthogonal decomposition and the Pythagorean identity (28).

For the KL statement, first note that rH=O(ε)r^{H}=O(\varepsilon) because it is a fixed linear map applied to \ell. The weighted KL objective differs from L(r;)-L(r;\ell) in (10) only by the constant

eρe{σ(e)logσ(e)+(1σ(e))logσ(e)},\sum_{e}\rho_{e}\{\sigma(\ell_{e})\log\sigma(\ell_{e})+(1-\sigma(\ell_{e}))\log\sigma(-\ell_{e})\},

which is independent of rr. Hence its centered minimizer is the same as the centered maximizer in Proposition 2, and that proposition localizes the minimizer in an O(ε)O(\varepsilon) neighborhood of zero. For |u|,|v|Cε|u|,|v|\leq C\varepsilon, a Taylor expansion of Bernoulli KL around (u,v)=(0,0)(u,v)=(0,0) in log-odds coordinates gives

DKL(Bernoulli(σ(u))Bernoulli(σ(v)))=18(uv)2+O(ε3),D_{\mathrm{KL}}\bigl(\operatorname{Bernoulli}(\sigma(u))\,\|\,\operatorname{Bernoulli}(\sigma(v))\bigr)=\frac{1}{8}(u-v)^{2}+O(\varepsilon^{3}), (31)

with a uniform remainder. Therefore

infr: 1r=0eρeDKL(Bernoulli(σ(e))Bernoulli(σ((Br)e)))\displaystyle\inf_{r:\,\mathbf{1}^{\top}r=0}\sum_{e}\rho_{e}D_{\mathrm{KL}}\bigl(\operatorname{Bernoulli}(\sigma(\ell_{e}))\,\|\,\operatorname{Bernoulli}(\sigma((Br)_{e}))\bigr) =18infr: 1r=0BrW2+O(ε3)\displaystyle=\frac{1}{8}\inf_{r:\,\mathbf{1}^{\top}r=0}\|\ell-Br\|_{W}^{2}+O(\varepsilon^{3})
=18cycW2+O(ε3),\displaystyle=\frac{1}{8}\|\ell^{\mathrm{cyc}}\|_{W}^{2}+O(\varepsilon^{3}),

where the last equality is the definition of the weighted projection residual. ∎