arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00251v1 [cs.AI] 31 Aug 2026

Hypotheses-Guided Self Distillation
for Continual Personalization

EunJeong Hwang ††thanks: Work done as an intern at Megagon Labs. Affiliation: University of British Columbia Email: ejhwang@cs.ubc.ca    Kushan Mitra Affiliation: Megagon Labs, USA Email: kushan@megagon.ai    Dan Zhang Affiliation: Megagon Labs, USA Email: danz@megagon.ai    Hannah Kim Affiliation: Megagon Labs, USA Email: hannah@megagon.ai    Estevam Hruschka Affiliation: Megagon Labs, USA Email: estevam@megagon.ai
Abstract

As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, latent, and noisy signals, with existing methods relying on raw interaction histories or costly reward-based optimization to manage personalization. We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation. Experiments across three personalization settings: online personalization, multi-session interactions, and implicit behavioral signals, show that HypReflect outperforms a range of baselines, including raw-history and incremental-update methods. We further demonstrate strong generalization to unseen users and cross-domain settings, along with stability across context budgets, reusable hypotheses, and more focused personalization. These results suggest a step towards reliable and scalable continual personalization through explicit, revisable user preference hypotheses.11 1 Code will be released upon acceptance.

Refer to caption
Figure 1: Preference hypotheses are generated across personalization settings from explicit feedback, implicit behavioral signals, and user profile information, along with confidence weights. They maintain explicit, revisable beliefs about the user and guide models to produce personalized responses.

1 Introduction

Conversational assistants that continuously adapt to individual users can provide increasingly effective and personalized support in diverse settings, including writing assistance (Mysore et al., 2024), conversational recommendation (Liang et al., 2024; Kim et al., 2025; Zhu et al., 2025), and education (Liu et al., 2025a). Personalization allows assistants to better address a user’s needs (Zhang et al., 2018; Salemi et al., 2024; Zhang et al., 2025; Zhao et al., 2025c), while continual adaptation enables long-term interactions with users by improving relevance, as assistants develop a richer understanding of users over time (Xu et al., 2022; Zhong et al., 2024; Li et al., 2025).

However, users rarely articulate their preferences in full and may not always be consciously aware of them, making continual adaptation inherently challenging. Figure 1 illustrates how evidence about a user’s preferences emerges incrementally from explicit feedback, implicit behavior, and user-profile information, requiring it to be consolidated into an explicit representation that can be revised as new evidence arrives. These signals are often sparse, noisy, and context dependent, making it difficult to distinguish stable preferences from task-specific requirements. Moreover, evidence may be distributed across long multi-turn and multi-session histories (Maharana et al., 2024; Wu et al., 2025), which language models often struggle to track consistently as conversations grow (Laban et al., 2025). The central problem is therefore to transform an evolving stream of interactions into a user representation that is both stable and revisable.

Existing approaches personalize LLMs by conditioning on user profiles, retrieved memories, or raw interaction histories (Packer et al., 2024; Zhao et al., 2025a; Li et al., 2026), making it increasingly costly to retain and retrieve past interactions over time. Other methods infer user traits or summaries through RL but require reward models, making continual optimization costly (Wan et al., 2025; Nam et al., 2026). Buening et al. (2026) applies self-distillation, improving scalability for continual personalization, but relies on raw interaction histories, leaving it unclear which preferences are learned and whether they generalize to future contexts.

In this work, we introduce HypReflect, a scalable framework for continual personalization that maintains preference hypotheses i.e. explicit, uncertainty-aware, and revisable beliefs about the user. HypReflect distills diverse user signals (e.g., Fig. 1) into preference hypotheses, reflectively consolidates useful user signals across interactions, and incorporates the refined hypotheses to guide personalized response generation through hypotheses-guided self-distillation.

We evaluate HypReflect across three personalization settings covering explicit feedback, multi-session interactions, and implicit behavioral choices, against a comprehensive suite of baselines, including self-distillation over raw interaction histories, incrementally updated summaries and hypotheses, and reflective prompting. Across these settings, HypReflect achieves relative improvements of up to 4.4%, 10.3%, and 5.0% respectively22 2 Maximum relative improvement per dataset with HypReflect+Sum in Table 1., while generalizing effectively to unseen users, shifted user characteristics, and cross-domain personalization. Further analyses demonstrate robustness across context budgets, hypothesis reusability, and more focused personalization. Overall, our findings show that reflective refinement enables stable and revisable preference hypotheses over long interaction histories, and that models benefit substantially from explicit, reusable user representations, enabling reliable and scalable continual personalization.

2 Problem Formulation

Continual Personalization

considers an assistant that repeatedly interacts with a user. At interaction tt, the assistant observes the preceding conversation history C<tC_{<t} and the current user turn utu_{t}, which may contain feedback ft−1f_{t-1} on the preceding assistant response, a new request qtq_{t}, or both. The assistant then generates a response:

yt∼πθ(⋅∣C<t,ut).y_{t}\sim\pi_{\theta}(\cdot\mid C_{<t},u_{t}).

As interactions accumulate, the history provides increasing signals about the user’s preferences. The goal therefore is to continually infer, maintain, and use this understanding to personalize future responses.

Self-Distillation for Continual Personalization.

Self-distillation uses the same model as both a student and a teacher, with the teacher conditioned on richer information. For continual personalization, self-distillation offers a simple and scalable alternative to costly reward optimization by treating the subsequent user turn ut+1={ft,qt+1}u_{t+1}=\{f_{t},q_{t+1}\} as privileged information when ground-truth personalized responses are unavailable (Buening et al., 2026). The student and teacher distributions are

πtS=πθ(⋅∣C<t,ut), πtT=πθ(⋅∣C<t,ut,ut+1),\pi_{t}^{\mathrm{S}}=\pi_{\theta}(\cdot\mid C_{<t},u_{t}),\text{ }\pi_{t}^{\mathrm{T}}=\pi_{\theta}(\cdot\mid C_{<t},u_{t},u_{t+1}),

and the teacher’s feedback-informed predictions are distilled into the student by minimizing

ℒSD​(θ)=∑tDKL​(πtS∥stopgrad⁡[πtT]).\mathcal{L}_{\mathrm{SD}}(\theta)=\sum_{t}D_{\mathrm{KL}}\left(\pi_{t}^{\mathrm{S}}\middle\|\operatorname{stopgrad}\left[\pi_{t}^{\mathrm{T}}\right]\right).

This enables learning from preferences revealed by subsequent interactions without an external teacher or reward model. However, directly internalizing preferences from raw interaction histories may not generalize well beyond the current interaction.

Preference Hypotheses Representation.

Hence, we introduce an explicit abstraction layer between raw interactions and response generation. We represent the model’s understanding of the user as a set of preference hypotheses HtH_{t} and formulate personalized response generation as

yt∼πθ(⋅∣C<t,qt,Ht).y_{t}\sim\pi_{\theta}(\cdot\mid C_{<t},q_{t},H_{t}).

This representation makes inferred preferences explicit, reusable, and revisable as new evidence arrives (Liu et al., 2026). The key challenge here is how to construct hypotheses from local interaction signals, refine and maintain them over long histories, and use them to guide self-distillation.

3 HypReflect: Continual User Modeling through Reflective Refinement

Refer to caption
Figure 2: Overview of HypReflect. Preference hypotheses are generated from context-bounded chunks (Steps 1–2) and reflectively refined over time (Step 3). The refined user state conditions both student and teacher during self-distillation, with the teacher seeing the next user turn as privileged information (Step 4). The resulting model produces continually personalized responses (Step 5).

To support continual personalization, we propose HypReflect, a framework that introduces an explicit preference hypothesis layer to better understand and respond to the user. Our framework consists of three stages. First, the model infers preference hypotheses from individual interaction chunks, capturing localized signals about the user (§3.1). Second, it recurrently refines these hypotheses by reflecting and consolidating user signals across chunks (§3.2). Finally, the refined hypotheses guide self-distillation, enabling the model to use subsequent user signals as privileged supervision for learning personalized response behavior (§3.3). We also consider a summary-augmented variant that jointly maintains summaries and hypotheses.

3.1 Inferring Preference Hypotheses from Interactions

Hypotheses representation.

To capture diverse interaction signals from the user (e.g., Fig. 1), we represent the model’s understanding of the user as a revisable set of preference hypotheses HtH_{t}, comprising mm multiple candidate hypotheses paired with verbalized confidence scores:

Ht={ht​1,…,ht​m},ht​j=(st​j,wt​j),H_{t}=\{h_{t1},\ldots,h_{tm}\},\qquad h_{tj}=(s_{tj},w_{tj}),

where st​js_{tj} is a natural-language statement describing a potential user preference and wt​j(∈{1,⋯,5})w_{tj}(\in\{1,\cdots,5\}) denotes the model’s confidence in that statement. Maintaining multiple hypotheses preserves alternative beliefs that can be revised as new evidence arrives, while confidence distinguishes well-supported preferences from uncertain ones. The natural-language representation makes the hypotheses interpretable, revisable, and directly usable for personalized response generation.

Inferring local hypotheses from interaction chunks.

Inferring preferences from the full interaction history is often impractical due to context-length limitations and, in online settings, because future interactions are not yet available. We therefore partition the observed interaction history into consecutive chunks {C1,…,CN}\{C_{1},\ldots,C_{N}\}, each containing the longest sequence of complete interactions within a token budget BB. For each chunk CiC_{i}, we infer a set of local hypotheses (Step 2 in Fig. 2) as

Ht=GenerateHypotheses​(Ct)={ht​j}j=1m.H_{t}=\textsc{GenerateHypotheses}(C_{t})=\{h_{tj}\}_{j=1}^{m}.

The model is instructed to infer reusable preferences grounded in the user’s instructions, corrections, feedback, and choices, while avoiding interaction-specific restatements and unsupported generalizations.

We additionally define a summary-augmented variant, where we generate a local summary SiS_{i} for each interaction chunk. Summaries preserve relevant conversational context from CiC_{i} that grounds the hypotheses, whereas HiH_{i} captures reusable preferences that are explicitly stated or inferred from user signals.

3.2 Maintaining User Understanding through Reflective Refinement

As interactions accumulate, the model must maintain preference hypotheses that (1) explain past interaction histories and (2) predict future user behavior. Simply using the most recent hypotheses discards earlier preferences, whereas repeatedly reconsidering the full interaction history is computationally infeasible. A refinement mechanism should therefore preserve long-term preferences while remaining revisable and bounded in computational context size.

A simple approach is incremental updating, which iteratively refines a single running user state. However, since each update operates on an already compressed representation, over time it is either prone to error accumulation, or it may have valid earlier preferences overwritten by recent observations. We empirically show that this leads to unstable long-term personalization (§5.1).

Reflective refinement across chunks.

To preserve evidence across multiple chunks, without relying exclusively on a recursively compressed state, we propose reflective refinement (Step 3 in Fig. 2). Rather than immediately merging each new local hypothesis set into a single running state, the model jointly reconsiders multiple independently generated local sets. By comparing hypotheses across interaction chunks, the model consolidates consistent evidence, resolves conflicts, merges redundant hypotheses, and revises beliefs that are no longer supported.

To keep refinement tractable, we maintain a bounded reflection window. Let HiH_{i} denote the local hypothesis set inferred from chunk CiC_{i}, Hi∗H_{i}^{*} the refined hypothesis set after observing the first ii chunks, and WW the reflection-window size. We construct the refinement input as

ℐi={[H1,…,Hi],i≤W,[Hi−W+1∗,Hi−W+2,…,Hi],i>W,\mathcal{I}_{i}=\begin{cases}[H_{1},\ldots,H_{i}],&i\leq W,\\[4.0pt] [H_{i-W+1}^{*},H_{i-W+2},\ldots,H_{i}],&i>W,\end{cases}\vskip 5.0pt

and update the maintained hypothesis set as Hi∗=ReflectAndRefine​(ℐi)H_{i}^{*}=\textsc{ReflectAndRefine}(\mathcal{I}_{i}).

When i≤Wi\leq W, the model reflects over all available local hypothesis sets. Thereafter, it replaces older local sets with their refined representation while retaining the most recent W−1W-1 local sets. This keeps the refinement input bounded without discarding earlier preference information.

For our summary-augmented variant, we apply the same windowed recurrent refinement procedure to a jointly maintained summary and hypothesis set, Zi=(Si,Hi)Z_{i}=(S_{i},H_{i}), for each chunk, producing Zi∗=(Si∗,Hi∗)Z_{i}^{*}=(S_{i}^{*},H_{i}^{*}). The refined summary Si∗S_{i}^{*} consolidates the context used to ground and revise the hypotheses, while Hi∗H_{i}^{*} maintains the refined preferences used for future personalization.

3.3 Hypothesis-Guided Self-Distillation

To effectively use the refined user state for personalized response generation, we incorporate it into self-distillation, training the model to translate maintained preferences into personalized responses. Building on the formulation in §2, we condition both the teacher and student on the same refined state, providing a shared understanding of the user and making the teacher’s targets learnable by the student. The teacher additionally observes the subsequent user turn ut+1=(ft,qt+1)u_{t+1}=(f_{t},q_{t+1}) and uses it as privileged hindsight to generate a preference-grounded response.

At interaction tt, given the preceding history C<tC_{<t}, current user turn utu_{t}, refined user hypotheses Ht∗H_{t}^{*}, and subsequent user turn ut+1u_{t+1}, the student and teacher distributions are

PS​(yt)\displaystyle P_{S}(y_{t}) =πθ(yt∣C<t,ut,Ht∗), PT(yt)=πθ(yt∣C<t,ut,Ht∗,ut+1).\displaystyle=\pi_{\theta}\left(y_{t}\mid C_{<t},u_{t},H_{t}^{*}\right),\text{ }P_{T}(y_{t})=\pi_{\theta}\left(y_{t}\mid C_{<t},u_{t},H_{t}^{*},u_{t+1}\right).

The teacher uses ut+1u_{t+1} to provide feedback-informed supervision. We then minimize the reverse KL divergence from the student to the detached teacher:

ℒdistill=𝔼t[DKL(PS(⋅)∥stopgrad[PT(⋅)])].\mathcal{L}_{\mathrm{distill}}=\mathbb{E}_{t}\left[D_{\mathrm{KL}}\left(P_{S}(\cdot)\,\|\,\operatorname{stopgrad}[P_{T}(\cdot)]\right)\right].

The objective is computed token-wise over the generated response (Step 4 in Fig. 2).

4 Experimental Setup

4.1 Datasets and Evaluation

To test whether models can infer reusable user preferences from diverse interaction signals and use them to personalize future responses, we evaluate three complementary settings: continual adaptation from explicit user feedback, multi-session personalization using conversational user profiles, and preference inference from implicit behavioral signals. See Appendix C.1 and C.3 more details.

HelpSteer2

(Wang et al., 2024). To evaluate continual adaptation from explicit user feedback, we construct an online personalization setting from HelpSteer2. Each user is associated with a set of writing-style preferences, and across multiple interactions, a user simulator provides feedback on one or two relevant preference dimensions after each model response.
Evaluation Models are trained over 250 interactions and evaluated every 50 interactions on held-out prompts using an LLM judge. We report pairwise win rates against the base model with respect to the target user’s preferences.

HiCupid

(Mok et al., 2025). To test personalization from user profile information, we use HiCupid, where user preferences and attributes are stated naturally across multi-session conversations. We train a model on QA interactions from 300 users and evaluate whether it can use previous user information to personalize future responses.
Evaluation We evaluate both seen-user and unseen-user generalization using held-out prompts and report pairwise win rates from an LLM judge33 3 We use deepseek-v4-flash as the LLM judge and verify its consistency with other judges in Appendix D. assessing personalization and response quality.

Flight Recommendation

(Qiu et al., 2026). To evaluate preference inference from implicit behavioral signals, we use a Flight Recommendation task. We utilize multi-turn recommendation trajectories from 624 users, with up to 25 interactions per user, where user choices provide partial evidence about latent preferences. These preferences may include non-obvious trade-offs, requiring models to infer user-specific behavior rather than rely on fixed assumptions.
Evaluation We report final-round preference prediction accuracy on the original flight task, held-out feature settings, and cross-domain hotel transfer.

4.2 Model Variants and Training Setup

(1) Base. Uses the vanilla language model without fine-tuning, conditioned on the relevant raw interaction history, to generate a response. (2) SD. Self-distillation is performed directly on the raw interaction history. When the history exceeds the context budget, only the most recent portion is retained; (3) SD-IncSum. The model is trained with an incrementally updated summary: a running summary SiS_{i} is maintained and refreshed after each chunk; (4) SD-IncHyp. Instead of a summary, the model maintains a running hypothesis set HiH_{i}, updated with the same incremental strategy; (5) HypReflect. Local hypotheses HiH_{i} are first generated from individual chunks and then consolidated into a refined set Hi∗H_{i}^{*} through the reflective refinement process described in §3.2; (6) HypReflect+Sum. Reflective refinement is applied to the paired memory state Zi=(Si,Hi)Z_{i}=(S_{i},H_{i}), yielding Zi∗=(Si∗,Hi∗)Z_{i}^{*}=(S_{i}^{*},H_{i}^{*}), to examine whether summaries supply complementary information during hypothesis refinement.

Training setup.

We use Qwen3.5-4B, Qwen3.5-9B (Yang et al., 2025), and Gemma4-4B (Team et al., 2026) as backbone LLMs. Additional hyperparameter details are provided in Appendix C.2.

5 Results

Table 1: Performance on three continual-personalization benchmarks covering continual adaptation with explicit user feedback (HelpSteer2), user-profile information across multiple sessions (HiCupid), and implicit behavioral signals (Flight). We report win rate against Base for HelpSteer2 (averaged over 50–250 interactions) and HiCupid, and accuracy for Flight.
HelpSteer2 (Win Rate ↑\uparrow) HiCupid (Win Rate ↑\uparrow) Flight Recommendation (Acc. ↑\uparrow)
Method Gemma4-4B Qwen3.5-4B Qwen3.5-9B Gemma4-4B Qwen3.5-4B Qwen3.5-9B Gemma4-4B Qwen3.5-4B Qwen3.5-9B
Base – – – – – – 40.9 40.3 41.7
SD 87.0 84.5 89.0 57.3 69.7 63.3 61.8 73.8 74.4
SD-IncSum 85.8 (−1.2↓)(-1.2\downarrow) 81.0 (−3.5↓)(-3.5\downarrow) 86.6 (−2.4↓)(-2.4\downarrow) 58.4 (+1.1↑)(+1.1\uparrow) 67.9 (−1.8↓)(-1.8\downarrow) 67.7 (+4.4↑)(+4.4\uparrow) 50.5 (−11.3↓)(-11.3\downarrow) 74.4 (+0.6↑)(+0.6\uparrow) 67.0 (−7.4↓)(-7.4\downarrow)
SD-IncHyp 89.6 (+2.6↑)(+2.6\uparrow) 81.9 (−2.6↓)(-2.6\downarrow) 88.2 (−0.8↓)(-0.8\downarrow) 60.7 (+3.4↑)(+3.4\uparrow) 71.6 (+1.9↑)(+1.9\uparrow) 66.2 (+2.9↑)(+2.9\uparrow) 57.3 (−4.5↓)(-4.5\downarrow) 74.8 (+1.0↑)(+1.0\uparrow) 75.4 (+1.0↑)(+1.0\uparrow)
HypReflect 86.3 (−0.7↓)(-0.7\downarrow) 88.9 (+4.4↑)(+4.4\uparrow) 90.3 (+1.3↑)(+1.3\uparrow) 61.1 (+3.8↑)(+3.8\uparrow) 71.1 (+1.4↑)(+1.4\uparrow) 67.7 (+4.4↑)(+4.4\uparrow) 62.3 (+0.5↑)(+0.5\uparrow) 74.0 (+0.2↑)(+0.2\uparrow) 76.0 (+1.6↑)(+1.6\uparrow)
HypReflect+Sum 90.0 (+3.0↑)(+3.0\uparrow) 88.2 (+3.7↑)(+3.7\uparrow) 92.1 (+3.1↑)(+3.1\uparrow) 61.9 (+4.6↑)(+4.6\uparrow) 71.8 (+2.1↑)(+2.1\uparrow) 69.8 (+6.5↑)(+6.5\uparrow) 64.9 (+3.1↑)(+3.1\uparrow) 74.2 (+0.4↑)(+0.4\uparrow) 75.3 (+0.9↑)(+0.9\uparrow)

5.1 Main Results

Preference hypotheses provide a stronger representation than summaries.

Table 1 compares performance across three personalization benchmarks. Under the incremental update strategy, SD-IncHyp outperforms SD-IncSum in eight of nine settings, showing that abstracting interaction signals into reusable preference beliefs is generally more effective than retaining them as summaries. HypReflect further improves over SD-IncHyp and outperforms SD in eight of nine settings, indicating that the effectiveness of preference hypotheses also depends on maintaining them accurately and consistently over time.

Reflective refinement makes personalization more reliable.

Furthermore, HypReflect outperforms SD in eight of nine settings and HypReflect+Sum improves over it in all nine. HypReflect+Sum also outperforms HypReflect in seven settings, suggesting that summaries provide complementary evidence omitted during hypothesis abstraction and ground refinement in both observed interactions and inferred preferences.

In contrast, SD-IncSum and SD-IncHyp underperform SD in six and three settings, respectively. We find that incremental updates can propagate earlier errors and allow recent observations to overwrite valid prior information, resulting in unstable supervision over time. In contrast, reflective refinement maintains more reliable preference hypotheses as interaction histories grow.

Table 2: Comparison of HypReflect against the base model and Base+RefHyp (base model reflective prompting approach).
HiCupid Flight Recommendation
Win Rate of HypReflect ↑\uparrow Accuracy ↑\uparrow
Model vs. Base vs. Base+RefHyp Base Base+RefHyp HypReflect
Gemma4-4B 61.1 59.7 40.9 47.9 62.3
Qwen3.5-4B 71.1 62.2 40.3 42.0 74.0
Qwen3.5-9B 67.7 57.6 41.7 46.8 76.0
Reflective hypotheses prompting yields consistent gains, which self-distillation further amplifies.

Table 2 demonstrates the effectiveness of reflective refinement both as an inference-time prompting method and when combined with self-distillation. We implement Base+RefHyp similarly to HypReflect, reflectively refining preference hypotheses to generate a response but without self-distillation. Even when used only through prompting, reflectively refining preference hypotheses improves the base model, with Base+RefHyp gaining an average of 4.6 accuracy points on Flight Recommendation and narrowing the gap to HypReflect on HiCupid, reducing the average win rate of HypReflect from 66.6% against Base to 59.8%. Combining reflective refinement with hypothesis-guided self-distillation yields substantially larger gains, as HypReflect outperforms Base+RefHyp by an average of 25.2 accuracy points on Flight Recommendation and achieves an average win rate of 59.8% against it on HiCupid. These results show that reflective refinement is effective even as inference-time prompting and combining it with self-distillation more effectively incorporates user preferences into personalized response behavior.

Table 3: Generalization results. Flight→\rightarrowFeatures: robustness across 2–8 features (with models trained on 4 features). Flight→\rightarrowHotel: cross-domain transfer. HiCupid Unseen: performance on unseen users.
Flight →\rightarrow Features (Acc. ↑\uparrow) Flight →\rightarrow Hotel (Acc. ↑\uparrow) HiCupid Unseen (Win Rate ↑\uparrow)
Method Gemma4-4B Qwen3.5-4B Qwen3.5-9B Gemma4-4B Qwen3.5-4B Qwen3.5-9B Gemma4-4B Qwen3.5-4B Qwen3.5-9B
SD 55.4 64.9 63.7 54.8 61.8 62.0 57.0 71.7 66.9
SD-IncSum 47.5 (−7.9↓)(-7.9\downarrow) 65.6 (+0.7↑)(+0.7\uparrow) 57.1 (−6.6↓)(-6.6\downarrow) 48.0 (−6.8↓)(-6.8\downarrow) 63.7 (+1.9↑)(+1.9\uparrow) 57.1 (−4.9↓)(-4.9\downarrow) 58.7 (+1.7↑)(+1.7\uparrow) 69.9 (−1.8↓)(-1.8\downarrow) 69.4 (+2.5↑)(+2.5\uparrow)
SD-IncHyp 53.4 (−2.0↓)(-2.0\downarrow) 64.8 (−0.1↓)(-0.1\downarrow) 65.5 (+1.8↑)(+1.8\uparrow) 52.8 (−2.0↓)(-2.0\downarrow) 62.3 (+0.5↑)(+0.5\uparrow) 62.9 (+0.9↑)(+0.9\uparrow) 62.1 (+5.1↑)(+5.1\uparrow) 71.8 (+0.1↑)(+0.1\uparrow) 68.0 (+1.1↑)(+1.1\uparrow)
HypReflect 55.6 (+0.2↑)(+0.2\uparrow) 65.8 (+0.9↑)(+0.9\uparrow) 66.0 (+2.3↑)(+2.3\uparrow) 55.7 (+0.9↑)(+0.9\uparrow) 63.6 (+1.8↑)(+1.8\uparrow) 63.8 (+1.8↑)(+1.8\uparrow) 62.9 (+5.9↑)(+5.9\uparrow) 71.6 (−0.1↓)(-0.1\downarrow) 68.1 (+1.2↑)(+1.2\uparrow)
HypReflect+Sum 57.7 (+2.3↑)(+2.3\uparrow) 65.0 (+0.1↑)(+0.1\uparrow) 65.9 (+2.2↑)(+2.2\uparrow) 56.5 (+1.7↑)(+1.7\uparrow) 64.1 (+2.3↑)(+2.3\uparrow) 63.7 (+1.7↑)(+1.7\uparrow) 62.6 (+5.6↑)(+5.6\uparrow) 73.3 (+1.6↑)(+1.6\uparrow) 70.3 (+3.4↑)(+3.4\uparrow)

5.2 Generalization Results

Reflective refinement improves robustness under distribution shift.

Table 3 reports results on three generalization settings. On HiCupid, we evaluate models on users unseen during training. On the Flight Recommendation dataset, we test models trained with 4 features on varying numbers of features (2–8) and a different Hotel domain.

Reflective methods remain the most robust. HypReflect+Sum improves over SD in all nine settings, often by substantial margins, and HypReflect outperforms SD in eight out of nine settings. In contrast, SD-IncSum produces large drops on the Flight transfer tasks. SD-IncHyp is more stable and improves performance in most cases, yet still underperforms the reflective approaches on several Flight settings.

These results show that reflective refinement produces higher-quality hypotheses, providing more reliable conditioning signals during self-distillation, allowing the model to learn a cleaner mapping from user behavior to underlying preferences.

5.3 Further Analysis

Table 4: Evaluation of the quality of the final user hypotheses generated by HypReflect, by transferring them to non-finetuned base models.
Model Flight (Acc.) HiCupid (Win Rate)
Base w/ Refined Hyp. Base w/ Refined Hyp.
Gemma4-4B 40.9 55.1+14.2\mathbf{55.1}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+14.2}} 50.0 53.7+3.7\mathbf{53.7}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+3.7}}
Qwen3.5-4B 40.3 48.8+8.5\mathbf{48.8}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+8.5}} 50.0 63.4+13.4\mathbf{63.4}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+13.4}}
Qwen3.5-9B 41.7 53.6+11.9\mathbf{53.6}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+11.9}} 50.0 62.8+12.8\mathbf{62.8}_{{\color[rgb]{0,0.5,0.5}\scriptscriptstyle+12.8}}
Reflective refinement produces reusable user hypotheses.

Table 4 presents our analysis on the quality of the user hypotheses learned by HypReflect through the hypothesis-transfer experiment. For each user, we take the final hypothesis generated by the HypReflect model and provide it to the corresponding non-finetuned base model for answer generation. During evaluation, each base model receives the hypothesis generated for the same user. The transferred hypotheses substantially improve performance across all models on both datasets, yielding average gains of 11.5 accuracy points on Flight Recommendation and 10.0 winrate points on HiCupid. These results indicate that our reflective refinement approach produces informative and reusable user hypotheses, suggesting that the resulting personalization benefits may extend beyond the model used to generate them.

Effect of reflection before consolidation.
Table 5: Performance of HypReflect with and without reflection during refinement.
Dataset Method Gemma4-4B Qwen3.5-4B
HiCupid HypReflect w/o refl. 72.8 83.8
HypReflect 74.2 84.5
Flight HypReflect w/o refl. 59.0 74.2
HypReflect 62.9 76.0

Table 5 presents an ablation on HiCupid and Flight Recommendation to isolate the effect of the reflection step. For HiCupid, we evaluate users with long interaction histories. Our refinement step first reflects on the inferred hypothesis sets ℐ\mathcal{I} and then consolidates them into an updated hypothesis set H∗H^{*} based on reflection. We compare direct consolidation of ℐ\mathcal{I} without reflection. Adding reflection before consolidation yields modest but consistent average gains of 1.0 and 2.8 points on HiCupid and Flight Recommendation, respectively, over direct consolidation of locally inferred hypothesis sets. Our qualitative analyses on both datasets in Appendix E.1 show that direct consolidation can retain broad, overlapping, or redundant preferences, whereas reflection before consolidation reorganizes the accumulated evidence into more specific and informative hypotheses. On Flight Recommendation, reflection also reevaluates recent behavioral evidence before consolidation, correcting misleading local inferences.

Hypothesis conditioning enables more focused personalization.

We further present qualitative analyses between SD and HypReflect in Appendix E.2. On HiCupid, explicit hypotheses help HypReflect emphasize the most relevant preference (e.g., narrative-driven explanations), while SD remains more generic. On Flight Recommendation, they help prioritize the strongest discriminative preference (e.g., shorter duration) leading HypReflect to select the correct option. These examples suggest that shared hypothesis conditioning enables the student to better capture the user’s style and preferences, leading to more targeted and personalized responses. This observation further suggests the shared abstraction layer may facilitate the distillation of personalized behaviors from the teacher to the student.

5.4 Effect of Context Size and Reflection Window

Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure 3: Win rate across user interactions under different context window sizes (2048, 4096) with Qwen3.5-4B and Gemma4-4B on HelpSteer2.
Reflective refinement stabilizes online personalization across context budgets.

Figure 3 presents the effect of the context budget BB on HelpSteer2 in an online setting, where performance is evaluated as the model incrementally receives new evidence about the user. Since responses can contain up to 512 tokens, we evaluate context budgets of 2048 and 4096 tokens. HypReflect adapts quickly and remains stable under both budgets. With 2048 tokens, SD and SD-IncHyp show unstable early performance, especially for Qwen3.5-4B; for Gemma4-4B, both initially perform well but degrade as interaction history grows. Increasing the budget to 4096 improves SD-IncHyp, but it still underperforms HypReflect, while SD again degrades during later interactions. These results suggest that reflective refinement improves robustness by reconsidering evidence across interaction chunks, whereas retaining more tokens alone does not ensure better online personalization.

Preference inference from implicit behavior benefits as context budgets grow.

Figure 4 examines the effect of context budget on Flight Recommendation, where user preferences are inferred from implicit behavioral signals. Flight uses fixed histories of short, structured choices over recurring attributes, so we evaluate smaller context budgets of 512, 1024, and 2048 tokens. Performance improves with larger context budgets, suggesting that additional context provides more evidence for inferring preferences from implicit behavioral signals. Across models and budgets, reflective refinement performs best in most settings, with clearer gains for Gemma4-4B and Qwen3.5-9B. Qwen3.5-4B, however, performs best with SD-IncHyp at larger budgets, suggesting that more context allows each user state to capture sufficient evidence, reducing the effects of noisy compression. These results demonstrate the robustness of reflective refinement across context budgets, supporting efficient continual personalization by reducing reliance on long, noisy interaction histories.

Refer to caption
Refer to caption
Refer to caption
Figure 4: Accuracy across different context lengths for difference models on Flight.
Effect of reflection window size.

Figure 10 compares different reflection window sizes on a subset of HelpSteer2. The window controls how many previous hypothesis sets are considered during refinement. Smaller windows are unstable early on but improve as evidence accumulates. A window of 5 enables quick adaptation and maintains stability, while increasing it to 10 yields little additional gain. We use a window of 5 across datasets and find it performs consistently well, indicating a reasonable default across datasets.

6 Related Work

Personalization in Conversational LLMs

Personalized LLMs adapt responses as users’ preferences, goals, and behaviors evolve. Prior work explores prompting, parameter adaptation, and alignment-based techniques (Zhang et al., 2025; Liu et al., 2025b; Hwang et al., 2023), while maintaining user representations through personas (Zhang et al., 2026a; Hwang et al., 2024), textual summaries (Nam et al., 2026), or RL-based profile updates (Zhao et al., 2025b). However, these methods typically target specific personalization setting rather than continually integrating diverse user signals. HypReflect studies continual personalization across diverse user signals, including feedback-driven, multi-session, and behavioral settings.

User Modeling

User modeling aims to infer and represent user preferences, goals, and behaviors for personalized interactions. Prior work models users through profiles, behavioral patterns, and predefined attributes (Purificato et al., 2024). More recent approaches treat user understanding as an evolving inference process by maintaining beliefs over user goals (Deng et al., 2026), refining hypotheses through exploration-exploitation objectives (Zhou et al., 2024), inferring communication styles (Garbacea and Tan, 2025), explicitly modeling mental-state hypotheses (Hwang et al., 2026), and updating them through Bayesian inverse planning (Zhang et al., 2026b). However, these approaches often model specific aspects of user state rather than a unified, revisable representation integrating heterogeneous evidence over time.

Feedback-Driven Learning and Adaptation

Recent work uses hindsight feedback to improve model behavior by distilling privileged information such as verified solutions (Zhao et al., 2026), combining self-distillation with policy optimization and external feedback (Hübotter et al., 2026), or leveraging subsequent user responses for alignment (Buening et al., 2026). Other methods adapt through interaction by refining user representations (Wan et al., 2025; Mehri et al., 2026) or modeling uncertainty over user hypotheses (Puri et al., 2026). However, these approaches often rely on task-specific feedback or costly optimization, limiting scalable long-horizon personalization. HypReflect instead uses user signals as privileged supervision while maintaining an evolving preference representation across interactions.

7 Conclusion

We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them over time, and incorporates them into the model via hypotheses-guided self-distillation. Our reflective refinement enables stable yet revisable user hypotheses over long interaction histories, providing a reliable user model that can guide self-distillation as new evidence accumulates. Across three personalization settings, we show HypReflect consistently improves continual personalization approaches based on raw histories or incremental update, along with stability across context budgets. Our results highlight the importance of maintaining explicit, revisable user hypotheses for achieving reliable and scalable continual personalization.

AI Use Statement

We used ChatGPT solely for language editing and proofreading during the preparation of this manuscript. Specifically, it was used to identify grammatical errors, improve sentence structure, enhance clarity and readability, and refine wording and phrasing. ChatGPT was not used to generate, develop, or substantively modify the research ideas, methodology, experimental design, analysis, results, or conclusions. All scientific content, interpretations, claims, technical decisions, code, and experiments were developed, implemented, and verified by the authors. The authors reviewed and approved all AI-assisted edits and remain fully responsible for the final manuscript.

Ethics Statement

Continual personalization systems rely on accumulating information from user interactions, which raises important privacy considerations. In real-world deployments, user data used to infer preferences should be collected with appropriate consent, handled securely, and subject to user control over storage, modification, and deletion. Although HypReflect maintains explicit preference hypotheses rather than directly exposing raw interaction histories, these representations may still encode sensitive information about users and should be treated as private user data.

Personalization systems also introduce risks when inferred preferences are inaccurate, outdated, or overly generalized from limited evidence. Incorrect user models may lead to responses that reinforce assumptions about users or reduce their ability to explore alternative behaviors. Practical deployments of systems like HypReflect should incorporate mechanisms for preference correction and transparency to ensure that inferred preferences remain aligned with users’ intended goals.

Personalization may further amplify sycophancy, as models can interpret user preferences or prior feedback as signals to agree with the user’s beliefs rather than merely adapt their responses. Over repeated interactions, this behavior may create a feedback loop in which the model increasingly reinforces the user’s existing views, including inaccurate, biased, or harmful beliefs. This risk is particularly relevant when inferred preferences concern opinions or values rather than benign stylistic or task-related choices. Personalized systems should therefore distinguish between adapting to user preferences and endorsing user claims, preserve uncertainty when evidence is limited, and avoid treating agreement or positive feedback as evidence of factual correctness. They should also retain the ability to provide corrective information or alternative perspectives when appropriate.

Reproducibility Statement

We are committed to ensuring the reproducibility of our results. Detailed descriptions of the experimental setup, including training, evaluation, and LLM-based judging procedures, as well as prompts, are provided in Appendix C and in Appendix F. We plan to release our code upon acceptance.

Acknowledgments

We thank Pouya Pezeshkpour, Farima Fatahi Bayat, Yanlin Feng, Seiji Maekawa, and other members of Megagon Labs for their insightful discussions and valuable feedback.

References

  • Buening et al. (2026) T. K. Buening, J. Hübotter, B. Pásztor, I. Shenfeld, G. Ramponi, and A. Krause Aligning language models from user interactions. In Continual Adaptation at Scale: Towards Sustainable AI Workshop, External Links: Link Cited by: §C.2, §C.3, §1, §2, §6.
  • Cachola et al. (2020) I. Cachola, K. Lo, A. Cohan, and D. Weld TLDR: extreme summarization of scientific documents. In Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu (Eds.), Online, pp. 4766–4777. External Links: Link, Document Cited by: §C.1.
  • Deng et al. (2026) M. Deng, Z. Li, X. Li, T. Zhu, Y. Zhao, Z. Guo, and W. Wang Uncertainty-aware clarification in llm agents with information gain. External Links: 2606.03135, Link Cited by: §6.
  • Garbacea and Tan (2025) C. Garbacea and C. Tan HyPerAlign: interpretable personalized llm alignment via hypothesis generation. External Links: 2505.00038, Link Cited by: §6.
  • Hübotter et al. (2026) J. Hübotter, F. Lübeck, L. Behric, A. Baumann, M. Bagatella, D. Marta, I. Hakimi, I. Shenfeld, T. K. Buening, C. Guestrin, and A. Krause Reinforcement learning via self-distillation. External Links: 2601.20802, Link Cited by: §6.
  • Hwang et al. (2023) E. Hwang, B. Majumder, and N. Tandon Aligning language models to user opinions. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 5906–5919. External Links: Link, Document Cited by: §6.
  • Hwang et al. (2024) E. Hwang, V. Shwartz, D. Gutfreund, and V. Thost A graph per persona: reasoning about subjective natural language descriptions. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 1928–1942. External Links: Link, Document Cited by: §6.
  • Hwang et al. (2026) E. Hwang, Y. Yin, G. Carenini, P. West, and V. Shwartz Infusing theory of mind into socially intelligent LLM agents. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 11327–11360. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §6.
  • Kim et al. (2025) T. Kim, J. Lee, S. Yoon, S. Kim, and D. Lee Towards personalized conversational sales agents: contextual user profiling for strategic action. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 5131–5154. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §1.
  • Laban et al. (2025) P. Laban, H. Hayashi, Y. Zhou, and J. Neville LLMs get lost in multi-turn conversation. External Links: 2505.06120, Link Cited by: §1.
  • Li et al. (2025) H. Li, C. Yang, A. Zhang, Y. Deng, X. Wang, and T. Chua Hello again! LLM-powered personalized agent for long-term dialogue. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 5259–5276. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §1.
  • Li et al. (2026) S. S. Li, B. Paranjape, K. Oktar, Z. Ma, G. Zhou, L. Guan, N. Zhang, S. Park, L. Chen, D. Yang, Y. Tsvetkov, and A. Celikyilmaz HorizonBench: long-horizon personalization with evolving preferences. External Links: 2604.17283, Link Cited by: §1.
  • Liang et al. (2024) T. Liang, C. Jin, L. Wang, W. Fan, C. Xia, K. Chen, and Y. Yin LLM-REDIAL: a large-scale dataset for conversational recommender systems created from user behaviors with LLMs. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 8926–8939. External Links: Link, Document Cited by: §1.
  • Liu et al. (2025a) B. Liu, J. Zhang, F. Lin, X. Jia, and M. Peng One size doesn’t fit all: a personalized conversational tutoring agent for mathematics instruction. In Companion Proceedings of the ACM on Web Conference 2025, WWW ’25, New York, NY, USA, pp. 2401–2410. External Links: ISBN 9798400713316, Link, Document Cited by: §1.
  • Liu et al. (2025b) J. Liu, Z. Qiu, Z. Li, Q. Dai, W. Yu, J. Zhu, M. Hu, M. Yang, T. Chua, and I. King A survey of personalized large language models: progress and future directions. External Links: 2502.11528, Link Cited by: §6.
  • Liu et al. (2026) Y. Liu, J. Guan, J. Li, W. Wu, J. Yang, J. Zhao, and G. Guo Text as a universal interface for transferable personalization. External Links: 2601.04963, Link Cited by: §2.
  • Maharana et al. (2024) A. Maharana, D. Lee, S. Tulyakov, M. Bansal, F. Barbieri, and Y. Fang Evaluating very long-term conversational memory of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 13851–13870. External Links: Link, Document Cited by: §1.
  • Mehri et al. (2026) S. Mehri, P. Kargupta, T. August, and D. Hakkani-Tür MultiSessionCollab: learning user preferences with memory to improve long-term collaboration. External Links: 2601.02702, Link Cited by: §6.
  • Mok et al. (2025) J. Mok, I. Kim, S. Park, and S. Yoon Exploring the potential of LLMs as personalized assistants: dataset, evaluation, and analysis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 10212–10239. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §C.1, §C.3, §4.1.
  • Mysore et al. (2024) S. Mysore, Z. Lu, M. Wan, L. Yang, B. Sarrafzadeh, S. Menezes, T. Baghaee, E. B. Gonzalez, J. Neville, and T. Safavi Pearl: personalizing large language model writing assistants with generation-calibrated retrievers. In Proceedings of the 1st Workshop on Customizable NLP: Progress and Challenges in Customizing NLP for a Domain, Application, Group, or Individual (CustomNLP4U), S. Kumar, V. Balachandran, C. Y. Park, W. Shi, S. A. Hayati, Y. Tsvetkov, N. Smith, H. Hajishirzi, D. Kang, and D. Jurgens (Eds.), Miami, Florida, USA, pp. 198–219. External Links: Link, Document Cited by: §1.
  • Nam et al. (2026) H. Nam, Y. Wan, M. Liu, P. Ahnn, J. Lian, and N. Jaques Learning to summarize user information for personalized reinforcement learning from human feedback. External Links: 2507.13579, Link Cited by: §1, §6.
  • Packer et al. (2024) C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez MemGPT: towards llms as operating systems. External Links: 2310.08560, Link Cited by: §1.
  • Puri et al. (2026) I. Puri, M. Damani, I. Shenfeld, M. Ghassemi, J. Andreas, and Y. Kim Reaching beyond the mode: rl for distributional reasoning in language models. External Links: 2603.24844, Link Cited by: §6.
  • Purificato et al. (2024) E. Purificato, L. Boratto, and E. W. D. Luca User modeling and user profiling: a comprehensive survey. External Links: 2402.09660, Link Cited by: §6.
  • Qiu et al. (2026) L. Qiu, F. Sha, K. Allen, Y. Kim, T. Linzen, and S. van Steenkiste Bayesian teaching enables probabilistic reasoning in large language models. Nature Communications. External Links: Link Cited by: §C.1, §4.1.
  • Salemi et al. (2024) A. Salemi, S. Mysore, M. Bendersky, and H. Zamani LaMP: when large language models meet personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7370–7392. External Links: Link, Document Cited by: §1.
  • Team et al. (2026) G. Team, S. E. Abd, V. Aggarwal, R. Algayres, A. Andreev, O. Bachem, I. Ballantyne, C. Brick, V. Cărbune, M. Casbon, et al. Gemma 4 technical report. arXiv preprint arXiv:2607.02770. External Links: Link Cited by: §4.2.
  • Wan et al. (2025) Y. Wan, J. Wu, M. Abdulhai, L. Shani, and N. Jaques Enhancing personalized multi-turn dialogue with curiosity reward. External Links: 2504.03206, Link Cited by: §1, §6.
  • Wang et al. (2024) Z. Wang, Y. Dong, O. Delalleau, J. Zeng, G. Shen, D. Egert, J. J. Zhang, M. N. Sreedhar, and O. Kuchaiev HelpSteer 2: open-source dataset for training top-performing reward models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §C.1, §4.1.
  • Wu et al. (2025) D. Wu, H. Wang, W. Yu, Y. Zhang, K. Chang, and D. Yu LongMemEval: benchmarking chat assistants on long-term interactive memory. External Links: 2410.10813, Link Cited by: §1.
  • Xu et al. (2022) J. Xu, A. Szlam, and J. Weston Beyond goldfish memory: long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 5180–5197. External Links: Link, Document Cited by: §1.
  • Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388. External Links: Link Cited by: §4.2.
  • Zhang et al. (2018) S. Zhang, E. Dinan, J. Urbanek, A. Szlam, D. Kiela, and J. Weston Personalizing dialogue agents: I have a dog, do you have pets too?. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych and Y. Miyao (Eds.), Melbourne, Australia, pp. 2204–2213. External Links: Link, Document Cited by: §1.
  • Zhang et al. (2026a) W. Zhang, X. Zhang, C. Zhang, L. Yang, J. Shang, Z. Wei, H. P. Zou, Z. Huang, Z. Wang, Y. Gao, X. Pan, L. Xiong, J. Liu, P. S. Yu, and X. Li PersonaAgent: bridging memory and action for personalized LLM agents. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 26421–26439. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §6.
  • Zhang et al. (2025) Z. Zhang, R. A. Rossi, B. Kveton, Y. Shao, D. Yang, H. Zamani, F. Dernoncourt, J. Barrow, T. Yu, S. Kim, R. Zhang, J. Gu, T. Derr, H. Chen, J. Wu, X. Chen, Z. Wang, S. Mitra, N. Lipka, N. Ahmed, and Y. Wang Personalization of large language models: a survey. External Links: 2411.00027, Link Cited by: §1, §6.
  • Zhang et al. (2026b) Z. Zhang, C. Jin, M. Y. Jia, S. Zhang, and T. Shu AutoToM: scaling model-based mental inference via automated agent modeling. External Links: 2502.15676, Link Cited by: §6.
  • Zhao et al. (2025a) S. Zhao, M. Hong, Y. Liu, D. Hazarika, and K. Lin Do llms recognize your preferences? evaluating personalized preference following in llms. External Links: 2502.09597, Link Cited by: §1.
  • Zhao et al. (2026) S. Zhao, Z. Xie, M. Liu, J. Huang, G. Pang, F. Chen, and A. Grover Self-distilled reasoner: on-policy self-distillation for large language models. External Links: 2601.18734, Link Cited by: §6.
  • Zhao et al. (2025b) W. Zhao, X. Sui, Y. Hu, J. Guo, H. Liu, B. Li, Y. Zhao, B. Qin, and T. Liu Teaching language models to evolve with users: dynamic profile modeling for personalized alignment. External Links: 2505.15456, Link Cited by: §6.
  • Zhao et al. (2025c) Z. Zhao, C. Vania, S. Kayal, N. Khan, S. B. Cohen, and E. Yilmaz PersonaLens: a benchmark for personalization evaluation in conversational AI assistants. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 18023–18055. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1.
  • Zhong et al. (2024) W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang MemoryBank: enhancing large language models with long-term memory. AAAI’24/IAAI’24/EAAI’24. External Links: ISBN 978-1-57735-887-9, Link, Document Cited by: §1.
  • Zhou et al. (2024) Y. Zhou, H. Liu, T. Srivastava, H. Mei, and C. Tan Hypothesis generation with large language models. In Proceedings of the 1st Workshop on NLP for Science (NLP4Science), L. Peled-Cohen, N. Calderon, S. Lissak, and R. Reichart (Eds.), Miami, FL, USA, pp. 117–139. External Links: Link, Document Cited by: §6.
  • Zhu et al. (2025) Y. Zhu, H. Steck, D. Liang, Y. He, N. Kallus, and J. Li LLM-based conversational recommendation agents with collaborative verbalized experience. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2207–2220. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §1.

Appendix A Limitations

Modeling diverse and evolving user preferences

HypReflect models user preferences by progressively refining evidence accumulated over interactions. In this work, we limit our experiments on three personalization settings, namely, explicit feedback, implicit behavioral signals, and conversational user profiles. However, users may express preferences through many other forms of interaction, including indirect corrections and contextual cues. Moreover, user preferences may change or evolve over time or depend on specific tasks and situations. Extending preference hypotheses to capture richer and more dynamic forms of user behavior remains an important direction.

Evaluation in real-world personalization settings.

Our evaluation spans three complementary personalization datasets that capture distinct modes of user evidence for continual personalization. While this provides a broad assessment across different interaction paradigms, it may not fully reflect the complexity of long-term human interactions. Future work should complement these evaluations with longitudinal user studies and real-world deployments.

Computational cost of continual refinement.

Maintaining explicit preference hypotheses introduces additional inference cost for hypothesis generation and reflective refinement. Although our bounded refinement strategy limits computation as interaction histories grow and avoids the cost of reward-model training and repeated policy optimization, continual personalization remains computationally demanding. Due to computational constraints, our experiments are limited to models up to 9B parameters. Exploring more efficient refinement strategies and evaluating stronger frontier models may lead to further improvements.

Preference hypothesis representation.

HypReflect assumes that user preferences can be represented as natural-language preference hypotheses with associated confidence scores. While this representation is interpretable, revisable, and directly usable for personalization, some preferences may be difficult to verbalize, interact in complex ways, or depend on latent contextual factors that are not easily captured through language alone. Future work could investigate richer representations that combine explicit preference hypotheses with structured, latent, or multimodal user representations.

Appendix B Methodology

B.1 Chunking

In the offline setting, we partition the complete history into consecutive, non-overlapping chunks C1,…,CNC_{1},\ldots,C_{N}, each containing the longest sequence of complete interactions that fits within a token budget BB. We then infer a set of local hypotheses from each chunk. In the online setting, interactions are accumulated sequentially until adding the next complete interaction would exceed BB. At that point, local hypotheses are inferred from the completed chunk, and accumulation begins for the next chunk. Before the first chunk is completed, the model uses an empty hypothesis set; afterward, it uses the latest maintained hypotheses while accumulating interactions for the next set of local hypotheses.

Appendix C Experimental Setting Details

C.1 Dataset Details

This section provides additional details about the construction of each evaluation benchmark. Figure 5 presents an example of each dataset’s user question, preferences, feedback (user’s follow up question) of HelpSteer2, HiCupid, and Flight recommendation datasets.

HelpSteer2 (Wang et al., 2024).

We use eight synthetic user profiles, where each profile is assigned a combination of three writing-style preferences adapted from TL;DR (Cachola et al., 2020). These preferences define the target response behavior for each user. At each interaction, the user simulator (deepseek-v4-flash) receives the generated response and provides feedback on one or two relevant preference dimensions. This creates a sparse feedback setting where the model must identify relevant preference signals over time. We train a separate model for each user profile.

HiCupid (Mok et al., 2025).

HiCupid contains personalized question-answering examples associated with individual users and their historical interactions. We train a single model jointly on 9K QA examples from 300 users. We use user’s profile data as a privileged information. For evaluation, we construct two settings: seen-user generalization, where the model receives previous interactions from users observed during training, and unseen-user generalization, where the model must personalize responses for users not encountered during training. The evaluation sets contain 1.8K prompts from seen users and 7.5K prompts from 250 held-out users.

Flight Recommendation (Qiu et al., 2026).

In Flight data, each user trajectory contains a sequence of recommendation interactions where user selections provide indirect evidence about underlying preferences. The benchmark contains 624 users and we use up to 25 interactions per user. User preferences may include non-obvious trade-offs, such as preferring higher prices or longer durations, requiring models to infer preferences from behavior rather than relying on standard assumptions. We evaluate across the original four-feature Flight setting, held-out feature settings with two to eight features, and transfer to the hotel recommendation domain.

Refer to caption
Figure 5: Example user question, preferences, and feedback (i.e., user’s follow up message) of Helpsteer2, HiCupid, and Flight recommendation datasets.

C.2 Training Hyperparameters

We initialize our configurations based on the implementation of Buening et al. (2026) and the corresponding dataset-specific codebases, and tune the learning rate and batch size for each dataset. We explore learning rates of {1×10−6,5×10−6,1×10−51\times 10^{-6},5\times 10^{-6},1\times 10^{-5}} and observe relatively consistent performance on HelpSteer and HiCupid, whereas Flight performs best with 1×10−51\times 10^{-5}. We also evaluate batch sizes of 2, 4, and 8 on Flight and 16 and 32 on HiCupid. HiCupid is relatively insensitive to batch size, while Flight benefits from smaller batches; we therefore use batch sizes of 32 and 4, respectively. Performance is otherwise stable across the tested hyperparameters, except for the maximum history/chunk and generation lengths, which are limited by GPU memory. When memory permits, larger history/chunk sizes generally improve performance by preserving more information from previous interactions. The final configurations are reported in Table 6. We used NVIDIA A800 GPUs, and each experiment can be run on a single GPU node. Training and evaluation take approximately 15 hours on HelpSteer2, up to 17 hours on Flight, and up to 10 hours on HiCupid.

Table 6: Training hyperparameters used for each dataset.
Hyperparameter HelpSteer Flight HiCupid
Number of epochs Online 1 1
Batch size 1 4 32
Learning rate 1×10−61\times 10^{-6} 1×10−51\times 10^{-5} 1×10−61\times 10^{-6}
LoRA rank 32 32 32
LoRA alpha 64 64 64
LoRA dropout 0 0 0
History/chunk size 4,096 1,024 1,024
Max generation length 512 400 512
Temperature 0.7 0.7 0.7
Top-pp 1.0 1.0 1.0
Top-kk 20 64 64
Reflection window size 5 5 5

C.3 Evaluation Details

We follow the original dataset papers for evaluation, including an LLM-based user simulator and pairwise LLM judging where applicable. We use DeepSeek-V4-Flash as the judge for HelpSteer and HiCupid, while Flight is evaluated through exact matching against the ground-truth choices. We additionally compare the judgments of DeepSeek-V4-Flash with those of DeepSeek-V4-Pro and Qwen3.7-Plus (see Table 8 and Appendix D) and observe consistent evaluation outcomes across the models. Given its comparable evaluation performance and lower inference cost, we select DeepSeek-V4-Flash as a practical balance between evaluation quality and cost. The complete evaluation settings are reported in Table 7. See Buening et al. (2026) (repo) for the LLM judge and user simulator prompts used for HelpSteer2, and Mok et al. (2025) (repo) for details of the LLM judge prompt used for HiCupid. Our code is also built based on these two repos. Running DeepSeek-V4-Flash as both the LLM judge and user simulator on HelpSteer2 costs approximately 10 USD and using it as the LLM judge on HiCupid costs approximately 10 USD.

Table 7: Evaluation settings used for each dataset.
Evaluation detail HelpSteer Flight HiCupid
Evaluation data 256 validation prompts 624 users / 20 questions per user 300 Seen users, 300 Unseen users
Evaluation frequency Every 50 interactions After training After training
Evaluator deepseek-v4-flash Exact answer matching deepseek-v4-flash
Primary metric Pairwise win rate Accuracy Pairwise win rate
Reference Frozen base model Ground-truth Flight Frozen base model

Appendix D LLM Judge Validity

We validate the use of deepseek-v4-flash as our primary LLM judge by comparing its evaluations with those of deepseek-v4-pro and qwen3.7-plus. Although the absolute scores vary across judges, all three produce the same ranking: HypReflect consistently outperforms SD-IncHyp, which in turn outperforms SD. This agreement indicates that our main conclusions are robust to the choice of LLM judge.

Table 8: Evaluation results using different LLM judges on the subset of HelpSteer2. All judges produce the same relative ranking across methods.
SD SD-IncHyp HypReflect
Deepseek-v4-flash 77.1 90.3 93.0
Deepseek-v4-pro 74.1 86.3 87.1
Qwen3.7-plus 71.8 83.4 86.4

Appendix E Additional results

E.1 Qualitative Examples for Reflection

Figures 6 and 7 presents some qualitative examples that compares between hypotheses consolidation without reflection and with reflection. We find that direct consolidation tends to transform local hypotheses into broad and partially overlapping meta-preferences, such as favoring actionable, multifaceted, or tailored advice. These abstractions omit concrete concepts from the original hypothesis sets and provide less distinctive guidance for personalization. Reflection instead explicitly identifies which hypotheses should be merged, retained, or removed, preserving specific preference anchors such as crowdfunding, local outreach, marketing, and the balance between project management and relationship building. This example suggests that reflection reduces vague generalization and redundancy while producing more specific and informative preference hypotheses.

User information and current query User profile: 35-year-old male; annual income of $45,000–$65,000; high-school physical-education teacher and cycling-team coach; motivated, supportive, and goal-oriented; positively uses crowdfunding platforms.
Current query: “What are some effective ways to fundraise for a school sports team?”
Hypotheses set 1 H1. The user may prefer comprehensive advice integrating diet, exercise (strength and cardiovascular), and sleep for health and fitness. [Weight: 4]
H2. The user may seek personalized financial advice focused on budgeting and saving, particularly within the $45,000–$65,000 income range. [Weight: 4]
H3. The user may value actionable, goal-oriented advice across domains like fitness, finance, and professional development. [Weight: 3]
H4. The user may be interested in guidance related to teaching, coaching, or leadership within sports, fitness, or team-management contexts. [Weight: 3]
Hypotheses set 2 H1. The user may be interested in advice related to community organizing and local initiatives, specifically concerning crowdfunding. [Weight: 4]
H2. The user may prefer detailed, actionable strategies for marketing and engagement, especially when targeting local audiences. [Weight: 4]
H3. The user may seek guidance on balancing operational tasks (project management) with outreach activities (promotion). [Weight: 3]
H4. The user may have an interest in general health and wellness topics, specifically for males in their mid-thirties. [Weight: 2]
H5. The user may value advice that focuses on building sustained community relationships beyond a single campaign. [Weight: 3]
Reflection The current set contains several distinct, high-priority clusters: comprehensive health and fitness advice, targeted financial guidance, community organizing and local marketing, and general actionable goal-setting. The fitness hypotheses should be merged into a single, well-defined preference anchor for health, while the community-organizing and marketing hypotheses should be consolidated around local engagement strategies. The generalized advice hypotheses serve as useful meta-preferences to maintain, but the specific finance, fitness, and community preferences are stronger anchors. Contradictory or very weak hypotheses should be removed, ensuring that the merged hypotheses retain specific entities such as “crowdfunding” and “mid-thirties male” where relevant.
Direct consolidation without reflection H1. The user may prefer advice that is highly actionable and practical, especially when related to project management, event promotion, or community organizing. [Weight: 4]
H2. The user may value multifaceted solutions that balance practical execution with relational aspects, such as managing team dynamics or handling criticism. [Weight: 4]
H3. The user may be interested in strategies that build long-term engagement and measure broader impact, particularly in community or educational contexts. [Weight: 4]
H4. The user may seek advice covering diverse domains, including health and fitness, professional development, and financial planning. [Weight: 3]
H5. When seeking guidance, the user may benefit from tailored advice that considers specific roles or demographic contexts. [Weight: 3]
Reflection-guided consolidation H1. The user may prefer comprehensive, actionable strategies encompassing diet, exercise (strength and cardiovascular), and sleep for health, potentially tailored for males in their mid-thirties. [Weight: 4]
H2. The user may seek personalized financial advice focused on budgeting and saving, particularly within the $45,000–$65,000 income range. [Weight: 4]
H3. The user may value guidance on local community engagement, including marketing, outreach, and crowdfunding initiatives. [Weight: 4]
H4. The user may prefer advice that helps balance operational tasks, such as project management, with external outreach and relationship building. [Weight: 3]
H5. The user may seek goal-oriented advice across domains such as fitness, finance, and professional development. [Weight: 3]
Figure 6: An example of hypothesis consolidation with and without explicit reflection.
User preference profile and current choice Hidden preference profile: Strong preference for later departures, longer durations, and higher prices; slight preference for more stops.
Current query: “Which flight is the best option?”
Flight 1: 06:00 AM; 30 min; 2 stops; $370.  Flight 2: 06:00 AM; 8 hr 18 min; 0 stops; $100.  Flight 3: 06:48 PM; 10 hr 15 min; 0 stops; $730.
Observed choice: Flight 3.
Hypotheses set 1 H1. Departure time = strong preference for later departures; duration = slight dislike for longer durations; stops = slight preference for more stops; price = slight dislike for higher prices. [Weight: 4]
H2. Departure time = strong preference for later departures; duration = slight preference for longer durations; stops = slight dislike for more stops; price = slight preference for higher prices. [Weight: 2]
H3. Departure time = strong preference for later departures. [Weight: 3]
Hypotheses set 2 H1. Departure time = slight preference for later departures; duration = slight preference for longer durations; stops = neutral; price = strong dislike for higher prices. [Weight: 2]
H2. Departure time = neutral; duration = neutral; stops = neutral; price = strong dislike for higher prices. [Weight: 4]
Reflection The recent interaction window provides repeated evidence that the user selects later, longer, and more expensive flights over earlier, shorter, and cheaper alternatives. This evidence directly conflicts with the hypotheses that infer a dislike for longer durations and higher prices. Evidence regarding stops is less consistent and should therefore be retained with lower confidence. Before consolidation, the conflicting hypotheses should be revised so that repeated behavioral evidence is prioritized over misleading local inferences.
Direct consolidation without reflection H1. Departure time = strong preference for later departures; duration = slight preference for longer durations; stops = slight dislike for more stops; price = slight dislike for higher prices. [Weight: 3]
H2. Departure time = strong preference for later departures; duration = slight dislike for longer durations; stops = strong dislike for more stops; price = strong dislike for higher prices. [Weight: 3]
H3. Departure time = strong preference for later departures; duration = slight preference for longer durations; stops = strong preference for more stops; price = slight dislike for higher prices. [Weight: 3]
Reflection-guided consolidation H1. Departure time = strong preference for later departures; duration = strong preference for longer durations; stops = slight preference for more stops; price = strong preference for higher prices. [Weight: 5]
H2. Departure time = strong preference for later departures; duration = strong preference for longer durations; stops = slight dislike for more stops; price = strong preference for higher prices. [Weight: 4]
H3. Departure time = strong preference for later departures; duration = slight preference for longer durations; stops = slight preference for more stops; price = slight preference for higher prices. [Weight: 3]
Figure 7: A qualitative Flight Recommendation example comparing direct and reflection-guided hypothesis consolidation. Direct consolidation propagates misleading inferences about duration and price, whereas reflection re-evaluates conflicting evidence across the recent interaction window and produces a top hypothesis that exactly matches the hidden preference profile.

E.2 Qualitative Examples of Final Responses: SD vs. HypReflect

Figures 8 and 9 present qualitative comparisons between the final responses generated by SD and HypReflect, together with the preference hypotheses used by HypReflect. In the HiCupid example, both methods provide generally relevant advice, but HypReflect more explicitly emphasizes narrative structure and personification, reflecting the highest-weight hypothesis inferred from the user’s interest in engaging, science-oriented content. In the Flight example, the hypotheses help HypReflect resolve competing preferences by prioritizing the user’s strong dislike of long travel durations. Because all three flights have the same price, HypReflect correctly selects Flight 2, which has the shortest duration, whereas SD selects Flight 3. These examples illustrate how explicit, weighted hypotheses can help the model identify the most relevant user preferences and prioritize them when generating personalized responses.

Dataset: HiCupid
Hidden user profile. 32-year-old man; sports journalist; income $40,000–$60,000; curious, analytical, and engaging; follows Mark Rober.
User query. What are some engaging ways to explain complex topics to a general audience? Hypotheses supplied to HypReflect.
H1: The user may prefer narrative-driven educational content that blends complex scientific concepts with relatable storytelling and personification (weight 5).
H2: The user may value high production quality while remaining open to cost-effective or free content-creation tools (weight 4).
H3: The user may prefer career guidance that balances existing skills with advanced opportunities in sports journalism (weight 4).
H4: The user may prefer learning methods that combine theoretical knowledge with hands-on application (weight 4).
H5: The user may prefer actionable health or financial guidance tailored to men in their early thirties (weight 3).
Final answer with HypReflect. To explain complex topics to a general audience, lean heavily into analogies and metaphors, as these create instant relatable connections. Frame the topic as a narrative, personifying concepts or structuring the explanation like a mystery to be solved. Furthermore, focus on the “why it matters”—connecting the science directly to the audience’s daily life or major societal issues makes the information immediately relevant and engaging. Final answer with SD. To explain complex topics to a general audience, employ strong analogies and metaphors that connect the abstract concept to everyday experiences. Visual demonstrations and simplified models are crucial for making the invisible visible. Additionally, framing the topic as a compelling story—highlighting the “why” and the real-world implications—keeps the audience hooked.
Figure 8: Final response comparison between SD and HypReflect on HiCupid
Dataset: Flight Recommendation
Hidden user profile. Slight preference for later departures; strong dislike for longer durations; slight preference for more stops; strong dislike for higher prices. User query. Which flight is the best option?
- Flight 1: 10:00 PM; 20 hr; 0 stops; $910.
- Flight 2: 07:36 AM; 8 hr 18 min; 2 stops; $910.
- Flight 3: 09:12 AM; 14 hr 9 min; 2 stops; $910.
Hypotheses supplied to HypReflect.
H1: Slight preference for later departures; strong dislike for longer durations; slight preference for more stops; strong dislike for higher prices (weight 4).
H2: Strong dislike for later departures, longer durations, more stops, and higher prices (weight 3).
H3: Slight preference for later departures; strong dislike for longer durations, more stops, and higher prices (weight 3).
User’s choice: Flight 2.
HypReflect answer: Flight 2 (correct).
SD answer: Flight 3 (incorrect).
Figure 9: Final response comparison between SD and HypReflect on Flight Recommendation
Refer to caption
Figure 10: Performance with HypReflect across different reflection window sizes on HelpSteer2

E.3 Reflection window size variation

Figure 10 compares performance under different reflection window sizes on a subset of the HelpSteer2 setting at every 50 interactions.

Appendix F Prompts

HypReflect follows the same four-stage procedure across datasets. First, the model generates local preference hypotheses from a recent chunk of interactions (Figure 11). Second, it reflects on the generated sets of local preference hypotheses to identify supported, conflicting, and less important hypotheses (Figure 12). Third, it produces a refined hypothesis set using a prompt nearly identical to that in Figure 11, augmented with the reflection traces and multiple sets of local hypotheses. Finally, the refined hypothesis set is provided to the model to guide the generation of the final personalized response (Figure 13).

Dataset-specific instantiations.

Although all datasets follow the same overall prompt structure, we slightly adapt the preference-extraction instructions and output format to reflect their distinct evidence types and personalization targets. For Flight, the history contains a flight-recommendation question, the user’s selected flight, and feedback. The hypotheses capture preferences over departure time, duration, number of stops, and price using a fixed set of directional labels. For HelpSteer2, the history contains a request, response, and user feedback. The hypotheses capture general response-style preferences, such as verbosity, tone, structure, and formatting, while excluding task-specific details. For HiCupid, the history consists of prior dialogues, and the hypotheses preserve relevant interests, constraints, roles, relationships, and other concrete personalization information. We retain at most three hypotheses for Flight and HelpSteer2 and five for HiCupid, reflecting the richer, multi-session nature of the latter. Summary-enhanced variants additionally produce an Evidence/summary note: containing the supporting observations.

[SYSTEM]
You are given a bounded portion of a user’s interaction history. Infer a compact set of persistent preference hypotheses that may help personalize future responses. The interaction history is evidence for preference inference only. Do not answer or continue any task contained in the history.
[PREFERENCE SPACE]
Infer preferences over <dataset-specific preference dimensions>.
[GUIDELINES]
- Use only evidence contained in the interaction history.
- Infer reusable preferences rather than merely describing the latest interaction.
- Express preferences as soft, conditional hypotheses.
- Avoid broad conclusions from weak or ambiguous evidence.
- Assign higher weights only when evidence is clear or repeated.
- Preserve or remove concrete details according to the dataset-specific personalization objective.
- Do not include unsupported preferences.
[CONVERSATION HISTORY]
<recent interaction chunk>
[TASK]
For each hypothesis, include a weight from 1 to 5 indicating how strongly the current interaction history supports it:
1 = weak inference from limited contextual evidence
2 = plausible inference with a clear but limited behavioral basis
3 = repeated behavioral evidence or a moderately clear stated preference
4 = directly stated preference or strong repeated behavioral evidence
5 = directly stated and repeatedly confirmed preference
Generate up to <maximum number> distinct preference hypotheses.
Return only:
hypothesis1: <dataset-specific preference statement> | weight: <1-5>
…
hypothesisN: <dataset-specific preference statement> | weight: <1-5>
Figure 11: local hypothesis-generation prompt structure.
[SYSTEM]
You are given several preference hypothesis sets inferred from different portions of a user’s interaction history, together with a previously refined hypothesis set when available. Reflect on how these hypotheses should be consolidated before a separate refinement step produces the updated preference memory. The hypotheses and summaries are evidence for preference analysis only. Do not answer or continue any task contained in them.
[GUIDELINES]
- Identify hypotheses that agree or describe the same preference.
- Identify conflicting, weak, ambiguous, unsupported, or overly broad hypotheses.
- Distinguish repeated evidence from isolated observations.
- Determine which hypotheses should be preserved, merged, narrowed, reweighted, or removed.
- Preserve dataset-specific information that remains useful for future personalization.
- Do not treat the reflection itself as additional evidence.
- Do not produce the final hypothesis set.
For each hypothesis, a weight from 1 to 5 indicates how strongly the current interaction history supports it:
1 = weak inference from limited contextual evidence
2 = plausible inference with a clear but limited behavioral basis
3 = repeated behavioral evidence or a moderately clear stated preference
4 = directly stated preference or strong repeated behavioral evidence
5 = directly stated and repeatedly confirmed preference
[LOCAL HYPOTHESIS SETS]
### Local hypothesis set max(N-W, 1)
<local hypotheses max(N-W, 1)>
…
### Local hypothesis set N
<local hypotheses N>
[TASK]
Write only a concise reflection that will guide the subsequent refinement step.
Figure 12: Hypotheses-reflection prompt across datasets.
[SYSTEM]
You are given hidden preference hypotheses inferred from the user’s previous interactions. You may also receive recent conversation history and an evidence summary. Use this information silently to personalize the response. The current request remains the primary task and must be answered directly and completely. Generate the assistant’s response to the current user request.
[GUIDELINES]
- Apply only hypotheses relevant to the current request.
- Give greater consideration to more strongly supported hypotheses.
- Treat every hypothesis as tentative rather than as a known fact.
- Prefer relevant and strongly supported hypotheses when hypotheses conflict.
- Ignore weak, irrelevant, or contradicted hypotheses.
- Do not force personalization when no hypothesis is relevant.
- Do not allow personalization to make the answer incomplete, incorrect, or unhelpful.
- Do not mention the hypotheses, evidence summary, interaction history, user profile, or personalization process.
- Follow the output format specified in system prompt.
For each hypothesis, include a weight from 1 to 5 indicating how strongly the current interaction history supports it:
1 = weak inference from limited contextual evidence
2 = plausible inference with a clear but limited behavioral basis
3 = repeated behavioral evidence or a moderately clear stated preference
4 = directly stated preference or strong repeated behavioral evidence
5 = directly stated and repeatedly confirmed preference
[RECENT CONVERSATION HISTORY]
<recent interaction history, if provided>
[HIDDEN USER PREFERENCE MEMORY]
<evidence summary, if provided>
<refined preference hypotheses>
[CURRENT USER REQUEST]
<current question or request>
[TASK]
Produce only the personalized response to the current request.
Figure 13: Prompt for answer generation conditioned on preference hypotheses.