Hypotheses-Guided Self Distillation
for Continual Personalization
Abstract
As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, latent, and noisy signals, with existing methods relying on raw interaction histories or costly reward-based optimization to manage personalization. We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation. Experiments across three personalization settings: online personalization, multi-session interactions, and implicit behavioral signals, show that HypReflect outperforms a range of baselines, including raw-history and incremental-update methods. We further demonstrate strong generalization to unseen users and cross-domain settings, along with stability across context budgets, reusable hypotheses, and more focused personalization. These results suggest a step towards reliable and scalable continual personalization through explicit, revisable user preference hypotheses.11 1 Code will be released upon acceptance.
1 Introduction
Conversational assistants that continuously adapt to individual users can provide increasingly effective and personalized support in diverse settings, including writing assistance (Mysore et al., 2024), conversational recommendation (Liang et al., 2024; Kim et al., 2025; Zhu et al., 2025), and education (Liu et al., 2025a). Personalization allows assistants to better address a user’s needs (Zhang et al., 2018; Salemi et al., 2024; Zhang et al., 2025; Zhao et al., 2025c), while continual adaptation enables long-term interactions with users by improving relevance, as assistants develop a richer understanding of users over time (Xu et al., 2022; Zhong et al., 2024; Li et al., 2025).
However, users rarely articulate their preferences in full and may not always be consciously aware of them, making continual adaptation inherently challenging. Figure 1 illustrates how evidence about a user’s preferences emerges incrementally from explicit feedback, implicit behavior, and user-profile information, requiring it to be consolidated into an explicit representation that can be revised as new evidence arrives. These signals are often sparse, noisy, and context dependent, making it difficult to distinguish stable preferences from task-specific requirements. Moreover, evidence may be distributed across long multi-turn and multi-session histories (Maharana et al., 2024; Wu et al., 2025), which language models often struggle to track consistently as conversations grow (Laban et al., 2025). The central problem is therefore to transform an evolving stream of interactions into a user representation that is both stable and revisable.
Existing approaches personalize LLMs by conditioning on user profiles, retrieved memories, or raw interaction histories (Packer et al., 2024; Zhao et al., 2025a; Li et al., 2026), making it increasingly costly to retain and retrieve past interactions over time. Other methods infer user traits or summaries through RL but require reward models, making continual optimization costly (Wan et al., 2025; Nam et al., 2026). Buening et al. (2026) applies self-distillation, improving scalability for continual personalization, but relies on raw interaction histories, leaving it unclear which preferences are learned and whether they generalize to future contexts.
In this work, we introduce HypReflect, a scalable framework for continual personalization that maintains preference hypotheses i.e. explicit, uncertainty-aware, and revisable beliefs about the user. HypReflect distills diverse user signals (e.g., Fig. 1) into preference hypotheses, reflectively consolidates useful user signals across interactions, and incorporates the refined hypotheses to guide personalized response generation through hypotheses-guided self-distillation.
We evaluate HypReflect across three personalization settings covering explicit feedback, multi-session interactions, and implicit behavioral choices, against a comprehensive suite of baselines, including self-distillation over raw interaction histories, incrementally updated summaries and hypotheses, and reflective prompting. Across these settings, HypReflect achieves relative improvements of up to 4.4%, 10.3%, and 5.0% respectively22 2 Maximum relative improvement per dataset with HypReflect+Sum in Table 1., while generalizing effectively to unseen users, shifted user characteristics, and cross-domain personalization. Further analyses demonstrate robustness across context budgets, hypothesis reusability, and more focused personalization. Overall, our findings show that reflective refinement enables stable and revisable preference hypotheses over long interaction histories, and that models benefit substantially from explicit, reusable user representations, enabling reliable and scalable continual personalization.
2 Problem Formulation
Continual Personalization
considers an assistant that repeatedly interacts with a user. At interaction , the assistant observes the preceding conversation history and the current user turn , which may contain feedback on the preceding assistant response, a new request , or both. The assistant then generates a response:
As interactions accumulate, the history provides increasing signals about the user’s preferences. The goal therefore is to continually infer, maintain, and use this understanding to personalize future responses.
Self-Distillation for Continual Personalization.
Self-distillation uses the same model as both a student and a teacher, with the teacher conditioned on richer information. For continual personalization, self-distillation offers a simple and scalable alternative to costly reward optimization by treating the subsequent user turn as privileged information when ground-truth personalized responses are unavailable (Buening et al., 2026). The student and teacher distributions are
and the teacher’s feedback-informed predictions are distilled into the student by minimizing
This enables learning from preferences revealed by subsequent interactions without an external teacher or reward model. However, directly internalizing preferences from raw interaction histories may not generalize well beyond the current interaction.
Preference Hypotheses Representation.
Hence, we introduce an explicit abstraction layer between raw interactions and response generation. We represent the model’s understanding of the user as a set of preference hypotheses and formulate personalized response generation as
This representation makes inferred preferences explicit, reusable, and revisable as new evidence arrives (Liu et al., 2026). The key challenge here is how to construct hypotheses from local interaction signals, refine and maintain them over long histories, and use them to guide self-distillation.
3 HypReflect: Continual User Modeling through Reflective Refinement
To support continual personalization, we propose HypReflect, a framework that introduces an explicit preference hypothesis layer to better understand and respond to the user. Our framework consists of three stages. First, the model infers preference hypotheses from individual interaction chunks, capturing localized signals about the user (§3.1). Second, it recurrently refines these hypotheses by reflecting and consolidating user signals across chunks (§3.2). Finally, the refined hypotheses guide self-distillation, enabling the model to use subsequent user signals as privileged supervision for learning personalized response behavior (§3.3). We also consider a summary-augmented variant that jointly maintains summaries and hypotheses.
3.1 Inferring Preference Hypotheses from Interactions
Hypotheses representation.
To capture diverse interaction signals from the user (e.g., Fig. 1), we represent the model’s understanding of the user as a revisable set of preference hypotheses , comprising multiple candidate hypotheses paired with verbalized confidence scores:
where is a natural-language statement describing a potential user preference and denotes the model’s confidence in that statement. Maintaining multiple hypotheses preserves alternative beliefs that can be revised as new evidence arrives, while confidence distinguishes well-supported preferences from uncertain ones. The natural-language representation makes the hypotheses interpretable, revisable, and directly usable for personalized response generation.
Inferring local hypotheses from interaction chunks.
Inferring preferences from the full interaction history is often impractical due to context-length limitations and, in online settings, because future interactions are not yet available. We therefore partition the observed interaction history into consecutive chunks , each containing the longest sequence of complete interactions within a token budget . For each chunk , we infer a set of local hypotheses (Step 2 in Fig. 2) as
The model is instructed to infer reusable preferences grounded in the user’s instructions, corrections, feedback, and choices, while avoiding interaction-specific restatements and unsupported generalizations.
We additionally define a summary-augmented variant, where we generate a local summary for each interaction chunk. Summaries preserve relevant conversational context from that grounds the hypotheses, whereas captures reusable preferences that are explicitly stated or inferred from user signals.
3.2 Maintaining User Understanding through Reflective Refinement
As interactions accumulate, the model must maintain preference hypotheses that (1) explain past interaction histories and (2) predict future user behavior. Simply using the most recent hypotheses discards earlier preferences, whereas repeatedly reconsidering the full interaction history is computationally infeasible. A refinement mechanism should therefore preserve long-term preferences while remaining revisable and bounded in computational context size.
A simple approach is incremental updating, which iteratively refines a single running user state. However, since each update operates on an already compressed representation, over time it is either prone to error accumulation, or it may have valid earlier preferences overwritten by recent observations. We empirically show that this leads to unstable long-term personalization (§5.1).
Reflective refinement across chunks.
To preserve evidence across multiple chunks, without relying exclusively on a recursively compressed state, we propose reflective refinement (Step 3 in Fig. 2). Rather than immediately merging each new local hypothesis set into a single running state, the model jointly reconsiders multiple independently generated local sets. By comparing hypotheses across interaction chunks, the model consolidates consistent evidence, resolves conflicts, merges redundant hypotheses, and revises beliefs that are no longer supported.
To keep refinement tractable, we maintain a bounded reflection window. Let denote the local hypothesis set inferred from chunk , the refined hypothesis set after observing the first chunks, and the reflection-window size. We construct the refinement input as
and update the maintained hypothesis set as .
When , the model reflects over all available local hypothesis sets. Thereafter, it replaces older local sets with their refined representation while retaining the most recent local sets. This keeps the refinement input bounded without discarding earlier preference information.
For our summary-augmented variant, we apply the same windowed recurrent refinement procedure to a jointly maintained summary and hypothesis set, , for each chunk, producing . The refined summary consolidates the context used to ground and revise the hypotheses, while maintains the refined preferences used for future personalization.
3.3 Hypothesis-Guided Self-Distillation
To effectively use the refined user state for personalized response generation, we incorporate it into self-distillation, training the model to translate maintained preferences into personalized responses. Building on the formulation in §2, we condition both the teacher and student on the same refined state, providing a shared understanding of the user and making the teacher’s targets learnable by the student. The teacher additionally observes the subsequent user turn and uses it as privileged hindsight to generate a preference-grounded response.
At interaction , given the preceding history , current user turn , refined user hypotheses , and subsequent user turn , the student and teacher distributions are
The teacher uses to provide feedback-informed supervision. We then minimize the reverse KL divergence from the student to the detached teacher:
The objective is computed token-wise over the generated response (Step 4 in Fig. 2).
4 Experimental Setup
4.1 Datasets and Evaluation
To test whether models can infer reusable user preferences from diverse interaction signals and use them to personalize future responses, we evaluate three complementary settings: continual adaptation from explicit user feedback, multi-session personalization using conversational user profiles, and preference inference from implicit behavioral signals. See Appendix C.1 and C.3 more details.
HelpSteer2
(Wang et al., 2024).
To evaluate continual adaptation from explicit user feedback, we construct an online personalization setting from HelpSteer2. Each user is associated with a set of writing-style preferences, and across multiple interactions, a user simulator provides feedback on one or two relevant preference dimensions after each model response.
Evaluation Models are trained over 250 interactions and evaluated every 50 interactions on held-out prompts using an LLM judge. We report pairwise win rates against the base model with respect to the target user’s preferences.
HiCupid
(Mok et al., 2025).
To test personalization from user profile information, we use HiCupid, where user preferences and attributes are stated naturally across multi-session conversations. We train a model on QA interactions from 300 users and evaluate whether it can use previous user information to personalize future responses.
Evaluation We evaluate both seen-user and unseen-user generalization using held-out prompts and report pairwise win rates from an LLM judge33
3
We use deepseek-v4-flash as the LLM judge and verify its consistency with other judges in Appendix D. assessing personalization and response quality.
Flight Recommendation
(Qiu et al., 2026).
To evaluate preference inference from implicit behavioral signals, we use a Flight Recommendation task. We utilize multi-turn recommendation trajectories from 624 users, with up to 25 interactions per user, where user choices provide partial evidence about latent preferences. These preferences may include non-obvious trade-offs, requiring models to infer user-specific behavior rather than rely on fixed assumptions.
Evaluation We report final-round preference prediction accuracy on the original flight task, held-out feature settings, and cross-domain hotel transfer.
4.2 Model Variants and Training Setup
(1) Base. Uses the vanilla language model without fine-tuning, conditioned on the relevant raw interaction history, to generate a response. (2) SD. Self-distillation is performed directly on the raw interaction history. When the history exceeds the context budget, only the most recent portion is retained; (3) SD-IncSum. The model is trained with an incrementally updated summary: a running summary is maintained and refreshed after each chunk; (4) SD-IncHyp. Instead of a summary, the model maintains a running hypothesis set , updated with the same incremental strategy; (5) HypReflect. Local hypotheses are first generated from individual chunks and then consolidated into a refined set through the reflective refinement process described in §3.2; (6) HypReflect+Sum. Reflective refinement is applied to the paired memory state , yielding , to examine whether summaries supply complementary information during hypothesis refinement.
Training setup.
We use Qwen3.5-4B, Qwen3.5-9B (Yang et al., 2025), and Gemma4-4B (Team et al., 2026) as backbone LLMs. Additional hyperparameter details are provided in Appendix C.2.
5 Results
| HelpSteer2 (Win Rate ) | HiCupid (Win Rate ) | Flight Recommendation (Acc. ) | |||||||
| Method | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B |
| Base | – | – | – | – | – | – | 40.9 | 40.3 | 41.7 |
| SD | 87.0 | 84.5 | 89.0 | 57.3 | 69.7 | 63.3 | 61.8 | 73.8 | 74.4 |
| SD-IncSum | 85.8 | 81.0 | 86.6 | 58.4 | 67.9 | 67.7 | 50.5 | 74.4 | 67.0 |
| SD-IncHyp | 89.6 | 81.9 | 88.2 | 60.7 | 71.6 | 66.2 | 57.3 | 74.8 | 75.4 |
| HypReflect | 86.3 | 88.9 | 90.3 | 61.1 | 71.1 | 67.7 | 62.3 | 74.0 | 76.0 |
| HypReflect+Sum | 90.0 | 88.2 | 92.1 | 61.9 | 71.8 | 69.8 | 64.9 | 74.2 | 75.3 |
5.1 Main Results
Preference hypotheses provide a stronger representation than summaries.
Table 1 compares performance across three personalization benchmarks. Under the incremental update strategy, SD-IncHyp outperforms SD-IncSum in eight of nine settings, showing that abstracting interaction signals into reusable preference beliefs is generally more effective than retaining them as summaries. HypReflect further improves over SD-IncHyp and outperforms SD in eight of nine settings, indicating that the effectiveness of preference hypotheses also depends on maintaining them accurately and consistently over time.
Reflective refinement makes personalization more reliable.
Furthermore, HypReflect outperforms SD in eight of nine settings and HypReflect+Sum improves over it in all nine. HypReflect+Sum also outperforms HypReflect in seven settings, suggesting that summaries provide complementary evidence omitted during hypothesis abstraction and ground refinement in both observed interactions and inferred preferences.
In contrast, SD-IncSum and SD-IncHyp underperform SD in six and three settings, respectively. We find that incremental updates can propagate earlier errors and allow recent observations to overwrite valid prior information, resulting in unstable supervision over time. In contrast, reflective refinement maintains more reliable preference hypotheses as interaction histories grow.
| HiCupid | Flight Recommendation | ||||
| Win Rate of HypReflect | Accuracy | ||||
| Model | vs. Base | vs. Base+RefHyp | Base | Base+RefHyp | HypReflect |
| Gemma4-4B | 61.1 | 59.7 | 40.9 | 47.9 | 62.3 |
| Qwen3.5-4B | 71.1 | 62.2 | 40.3 | 42.0 | 74.0 |
| Qwen3.5-9B | 67.7 | 57.6 | 41.7 | 46.8 | 76.0 |
Reflective hypotheses prompting yields consistent gains, which self-distillation further amplifies.
Table 2 demonstrates the effectiveness of reflective refinement both as an inference-time prompting method and when combined with self-distillation. We implement Base+RefHyp similarly to HypReflect, reflectively refining preference hypotheses to generate a response but without self-distillation. Even when used only through prompting, reflectively refining preference hypotheses improves the base model, with Base+RefHyp gaining an average of 4.6 accuracy points on Flight Recommendation and narrowing the gap to HypReflect on HiCupid, reducing the average win rate of HypReflect from 66.6% against Base to 59.8%. Combining reflective refinement with hypothesis-guided self-distillation yields substantially larger gains, as HypReflect outperforms Base+RefHyp by an average of 25.2 accuracy points on Flight Recommendation and achieves an average win rate of 59.8% against it on HiCupid. These results show that reflective refinement is effective even as inference-time prompting and combining it with self-distillation more effectively incorporates user preferences into personalized response behavior.
| Flight Features (Acc. ) | Flight Hotel (Acc. ) | HiCupid Unseen (Win Rate ) | |||||||
| Method | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B | Gemma4-4B | Qwen3.5-4B | Qwen3.5-9B |
| SD | 55.4 | 64.9 | 63.7 | 54.8 | 61.8 | 62.0 | 57.0 | 71.7 | 66.9 |
| SD-IncSum | 47.5 | 65.6 | 57.1 | 48.0 | 63.7 | 57.1 | 58.7 | 69.9 | 69.4 |
| SD-IncHyp | 53.4 | 64.8 | 65.5 | 52.8 | 62.3 | 62.9 | 62.1 | 71.8 | 68.0 |
| HypReflect | 55.6 | 65.8 | 66.0 | 55.7 | 63.6 | 63.8 | 62.9 | 71.6 | 68.1 |
| HypReflect+Sum | 57.7 | 65.0 | 65.9 | 56.5 | 64.1 | 63.7 | 62.6 | 73.3 | 70.3 |
5.2 Generalization Results
Reflective refinement improves robustness under distribution shift.
Table 3 reports results on three generalization settings. On HiCupid, we evaluate models on users unseen during training. On the Flight Recommendation dataset, we test models trained with 4 features on varying numbers of features (2–8) and a different Hotel domain.
Reflective methods remain the most robust. HypReflect+Sum improves over SD in all nine settings, often by substantial margins, and HypReflect outperforms SD in eight out of nine settings. In contrast, SD-IncSum produces large drops on the Flight transfer tasks. SD-IncHyp is more stable and improves performance in most cases, yet still underperforms the reflective approaches on several Flight settings.
These results show that reflective refinement produces higher-quality hypotheses, providing more reliable conditioning signals during self-distillation, allowing the model to learn a cleaner mapping from user behavior to underlying preferences.
5.3 Further Analysis
| Model | Flight (Acc.) | HiCupid (Win Rate) | ||
|---|---|---|---|---|
| Base | w/ Refined Hyp. | Base | w/ Refined Hyp. | |
| Gemma4-4B | 40.9 | 50.0 | ||
| Qwen3.5-4B | 40.3 | 50.0 | ||
| Qwen3.5-9B | 41.7 | 50.0 | ||
Reflective refinement produces reusable user hypotheses.
Table 4 presents our analysis on the quality of the user hypotheses learned by HypReflect through the hypothesis-transfer experiment. For each user, we take the final hypothesis generated by the HypReflect model and provide it to the corresponding non-finetuned base model for answer generation. During evaluation, each base model receives the hypothesis generated for the same user. The transferred hypotheses substantially improve performance across all models on both datasets, yielding average gains of 11.5 accuracy points on Flight Recommendation and 10.0 winrate points on HiCupid. These results indicate that our reflective refinement approach produces informative and reusable user hypotheses, suggesting that the resulting personalization benefits may extend beyond the model used to generate them.
Effect of reflection before consolidation.
| Dataset | Method | Gemma4-4B | Qwen3.5-4B |
|---|---|---|---|
| HiCupid | HypReflect w/o refl. | 72.8 | 83.8 |
| HypReflect | 74.2 | 84.5 | |
| Flight | HypReflect w/o refl. | 59.0 | 74.2 |
| HypReflect | 62.9 | 76.0 |
Table 5 presents an ablation on HiCupid and Flight Recommendation to isolate the effect of the reflection step. For HiCupid, we evaluate users with long interaction histories. Our refinement step first reflects on the inferred hypothesis sets and then consolidates them into an updated hypothesis set based on reflection. We compare direct consolidation of without reflection. Adding reflection before consolidation yields modest but consistent average gains of 1.0 and 2.8 points on HiCupid and Flight Recommendation, respectively, over direct consolidation of locally inferred hypothesis sets. Our qualitative analyses on both datasets in Appendix E.1 show that direct consolidation can retain broad, overlapping, or redundant preferences, whereas reflection before consolidation reorganizes the accumulated evidence into more specific and informative hypotheses. On Flight Recommendation, reflection also reevaluates recent behavioral evidence before consolidation, correcting misleading local inferences.
Hypothesis conditioning enables more focused personalization.
We further present qualitative analyses between SD and HypReflect in Appendix E.2. On HiCupid, explicit hypotheses help HypReflect emphasize the most relevant preference (e.g., narrative-driven explanations), while SD remains more generic. On Flight Recommendation, they help prioritize the strongest discriminative preference (e.g., shorter duration) leading HypReflect to select the correct option. These examples suggest that shared hypothesis conditioning enables the student to better capture the user’s style and preferences, leading to more targeted and personalized responses. This observation further suggests the shared abstraction layer may facilitate the distillation of personalized behaviors from the teacher to the student.
5.4 Effect of Context Size and Reflection Window




Reflective refinement stabilizes online personalization across context budgets.
Figure 3 presents the effect of the context budget on HelpSteer2 in an online setting, where performance is evaluated as the model incrementally receives new evidence about the user. Since responses can contain up to 512 tokens, we evaluate context budgets of 2048 and 4096 tokens. HypReflect adapts quickly and remains stable under both budgets. With 2048 tokens, SD and SD-IncHyp show unstable early performance, especially for Qwen3.5-4B; for Gemma4-4B, both initially perform well but degrade as interaction history grows. Increasing the budget to 4096 improves SD-IncHyp, but it still underperforms HypReflect, while SD again degrades during later interactions. These results suggest that reflective refinement improves robustness by reconsidering evidence across interaction chunks, whereas retaining more tokens alone does not ensure better online personalization.
Preference inference from implicit behavior benefits as context budgets grow.
Figure 4 examines the effect of context budget on Flight Recommendation, where user preferences are inferred from implicit behavioral signals. Flight uses fixed histories of short, structured choices over recurring attributes, so we evaluate smaller context budgets of 512, 1024, and 2048 tokens. Performance improves with larger context budgets, suggesting that additional context provides more evidence for inferring preferences from implicit behavioral signals. Across models and budgets, reflective refinement performs best in most settings, with clearer gains for Gemma4-4B and Qwen3.5-9B. Qwen3.5-4B, however, performs best with SD-IncHyp at larger budgets, suggesting that more context allows each user state to capture sufficient evidence, reducing the effects of noisy compression. These results demonstrate the robustness of reflective refinement across context budgets, supporting efficient continual personalization by reducing reliance on long, noisy interaction histories.



Effect of reflection window size.
Figure 10 compares different reflection window sizes on a subset of HelpSteer2. The window controls how many previous hypothesis sets are considered during refinement. Smaller windows are unstable early on but improve as evidence accumulates. A window of 5 enables quick adaptation and maintains stability, while increasing it to 10 yields little additional gain. We use a window of 5 across datasets and find it performs consistently well, indicating a reasonable default across datasets.
6 Related Work
Personalization in Conversational LLMs
Personalized LLMs adapt responses as users’ preferences, goals, and behaviors evolve. Prior work explores prompting, parameter adaptation, and alignment-based techniques (Zhang et al., 2025; Liu et al., 2025b; Hwang et al., 2023), while maintaining user representations through personas (Zhang et al., 2026a; Hwang et al., 2024), textual summaries (Nam et al., 2026), or RL-based profile updates (Zhao et al., 2025b). However, these methods typically target specific personalization setting rather than continually integrating diverse user signals. HypReflect studies continual personalization across diverse user signals, including feedback-driven, multi-session, and behavioral settings.
User Modeling
User modeling aims to infer and represent user preferences, goals, and behaviors for personalized interactions. Prior work models users through profiles, behavioral patterns, and predefined attributes (Purificato et al., 2024). More recent approaches treat user understanding as an evolving inference process by maintaining beliefs over user goals (Deng et al., 2026), refining hypotheses through exploration-exploitation objectives (Zhou et al., 2024), inferring communication styles (Garbacea and Tan, 2025), explicitly modeling mental-state hypotheses (Hwang et al., 2026), and updating them through Bayesian inverse planning (Zhang et al., 2026b). However, these approaches often model specific aspects of user state rather than a unified, revisable representation integrating heterogeneous evidence over time.
Feedback-Driven Learning and Adaptation
Recent work uses hindsight feedback to improve model behavior by distilling privileged information such as verified solutions (Zhao et al., 2026), combining self-distillation with policy optimization and external feedback (Hübotter et al., 2026), or leveraging subsequent user responses for alignment (Buening et al., 2026). Other methods adapt through interaction by refining user representations (Wan et al., 2025; Mehri et al., 2026) or modeling uncertainty over user hypotheses (Puri et al., 2026). However, these approaches often rely on task-specific feedback or costly optimization, limiting scalable long-horizon personalization. HypReflect instead uses user signals as privileged supervision while maintaining an evolving preference representation across interactions.
7 Conclusion
We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them over time, and incorporates them into the model via hypotheses-guided self-distillation. Our reflective refinement enables stable yet revisable user hypotheses over long interaction histories, providing a reliable user model that can guide self-distillation as new evidence accumulates. Across three personalization settings, we show HypReflect consistently improves continual personalization approaches based on raw histories or incremental update, along with stability across context budgets. Our results highlight the importance of maintaining explicit, revisable user hypotheses for achieving reliable and scalable continual personalization.
AI Use Statement
We used ChatGPT solely for language editing and proofreading during the preparation of this manuscript. Specifically, it was used to identify grammatical errors, improve sentence structure, enhance clarity and readability, and refine wording and phrasing. ChatGPT was not used to generate, develop, or substantively modify the research ideas, methodology, experimental design, analysis, results, or conclusions. All scientific content, interpretations, claims, technical decisions, code, and experiments were developed, implemented, and verified by the authors. The authors reviewed and approved all AI-assisted edits and remain fully responsible for the final manuscript.
Ethics Statement
Continual personalization systems rely on accumulating information from user interactions, which raises important privacy considerations. In real-world deployments, user data used to infer preferences should be collected with appropriate consent, handled securely, and subject to user control over storage, modification, and deletion. Although HypReflect maintains explicit preference hypotheses rather than directly exposing raw interaction histories, these representations may still encode sensitive information about users and should be treated as private user data.
Personalization systems also introduce risks when inferred preferences are inaccurate, outdated, or overly generalized from limited evidence. Incorrect user models may lead to responses that reinforce assumptions about users or reduce their ability to explore alternative behaviors. Practical deployments of systems like HypReflect should incorporate mechanisms for preference correction and transparency to ensure that inferred preferences remain aligned with users’ intended goals.
Personalization may further amplify sycophancy, as models can interpret user preferences or prior feedback as signals to agree with the user’s beliefs rather than merely adapt their responses. Over repeated interactions, this behavior may create a feedback loop in which the model increasingly reinforces the user’s existing views, including inaccurate, biased, or harmful beliefs. This risk is particularly relevant when inferred preferences concern opinions or values rather than benign stylistic or task-related choices. Personalized systems should therefore distinguish between adapting to user preferences and endorsing user claims, preserve uncertainty when evidence is limited, and avoid treating agreement or positive feedback as evidence of factual correctness. They should also retain the ability to provide corrective information or alternative perspectives when appropriate.
Reproducibility Statement
We are committed to ensuring the reproducibility of our results. Detailed descriptions of the experimental setup, including training, evaluation, and LLM-based judging procedures, as well as prompts, are provided in Appendix C and in Appendix F. We plan to release our code upon acceptance.
Acknowledgments
We thank Pouya Pezeshkpour, Farima Fatahi Bayat, Yanlin Feng, Seiji Maekawa, and other members of Megagon Labs for their insightful discussions and valuable feedback.
References
- Aligning language models from user interactions. In Continual Adaptation at Scale: Towards Sustainable AI Workshop, External Links: Link Cited by: §C.2, §C.3, §1, §2, §6.
- TLDR: extreme summarization of scientific documents. In Findings of the Association for Computational Linguistics: EMNLP 2020, T. Cohn, Y. He, and Y. Liu (Eds.), Online, pp. 4766–4777. External Links: Link, Document Cited by: §C.1.
- Uncertainty-aware clarification in llm agents with information gain. External Links: 2606.03135, Link Cited by: §6.
- HyPerAlign: interpretable personalized llm alignment via hypothesis generation. External Links: 2505.00038, Link Cited by: §6.
- Reinforcement learning via self-distillation. External Links: 2601.20802, Link Cited by: §6.
- Aligning language models to user opinions. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 5906–5919. External Links: Link, Document Cited by: §6.
- A graph per persona: reasoning about subjective natural language descriptions. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 1928–1942. External Links: Link, Document Cited by: §6.
- Infusing theory of mind into socially intelligent LLM agents. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 11327–11360. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §6.
- Towards personalized conversational sales agents: contextual user profiling for strategic action. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 5131–5154. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §1.
- LLMs get lost in multi-turn conversation. External Links: 2505.06120, Link Cited by: §1.
- Hello again! LLM-powered personalized agent for long-term dialogue. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 5259–5276. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §1.
- HorizonBench: long-horizon personalization with evolving preferences. External Links: 2604.17283, Link Cited by: §1.
- LLM-REDIAL: a large-scale dataset for conversational recommender systems created from user behaviors with LLMs. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 8926–8939. External Links: Link, Document Cited by: §1.
- One size doesn’t fit all: a personalized conversational tutoring agent for mathematics instruction. In Companion Proceedings of the ACM on Web Conference 2025, WWW ’25, New York, NY, USA, pp. 2401–2410. External Links: ISBN 9798400713316, Link, Document Cited by: §1.
- A survey of personalized large language models: progress and future directions. External Links: 2502.11528, Link Cited by: §6.
- Text as a universal interface for transferable personalization. External Links: 2601.04963, Link Cited by: §2.
- Evaluating very long-term conversational memory of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 13851–13870. External Links: Link, Document Cited by: §1.
- MultiSessionCollab: learning user preferences with memory to improve long-term collaboration. External Links: 2601.02702, Link Cited by: §6.
- Exploring the potential of LLMs as personalized assistants: dataset, evaluation, and analysis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 10212–10239. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §C.1, §C.3, §4.1.
- Pearl: personalizing large language model writing assistants with generation-calibrated retrievers. In Proceedings of the 1st Workshop on Customizable NLP: Progress and Challenges in Customizing NLP for a Domain, Application, Group, or Individual (CustomNLP4U), S. Kumar, V. Balachandran, C. Y. Park, W. Shi, S. A. Hayati, Y. Tsvetkov, N. Smith, H. Hajishirzi, D. Kang, and D. Jurgens (Eds.), Miami, Florida, USA, pp. 198–219. External Links: Link, Document Cited by: §1.
- Learning to summarize user information for personalized reinforcement learning from human feedback. External Links: 2507.13579, Link Cited by: §1, §6.
- MemGPT: towards llms as operating systems. External Links: 2310.08560, Link Cited by: §1.
- Reaching beyond the mode: rl for distributional reasoning in language models. External Links: 2603.24844, Link Cited by: §6.
- User modeling and user profiling: a comprehensive survey. External Links: 2402.09660, Link Cited by: §6.
- Bayesian teaching enables probabilistic reasoning in large language models. Nature Communications. External Links: Link Cited by: §C.1, §4.1.
- LaMP: when large language models meet personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 7370–7392. External Links: Link, Document Cited by: §1.
- Gemma 4 technical report. arXiv preprint arXiv:2607.02770. External Links: Link Cited by: §4.2.
- Enhancing personalized multi-turn dialogue with curiosity reward. External Links: 2504.03206, Link Cited by: §1, §6.
- HelpSteer 2: open-source dataset for training top-performing reward models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §C.1, §4.1.
- LongMemEval: benchmarking chat assistants on long-term interactive memory. External Links: 2410.10813, Link Cited by: §1.
- Beyond goldfish memory: long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 5180–5197. External Links: Link, Document Cited by: §1.
- Qwen3 technical report. arXiv preprint arXiv:2505.09388. External Links: Link Cited by: §4.2.
- Personalizing dialogue agents: I have a dog, do you have pets too?. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych and Y. Miyao (Eds.), Melbourne, Australia, pp. 2204–2213. External Links: Link, Document Cited by: §1.
- PersonaAgent: bridging memory and action for personalized LLM agents. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 26421–26439. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §6.
- Personalization of large language models: a survey. External Links: 2411.00027, Link Cited by: §1, §6.
- AutoToM: scaling model-based mental inference via automated agent modeling. External Links: 2502.15676, Link Cited by: §6.
- Do llms recognize your preferences? evaluating personalized preference following in llms. External Links: 2502.09597, Link Cited by: §1.
- Self-distilled reasoner: on-policy self-distillation for large language models. External Links: 2601.18734, Link Cited by: §6.
- Teaching language models to evolve with users: dynamic profile modeling for personalized alignment. External Links: 2505.15456, Link Cited by: §6.
- PersonaLens: a benchmark for personalization evaluation in conversational AI assistants. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 18023–18055. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1.
- MemoryBank: enhancing large language models with long-term memory. AAAI’24/IAAI’24/EAAI’24. External Links: ISBN 978-1-57735-887-9, Link, Document Cited by: §1.
- Hypothesis generation with large language models. In Proceedings of the 1st Workshop on NLP for Science (NLP4Science), L. Peled-Cohen, N. Calderon, S. Lissak, and R. Reichart (Eds.), Miami, FL, USA, pp. 117–139. External Links: Link, Document Cited by: §6.
- LLM-based conversational recommendation agents with collaborative verbalized experience. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 2207–2220. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §1.
Appendix A Limitations
Modeling diverse and evolving user preferences
HypReflect models user preferences by progressively refining evidence accumulated over interactions. In this work, we limit our experiments on three personalization settings, namely, explicit feedback, implicit behavioral signals, and conversational user profiles. However, users may express preferences through many other forms of interaction, including indirect corrections and contextual cues. Moreover, user preferences may change or evolve over time or depend on specific tasks and situations. Extending preference hypotheses to capture richer and more dynamic forms of user behavior remains an important direction.
Evaluation in real-world personalization settings.
Our evaluation spans three complementary personalization datasets that capture distinct modes of user evidence for continual personalization. While this provides a broad assessment across different interaction paradigms, it may not fully reflect the complexity of long-term human interactions. Future work should complement these evaluations with longitudinal user studies and real-world deployments.
Computational cost of continual refinement.
Maintaining explicit preference hypotheses introduces additional inference cost for hypothesis generation and reflective refinement. Although our bounded refinement strategy limits computation as interaction histories grow and avoids the cost of reward-model training and repeated policy optimization, continual personalization remains computationally demanding. Due to computational constraints, our experiments are limited to models up to 9B parameters. Exploring more efficient refinement strategies and evaluating stronger frontier models may lead to further improvements.
Preference hypothesis representation.
HypReflect assumes that user preferences can be represented as natural-language preference hypotheses with associated confidence scores. While this representation is interpretable, revisable, and directly usable for personalization, some preferences may be difficult to verbalize, interact in complex ways, or depend on latent contextual factors that are not easily captured through language alone. Future work could investigate richer representations that combine explicit preference hypotheses with structured, latent, or multimodal user representations.
Appendix B Methodology
B.1 Chunking
In the offline setting, we partition the complete history into consecutive, non-overlapping chunks , each containing the longest sequence of complete interactions that fits within a token budget . We then infer a set of local hypotheses from each chunk. In the online setting, interactions are accumulated sequentially until adding the next complete interaction would exceed . At that point, local hypotheses are inferred from the completed chunk, and accumulation begins for the next chunk. Before the first chunk is completed, the model uses an empty hypothesis set; afterward, it uses the latest maintained hypotheses while accumulating interactions for the next set of local hypotheses.
Appendix C Experimental Setting Details
C.1 Dataset Details
This section provides additional details about the construction of each evaluation benchmark. Figure 5 presents an example of each dataset’s user question, preferences, feedback (user’s follow up question) of HelpSteer2, HiCupid, and Flight recommendation datasets.
HelpSteer2 (Wang et al., 2024).
We use eight synthetic user profiles, where each profile is assigned a combination of three writing-style preferences adapted from TL;DR (Cachola et al., 2020). These preferences define the target response behavior for each user. At each interaction, the user simulator (deepseek-v4-flash) receives the generated response and provides feedback on one or two relevant preference dimensions. This creates a sparse feedback setting where the model must identify relevant preference signals over time. We train a separate model for each user profile.
HiCupid (Mok et al., 2025).
HiCupid contains personalized question-answering examples associated with individual users and their historical interactions. We train a single model jointly on 9K QA examples from 300 users. We use user’s profile data as a privileged information. For evaluation, we construct two settings: seen-user generalization, where the model receives previous interactions from users observed during training, and unseen-user generalization, where the model must personalize responses for users not encountered during training. The evaluation sets contain 1.8K prompts from seen users and 7.5K prompts from 250 held-out users.
Flight Recommendation (Qiu et al., 2026).
In Flight data, each user trajectory contains a sequence of recommendation interactions where user selections provide indirect evidence about underlying preferences. The benchmark contains 624 users and we use up to 25 interactions per user. User preferences may include non-obvious trade-offs, such as preferring higher prices or longer durations, requiring models to infer preferences from behavior rather than relying on standard assumptions. We evaluate across the original four-feature Flight setting, held-out feature settings with two to eight features, and transfer to the hotel recommendation domain.
C.2 Training Hyperparameters
We initialize our configurations based on the implementation of Buening et al. (2026) and the corresponding dataset-specific codebases, and tune the learning rate and batch size for each dataset. We explore learning rates of {} and observe relatively consistent performance on HelpSteer and HiCupid, whereas Flight performs best with . We also evaluate batch sizes of 2, 4, and 8 on Flight and 16 and 32 on HiCupid. HiCupid is relatively insensitive to batch size, while Flight benefits from smaller batches; we therefore use batch sizes of 32 and 4, respectively. Performance is otherwise stable across the tested hyperparameters, except for the maximum history/chunk and generation lengths, which are limited by GPU memory. When memory permits, larger history/chunk sizes generally improve performance by preserving more information from previous interactions. The final configurations are reported in Table 6. We used NVIDIA A800 GPUs, and each experiment can be run on a single GPU node. Training and evaluation take approximately 15 hours on HelpSteer2, up to 17 hours on Flight, and up to 10 hours on HiCupid.
| Hyperparameter | HelpSteer | Flight | HiCupid |
|---|---|---|---|
| Number of epochs | Online | 1 | 1 |
| Batch size | 1 | 4 | 32 |
| Learning rate | |||
| LoRA rank | 32 | 32 | 32 |
| LoRA alpha | 64 | 64 | 64 |
| LoRA dropout | 0 | 0 | 0 |
| History/chunk size | 4,096 | 1,024 | 1,024 |
| Max generation length | 512 | 400 | 512 |
| Temperature | 0.7 | 0.7 | 0.7 |
| Top- | 1.0 | 1.0 | 1.0 |
| Top- | 20 | 64 | 64 |
| Reflection window size | 5 | 5 | 5 |
C.3 Evaluation Details
We follow the original dataset papers for evaluation, including an LLM-based user simulator and pairwise LLM judging where applicable. We use DeepSeek-V4-Flash as the judge for HelpSteer and HiCupid, while Flight is evaluated through exact matching against the ground-truth choices. We additionally compare the judgments of DeepSeek-V4-Flash with those of DeepSeek-V4-Pro and Qwen3.7-Plus (see Table 8 and Appendix D) and observe consistent evaluation outcomes across the models. Given its comparable evaluation performance and lower inference cost, we select DeepSeek-V4-Flash as a practical balance between evaluation quality and cost. The complete evaluation settings are reported in Table 7. See Buening et al. (2026) (repo) for the LLM judge and user simulator prompts used for HelpSteer2, and Mok et al. (2025) (repo) for details of the LLM judge prompt used for HiCupid. Our code is also built based on these two repos. Running DeepSeek-V4-Flash as both the LLM judge and user simulator on HelpSteer2 costs approximately 10 USD and using it as the LLM judge on HiCupid costs approximately 10 USD.
| Evaluation detail | HelpSteer | Flight | HiCupid |
|---|---|---|---|
| Evaluation data | 256 validation prompts | 624 users / 20 questions per user | 300 Seen users, 300 Unseen users |
| Evaluation frequency | Every 50 interactions | After training | After training |
| Evaluator | deepseek-v4-flash | Exact answer matching | deepseek-v4-flash |
| Primary metric | Pairwise win rate | Accuracy | Pairwise win rate |
| Reference | Frozen base model | Ground-truth Flight | Frozen base model |
Appendix D LLM Judge Validity
We validate the use of deepseek-v4-flash as our primary LLM judge by comparing its evaluations with those of deepseek-v4-pro and qwen3.7-plus. Although the absolute scores vary across judges, all three produce the same ranking: HypReflect consistently outperforms SD-IncHyp, which in turn outperforms SD. This agreement indicates that our main conclusions are robust to the choice of LLM judge.
| SD | SD-IncHyp | HypReflect | |
|---|---|---|---|
| Deepseek-v4-flash | 77.1 | 90.3 | 93.0 |
| Deepseek-v4-pro | 74.1 | 86.3 | 87.1 |
| Qwen3.7-plus | 71.8 | 83.4 | 86.4 |
Appendix E Additional results
E.1 Qualitative Examples for Reflection
Figures 6 and 7 presents some qualitative examples that compares between hypotheses consolidation without reflection and with reflection. We find that direct consolidation tends to transform local hypotheses into broad and partially overlapping meta-preferences, such as favoring actionable, multifaceted, or tailored advice. These abstractions omit concrete concepts from the original hypothesis sets and provide less distinctive guidance for personalization. Reflection instead explicitly identifies which hypotheses should be merged, retained, or removed, preserving specific preference anchors such as crowdfunding, local outreach, marketing, and the balance between project management and relationship building. This example suggests that reflection reduces vague generalization and redundancy while producing more specific and informative preference hypotheses.
E.2 Qualitative Examples of Final Responses: SD vs. HypReflect
Figures 8 and 9 present qualitative comparisons between the final responses generated by SD and HypReflect, together with the preference hypotheses used by HypReflect. In the HiCupid example, both methods provide generally relevant advice, but HypReflect more explicitly emphasizes narrative structure and personification, reflecting the highest-weight hypothesis inferred from the user’s interest in engaging, science-oriented content. In the Flight example, the hypotheses help HypReflect resolve competing preferences by prioritizing the user’s strong dislike of long travel durations. Because all three flights have the same price, HypReflect correctly selects Flight 2, which has the shortest duration, whereas SD selects Flight 3. These examples illustrate how explicit, weighted hypotheses can help the model identify the most relevant user preferences and prioritize them when generating personalized responses.
E.3 Reflection window size variation
Figure 10 compares performance under different reflection window sizes on a subset of the HelpSteer2 setting at every 50 interactions.
Appendix F Prompts
HypReflect follows the same four-stage procedure across datasets. First, the model generates local preference hypotheses from a recent chunk of interactions (Figure 11). Second, it reflects on the generated sets of local preference hypotheses to identify supported, conflicting, and less important hypotheses (Figure 12). Third, it produces a refined hypothesis set using a prompt nearly identical to that in Figure 11, augmented with the reflection traces and multiple sets of local hypotheses. Finally, the refined hypothesis set is provided to the model to guide the generation of the final personalized response (Figure 13).
Dataset-specific instantiations.
Although all datasets follow the same overall prompt structure, we slightly adapt the preference-extraction instructions and output format to reflect their distinct evidence types and personalization targets. For Flight, the history contains a flight-recommendation question, the user’s selected flight, and feedback. The hypotheses capture preferences over departure time, duration, number of stops, and price using a fixed set of directional labels. For HelpSteer2, the history contains a request, response, and user feedback. The hypotheses capture general response-style preferences, such as verbosity, tone, structure, and formatting, while excluding task-specific details. For HiCupid, the history consists of prior dialogues, and the hypotheses preserve relevant interests, constraints, roles, relationships, and other concrete personalization information. We retain at most three hypotheses for Flight and HelpSteer2 and five for HiCupid, reflecting the richer, multi-session nature of the latter. Summary-enhanced variants additionally produce an Evidence/summary note: containing the supporting observations.