PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces
Abstract
Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users’ styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large user populations. We propose PLUME (Personalized Low-Rank Adaptation through User Modulation and Shared Subspace), a lightweight framework that achieves efficient and expressive per-user adaptation by leveraging a shared task-specific subspace. Specifically, PLUME first learns a global task subspace from aggregated user data. Personalization is then achieved by training only a lightweight small square matrix within this subspace, enabling each user to obtain a tailored model while keeping shared components fixed. Cross-layer shared parameters and rank-1 residual terms are further introduced to significantly reduce redundancy while maintaining expressiveness. Experiments on multiple personalized text generation benchmarks demonstrate that PLUME achieves comparable or superior performance to strong baselines, while reducing per-user parameters by over 95%. These results establish shared-subspace modulation with minimal residuals as a scalable and semantically grounded approach to LLM personalization. Our code is available at https://github.com/zhouhao0218/PLUME
1 Introduction
General-purpose large language models (LLMs) have achieved remarkable success across a wide range of natural language tasks. By pre-training on web-scale text corpora, modern LLMs such as GPT, LLaMA, Gemini, and their variants Zhao et al. (2023) acquire strong general-purpose language understanding and generation capabilities, enabling impressive performance spanning QA dialogue, machine translation, and logical reasoning Xu et al. (2024); Abbasiantaeb et al. (2024); Ferrag et al. (2025). However, these models are typically one-size-fits-all, treating all users the same and aiming to generally fit for tasks. This has spurred growing interest in personalized LLM, where an LLM’s outputs are tailored to an individual user’s style, preferences, or context towards an individual’s specific domain, task, and environment Fan et al. (2024); Zhao et al. (2023). Indeed, user-level personalization is increasingly viewed as crucial in applications like web-based QA assistants Wu et al. (2025), education Chu et al. (2025), and healthcare Bajwa et al. (2021). Among these diverse applications, personalized text generation represents a critical frontier in LLM research. Users increasingly demand AI systems that reflect their individuality rather than merely produce what they intend to say.
In response to these demands, a substantial wave of research towards personalized LLMs has emerged. Broadly speaking, existing approaches can be divided into two paradigms. The first category, prompt-based approaches Salemi et al. (2023); Liu et al. (2021); Wang et al. (2023); Kang et al. (2023); Qiu et al. (2025), attempts to achieve personalization through carefully constructed prompts that are built upon retrieved user history, profiles, or interaction logs. While such designs offer simplicity and interpretability, they heavily rely on explicit and high-quality user signals and are constrained by limited context length. The second category, fine-tuning-based approaches, directly modifies model parameters, hoping to encode personalized information in LLM itself. Majority of studies Zhang et al. (2024a); Zhao et al. (2025) choose to fine-tune a universal model across all user data, hoping to implicitly capture individual preferences deeply mined in different individuals’ data via a single run; however, such models tend to blur user distinctions and underperform in fine-grained personalization.
More recently, a small but growing body of work has explored the per-user model paradigm Tan et al. (2024); Bu et al. (2025), in which each individual maintains a dedicated model or adapter. Despite its advantages in personalization peroformance, this paradigm suffers from storage overhead when scaling to thousands or millions of users. Moreover, given the limited personal writing data, these adapters are at high risk of overfitting. This leads to a central challenge for personalized language modeling in writing assistance: How to achieve user-level adaptability that captures stylistic nuance and expressiveness without compromising storage efficiency?
To bridge this gap, we propose PLUME (Personalized Low-rank Adaptation through User Modulation and Shared Subspace), a lightweight yet expressive framework for personalized LLM fine-tuning in writing tasks. PLUME first performs global task adaptation by training a non-personalized LoRA module on the aggregated data from all users. The resulting low-rank LoRA matrices capture task-specific knowledge and establish a compact shared latent subspace that serves as a global foundation for subsequent personalization. Building upon this shared representation, PLUME introduces an User-Conditioned Subspace Mixer (USM), which is inserted between the pre-trained LoRA factors to modulate the shared subspace according to each user’s unique preferences. To further reduce the parameter redundancy, PLUME incorporates a Personalized Cross-Layer Shared (PCLS) module shared across layers, thereby eliminating redundant per-layer parameterization within each personalized model. Finally, a lightweight rank-1 residual component (Resid) is attached to provide fine-grained layer-specific correction with negligible cost. Through the principled composition of these components, PLUME achieves highly expressive personalization while significanly reducing per-user trainable parameters. Our key contributions are as follows:
- •
New Problem formulation We introduce the problem such that for the personalized text generation task, how to efficiently compress per-user PEFT parameters while maintaining both personalization effectiveness and model expressiveness, which is critical for practical personalization training.
- •
PLUME Framework We show that classic full LoRA is redundant under personalization settings and propose PLUME, a novel and lightweight personalization framework that alleviates the inefficiency of relatively heavy per-user LoRA adapters. PLUME significantly reduces per-user parameter cost while preserving personalization performance and expressive capacity.
- •
Extensive Empirical Validation and Analysis: We conduct comprehensive experiments on five different tasks from LaMP and LongLaMP benchmarks, covering both short and long form content generation. Experiment results show that our framework could robustly reduce the per-user parameter space to less than 5% of standard LoRA while maintaining or even improving personalization performance.
2 Problem and Preliminaries
2.1 Problem Formulation
We consider the task of personalized text generation, where the goal is to find a unique adapter for each individual to generate user-specific textual responses conditioned on both the current query and contextual information.
Data Construction. Let there be users denoted by . Each user has pairs of history of query–response:
where represents the -th query describing a specific task or instruction, and denotes the corresponding ground-truth response. Each response is tokenized as:
where denotes the -th token and its token-wise length.
To both emulate the real deployment regime and conform to the established best practice of few-shot in-context prompting, for each query , we retrieve its top- similar historical queries from the same user and pair them with their corresponding responses to construct the personalized context:
where could be any function that measures query-level similarity. Each instance is therefore represented as a triplet .
Training Objective. We adopt Decoder-Only LLMs for training where a pre-trained large language model parameterized by serves as the shared backbone, and each user is associated with a lightweight personalized adapter parametrized by . The full model for user is denoted as , which conditions generation on both the shared backbone and user-specific adaptation.
For user , the model is trained to predict each token in the user’s response given all previous tokens, the current query, and the retrieved context:
| (1) |
where represents auxiliary information from other users (Note: We formulate a general LLM personalization task in Eq.1, where could be in any form, such as other user’s textual histories. In this work, we encode into global Space and as a learned task-specific representation shared by all users). During optimization, is shared or fixed, while is optimized to minimize the user-specific loss:
| (2) |
Inference. At test time, given a new query for user , we retrieve its top- similar query–response pairs from to form . The shared backbone combined with the learned personalized parameters is then used to generate the predicted response:
and the output is compared with the ground-truth using standard evaluation metrics.
| Symbol | Description |
| User index, . | |
| Query–response pair index for user . | |
| Token index within a response sequence. | |
| The -th query of user . | |
| Ground-truth response corresponding to . | |
| Set of all historical pairs for user . | |
| Top- similar pairs retrieved from . | |
| Auxiliary information (optionally) derived from other users’ histories. | |
| Parameters of the shared base language model. | |
| Personalized parameters (adapter) for user . | |
| Personalized LLM combining shared backbone and user’s PEFT module. | |
| Subscript indicating user-specific parameters. | |
| Superscript indicating the -th layer in the model. |
2.2 One-PEFT-per-User Personalization Paradigm
We follow the One PEFT per User (OPPU Tan et al. (2024)) training framework, which assigns an independent PEFT module to each user for personalized fine-tuning.
Concretely, OPPU first ignores user differences and trains a task-specific global LoRA adapter using the union of all users’ training data , resulting in a shared adapter . This stage corresponds to optimizing
| (3) |
The Second step simply follows Eq 2, but with merged model .
If we further examine this paradigm through the lens of a specific target module, e.g., , the resulting parameter composition becomes more explicit. Typically, the effective weight matrix on a layer can be expressed as:
| (4) |
where denotes the shared base model weight, represents the global LoRA adapter trained from the aggregated non-personalized data , and denotes the user-specific LoRA adapter paramterized from .
Vanilla OPPU provides a straightforward way to incorporate user-specific knowledge, nevertheless, it faces two major limitations. First, maintaining a separate LoRA module for every user leads to prohibitively high storage and memory costs, especially when the number of users is large. Second, since each user’s training data can vary drastically in scale and quality, independently optimizing a full-rank LoRA for every user often causes overfitting or unstable personalization performance. As illustrated in Figure 2(a), the performance of OPPU drops significantly when the individual LoRA rank is reduced.
These issues highlight the need for a more parameter-efficient and stable approach to capture user-specific preferences without requiring a full LoRA module per user.
2.3 Notation Summary
Table 1 summarizes the key symbols used in this paper, covering both the preceding problem formulation and the following method sections.
3 Method
In this section, we propose PLUME, a novel framework designed to reduce redundant personalized parameters while preserving personalization expressivity. Figure 1 illustrates the overall architecture of PLUME. We will start from our key observation, comparison with vanilla OPPU and progressively builds our efficient yet effective model.
3.1 PLUME Framework
Revisiting Task Space. From a linear algebraic perspective, the task-specific LoRA update on layer can be expressed as , where projects inputs into a low-dimensional task subspace, and expands the projected representation back to the model’s hidden space. Hence, for any input , the LoRA output always resides in the column space . OPPU extends this by assigning each user an individual adapter , which effectively introduces a new output subspace basis specific to each user. From this view, OPPU increases the expressive capacity of the model by expanding the dimensional coverage of across users, thereby enhancing personalization.
However, this completely “separates” the shared task space and the augmented individual space , leaving the shared adapter only learns how to project tokens into the task subspace, while not being able to control how different users behave within the collaboratively trained existing shared task subspace . This motivates us to allow user-dependent modulation of the shared subspace.
User-Conditioned Subspace Mixer (). Intuitively, since users may differ in style, intent, or preference, the way they combine or emphasize task-subspace directions should vary. To capture user-specific task space utilization, we insert a lightweight matrix between and :
| (5) |
We term a User-Conditioned Subspace Mixer (USM). It can be interpreted as a re-indexing or re-weighting operator that re-combines latent task directions inside the shared LoRA subspace. Thus, enables personalized modulation of how a user exploits the common task subspace , offering additional expressivity with negligible parameter cost.
PLUME. Before we move on to a more compact model, we first assess how much expressivity the USM alone provides, we introduce a lightweight per-layer residual term aiming to approximate the representational power of high-rank OPPU:
| (6) |
Here follow the same formulation as , but with a reduced individual rank . Empirically, we find that when , PLUME already achieves performance comparable to OPPU with rank 64, as shown in Fig 2(a). Interestingly, further increasing provides no additional gains—indicating that once the shared task subspace is efficiently utilized via , a large individual subspace becomes redundant and inefficient, leading to overfitting. This observation suggests that there may still exist parameter redundancy even within the reduced adapters. Thus, we further push PLUME to the extreme—seeking the most compact yet expressive form of individual representation. As illustrated in Fig 2(b), our proposed novel shared subspace mechanism enables PLUME to reduce the individual parameter while maintaining highly compatible performance.
3.2 Parameter Reduction via Shared Subspace
CKA Analysis for Layer-Wise Representation Similarity for Personalized Model Inspired by prior work He et al. (2025); Kopiczko et al. (2023); Zhou et al. () that sharing weights across layers can enhance model expressiveness while reducing parameter usage, we investigate whether different layers in personalized LoRA models actually learn similar user-specific representations. To this end, we perform a Centered Kernel Alignment (CKA), a widely used approach for comparing neural representations across layers and models Kornblith et al. (2019); Liu et al. (2025b) to quantify layer-wise representational similarity across the personalized adapters (See Appendix A.3 for more about CKA).
As shown in Figure 3, OPPU-trained adapters exhibit remarkably high inter-layer CKA values (often exceeding ), indicating that the personalized LoRA subspaces learned by different layers are strongly aligned. This suggests that the layer-wise residuals tend to encode similar directions of user-specific variation, leading to redundant parameterization across depth. Motivated by this finding, we hypothesize that a shared subspace across layers could capture the dominant personalized factors more efficiently.
PLUME-s. Building on this insight, we observe that the per-layer residuals often act similarly to layer-specific biases that adjust the shared projection directions. Accordingly, we decompose into two functional components: (a) a Personalized Cross-Layer Shared Subspace (PCLS), ) capturing cross-layer user-specific residuals in a low-rank shared form, and (b) a rank-1 layer-wise Residual (Resid), that enables fine-grained local adaptation. The resulting formulation becomes:
| (7) |
Here with rank parameterize the Layer-Shared Personalized Subspace reused across all layers, while serve as the rank-1 layer refinements for fine adjustment. Extensive experiments demonstrate that leveraging a shared residual subspace allows substantial reduction of individual residual ranks without compromising performance, and can even outperform OPPU in certain settings (see Sec. 4).
Summary and Parameter Efficiency. Let be the number of layers and the global LoRA rank, and the embedding vector dimension. In conclusion, PLUME adaptively unifies four complementary components: a global task-specific adapter shared by all users, and three lightweight personalized modules. (1) The User-Conditioned Subspace Mixer (USM) adaptively reuses the global task space with parameters. (2) The Personalized Cross-Layer-Shared Personalized Subspace (PCLS) captures cross-layer personalization compactly with parameters. (3) The rank-1 personalized residuals restore fine local flexibility using parameters. In total, the per-user complexity is compared to for OPPU’s layer-wise independent adapters. When , the parameter reduction becomes particularly significant, see Table 3 for practical analysis.
4 Experiments
| Type | Method | Long Content Generation | Short Content Generation | |||||||||||||
| Abstract Generation | Product Review | Topic Writing | News Headline | Scholarly Title | ||||||||||||
| R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | ||
| Non-Personalized | BASE | 0.3497 | 0.1688 | 0.2408 | 0.3397 | 0.1389 | 0.2335 | 0.2892 | 0.1238 | 0.2050 | 0.1313 | 0.1158 | 0.0962 | 0.3905 | 0.3165 | 0.4084 |
| LoRA | 0.3491 | 0.2036 | 0.2456 | 0.3877 | 0.2323 | 0.2707 | 0.2655 | 0.1378 | 0.1732 | 0.2188 | 0.2010 | 0.1939 | 0.4648 | 0.4138 | 0.4228 | |
| PiSSA | 0.3523 | 0.1995 | 0.2536 | 0.3964 | 0.2326 | 0.2826 | 0.2854 | 0.1393 | 0.1927 | 0.2230 | 0.2056 | 0.1999 | 0.4801 | 0.4266 | 0.4351 | |
| AdaLoRA | 0.3503 | 0.2089 | 0.2395 | 0.3419 | 0.2056 | 0.2323 | 0.2279 | 0.1212 | 0.1530 | 0.2299 | 0.2113 | 0.1892 | 0.4621 | 0.4203 | 0.3940 | |
| QLoRA | 0.3450 | 0.2001 | 0.2394 | 0.3940 | 0.2339 | 0.2769 | 0.2697 | 0.1399 | 0.1762 | 0.2224 | 0.2037 | 0.1958 | 0.4628 | 0.4101 | 0.4196 | |
| LoRA-One | 0.3780 | 0.2317 | 0.2750 | 0.3800 | 0.2320 | 0.2680 | 0.2930 | 0.1410 | 0.2110 | – | – | – | 0.4480 | 0.4030 | 0.4070 | |
| Personalized | RAG | 0.3512 | 0.1743 | 0.2567 | 0.3389 | 0.1454 | 0.2336 | 0.2936 | 0.1264 | 0.2199 | 0.1470 | 0.1304 | 0.1047 | 0.3982 | 0.3229 | 0.4050 |
| OPPU | 0.4147 | 0.2374 | 0.2824 | 0.4416 | 0.2481 | 0.3165 | 0.3191 | 0.1510 | 0.2126 | 0.2383 | 0.2178 | 0.2008 | 0.5146 | 0.4510 | 0.4365 | |
| CoPE | 0.3779 | 0.2247 | 0.2530 | 0.3600 | 0.2385 | 0.2645 | 0.2303 | 0.1386 | 0.1754 | 0.2324 | 0.2101 | 0.1963 | 0.4741 | 0.4076 | 0.4133 | |
| PLUME | 0.4168 | 0.2395 | 0.2839 | 0.4423 | 0.2460 | 0.3161 | 0.3217 | 0.1534 | 0.2172 | 0.2150 | 0.1974 | 0.1778 | 0.5075 | 0.4459 | 0.4246 | |
| PLUME-s | 0.4117 | 0.2322 | 0.2803 | 0.4369 | 0.2434 | 0.3107 | 0.3176 | 0.1517 | 0.2131 | 0.2427 | 0.2219 | 0.2039 | 0.5177 | 0.4537 | 0.4359 | |
| Method | #Params | % of OPPU | Memory (GB) | Runtime |
| OPPU | 167,772,160 | 100% | 0.6250 | 4h53m |
| CoPE | 167,772,160 | 100% | 0.1016 | 4h56m |
| PLUME | 11,403,264 | 6.80% | 0.0425 | 4h30m |
| PLUME-s | 3,702,784 | 3.11% | 0.0156 | 4h42m |
In this section, we conduct experiments to systematically investigate the following research questions:
RQ1: How does PLUME perform against non-personalized and personalized PEFT benchmarks in various personalized text generation tasks?
RQ2: How much efficiency gain does PLUME achieve?
RQ3: How do the different components in PLUME affect its performance effectiveness?
RQ4: How does PLUME behave under varying degrees of parameter reduction, and what is the resulting trade-off between model compactness and personalization performance?
4.1 Experimental Setup.
Datasets We adopt the widely used benchmarks LongLaMP Kumar et al. (2024) and LaMP Salemi et al. (2023) to evaluate our model’s ability across both short-form and long-form content personalized content generation tasks. Specifically, we select 3 datasets: Abstract Generation, Product Review, and Topic Writing from LongLaMP, and 2 datasets: News Headline generation, Scholarly Title generation from LaMP. We leave details on data preprocessing and data statistics to Appendix B
Baselines. We compare our methods against a series of strong parameter-efficient fine-tuning baselines under both non-personalized training and personalized training settings. For the non-personalized setting, we compare against the base model, and LoRA Hu et al. (2021), AdaLoRA Zhang et al. (2023), PiSSA Meng et al. (2024), and QLoRA Dettmers et al. (2023) and LoRA-One Zhang et al. (2025b) which all adapt model parameters through a shared low-rank adapter across users. For the personalized setting, we consider retrival-based approach such as RAG Tan et al. (2024), training-based baseline OPPU Tan et al. (2024), and a more recent work CoPE Bu et al. (2025). Finally, we evaluate our two variants: PLUME and PLUME-s.
Implementation To ensure fairness, all models are evaluated with a consistent LoRA configuration (rank = 64) across both non-personalized and personalized settings, including the OPPU and CoPE baselines. To demonstrate the generalizability and robustness of our approach, we evaluate PLUME on two distinct open-source large language models: LLaMA2-7B Touvron et al. (2023) and Mistral-7B-Instruct-v0.2 Jiang et al. (2023). As illustrated in Figure 4(b) in the later discussion, PLUME already reaches competitive results with a very small rank, so we report PLUME results with residual rank , and PLUME-s with residual rank and shared rank as low as . Due to space constraints, we present the results using the Mistral-7B backbone in the main text and leave more experiment results and details in Appendix C.
Evaluation Metrics In accordance with established practices in prior work Tan et al. (2024); Kumar et al. (2024), we employ a standard set of automatic metrics ROUGE-1, ROUGE-L Lin (2004), and METEOR Banerjee and Lavie (2005) to quantitatively assess the lexical overlap, fluency, and semantic correspondence between generated outputs and reference responses (See more details in Appendix C.1)
4.2 Main Results
Overall Perfomance (RQ1) Table 2 reports the overall performance across five generation tasks. Overall, personalized methods outperform non-personalized ones across all metrics, confirming the advantage of user-specific adaptation. Within the personalized group, despite using far fewer parameters than strong baselines such as OPPU and CoPE, the PLUME series delivers comparable or better results than other baseline methods. For instance, PLUME attains the highest ROUGE-L on Abstract Generation (0.2395) and comparable performance on other long-text tasks, demonstrating that the personalized module USM would substantially enhances expressiveness. When incorporating the shared subspace, PLUME-s performance slightly improves, e.g., ROUGE-1 rises from 0.2150 to 0.2427 on News Headline Generation Tasks. This indicates that the shared subspace not only compresses parameters but also mitigates redundancy and overfitting issues observed in OPPU.
In summary, PLUME achieves a favorable balance between expressiveness and efficiency. It retains nearly identical generation quality to strong personalized baselines OPPU and CoPE while using only a fraction of their parameters, and even surpasses them on several short-text tasks.
Efficiency Comparison (RQ2) One of our main contributions is that PLUM substantially improves parameter efficiency for LLM personalization adapters. Here, we quantify the extent to which this compression frontier can be pushed under practical setting. As in Table 3, compared to OPPU, which requires 168M parameters per user, PLUME and PLUME-s reduce the cost to only 6.8% and 3.1% respectively, achieving over – compression while preserving personalization quality and requiring less training time.
Ablation Study: Effectiveness of Components (RQ3)
Figure 4(a) compares the effects of three core modules: USM, PCLS, and Resid on both long and short text generation tasks.
Across both tasks, removing any single module consistently leads to a drop in ROUGE-L and METEOR, suggesting that all three work jointly to enhance expressiveness, coherence, and efficiency. This consistent trend highlights that the modules are complementary rather than redundant. The relative impact of each module remains consistent across both tasks (Resid > PCLS > USM). Moreover, the effect of the Resid module is more pronounced in long-form generation tasks. Notably, even though Resid employs only a rank-1 update, it still contributes sufficiently to representational richness, showing that lightweight residual paths can complement USM’s user-specific modulation effectively. These results demonstrate that each component is indispensable for achieving a strong balance between personalization, coherence, and parameter efficiency.
Parameter Efficiency VS Model Performance (RQ4)
Sensitivity of residual rank.
To investigate the sensitivity of residual dimensionality, we vary the rank of the residual module. As shown in Figure 4(b), performance on Abstract Generation (left) first increases with larger residual ranks and peaks around rank 4, after which both ROUGE-L and METEOR scores gradually decline. This trend indicates that a small residual rank is sufficient to capture user-specific variations, while larger ranks introduce redundancy and overfitting. In contrast, the Scholarly Title task (right) shows much smaller variance across ranks, with performance already strong at rank 1 and exhibiting a slight downward trend thereafter. This validates our conjecture that short-form generation requires fewer personalized parameters and benefits more from PLUME’s shared-subspace regularization.
Sensitivity on PCLS rank. To further examine the influence of the shared subspace capacity, we vary the rank of PCLS module. As illustrated in Figure 4(c), the results show an opposite trend between the two settings. For Abstract Generation, performance steadily improves as the shared rank increases, indicating that long-form generation benefits from a richer shared subspace capable of modeling broader contextual dependencies and semantic consistency across layers. In contrast, the Scholarly Title task exhibits an inverted trend—performance peaks at a small shared rank and then declines as the rank grows—suggesting that short-form generation requires only limited shared capacity. These contrasting patterns highlight PLUME’s flexibility that its shared subspace can effectively balance generalization and personalization.
5 Conclusion
This paper investigates how to achieve fine-grained user adaptation with better parameter efficiency in the personalized text generation task. To tackle this, we proposed PLUME, a novel framework for personalized LLMs through low-rank user modulation and shared subspaces. The proposed design comprises of the User-Conditioned Subspace Mixer (USM), the cross-layer shared personalized subspace (PCLS), and the rank-1 residuals (Resid) and demonstrates that rich personalization can be achieved with only a fraction of the parameters required by conventional per-user adapters such as OPPU. Comprehensive experiments across five personalized text-generation benchmarks with two different foundation models as backbone show that PLUME consistently matches or exceeds existing personalized LoRA variants while reducing per-user parameters by over 95%.
Overall, PLUME provides a more effective and highly compact user representation for personalized LLMs, and we believe it establishes a new research pathway for advancing LLM personalization.
References
- Let the llms talk: simulating human-to-human conversational qa via zero-shot llm-to-llm interactions. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pp. 8–17. Cited by: §1.
- Artificial intelligence in healthcare: transforming the practice of medicine. Future healthcare journal 8 (2), pp. e188–e194. Cited by: §1.
- METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72. Cited by: §C.1, §4.1.
- Personalized llm decoding via contrasting personal preference. arXiv preprint arXiv:2506.12109. Cited by: §A.2, Appendix B, §1, §4.1.
- Llm agents for education: advances and applications. arXiv preprint arXiv:2503.11733. Cited by: §1.
- Uncovering chatgpt’s capabilities in recommender systems. In Proceedings of the 17th ACM Conference on Recommender Systems, pp. 1126–1132. Cited by: §A.2.
- Qlora: efficient finetuning of quantized llms. Advances in neural information processing systems 36, pp. 10088–10115. Cited by: §A.1, §4.1.
- A survey on rag meeting llms: towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pp. 6491–6501. Cited by: §1.
- From llm reasoning to autonomous ai agents: a comprehensive review. arXiv preprint arXiv:2504.19678. Cited by: §1.
- RaSA: rank-sharing low-rank adaptation. arXiv preprint arXiv:2503.12576. Cited by: §3.2.
- Lora: low-rank adaptation of large language models. arxiv 2021. arXiv preprint arXiv:2106.09685 10. Cited by: §A.1, §4.1.
- Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118. Cited by: Appendix B.
- Mistral 7b. External Links: 2310.06825 Cited by: §4.1.
- Do llms understand user preferences? evaluating llms on user rating prediction. arXiv preprint arXiv:2305.06474. Cited by: §A.2, §1.
- Customizing language models with instance-wise lora for sequential recommendation. Advances in Neural Information Processing Systems 37, pp. 113072–113095. Cited by: §A.2.
- Vera: vector-based random matrix adaptation. arXiv preprint arXiv:2310.11454. Cited by: §3.2.
- Similarity of neural network representations revisited. In International conference on machine learning, pp. 3519–3529. Cited by: §3.2.
- Longlamp: a benchmark for personalized long-form text generation. arXiv preprint arXiv:2407.11016. Cited by: §4.1, §4.1.
- Learning to rewrite prompts for personalized text generation. In Proceedings of the ACM Web Conference 2024, pp. 3367–3378. Cited by: §A.2.
- Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81. Cited by: §C.1, §C.1, §4.1.
- A survey of personalized large language models: progress and future directions. arXiv preprint arXiv:2502.11528. Cited by: §A.2.
- Dora: weight-decomposed low-rank adaptation. In Forty-first International Conference on Machine Learning, Cited by: §A.1.
- Spectral insights into data-oblivious critical layers in large language models. arXiv preprint arXiv:2506.00382. Cited by: §A.3, §3.2.
- Pre-trained language model for web-scale retrieval in baidu search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 3365–3375. Cited by: §A.2, §1.
- Pissa: principal singular values and singular vectors adaptation of large language models. Advances in Neural Information Processing Systems 37, pp. 121038–121072. Cited by: §A.1, §4.1.
- Pearl: personalizing large language model writing assistants with generation-calibrated retrievers. arXiv preprint arXiv:2311.09180. Cited by: §A.2.
- Memorag: boosting long context processing with global memory-enhanced retrieval augmentation. In Proceedings of the ACM on Web Conference 2025, pp. 2366–2377. Cited by: §A.2.
- Measuring what makes you unique: difference-aware user modeling for enhancing llm personalization. arXiv preprint arXiv:2503.02450. Cited by: §A.2, §1.
- Lamp: when large language models meet personalization. arXiv preprint arXiv:2304.11406. Cited by: §A.2, §A.2, §1, §4.1.
- Persona-db: efficient large language model personalization for response prediction with collaborative data refinement. arXiv preprint arXiv:2402.11060. Cited by: §A.2.
- Aligning large language models with implicit preferences from user-generated content. arXiv preprint arXiv:2506.04463. Cited by: §A.2.
- Democratizing large language models via personalized parameter-efficient fine-tuning. arXiv preprint arXiv:2402.04401. Cited by: §A.2, Appendix B, §1, §2.2, §4.1, §4.1.
- Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §4.1.
- Learning personalized alignment for evaluating open-ended text generation. arXiv preprint arXiv:2310.03304. Cited by: §A.2, §1.
- KaSA: knowledge-aware singular-value adaptation of large language models. arXiv preprint arXiv:2412.06071. Cited by: §A.1.
- Milora: harnessing minor singular components for parameter-efficient llm finetuning. arXiv preprint arXiv:2406.09044. Cited by: §A.1.
- Personalized large language models. In 2024 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 511–520. Cited by: §A.2.
- Webwalker: benchmarking llms in web traversal. arXiv preprint arXiv:2501.07572. Cited by: §1.
- Contrastive preference optimization: pushing the boundaries of llm performance in machine translation. arXiv preprint arXiv:2401.08417. Cited by: §1.
- Personalized llm response generation with parameterized memory injection. arXiv preprint arXiv:2404.03565. Cited by: §A.2, §1.
- Adalora: adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512. Cited by: §A.1, §4.1.
- PRIME: large language model personalization with cognitive memory and thought processes. arXiv preprint arXiv:2507.04607. Cited by: §A.2.
- Personalized lora for human-centered text understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 19588–19596. Cited by: §A.2.
- LoRA-one: one-step full gradient could suffice for fine-tuning large language models, provably and efficiently. arXiv preprint arXiv:2502.01235. Cited by: §4.1.
- A survey of large language models. arXiv preprint arXiv:2303.18223 1 (2). Cited by: §1.
- NextQuill: causal preference modeling for enhancing llm personalization. arXiv preprint arXiv:2506.02368. Cited by: §1.
- [47] BSLoRA: enhancing the parameter efficiency of lora with intra-layer and inter-layer sharing. In Forty-second International Conference on Machine Learning, Cited by: §3.2.
- Lifelong personalized low-rank adaptation of large language models for recommendation. arXiv preprint arXiv:2408.03533. Cited by: §A.2.
Appendix A Related Work
A.1 LoRA Based Fine Tuning
LoRA (Low-Rank Adaptation) (Hu et al., 2021) is one of the most widely adopted parameter-efficient fine-tuning (PEFT) techniques for adapting large language models (LLMs) with minimal trainable parameters. Instead of updating the full weight matrix, LoRA introduces a low-rank update , where and with , thus reducing the number of trainable parameters from to . Due to its effectiveness and efficiency, LoRA has become one of the most effective and practical approaches in personalized LLM adaptation tasks. Building upon this foundation, several variants have been developed to improve flexibility, stability, and efficiency. AdaLoRA (Zhang et al., 2023) dynamically allocates ranks through orthogonal regularization, DoRA (Liu et al., 2024) decouples direction and magnitude learning to stabilize optimization, and QLoRA (Dettmers et al., 2023) integrates quantization for memory-efficient adaptation of large-scale models. A complementary line of work leverages singular value decomposition (SVD) for structured adaptation: PiSSA (Meng et al., 2024) initializes LoRA adapters with principal singular components for faster convergence, while MiLoRA (Wang et al., 2024b) and KASA (Wang et al., 2024a) focus on minor singular directions to enhance generalization. Collectively, when personalization requires fine-tuning, LoRA and its variants serve as the most effective and widely adopted PEFT solutions, enabling scalable adaptation of LLMs to individual users while preserving efficiency and generalization.
A.2 LLM personalization
The problem of adapting large language models to individual users has been extensively studied, yielding a diverse landscape of personalization strategies. Typically, these paradigms could be classified into two categories: Prompt-based and Adapter-based approaches.
Prompt-Based Approaches Among the earliest and most lightweight approaches are prompt-based methods, which target to extract and encode user-specific information into handcrafted or learned prompts to guide model behavior without modifying LLM parameters Liu et al. (2025a); Liu et al. (2021). In light of this, various prompting techniques, such as CoT and in-context learning, have been employed to provide a summarized or sampled user behavior history Wang et al. (2023); Kang et al. (2023). For instance, Dai et al Dai et al. (2023) leverages sampled user purchase history to guide LLMs to make limited-size personalized item recommendations, while DPL Qiu et al. (2025) improves personalization by encoding extracted inter-user comparisons into personalized prompts. However, such prompt-based approaches are fundamentally constrained by the model’s finite context window, and further suffer from the increasing difficulty of extracting effective information as user data grows. To mitigate these limitations, recent work has turned to retrieval-augmented personalized prompting, which dynamically retrieves salient user records from a long-term memory to populate the prompt, obviating the need for exhaustive history inclusion Salemi et al. (2023); Qian et al. (2025); Sun et al. (2024). As an example, Pearl Mysore et al. (2023) leverages a retriever calibrated to the generation objective to select historical user-authored documents that properly enhance and augment the prompt. More recently, PRIME Zhang et al. (2025a) proposes to enhance LLM personalization via episodic and semantic memory mechanisms, aiming to achieve retaining and updating memory for more efficient individual information retrieval.
While being conceptually simple and highly interpretable, prompt-based personalization relies on strong, explicit user signals in historical data Tan et al. (2025). In content generation tasks, where personalization hinges on implicit factors such as stylistic preferences or individual personality traits, these approaches often struggle to reliably capture users’ intent and preferences, leading to unstable and degraded performance.
Adapter-Based Approaches Adapter-based approaches typically adapt PEFT methods and could be divided into two types, with the first category training all users with a shared model Zhu et al. (2024); Li et al. (2024); Existing work such as LM-P Woźniak et al. (2024), PLoRAZhang et al. (2024b) and MiLP Zhang et al. (2024a) all fall in this research line. Recent advancements like iLoRA Kong et al. (2024) and RecLoRA Zhu et al. (2024) have incorporated the Mixture of Experts (MoE) structure to model the diverse range of user preferences and behaviors. The second category of methods further enhances personalization performance by assigning each user a PEFT-trained model. For example, OPPU Tan et al. (2024) equips each user with a LoRA module and trains it on the user’s individual data, achieving SOTA results on personalized classification and short content generation tasks on LaMP Salemi et al. (2023) dataset. Further, CoPE Bu et al. (2025) demonstrates the effectiveness of the strategy on long content generation tasks by augmenting training data with negative sampling together with contrastive loss.
Despite its promising results, the main limitation of OPPU lies in its parameter growth. By assigning an individual PEFT module to each user, the total number of parameters increases rapidly with the number of users, resulting in considerable storage overhead. In addition, since each user typically has only a small amount of data, training a conventional PEFT module (such as LoRA) on such limited data often leads to overfitting and redundant parameters. Hence, reducing user-specific parameters is both urgent and essential for effective personalized LLMs. This work first introduces this pivotal challenge, and proposes our solution–PLUME, which decomposes traditional LoRA into ultra-lightweight modules through a shared task-specific parameter space and a layer-wise sharing mechanism, achieving competitive personalization performance with minimal parameter redundancy.
A.3 Details of Centered Kernel Alignment (CKA)
We use linear Centered Kernel Alignment (CKA) Liu et al. (2025b) to measure representation similarity between LoRA parameters across layers. CKA provides a scale-invariant and rotation-invariant similarity measure, making it suitable for comparing learned adapter representations.
Specifically, given two column-centered and vectorized LoRA weight matrices and , the linear CKA is defined as: Given two column-centered vectorized LoRA matrices from different layers of the same user model, the linear CKA is defined as:
Appendix B Data Processing Detail and Data Statistics
In preparing the data, we follow the general setup of prior frameworks such as OPPU Tan et al. (2024) and CoPE Bu et al. (2025) and select 200 users with sufficient interaction histories as our evaluation cohort. For each user, we aggregate all historical interactions and then partition them into training, validation, and test subsets with an 8:1:1 ratio based on temporal ordering. When a personalization prompt requires leveraging user history as examples, we restrict retrieval to the training portion only, ensuring a realistic personalization setup without test leakage. In this case, we employ Contriever Izacard et al. (2021) to retrieve the top-K most relevant history entries, which are then included as few-shot demonstrations for generating the target content. This design allows us to assess both the model’s raw personalization ability and its robustness to retrieved history length and quality. Dataset statistics and splits are provided in Table 4.
| Dataset | AG | PR | TW | NH | ST |
| # Questions | 14,065 | 6,389 | 5,138 | 33,072 | 15,202 |
| Avg Q Length | 332.61 | 881.90 | 585.68 | 188.36 | 509.99 |
| Avg Target Length | 181.44 | 416.73 | 324.66 | 14.95 | 14.54 |
| Type | Method | Long Content Generation | Short Content Generation | |||||||||||||
| Abstract Generation | Product Review | Topic Writing | News Headline | Scholarly Title | ||||||||||||
| R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | R-1 | R-L | MTR | ||
| Non-Personalized | BASE | 0.2050 | 0.1092 | 0.1286 | 0.3248 | 0.1390 | 0.2011 | 0.2710 | 0.1204 | 0.1727 | 0.1320 | 0.1157 | 0.0831 | 0.4379 | 0.3556 | 0.4192 |
| LoRA | 0.3358 | 0.1983 | 0.2279 | 0.3602 | 0.2204 | 0.2489 | 0.2465 | 0.1331 | 0.1585 | 0.2106 | 0.1948 | 0.1786 | 0.4474 | 0.4004 | 0.3849 | |
| PiSSA | 0.3378 | 0.1910 | 0.2374 | 0.3833 | 0.2276 | 0.2687 | 0.2681 | 0.1359 | 0.1768 | 0.2140 | 0.1970 | 0.1886 | 0.4718 | 0.4207 | 0.4263 | |
| AdaLoRA | 0.3403 | 0.1989 | 0.2328 | 0.2713 | 0.1336 | 0.1595 | 0.2135 | 0.1095 | 0.1325 | 0.2127 | 0.1963 | 0.1718 | 0.4445 | 0.4029 | 0.3744 | |
| QLoRA | 0.3260 | 0.1934 | 0.2157 | 0.3668 | 0.2238 | 0.2563 | 0.2401 | 0.1291 | 0.1515 | 0.2062 | 0.1907 | 0.1735 | 0.4482 | 0.4030 | 0.3835 | |
| LoRA-One | 0.4080 | 0.2700 | 0.3060 | 0.3750 | 0.2180 | 0.2660 | 0.2720 | 0.1340 | 0.1830 | 0.2150 | 0.1980 | 0.1900 | 0.4700 | 0.4190 | 0.4230 | |
| Personalized | RAG | 0.3549 | 0.1808 | 0.2435 | 0.3268 | 0.1399 | 0.2111 | 0.2435 | 0.1124 | 0.1607 | 0.1322 | 0.1157 | 0.0941 | 0.4231 | 0.3461 | 0.3881 |
| OPPU | 0.3933 | 0.2214 | 0.2614 | 0.4156 | 0.2369 | 0.2872 | 0.2878 | 0.1428 | 0.1827 | 0.2285 | 0.2091 | 0.1900 | 0.4955 | 0.4364 | 0.4050 | |
| CoPE | 0.3320 | 0.2071 | 0.2731 | 0.3263 | 0.2125 | 0.2644 | 0.1607 | 0.1041 | 0.1370 | 0.1779 | 0.1589 | 0.1928 | 0.4030 | 0.3485 | 0.3920 | |
| PLUME | 0.3962 | 0.2225 | 0.2664 | 0.4209 | 0.2393 | 0.2947 | 0.2944 | 0.1439 | 0.1872 | 0.2291 | 0.2100 | 0.1916 | 0.5103 | 0.4487 | 0.4231 | |
| PLUME-s | 0.3865 | 0.2169 | 0.2543 | 0.4142 | 0.2397 | 0.2890 | 0.2790 | 0.1406 | 0.1742 | 0.2336 | 0.2155 | 0.1940 | 0.5107 | 0.4512 | 0.4243 | |
| # Target Token | OPPU | PLUME | PLUME-s |
| Bin 1: 79–125 | 120 | 119 ( -0.83% ) | 123 ( +0.2.5% ) |
| Bin 2: 126–140 | 133 | 130 ( -2.20% ) | 135 ( +1.50% ) |
| Bin 3: 141–155 | 140 | 139 ( -0.71% ) | 141 ( +0.71% ) |
| Bin 4: 156–170 | 151 | 150 ( -0.67% ) | 155 ( +2.64% ) |
| Bin 5: 171–245 | 156 | 158 ( +1.28% ) | 166 ( +6.41% ) |
Appendix C More Experiment
C.1 Details of Evaluation Metrics
To improve clarity with standard text generation benchmarks, we provide formal definitions of the automatic evaluation metrics used in our experiments. Let denote the generated text and denote the reference text.
ROUGE-1.
ROUGE-1 Lin (2004) measures unigram (1-gram) overlap between the generated and reference texts, serving as a proxy for lexical content coverage. Let and denote the frequency of unigram in and , respectively. The unigram overlap is defined as:
| (8) |
The recall form of ROUGE-1 is computed as:
| (9) |
In practice, we report the F1-score to balance precision and recall:
| (10) |
where and denote unigram precision and recall.
ROUGE-L.
ROUGE-L Lin (2004) is based on the Longest Common Subsequence (LCS) between and , capturing sentence-level structural similarity without requiring consecutive n-gram matches. Let denote the length of the longest common subsequence. The recall and precision are defined as:
| (11) |
The ROUGE-L F-measure is:
| (12) |
where is typically set to favor recall.
METEOR.
METEOR (Metric for Evaluation of Translation with Explicit ORdering) Banerjee and Lavie (2005) extends lexical overlap metrics by incorporating exact matches, stem matches, synonym matches, and a fragmentation penalty to account for word ordering. Let denote the number of matched unigrams between and . Precision and recall are defined as:
| (13) |
METEOR computes a weighted harmonic mean emphasizing recall:
| (14) |
To penalize fragmented matches, a penalty term is introduced:
| (15) |
where is the number of matched chunks and are hyperparameters. The final METEOR score is:
| (16) |
Overall, ROUGE-1 measures lexical overlap, ROUGE-L captures structural similarity via sequence alignment, and METEOR incorporates semantic matching and ordering penalties. Together, they provide complementary perspectives for evaluating personalized text generation quality.
C.2 Training Details and Reproducibility
For fair comparison, we conduct all experiments with 5 epochs with the AdamW optimizer. To avoid randomness, we set the generation temperature to be 0. For hyper-parameters, we tried different combinations and report the best results. Specifically, learning rate is chosen from ; component coefficient and from . In sensitivity study, we investigate the influence of LoRA rank from . All experiments were conducted on a single cluster node equipped with a Dell PowerEdge C6620 and NVIDIA H100 GPUs with 94 GB of memory.
C.3 Llama2-7B Results
Table 5 reports the results using LLaMA-2-7b-chat as the backbone model. Both PLUME variants outperform prior personalized baselines. PLUME achieves slightly higher scores on several tasks, while PLUME-s attains the best overall balance between performance and parameter efficiency. Together with the Mistral-7B results in Table 2, these findings demonstrate the robustness of our method across different backbone models.
Appendix D More Analysis
D.1 Sensitivity on Number of Retrieved History interactions
From Figure 5, we observe contrasting effects of increasing the number of in-prompt reference examples on personalized long-form content generation (e.g., Abstract Generation) versus short-form generation (e.g., Scholarly Title Generation) tasks. For Abstract Generation, adding examples yields small but consistent gains and then plateaus: performance nudges up from about 1 example to 3–4 examples, suggesting that extra stylistic cues help open-ended expansion without overwhelming the model. In contrast, for Scholarly Title Generation, quality peaks early and then declines: 1–2 examples give the best scores, while 3–4 examples slightly hurt, likely due to prompt dilution and competing keyword signals in a short, constrained output space. These findings indicate that long-form personalized generation benefits from more reference examples, whereas short-form generation achieves optimal performance with only 1–2 examples, providing practical guidance for designing in-context learning strategies tailored to different generation tasks.
D.2 Performance Variation by Output Length
Figure 6 reports performance on the Abstract Generation task when users are grouped by the average number of ground-truth tokens. Both PLUME and PLUME-s consistently outperform OPPU across most length bins. A consistent trend appears across the OPPU, PLUME and PLUME-s models: ROUGE-1 and METEOR peak in the Bin 2 (token length 126–140) and then decline as ground-truth length increases. This trend highlights a critical characteristic in abstract generation tasks: there exists an optimal output length range for maintaining high summarization quality. When the ground-truth length exceeds a certain threshold—approximately 140 tokens—both lexical overlap (ROUGE-1) and semantic similarity (METEOR) degrade significantly, even though the models gradually increase their output length. From Table 6, while the models generate progressively longer outputs with increasing bin size, they remain shorter and more conservative than the ground-truth abstracts. Longer ground-truth abstracts contain more entities and clause structure, so partial coverage and paraphrasing are penalized more strongly, revealing capacity limits of the per user LoRA adapters together with deterministic decoding. These results suggest using length aware decoding, explicit coverage objectives for salient terms and entities, and lightweight hierarchical planning to better match long ground-truth targets.
Appendix E Case Study
To qualitatively illustrate the generation result of our approach over baseline approaches, we present a few representative examples from abstract generation and scholarly title generation tasks.
Prompt Template
You are an academic researcher.
Your task is to generate an academic-style abstract that matches the author’s writing style based on the paper’s abstract.
Here are reference examples (for style/tone ONLY; do not copy them): <REFERENCE_ABSTRACT> is the abstract of <REFERENCE_TITLE>.
Generate a NEW abstract for the paper titled: <TARGET_TITLE>
Constraints:
•
DO NOT copy any sentences from the reference abstracts.
•
Length:150–250 words; formal, concise, and self-contained.
•
Suggested structure: background/objective method data/pipeline results/impact (optional) deployment/cloud aspects.
•
Write only the abstract text, without headings or extra commentary.
| Prompt | You are user #009, and a academic researcher. Your task is to generate an academic-style abstract that matches the author’s writing style based on the paper’s abstract. Here are reference examples (for style/tone ONLY; do not copy them): “The User Centric Smart Card Ownership Model (UCOM) provides an open and dynamic smart card environment enabling cardholders to request installation/deletion of an application to which they are entitled. As in this model, smart cards are not under the control of a centralised authority; hence, it is difficult for an application provider to ascertain their trustworthiness. At present, proposed secure channel protocols for the smart card environment do not provide adequate assurance required by the UCOM. In this paper, we explore the reasons behind their failure to meet the UCOM requirements and then propose a secure and trusted channel protocol that meets them. In addition, the proposed protocol is also suitable to GlobalPlatform’s consumer-centric smart cards. A comparison of the proposed protocol with existing smart card and selected Internet protocols is provided. Then we analyse the protocol with the CasperFDR tool. Finally, we detail the implementation and the performance measurement.” is the abstract of “A Secure and Trusted Channel Protocol for the User Centric Smart Card Ownership Model.” Generate a NEW abstract for the paper titled: “Coopetitive architecture to support a dynamic and scalable NFC based mobile services architecture.” Constraints: - DO NOT copy any sentences from the reference abstracts. - Length:150–250 words; formal, concise, and self-contained. - Suggested structure: background/objective method data/pipeline results/impact (optional) deployment/cloud aspects. Write only the abstract text, without headings or extra commentary. ### Response: |
| Version | Generated Abstract Output |