arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.04715v1 [cs.AI] 04 Sep 2026
\keepXColumns

PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

Xinyu Li, Hao Zhou, Jianfeng Zhu, Julina Maharjan, Ruixin Guo, Feodor Dragan, Ruoming Jin Department of Computer Science Kent State University Kent, OH 44242, USA {xli74,hzhou6,jzhu10,jmaharja,rguo5,fdragan,rjin1}@kent.edu
Abstract

Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users’ styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large user populations. We propose PLUME (Personalized Low-Rank Adaptation through User Modulation and Shared Subspace), a lightweight framework that achieves efficient and expressive per-user adaptation by leveraging a shared task-specific subspace. Specifically, PLUME first learns a global task subspace from aggregated user data. Personalization is then achieved by training only a lightweight small square matrix within this subspace, enabling each user to obtain a tailored model while keeping shared components fixed. Cross-layer shared parameters and rank-1 residual terms are further introduced to significantly reduce redundancy while maintaining expressiveness. Experiments on multiple personalized text generation benchmarks demonstrate that PLUME achieves comparable or superior performance to strong baselines, while reducing per-user parameters by over 95%. These results establish shared-subspace modulation with minimal residuals as a scalable and semantically grounded approach to LLM personalization. Our code is available at https://github.com/zhouhao0218/PLUME

1 Introduction

General-purpose large language models (LLMs) have achieved remarkable success across a wide range of natural language tasks. By pre-training on web-scale text corpora, modern LLMs such as GPT, LLaMA, Gemini, and their variants  Zhao et al. (2023) acquire strong general-purpose language understanding and generation capabilities, enabling impressive performance spanning QA dialogue, machine translation, and logical reasoning  Xu et al. (2024); Abbasiantaeb et al. (2024); Ferrag et al. (2025). However, these models are typically one-size-fits-all, treating all users the same and aiming to generally fit for tasks. This has spurred growing interest in personalized LLM, where an LLM’s outputs are tailored to an individual user’s style, preferences, or context towards an individual’s specific domain, task, and environment  Fan et al. (2024); Zhao et al. (2023). Indeed, user-level personalization is increasingly viewed as crucial in applications like web-based QA assistants Wu et al. (2025), education  Chu et al. (2025), and healthcare  Bajwa et al. (2021). Among these diverse applications, personalized text generation represents a critical frontier in LLM research. Users increasingly demand AI systems that reflect their individuality rather than merely produce what they intend to say.

In response to these demands, a substantial wave of research towards personalized LLMs has emerged. Broadly speaking, existing approaches can be divided into two paradigms. The first category, prompt-based approaches  Salemi et al. (2023); Liu et al. (2021); Wang et al. (2023); Kang et al. (2023); Qiu et al. (2025), attempts to achieve personalization through carefully constructed prompts that are built upon retrieved user history, profiles, or interaction logs. While such designs offer simplicity and interpretability, they heavily rely on explicit and high-quality user signals and are constrained by limited context length. The second category, fine-tuning-based approaches, directly modifies model parameters, hoping to encode personalized information in LLM itself. Majority of studies  Zhang et al. (2024a); Zhao et al. (2025) choose to fine-tune a universal model across all user data, hoping to implicitly capture individual preferences deeply mined in different individuals’ data via a single run; however, such models tend to blur user distinctions and underperform in fine-grained personalization.

More recently, a small but growing body of work has explored the per-user model paradigm Tan et al. (2024); Bu et al. (2025), in which each individual maintains a dedicated model or adapter. Despite its advantages in personalization peroformance, this paradigm suffers from storage overhead when scaling to thousands or millions of users. Moreover, given the limited personal writing data, these adapters are at high risk of overfitting. This leads to a central challenge for personalized language modeling in writing assistance: How to achieve user-level adaptability that captures stylistic nuance and expressiveness without compromising storage efficiency?

To bridge this gap, we propose PLUME (Personalized Low-rank Adaptation through User Modulation and Shared Subspace), a lightweight yet expressive framework for personalized LLM fine-tuning in writing tasks. PLUME first performs global task adaptation by training a non-personalized LoRA module on the aggregated data from all users. The resulting low-rank LoRA matrices capture task-specific knowledge and establish a compact shared latent subspace that serves as a global foundation for subsequent personalization. Building upon this shared representation, PLUME introduces an User-Conditioned Subspace Mixer (USM), which is inserted between the pre-trained LoRA factors to modulate the shared subspace according to each user’s unique preferences. To further reduce the parameter redundancy, PLUME incorporates a Personalized Cross-Layer Shared (PCLS) module shared across layers, thereby eliminating redundant per-layer parameterization within each personalized model. Finally, a lightweight rank-1 residual component (Resid) is attached to provide fine-grained layer-specific correction with negligible cost. Through the principled composition of these components, PLUME achieves highly expressive personalization while significanly reducing per-user trainable parameters. Our key contributions are as follows:

  • •

    New Problem formulation We introduce the problem such that for the personalized text generation task, how to efficiently compress per-user PEFT parameters while maintaining both personalization effectiveness and model expressiveness, which is critical for practical personalization training.

  • •

    PLUME Framework We show that classic full LoRA is redundant under personalization settings and propose PLUME, a novel and lightweight personalization framework that alleviates the inefficiency of relatively heavy per-user LoRA adapters. PLUME significantly reduces per-user parameter cost while preserving personalization performance and expressive capacity.

  • •

    Extensive Empirical Validation and Analysis: We conduct comprehensive experiments on five different tasks from LaMP and LongLaMP benchmarks, covering both short and long form content generation. Experiment results show that our framework could robustly reduce the per-user parameter space to less than 5% of standard LoRA while maintaining or even improving personalization performance.

Refer to caption
Figure 1: Overall System Architecture. PLUME consists of four components: (i) globally task-specific non-personalized LoRA modules (ii) User-Conditioned Subspace Mixer (USM): cu(l)c_{u}^{(l)} between global LoRA factors, (iii) Personalized Cross-Layer Shared Subspace (PCLS): (𝒜u,ℬu)(\mathcal{A}_{u},\mathcal{B}_{u}), and (iv) Layer-aware rank-1 residuals αu(l)​βu(l)\alpha_{u}^{(l)}\beta_{u}^{(l)}. Through the principled composition of these complementary modules, PLUME achieves personalized parameter compression while maintaining competitive or superior performance.

2 Problem and Preliminaries

2.1 Problem Formulation

We consider the task of personalized text generation, where the goal is to find a unique adapter for each individual to generate user-specific textual responses conditioned on both the current query and contextual information.

Data Construction. Let there be NN users denoted by {u1,u2,…,uN}\{u_{1},u_{2},\dots,u_{N}\}. Each user uiu_{i} has MiM_{i} pairs of history of query–response:

ℋi={(qi,1,ai,1),(qi,2,ai,2),…,(qi,Mi,ai,Mi)},\mathcal{H}_{i}=\{(q_{i,1},a_{i,1}),(q_{i,2},a_{i,2}),\dots,(q_{i,M_{i}},a_{i,M_{i}})\},

where qi,jq_{i,j} represents the jj-th query describing a specific task or instruction, and ai,ja_{i,j} denotes the corresponding ground-truth response. Each response is tokenized as:

ai,j=(ai,j(1),ai,j(2),…,ai,j(Ti,j)),a_{i,j}=(a_{i,j}^{(1)},a_{i,j}^{(2)},\dots,a_{i,j}^{(T_{i,j})}),

where ai,j(t)a_{i,j}^{(t)} denotes the tt-th token and Ti,jT_{i,j} its token-wise length.

To both emulate the real deployment regime and conform to the established best practice of few-shot in-context prompting, for each query qi,jq_{i,j}, we retrieve its top-kk similar historical queries 𝒯i,j=TopK​(Sim​(qi,j,qi,⋅))\mathcal{T}_{i,j}=\text{TopK}\big(\text{Sim}(q_{i,j},q_{i,\cdot})\big) from the same user and pair them with their corresponding responses to construct the personalized context:

𝒞⁡(qi,j)={(qi,t,ai,t)∣t∈𝒯i,j}\mathcal{C}(q_{i,j})=\{(q_{i,t},a_{i,t})\mid t\in\mathcal{T}_{i,j}\}

where Sim​(⋅)\text{Sim}(\cdot) could be any function that measures query-level similarity. Each instance is therefore represented as a triplet (qi,j,𝒞⁡(qi,j),ai,j)(q_{i,j},\mathcal{C}(q_{i,j}),a_{i,j}).

Training Objective. We adopt Decoder-Only LLMs for training where a pre-trained large language model parameterized by Θ\Theta serves as the shared backbone, and each user uiu_{i} is associated with a lightweight personalized adapter parametrized by θi\theta_{i}. The full model for user uiu_{i} is denoted as fΘ,θif_{\Theta,\theta_{i}}, which conditions generation on both the shared backbone and user-specific adaptation.

For user uiu_{i}, the model is trained to predict each token in the user’s response given all previous tokens, the current query, and the retrieved context:

ℒtrain(Θ,θi)=−∑j=1Mi∑t=1Ti,jlogPfΘ,θi​(ai,j(t)∣ai,j(<t),qi,j,𝒞⁡(qi,j),𝒜other),\begin{split}&\mathcal{L}_{\text{train}}(\Theta,\theta_{i})=-\sum_{j=1}^{M_{i}}\sum_{t=1}^{T_{i,j}}\\ \log&P_{f_{\Theta,\theta_{i}}}\big(a_{i,j}^{(t)}\mid a_{i,j}^{(<t)},q_{i,j},\mathcal{C}(q_{i,j}),\mathcal{A}_{\text{other}}\big),\end{split} (1)

where 𝒜other\mathcal{A}_{\text{other}} represents auxiliary information from other users (Note: We formulate a general LLM personalization task in Eq.1, where 𝒜other\mathcal{A}_{\text{other}} could be in any form, such as other user’s textual histories. In this work, we encode 𝒜other\mathcal{A}_{\text{other}} into global Space A(l)A^{(l)} and B(l)B^{(l)} as a learned task-specific representation shared by all users). During optimization, Θ\Theta is shared or fixed, while θi\theta_{i} is optimized to minimize the user-specific loss:

θi∗=argminθiℒtrain(Θ,θi),i=1,…,N.\theta_{i}^{*}=\arg\min_{\theta_{i}}\mathcal{L}_{\text{train}}(\Theta,\theta_{i}),\quad i=1,\dots,N. (2)

Inference. At test time, given a new query qitestq_{i}^{\text{test}} for user uiu_{i}, we retrieve its top-kk similar query–response pairs from ℋi\mathcal{H}_{i} to form 𝒞⁡(qitest)\mathcal{C}(q_{i}^{\text{test}}). The shared backbone combined with the learned personalized parameters is then used to generate the predicted response:

a^itest=fΘ,θi∗​(qitest,𝒞⁡(qitest),𝒜other),\hat{a}_{i}^{\text{test}}=f_{\Theta,\theta_{i}^{*}}\big(q_{i}^{\text{test}},\mathcal{C}(q_{i}^{\text{test}}),\mathcal{A}_{\text{other}}\big),

and the output a^itest\hat{a}_{i}^{\text{test}} is compared with the ground-truth aitesta_{i}^{\text{test}} using standard evaluation metrics.

Symbol Description
ii User index, i=1,…,Ni=1,\dots,N.
jj Query–response pair index for user uiu_{i}.
tt Token index within a response sequence.
qi,jq_{i,j} The jj-th query of user uiu_{i}.
ai,ja_{i,j} Ground-truth response corresponding to qi,jq_{i,j}.
ℋi\mathcal{H}_{i} Set of all historical (q,a)(q,a) pairs for user uiu_{i}.
𝒞⁡(qi,j)\mathcal{C}(q_{i,j}) Top-kk similar (q,a)(q,a) pairs retrieved from ℋi\mathcal{H}_{i}.
𝒜other\mathcal{A}_{\text{other}} Auxiliary information (optionally) derived from other users’ histories.
Θ\Theta Parameters of the shared base language model.
θi\theta_{i} Personalized parameters (adapter) for user uiu_{i}.
fΘ,θif_{\Theta,\theta_{i}} Personalized LLM combining shared backbone and user’s PEFT module.
(⋅)u(\cdot)_{u} Subscript indicating user-specific parameters.
(⋅)(l)(\cdot)^{(l)} Superscript indicating the ll-th layer in the model.
Table 1: Notation Summary.

2.2 One-PEFT-per-User Personalization Paradigm

(a)
(b)
Figure 2: (a) OPPU requires a high rank to maintain performance. (b) PLUME achieves comparable results even at low ranks by using layer-wise shared parameters and lightweight user-specific modules.

We follow the One PEFT per User (OPPU Tan et al. (2024)) training framework, which assigns an independent PEFT module to each user for personalized fine-tuning.

Concretely, OPPU first ignores user differences and trains a task-specific global LoRA adapter using the union of all users’ training data ⋃iℋi\bigcup_{i}\mathcal{H}_{i}, resulting in a shared adapter Δ​Θ\Delta\Theta. This stage corresponds to optimizing

Δ​Θ∗=arg⁡min⁡∑i=1NΔ​Θ⁡ℒtrain​(Θ,Δ​Θ,ℋi),\Delta\Theta^{*}=\arg\min_{\Delta\Theta}\sum_{i=1}^{N}\mathcal{L}_{\text{train}}(\Theta,\Delta\Theta;\mathcal{H}_{i}), (3)

The Second step simply follows Eq  2, but with merged model Θ←Θ+Δ​Θ∗\Theta\leftarrow\Theta+\Delta\Theta^{*}.

If we further examine this paradigm through the lens of a specific target module, e.g., qprojq_{\text{proj}}, the resulting parameter composition becomes more explicit. Typically, the effective weight matrix on a layer can be expressed as:

Wu(l)=W0(l)+s​B(l)​A(l)+s′​Bu(l)​Au(l),W_{u}^{(l)}=W_{0}^{(l)}+s\,B^{(l)}A^{(l)}+s^{\prime}\,B_{u}^{(l)}A_{u}^{(l)}, (4)

where W0(l)W_{0}^{(l)} denotes the shared base model weight, (A(l),B(l))(A^{(l)},B^{(l)}) represents the global LoRA adapter trained from the aggregated non-personalized data ⋃iℋi\bigcup_{i}\mathcal{H}_{i}, and (Au(l),Bu(l))(A^{(l)}_{u},B^{(l)}_{u}) denotes the user-specific LoRA adapter paramterized from θi\theta_{i}.

Vanilla OPPU provides a straightforward way to incorporate user-specific knowledge, nevertheless, it faces two major limitations. First, maintaining a separate LoRA module for every user leads to prohibitively high storage and memory costs, especially when the number of users is large. Second, since each user’s training data can vary drastically in scale and quality, independently optimizing a full-rank LoRA for every user often causes overfitting or unstable personalization performance. As illustrated in Figure 2(a), the performance of OPPU drops significantly when the individual LoRA rank is reduced.

These issues highlight the need for a more parameter-efficient and stable approach to capture user-specific preferences without requiring a full LoRA module per user.

2.3 Notation Summary

Table 1 summarizes the key symbols used in this paper, covering both the preceding problem formulation and the following method sections.

3 Method

In this section, we propose PLUME, a novel framework designed to reduce redundant personalized parameters while preserving personalization expressivity. Figure 1 illustrates the overall architecture of PLUME. We will start from our key observation, comparison with vanilla OPPU and progressively builds our efficient yet effective model.

3.1 PLUME Framework

Revisiting Task Space. From a linear algebraic perspective, the task-specific LoRA update on layer ll can be expressed as Δ​W(l)=s​B(l)​A(l)\Delta W^{(l)}=s\,B^{(l)}A^{(l)}, where A(l)A^{(l)} projects inputs into a low-dimensional task subspace, and B(l)B^{(l)} expands the projected representation back to the model’s hidden space. Hence, for any input xx, the LoRA output Δ​W(l)​x=s​B(l)​A(l)​x\Delta W^{(l)}x=s\,B^{(l)}A^{(l)}x always resides in the column space col⁡(B(l))\mathrm{col}(B^{(l)}). OPPU extends this by assigning each user an individual adapter (Au(l),Bu(l))(A_{u}^{(l)},B_{u}^{(l)}), which effectively introduces a new output subspace basis col⁡(Bu(l))\mathrm{col}(B^{(l)}_{u}) specific to each user. From this view, OPPU increases the expressive capacity of the model by expanding the dimensional coverage of col⁡(B(l))\mathrm{col}(B^{(l)}) across users, thereby enhancing personalization.

However, this completely “separates” the shared task space col⁡(B(l))\mathrm{col}(B^{(l)}) and the augmented individual space col⁡(Bu(l))\mathrm{col}(B^{(l)}_{u}), leaving the shared adapter only learns how to project tokens into the task subspace, while not being able to control how different users behave within the collaboratively trained existing shared task subspace col⁡(B(l))\mathrm{col}(B^{(l)}). This motivates us to allow user-dependent modulation of the shared subspace.

User-Conditioned Subspace Mixer (cuc_{u}). Intuitively, since users may differ in style, intent, or preference, the way they combine or emphasize task-subspace directions should vary. To capture user-specific task space utilization, we insert a lightweight matrix cu(l)∈ℝrg×rgc_{u}^{(l)}\in\mathbb{R}^{r_{g}\times r_{g}} between A(l)A^{(l)} and B(l)B^{(l)}:

Δ​Wu(l)=s​B(l)​cu(l)​A(l).\Delta W_{u}^{(l)}=s\,B^{(l)}c_{u}^{(l)}A^{(l)}. (5)

We term cu(l)c_{u}^{(l)} a User-Conditioned Subspace Mixer (USM). It can be interpreted as a re-indexing or re-weighting operator that re-combines latent task directions inside the shared LoRA subspace. Thus, cu(l)c_{u}^{(l)} enables personalized modulation of how a user exploits the common task subspace col⁡(B(l))\mathrm{col}(B^{(l)}), offering additional expressivity with negligible parameter cost.

PLUME. Before we move on to a more compact model, we first assess how much expressivity the USM alone provides, we introduce a lightweight per-layer residual term aiming to approximate the representational power of high-rank OPPU:

Wu(l)=W0(l)+s​B(l)​cu(l)​A(l)+s′​bu(l)​au(l).W_{u}^{(l)}=W_{0}^{(l)}+s\,B^{(l)}c_{u}^{(l)}A^{(l)}+s^{\prime}\,b_{u}^{(l)}a_{u}^{(l)}. (6)

Here (au(l),bu(l))(a_{u}^{(l)},b_{u}^{(l)}) follow the same formulation as (Au(l),Bu(l))(A_{u}^{(l)},B_{u}^{(l)}), but with a reduced individual rank rindr_{\text{ind}}. Empirically, we find that when rind≈4r_{\text{ind}}\!\approx\!4, PLUME already achieves performance comparable to OPPU with rank 64, as shown in Fig 2(a). Interestingly, further increasing rindr_{\text{ind}} provides no additional gains—indicating that once the shared task subspace is efficiently utilized via cu(l)c_{u}^{(l)}, a large individual subspace becomes redundant and inefficient, leading to overfitting. This observation suggests that there may still exist parameter redundancy even within the reduced adapters. Thus, we further push PLUME to the extreme—seeking the most compact yet expressive form of individual representation. As illustrated in Fig 2(b), our proposed novel shared subspace mechanism enables PLUME to reduce the individual parameter while maintaining highly compatible performance.

Refer to caption
Figure 3: CKA analysis of a user’s LoRA parameters. (a) Pairwise CKA across layers; (b) Average CKA similarity per layer. Strong inter-layer similarity reveals significant redundancy, supporting PLUME’s shared-layer design.

3.2 Parameter Reduction via Shared Subspace

CKA Analysis for Layer-Wise Representation Similarity for Personalized Model Inspired by prior work He et al. (2025); Kopiczko et al. (2023); Zhou et al. () that sharing weights across layers can enhance model expressiveness while reducing parameter usage, we investigate whether different layers in personalized LoRA models actually learn similar user-specific representations. To this end, we perform a Centered Kernel Alignment (CKA), a widely used approach for comparing neural representations across layers and models  Kornblith et al. (2019); Liu et al. (2025b) to quantify layer-wise representational similarity across the personalized adapters (See Appendix A.3 for more about CKA).

As shown in Figure 3, OPPU-trained adapters exhibit remarkably high inter-layer CKA values (often exceeding 0.90.9), indicating that the personalized LoRA subspaces learned by different layers are strongly aligned. This suggests that the layer-wise residuals {au(l),bu(l)}\{a_{u}^{(l)},b_{u}^{(l)}\} tend to encode similar directions of user-specific variation, leading to redundant parameterization across depth. Motivated by this finding, we hypothesize that a shared subspace across layers could capture the dominant personalized factors more efficiently.

PLUME-s. Building on this insight, we observe that the per-layer residuals (au(l),bu(l))(a_{u}^{(l)},b_{u}^{(l)}) often act similarly to layer-specific biases that adjust the shared projection directions. Accordingly, we decompose (au(l),bu(l))(a_{u}^{(l)},b_{u}^{(l)}) into two functional components: (a) a Personalized Cross-Layer Shared Subspace (PCLS), (𝒜u,ℬu)(\mathcal{A}_{u},\mathcal{B}_{u})) capturing cross-layer user-specific residuals in a low-rank shared form, and (b) a rank-1 layer-wise Residual (Resid), (αu(l),βu(l))(\alpha_{u}^{(l)},\beta_{u}^{(l)}) that enables fine-grained local adaptation. The resulting formulation becomes:

Wu(l)=W0(l)+s​B(l)​cu(l)​A(l)+s′​ℬu​𝒜u+βu(l)​αu(l).W_{u}^{(l)}=W_{0}^{(l)}+s\,B^{(l)}c_{u}^{(l)}A^{(l)}+s^{\prime}\,\mathcal{B}_{u}\mathcal{A}_{u}+\beta_{u}^{(l)}\alpha_{u}^{(l)}. (7)

Here (𝒜u,ℬu)(\mathcal{A}_{u},\mathcal{B}_{u}) with rank rshr_{\textbf{sh}} parameterize the Layer-Shared Personalized Subspace reused across all layers, while (αu(l),βu(l))(\alpha_{u}^{(l)},\beta_{u}^{(l)}) serve as the rank-1 layer refinements for fine adjustment. Extensive experiments demonstrate that leveraging a shared residual subspace allows substantial reduction of individual residual ranks without compromising performance, and can even outperform OPPU in certain settings (see Sec. 4).

Summary and Parameter Efficiency. Let LL be the number of layers and rgr_{g} the global LoRA rank, and dd the embedding vector dimension. In conclusion, PLUME adaptively unifies four complementary components: a global task-specific adapter (A(l),B(l))(A^{(l)},B^{(l)}) shared by all users, and three lightweight personalized modules. (1) The User-Conditioned Subspace Mixer (USM) cu(l)c_{u}^{(l)} adaptively reuses the global task space with O⁡(L​rg2)O(Lr_{g}^{2}) parameters. (2) The Personalized Cross-Layer-Shared Personalized Subspace (PCLS) (𝒜u,ℬu)(\mathcal{A}_{u},\mathcal{B}_{u}) captures cross-layer personalization compactly with O⁡(rsh​d)O(r_{\text{sh}}d) parameters. (3) The rank-1 personalized residuals (αu(l),βu(l))(\alpha_{u}^{(l)},\beta_{u}^{(l)}) restore fine local flexibility using O⁡(L​d)O(Ld) parameters. In total, the per-user complexity is O⁡(L​rg2+rsh​d+L​d),O(Lr_{g}^{2}+r_{\text{sh}}d+Ld), compared to O⁡(L​d​roppu)O(Ldr_{\text{oppu}}) for OPPU’s layer-wise independent adapters. When rsh≪roppur_{\text{sh}}\!\ll\!r_{\text{oppu}}, the parameter reduction becomes particularly significant, see Table 3 for practical analysis.

4 Experiments

Type Method Long Content Generation Short Content Generation
Abstract Generation Product Review Topic Writing News Headline Scholarly Title
R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR
Non-Personalized BASE 0.3497 0.1688 0.2408 0.3397 0.1389 0.2335 0.2892 0.1238 0.2050 0.1313 0.1158 0.0962 0.3905 0.3165 0.4084
LoRA 0.3491 0.2036 0.2456 0.3877 0.2323 0.2707 0.2655 0.1378 0.1732 0.2188 0.2010 0.1939 0.4648 0.4138 0.4228
PiSSA 0.3523 0.1995 0.2536 0.3964 0.2326 0.2826 0.2854 0.1393 0.1927 0.2230 0.2056 0.1999 0.4801 0.4266 0.4351
AdaLoRA 0.3503 0.2089 0.2395 0.3419 0.2056 0.2323 0.2279 0.1212 0.1530 0.2299 0.2113 0.1892 0.4621 0.4203 0.3940
QLoRA 0.3450 0.2001 0.2394 0.3940 0.2339 0.2769 0.2697 0.1399 0.1762 0.2224 0.2037 0.1958 0.4628 0.4101 0.4196
LoRA-One 0.3780 0.2317 0.2750 0.3800 0.2320 0.2680 0.2930 0.1410 0.2110 – – – 0.4480 0.4030 0.4070
Personalized RAG 0.3512 0.1743 0.2567 0.3389 0.1454 0.2336 0.2936 0.1264 0.2199 0.1470 0.1304 0.1047 0.3982 0.3229 0.4050
OPPU 0.4147 0.2374 0.2824 0.4416 0.2481 0.3165 0.3191 0.1510 0.2126 0.2383 0.2178 0.2008 0.5146 0.4510 0.4365
CoPE 0.3779 0.2247 0.2530 0.3600 0.2385 0.2645 0.2303 0.1386 0.1754 0.2324 0.2101 0.1963 0.4741 0.4076 0.4133
PLUME 0.4168 0.2395 0.2839 0.4423 0.2460 0.3161 0.3217 0.1534 0.2172 0.2150 0.1974 0.1778 0.5075 0.4459 0.4246
PLUME-s 0.4117 0.2322 0.2803 0.4369 0.2434 0.3107 0.3176 0.1517 0.2131 0.2427 0.2219 0.2039 0.5177 0.4537 0.4359
Table 2: Performance comparison on Mistral-7B across five personalized text generation tasks. Bold numbers indicate the best results within each task. PLUME refers to Eq 6, and PLUME-s refers to Eq 7. Results of LoRA-One on the News Headline task are abnormally low and thus excluded from reporting.
Method #Params % of OPPU Memory (GB) Runtime
OPPU 167,772,160 100% 0.6250 4h53m
CoPE 167,772,160 100% 0.1016 4h56m
PLUME 11,403,264 6.80% 0.0425 4h30m
PLUME-s 3,702,784 3.11% 0.0156 4h42m
Table 3: Per-user parameter cost and running time comparison on Mistral-7B backbone. PLUME achieves substantial parameter savings while maintaining personalization capacity.

In this section, we conduct experiments to systematically investigate the following research questions:
RQ1: How does PLUME perform against non-personalized and personalized PEFT benchmarks in various personalized text generation tasks?
RQ2: How much efficiency gain does PLUME achieve?
RQ3: How do the different components in PLUME affect its performance effectiveness?
RQ4: How does PLUME behave under varying degrees of parameter reduction, and what is the resulting trade-off between model compactness and personalization performance?

4.1 Experimental Setup.

Datasets We adopt the widely used benchmarks LongLaMP Kumar et al. (2024) and LaMP Salemi et al. (2023) to evaluate our model’s ability across both short-form and long-form content personalized content generation tasks. Specifically, we select 3 datasets: Abstract Generation, Product Review, and Topic Writing from LongLaMP, and 2 datasets: News Headline generation, Scholarly Title generation from LaMP. We leave details on data preprocessing and data statistics to Appendix B

Baselines. We compare our methods against a series of strong parameter-efficient fine-tuning baselines under both non-personalized training and personalized training settings. For the non-personalized setting, we compare against the base model, and LoRA Hu et al. (2021), AdaLoRA Zhang et al. (2023), PiSSA Meng et al. (2024), and QLoRA Dettmers et al. (2023) and LoRA-One Zhang et al. (2025b) which all adapt model parameters through a shared low-rank adapter across users. For the personalized setting, we consider retrival-based approach such as RAG Tan et al. (2024), training-based baseline OPPU Tan et al. (2024), and a more recent work CoPE Bu et al. (2025). Finally, we evaluate our two variants: PLUME and PLUME-s.

Implementation To ensure fairness, all models are evaluated with a consistent LoRA configuration (rank = 64) across both non-personalized and personalized settings, including the OPPU and CoPE baselines. To demonstrate the generalizability and robustness of our approach, we evaluate PLUME on two distinct open-source large language models: LLaMA2-7B Touvron et al. (2023) and Mistral-7B-Instruct-v0.2 Jiang et al. (2023). As illustrated in Figure 4(b) in the later discussion, PLUME already reaches competitive results with a very small rank, so we report PLUME results with residual rank 44, and PLUME-s with residual rank 11 and shared rank as low as 88. Due to space constraints, we present the results using the Mistral-7B backbone in the main text and leave more experiment results and details in Appendix C.

Evaluation Metrics In accordance with established practices in prior work Tan et al. (2024); Kumar et al. (2024), we employ a standard set of automatic metrics ROUGE-1, ROUGE-L Lin (2004), and METEOR Banerjee and Lavie (2005) to quantitatively assess the lexical overlap, fluency, and semantic correspondence between generated outputs and reference responses (See more details in Appendix C.1)

4.2 Main Results

Overall Perfomance (RQ1) Table 2 reports the overall performance across five generation tasks. Overall, personalized methods outperform non-personalized ones across all metrics, confirming the advantage of user-specific adaptation. Within the personalized group, despite using far fewer parameters than strong baselines such as OPPU and CoPE, the PLUME series delivers comparable or better results than other baseline methods. For instance, PLUME attains the highest ROUGE-L on Abstract Generation (0.2395) and comparable performance on other long-text tasks, demonstrating that the personalized module USM would substantially enhances expressiveness. When incorporating the shared subspace, PLUME-s performance slightly improves, e.g., ROUGE-1 rises from 0.2150 to 0.2427 on News Headline Generation Tasks. This indicates that the shared subspace not only compresses parameters but also mitigates redundancy and overfitting issues observed in OPPU.

In summary, PLUME achieves a favorable balance between expressiveness and efficiency. It retains nearly identical generation quality to strong personalized baselines OPPU and CoPE while using only a fraction of their parameters, and even surpasses them on several short-text tasks.

Efficiency Comparison (RQ2) One of our main contributions is that PLUM substantially improves parameter efficiency for LLM personalization adapters. Here, we quantify the extent to which this compression frontier can be pushed under practical setting. As in Table 3, compared to OPPU, which requires 168M parameters per user, PLUME and PLUME-s reduce the cost to only 6.8% and 3.1% respectively, achieving over 15×15\times–30×30\times compression while preserving personalization quality and requiring less training time.

Ablation Study: Effectiveness of Components (RQ3)

(a)
(b)
(c)
Figure 4: (a) Component effectiveness. (b) Sensitivity on Resid Rank. (c) Sensitivity on PCLS Rank.

Figure 4(a) compares the effects of three core modules: USM, PCLS, and Resid on both long and short text generation tasks.

Across both tasks, removing any single module consistently leads to a drop in ROUGE-L and METEOR, suggesting that all three work jointly to enhance expressiveness, coherence, and efficiency. This consistent trend highlights that the modules are complementary rather than redundant. The relative impact of each module remains consistent across both tasks (Resid > PCLS > USM). Moreover, the effect of the Resid module is more pronounced in long-form generation tasks. Notably, even though Resid employs only a rank-1 update, it still contributes sufficiently to representational richness, showing that lightweight residual paths can complement USM’s user-specific modulation effectively. These results demonstrate that each component is indispensable for achieving a strong balance between personalization, coherence, and parameter efficiency.

Parameter Efficiency VS Model Performance (RQ4)
Sensitivity of residual rank. To investigate the sensitivity of residual dimensionality, we vary the rank of the residual module. As shown in Figure 4(b), performance on Abstract Generation (left) first increases with larger residual ranks and peaks around rank 4, after which both ROUGE-L and METEOR scores gradually decline. This trend indicates that a small residual rank is sufficient to capture user-specific variations, while larger ranks introduce redundancy and overfitting. In contrast, the Scholarly Title task (right) shows much smaller variance across ranks, with performance already strong at rank 1 and exhibiting a slight downward trend thereafter. This validates our conjecture that short-form generation requires fewer personalized parameters and benefits more from PLUME’s shared-subspace regularization.

Sensitivity on PCLS rank. To further examine the influence of the shared subspace capacity, we vary the rank of PCLS module. As illustrated in Figure 4(c), the results show an opposite trend between the two settings. For Abstract Generation, performance steadily improves as the shared rank increases, indicating that long-form generation benefits from a richer shared subspace capable of modeling broader contextual dependencies and semantic consistency across layers. In contrast, the Scholarly Title task exhibits an inverted trend—performance peaks at a small shared rank and then declines as the rank grows—suggesting that short-form generation requires only limited shared capacity. These contrasting patterns highlight PLUME’s flexibility that its shared subspace can effectively balance generalization and personalization.

5 Conclusion

This paper investigates how to achieve fine-grained user adaptation with better parameter efficiency in the personalized text generation task. To tackle this, we proposed PLUME, a novel framework for personalized LLMs through low-rank user modulation and shared subspaces. The proposed design comprises of the User-Conditioned Subspace Mixer (USM), the cross-layer shared personalized subspace (PCLS), and the rank-1 residuals (Resid) and demonstrates that rich personalization can be achieved with only a fraction of the parameters required by conventional per-user adapters such as OPPU. Comprehensive experiments across five personalized text-generation benchmarks with two different foundation models as backbone show that PLUME consistently matches or exceeds existing personalized LoRA variants while reducing per-user parameters by over 95%.

Overall, PLUME provides a more effective and highly compact user representation for personalized LLMs, and we believe it establishes a new research pathway for advancing LLM personalization.

References

  • Abbasiantaeb et al. (2024) Z. Abbasiantaeb, Y. Yuan, E. Kanoulas, and M. Aliannejadi Let the llms talk: simulating human-to-human conversational qa via zero-shot llm-to-llm interactions. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pp. 8–17. Cited by: §1.
  • Bajwa et al. (2021) J. Bajwa, U. Munir, A. Nori, and B. Williams Artificial intelligence in healthcare: transforming the practice of medicine. Future healthcare journal 8 (2), pp. e188–e194. Cited by: §1.
  • Banerjee and Lavie (2005) S. Banerjee and A. Lavie METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72. Cited by: §C.1, §4.1.
  • Bu et al. (2025) H. Bu, C. Jung, M. Kang, and J. Kim Personalized llm decoding via contrasting personal preference. arXiv preprint arXiv:2506.12109. Cited by: §A.2, Appendix B, §1, §4.1.
  • Chu et al. (2025) Z. Chu, S. Wang, J. Xie, T. Zhu, Y. Yan, J. Ye, A. Zhong, X. Hu, J. Liang, P. S. Yu, et al. Llm agents for education: advances and applications. arXiv preprint arXiv:2503.11733. Cited by: §1.
  • Dai et al. (2023) S. Dai, N. Shao, H. Zhao, W. Yu, Z. Si, C. Xu, Z. Sun, X. Zhang, and J. Xu Uncovering chatgpt’s capabilities in recommender systems. In Proceedings of the 17th ACM Conference on Recommender Systems, pp. 1126–1132. Cited by: §A.2.
  • Dettmers et al. (2023) T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer Qlora: efficient finetuning of quantized llms. Advances in neural information processing systems 36, pp. 10088–10115. Cited by: §A.1, §4.1.
  • Fan et al. (2024) W. Fan, Y. Ding, L. Ning, S. Wang, H. Li, D. Yin, T. Chua, and Q. Li A survey on rag meeting llms: towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pp. 6491–6501. Cited by: §1.
  • Ferrag et al. (2025) M. A. Ferrag, N. Tihanyi, and M. Debbah From llm reasoning to autonomous ai agents: a comprehensive review. arXiv preprint arXiv:2504.19678. Cited by: §1.
  • He et al. (2025) Z. He, Z. Tu, X. Wang, X. Chen, Z. Wang, J. Xu, T. Liang, W. Jiao, Z. Zhang, and R. Wang RaSA: rank-sharing low-rank adaptation. arXiv preprint arXiv:2503.12576. Cited by: §3.2.
  • Hu et al. (2021) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen Lora: low-rank adaptation of large language models. arxiv 2021. arXiv preprint arXiv:2106.09685 10. Cited by: §A.1, §4.1.
  • Izacard et al. (2021) G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118. Cited by: Appendix B.
  • Jiang et al. (2023) A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed Mistral 7b. External Links: 2310.06825 Cited by: §4.1.
  • Kang et al. (2023) W. Kang, J. Ni, N. Mehta, M. Sathiamoorthy, L. Hong, E. Chi, and D. Z. Cheng Do llms understand user preferences? evaluating llms on user rating prediction. arXiv preprint arXiv:2305.06474. Cited by: §A.2, §1.
  • Kong et al. (2024) X. Kong, J. Wu, A. Zhang, L. Sheng, H. Lin, X. Wang, and X. He Customizing language models with instance-wise lora for sequential recommendation. Advances in Neural Information Processing Systems 37, pp. 113072–113095. Cited by: §A.2.
  • Kopiczko et al. (2023) D. J. Kopiczko, T. Blankevoort, and Y. M. Asano Vera: vector-based random matrix adaptation. arXiv preprint arXiv:2310.11454. Cited by: §3.2.
  • Kornblith et al. (2019) S. Kornblith, M. Norouzi, H. Lee, and G. Hinton Similarity of neural network representations revisited. In International conference on machine learning, pp. 3519–3529. Cited by: §3.2.
  • Kumar et al. (2024) I. Kumar, S. Viswanathan, S. Yerra, A. Salemi, R. A. Rossi, F. Dernoncourt, H. Deilamsalehy, X. Chen, R. Zhang, S. Agarwal, et al. Longlamp: a benchmark for personalized long-form text generation. arXiv preprint arXiv:2407.11016. Cited by: §4.1, §4.1.
  • Li et al. (2024) C. Li, M. Zhang, Q. Mei, W. Kong, and M. Bendersky Learning to rewrite prompts for personalized text generation. In Proceedings of the ACM Web Conference 2024, pp. 3367–3378. Cited by: §A.2.
  • Lin (2004) C. Lin Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81. Cited by: §C.1, §C.1, §4.1.
  • Liu et al. (2025a) J. Liu, Z. Qiu, Z. Li, Q. Dai, J. Zhu, M. Hu, M. Yang, and I. King A survey of personalized large language models: progress and future directions. arXiv preprint arXiv:2502.11528. Cited by: §A.2.
  • Liu et al. (2024) S. Liu, C. Wang, H. Yin, P. Molchanov, Y. F. Wang, K. Cheng, and M. Chen Dora: weight-decomposed low-rank adaptation. In Forty-first International Conference on Machine Learning, Cited by: §A.1.
  • Liu et al. (2025b) X. Liu, L. Hsiung, Y. Yang, and Y. Yan Spectral insights into data-oblivious critical layers in large language models. arXiv preprint arXiv:2506.00382. Cited by: §A.3, §3.2.
  • Liu et al. (2021) Y. Liu, W. Lu, S. Cheng, D. Shi, S. Wang, Z. Cheng, and D. Yin Pre-trained language model for web-scale retrieval in baidu search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 3365–3375. Cited by: §A.2, §1.
  • Meng et al. (2024) F. Meng, Z. Wang, and M. Zhang Pissa: principal singular values and singular vectors adaptation of large language models. Advances in Neural Information Processing Systems 37, pp. 121038–121072. Cited by: §A.1, §4.1.
  • Mysore et al. (2023) S. Mysore, Z. Lu, M. Wan, L. Yang, S. Menezes, T. Baghaee, E. B. Gonzalez, J. Neville, and T. Safavi Pearl: personalizing large language model writing assistants with generation-calibrated retrievers. arXiv preprint arXiv:2311.09180. Cited by: §A.2.
  • Qian et al. (2025) H. Qian, Z. Liu, P. Zhang, K. Mao, D. Lian, Z. Dou, and T. Huang Memorag: boosting long context processing with global memory-enhanced retrieval augmentation. In Proceedings of the ACM on Web Conference 2025, pp. 2366–2377. Cited by: §A.2.
  • Qiu et al. (2025) Y. Qiu, X. Zhao, Y. Zhang, Y. Bai, W. Wang, H. Cheng, F. Feng, and T. Chua Measuring what makes you unique: difference-aware user modeling for enhancing llm personalization. arXiv preprint arXiv:2503.02450. Cited by: §A.2, §1.
  • Salemi et al. (2023) A. Salemi, S. Mysore, M. Bendersky, and H. Zamani Lamp: when large language models meet personalization. arXiv preprint arXiv:2304.11406. Cited by: §A.2, §A.2, §1, §4.1.
  • Sun et al. (2024) C. Sun, K. Yang, R. G. Reddy, Y. R. Fung, H. P. Chan, K. Small, C. Zhai, and H. Ji Persona-db: efficient large language model personalization for response prediction with collaborative data refinement. arXiv preprint arXiv:2402.11060. Cited by: §A.2.
  • Tan et al. (2025) Z. Tan, Z. Li, T. Liu, H. Wang, H. Yun, M. Zeng, P. Chen, Z. Zhang, Y. Gao, R. Wang, et al. Aligning large language models with implicit preferences from user-generated content. arXiv preprint arXiv:2506.04463. Cited by: §A.2.
  • Tan et al. (2024) Z. Tan, Q. Zeng, Y. Tian, Z. Liu, B. Yin, and M. Jiang Democratizing large language models via personalized parameter-efficient fine-tuning. arXiv preprint arXiv:2402.04401. Cited by: §A.2, Appendix B, §1, §2.2, §4.1, §4.1.
  • Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §4.1.
  • Wang et al. (2023) D. Wang, K. Yang, H. Zhu, X. Yang, A. Cohen, L. Li, and Y. Tian Learning personalized alignment for evaluating open-ended text generation. arXiv preprint arXiv:2310.03304. Cited by: §A.2, §1.
  • Wang et al. (2024a) F. Wang, J. Jiang, C. Park, S. Kim, and J. Tang KaSA: knowledge-aware singular-value adaptation of large language models. arXiv preprint arXiv:2412.06071. Cited by: §A.1.
  • Wang et al. (2024b) H. Wang, Y. Li, S. Wang, G. Chen, and Y. Chen Milora: harnessing minor singular components for parameter-efficient llm finetuning. arXiv preprint arXiv:2406.09044. Cited by: §A.1.
  • Woźniak et al. (2024) S. Woźniak, B. Koptyra, A. Janz, P. Kazienko, and J. Kocoń Personalized large language models. In 2024 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 511–520. Cited by: §A.2.
  • Wu et al. (2025) J. Wu, W. Yin, Y. Jiang, Z. Wang, Z. Xi, R. Fang, L. Zhang, Y. He, D. Zhou, P. Xie, et al. Webwalker: benchmarking llms in web traversal. arXiv preprint arXiv:2501.07572. Cited by: §1.
  • Xu et al. (2024) H. Xu, A. Sharaf, Y. Chen, W. Tan, L. Shen, B. Van Durme, K. Murray, and Y. J. Kim Contrastive preference optimization: pushing the boundaries of llm performance in machine translation. arXiv preprint arXiv:2401.08417. Cited by: §1.
  • Zhang et al. (2024a) K. Zhang, Y. Kim, and X. Liu Personalized llm response generation with parameterized memory injection. arXiv preprint arXiv:2404.03565. Cited by: §A.2, §1.
  • Zhang et al. (2023) Q. Zhang, M. Chen, A. Bukharin, N. Karampatziakis, P. He, Y. Cheng, W. Chen, and T. Zhao Adalora: adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512. Cited by: §A.1, §4.1.
  • Zhang et al. (2025a) X. F. Zhang, N. Beauchamp, and L. Wang PRIME: large language model personalization with cognitive memory and thought processes. arXiv preprint arXiv:2507.04607. Cited by: §A.2.
  • Zhang et al. (2024b) Y. Zhang, J. Wang, L. Yu, D. Xu, and X. Zhang Personalized lora for human-centered text understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 19588–19596. Cited by: §A.2.
  • Zhang et al. (2025b) Y. Zhang, F. Liu, and Y. Chen LoRA-one: one-step full gradient could suffice for fine-tuning large language models, provably and efficiently. arXiv preprint arXiv:2502.01235. Cited by: §4.1.
  • Zhao et al. (2023) W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223 1 (2). Cited by: §1.
  • Zhao et al. (2025) X. Zhao, J. You, Y. Zhang, W. Wang, H. Cheng, F. Feng, S. Ng, and T. Chua NextQuill: causal preference modeling for enhancing llm personalization. arXiv preprint arXiv:2506.02368. Cited by: §1.
  • [47] Y. Zhou, R. Li, C. Zhou, F. Yang, and A. Pan BSLoRA: enhancing the parameter efficiency of lora with intra-layer and inter-layer sharing. In Forty-second International Conference on Machine Learning, Cited by: §3.2.
  • Zhu et al. (2024) J. Zhu, J. Lin, X. Dai, B. Chen, R. Shan, J. Zhu, R. Tang, Y. Yu, and W. Zhang Lifelong personalized low-rank adaptation of large language models for recommendation. arXiv preprint arXiv:2408.03533. Cited by: §A.2.

Appendix A Related Work

A.1 LoRA Based Fine Tuning

LoRA (Low-Rank Adaptation) (Hu et al., 2021) is one of the most widely adopted parameter-efficient fine-tuning (PEFT) techniques for adapting large language models (LLMs) with minimal trainable parameters. Instead of updating the full weight matrix, LoRA introduces a low-rank update Δ​W=s​A​B\Delta W=sAB, where A∈ℝd×rA\in\mathbb{R}^{d\times r} and B∈ℝr×kB\in\mathbb{R}^{r\times k} with r≪min⁡(d,k)r\ll\min({d,k}), thus reducing the number of trainable parameters from O⁡(d​k)O(dk) to O⁡(r⁡(d+k))O(r(d+k)). Due to its effectiveness and efficiency, LoRA has become one of the most effective and practical approaches in personalized LLM adaptation tasks. Building upon this foundation, several variants have been developed to improve flexibility, stability, and efficiency. AdaLoRA (Zhang et al., 2023) dynamically allocates ranks through orthogonal regularization, DoRA (Liu et al., 2024) decouples direction and magnitude learning to stabilize optimization, and QLoRA (Dettmers et al., 2023) integrates quantization for memory-efficient adaptation of large-scale models. A complementary line of work leverages singular value decomposition (SVD) for structured adaptation: PiSSA (Meng et al., 2024) initializes LoRA adapters with principal singular components for faster convergence, while MiLoRA (Wang et al., 2024b) and KASA (Wang et al., 2024a) focus on minor singular directions to enhance generalization. Collectively, when personalization requires fine-tuning, LoRA and its variants serve as the most effective and widely adopted PEFT solutions, enabling scalable adaptation of LLMs to individual users while preserving efficiency and generalization.

A.2 LLM personalization

The problem of adapting large language models to individual users has been extensively studied, yielding a diverse landscape of personalization strategies. Typically, these paradigms could be classified into two categories: Prompt-based and Adapter-based approaches.

Prompt-Based Approaches Among the earliest and most lightweight approaches are prompt-based methods, which target to extract and encode user-specific information into handcrafted or learned prompts to guide model behavior without modifying LLM parameters Liu et al. (2025a); Liu et al. (2021). In light of this, various prompting techniques, such as CoT and in-context learning, have been employed to provide a summarized or sampled user behavior history  Wang et al. (2023); Kang et al. (2023). For instance, Dai et al Dai et al. (2023) leverages sampled user purchase history to guide LLMs to make limited-size personalized item recommendations, while DPL Qiu et al. (2025) improves personalization by encoding extracted inter-user comparisons into personalized prompts. However, such prompt-based approaches are fundamentally constrained by the model’s finite context window, and further suffer from the increasing difficulty of extracting effective information as user data grows. To mitigate these limitations, recent work has turned to retrieval-augmented personalized prompting, which dynamically retrieves salient user records from a long-term memory to populate the prompt, obviating the need for exhaustive history inclusion Salemi et al. (2023); Qian et al. (2025); Sun et al. (2024). As an example, Pearl Mysore et al. (2023) leverages a retriever calibrated to the generation objective to select historical user-authored documents that properly enhance and augment the prompt. More recently, PRIME Zhang et al. (2025a) proposes to enhance LLM personalization via episodic and semantic memory mechanisms, aiming to achieve retaining and updating memory for more efficient individual information retrieval.

While being conceptually simple and highly interpretable, prompt-based personalization relies on strong, explicit user signals in historical data  Tan et al. (2025). In content generation tasks, where personalization hinges on implicit factors such as stylistic preferences or individual personality traits, these approaches often struggle to reliably capture users’ intent and preferences, leading to unstable and degraded performance.

Adapter-Based Approaches Adapter-based approaches typically adapt PEFT methods and could be divided into two types, with the first category training all users with a shared model Zhu et al. (2024); Li et al. (2024); Existing work such as LM-P Woźniak et al. (2024), PLoRAZhang et al. (2024b) and MiLP Zhang et al. (2024a) all fall in this research line. Recent advancements like iLoRA Kong et al. (2024) and RecLoRA Zhu et al. (2024) have incorporated the Mixture of Experts (MoE) structure to model the diverse range of user preferences and behaviors. The second category of methods further enhances personalization performance by assigning each user a PEFT-trained model. For example, OPPU Tan et al. (2024) equips each user with a LoRA module and trains it on the user’s individual data, achieving SOTA results on personalized classification and short content generation tasks on LaMP Salemi et al. (2023) dataset. Further, CoPE Bu et al. (2025) demonstrates the effectiveness of the strategy on long content generation tasks by augmenting training data with negative sampling together with contrastive loss.

Despite its promising results, the main limitation of OPPU lies in its parameter growth. By assigning an individual PEFT module to each user, the total number of parameters increases rapidly with the number of users, resulting in considerable storage overhead. In addition, since each user typically has only a small amount of data, training a conventional PEFT module (such as LoRA) on such limited data often leads to overfitting and redundant parameters. Hence, reducing user-specific parameters is both urgent and essential for effective personalized LLMs. This work first introduces this pivotal challenge, and proposes our solution–PLUME, which decomposes traditional LoRA into ultra-lightweight modules through a shared task-specific parameter space and a layer-wise sharing mechanism, achieving competitive personalization performance with minimal parameter redundancy.

A.3 Details of Centered Kernel Alignment (CKA)

We use linear Centered Kernel Alignment (CKA) Liu et al. (2025b) to measure representation similarity between LoRA parameters across layers. CKA provides a scale-invariant and rotation-invariant similarity measure, making it suitable for comparing learned adapter representations.

Specifically, given two column-centered and vectorized LoRA weight matrices XX and YY, the linear CKA is defined as: Given two column-centered vectorized LoRA matrices from different layers of the same user model, the linear CKA is defined as:

CKA⁡(X,Y)=‖X⊤​Y‖F2‖X⊤​X‖F⋅‖Y⊤​Y‖F.\mathrm{CKA}(X,Y)=\frac{\|X^{\top}Y\|_{F}^{2}}{\|X^{\top}X\|_{F}\;\cdot\;\|Y^{\top}Y\|_{F}}.

Appendix B Data Processing Detail and Data Statistics

In preparing the data, we follow the general setup of prior frameworks such as OPPU Tan et al. (2024) and CoPE Bu et al. (2025) and select 200 users with sufficient interaction histories as our evaluation cohort. For each user, we aggregate all historical interactions and then partition them into training, validation, and test subsets with an 8:1:1 ratio based on temporal ordering. When a personalization prompt requires leveraging user history as examples, we restrict retrieval to the training portion only, ensuring a realistic personalization setup without test leakage. In this case, we employ Contriever Izacard et al. (2021) to retrieve the top-K most relevant history entries, which are then included as few-shot demonstrations for generating the target content. This design allows us to assess both the model’s raw personalization ability and its robustness to retrieved history length and quality. Dataset statistics and splits are provided in Table 4.

Dataset AG PR TW NH ST
# Questions 14,065 6,389 5,138 33,072 15,202
Avg Q Length 332.61 881.90 585.68 188.36 509.99
Avg Target Length 181.44 416.73 324.66 14.95 14.54
Table 4: Dataset statistics for all tasks. AG = Abstract Generation, PR = Product Review, TW = Topic Writing, NH = News Headline, ST = Scholarly Title.
Type Method Long Content Generation Short Content Generation
Abstract Generation Product Review Topic Writing News Headline Scholarly Title
R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR R-1 R-L MTR
Non-Personalized BASE 0.2050 0.1092 0.1286 0.3248 0.1390 0.2011 0.2710 0.1204 0.1727 0.1320 0.1157 0.0831 0.4379 0.3556 0.4192
LoRA 0.3358 0.1983 0.2279 0.3602 0.2204 0.2489 0.2465 0.1331 0.1585 0.2106 0.1948 0.1786 0.4474 0.4004 0.3849
PiSSA 0.3378 0.1910 0.2374 0.3833 0.2276 0.2687 0.2681 0.1359 0.1768 0.2140 0.1970 0.1886 0.4718 0.4207 0.4263
AdaLoRA 0.3403 0.1989 0.2328 0.2713 0.1336 0.1595 0.2135 0.1095 0.1325 0.2127 0.1963 0.1718 0.4445 0.4029 0.3744
QLoRA 0.3260 0.1934 0.2157 0.3668 0.2238 0.2563 0.2401 0.1291 0.1515 0.2062 0.1907 0.1735 0.4482 0.4030 0.3835
LoRA-One 0.4080 0.2700 0.3060 0.3750 0.2180 0.2660 0.2720 0.1340 0.1830 0.2150 0.1980 0.1900 0.4700 0.4190 0.4230
Personalized RAG 0.3549 0.1808 0.2435 0.3268 0.1399 0.2111 0.2435 0.1124 0.1607 0.1322 0.1157 0.0941 0.4231 0.3461 0.3881
OPPU 0.3933 0.2214 0.2614 0.4156 0.2369 0.2872 0.2878 0.1428 0.1827 0.2285 0.2091 0.1900 0.4955 0.4364 0.4050
CoPE 0.3320 0.2071 0.2731 0.3263 0.2125 0.2644 0.1607 0.1041 0.1370 0.1779 0.1589 0.1928 0.4030 0.3485 0.3920
PLUME 0.3962 0.2225 0.2664 0.4209 0.2393 0.2947 0.2944 0.1439 0.1872 0.2291 0.2100 0.1916 0.5103 0.4487 0.4231
PLUME-s 0.3865 0.2169 0.2543 0.4142 0.2397 0.2890 0.2790 0.1406 0.1742 0.2336 0.2155 0.1940 0.5107 0.4512 0.4243
Table 5: Performance comparison on LLaMA-2-7B across five personalized text generation tasks. Bold numbers indicate the best results within each task.
# Target Token OPPU PLUME PLUME-s
Bin 1: 79–125 120 119 ( -0.83% ) 123 ( +0.2.5% )
Bin 2: 126–140 133 130 ( -2.20% ) 135 ( +1.50% )
Bin 3: 141–155 140 139 ( -0.71% ) 141 ( +0.71% )
Bin 4: 156–170 151 150 ( -0.67% ) 155 ( +2.64% )
Bin 5: 171–245 156 158 ( +1.28% ) 166 ( +6.41% )
Table 6: Average generated tokens by configuration across ground-truth length bins. Percentages for PLUME and PLUME-s indicate change relative to OPPU.

Appendix C More Experiment

C.1 Details of Evaluation Metrics

To improve clarity with standard text generation benchmarks, we provide formal definitions of the automatic evaluation metrics used in our experiments. Let GG denote the generated text and RR denote the reference text.

ROUGE-1.

ROUGE-1 Lin (2004) measures unigram (1-gram) overlap between the generated and reference texts, serving as a proxy for lexical content coverage. Let CountG​(w)\text{Count}_{G}(w) and CountR​(w)\text{Count}_{R}(w) denote the frequency of unigram ww in GG and RR, respectively. The unigram overlap is defined as:

Overlap​(w)=min⁡(CountG​(w),CountR​(w)).\text{Overlap}(w)=\min(\text{Count}_{G}(w),\text{Count}_{R}(w)). (8)

The recall form of ROUGE-1 is computed as:

ROUGE-1=∑w∈ROverlap​(w)∑w∈RCountR​(w).\text{ROUGE-1}=\frac{\sum_{w\in R}\text{Overlap}(w)}{\sum_{w\in R}\text{Count}_{R}(w)}. (9)

In practice, we report the F1-score to balance precision and recall:

F1=2​P​RP+R,\text{F1}=\frac{2PR}{P+R}, (10)

where PP and RR denote unigram precision and recall.

ROUGE-L.

ROUGE-L Lin (2004) is based on the Longest Common Subsequence (LCS) between GG and RR, capturing sentence-level structural similarity without requiring consecutive n-gram matches. Let LCS​(G,R)\text{LCS}(G,R) denote the length of the longest common subsequence. The recall and precision are defined as:

RL​C​S=LCS​(G,R)|R|,PL​C​S=LCS​(G,R)|G|.R_{LCS}=\frac{\text{LCS}(G,R)}{|R|},\quad P_{LCS}=\frac{\text{LCS}(G,R)}{|G|}. (11)

The ROUGE-L F-measure is:

ROUGE-L=(1+β2)​RL​C​S​PL​C​SRL​C​S+β2​PL​C​S,\text{ROUGE-L}=\frac{(1+\beta^{2})R_{LCS}P_{LCS}}{R_{LCS}+\beta^{2}P_{LCS}}, (12)

where β\beta is typically set to favor recall.

METEOR.

METEOR (Metric for Evaluation of Translation with Explicit ORdering) Banerjee and Lavie (2005) extends lexical overlap metrics by incorporating exact matches, stem matches, synonym matches, and a fragmentation penalty to account for word ordering. Let mm denote the number of matched unigrams between GG and RR. Precision and recall are defined as:

P=m|G|,R=m|R|.P=\frac{m}{|G|},\quad R=\frac{m}{|R|}. (13)

METEOR computes a weighted harmonic mean emphasizing recall:

Fmean=10​P​RR+9​P.F_{\text{mean}}=\frac{10PR}{R+9P}. (14)

To penalize fragmented matches, a penalty term is introduced:

Penalty=γ​(cm)θ,\text{Penalty}=\gamma\left(\frac{c}{m}\right)^{\theta}, (15)

where cc is the number of matched chunks and γ,θ\gamma,\theta are hyperparameters. The final METEOR score is:

METEOR=Fmean⋅(1−Penalty).\text{METEOR}=F_{\text{mean}}\cdot(1-\text{Penalty}). (16)

Overall, ROUGE-1 measures lexical overlap, ROUGE-L captures structural similarity via sequence alignment, and METEOR incorporates semantic matching and ordering penalties. Together, they provide complementary perspectives for evaluating personalized text generation quality.

C.2 Training Details and Reproducibility

For fair comparison, we conduct all experiments with 5 epochs with the AdamW optimizer. To avoid randomness, we set the generation temperature to be 0. For hyper-parameters, we tried different combinations and report the best results. Specifically, learning rate is chosen from {3​e−4,1​e−4,5​e−5,1​e−5,5​e−6,1​e−6}\{3e^{-4},1e^{-4},5e^{-5},1e^{-5},5e^{-6},1e^{-6}\}; component coefficient ss and s′s^{\prime} from {0.01,0.1,0.5,1,5,10.0,20.0,50.0}\{0.01,0.1,0.5,1,5,10.0,20.0,50.0\}. In sensitivity study, we investigate the influence of LoRA rank from {1,2,3,4,5,6,7,8,16,32,64,128}\{1,2,3,4,5,6,7,8,16,32,64,128\}. All experiments were conducted on a single cluster node equipped with a Dell PowerEdge C6620 and NVIDIA H100 GPUs with 94 GB of memory.

C.3 Llama2-7B Results

Table 5 reports the results using LLaMA-2-7b-chat as the backbone model. Both PLUME variants outperform prior personalized baselines. PLUME achieves slightly higher scores on several tasks, while PLUME-s attains the best overall balance between performance and parameter efficiency. Together with the Mistral-7B results in Table 2, these findings demonstrate the robustness of our method across different backbone models.

Appendix D More Analysis

D.1 Sensitivity on Number of Retrieved History interactions

Figure 5: Performance variations on different numbers of retrieved interaction history.

From Figure  5, we observe contrasting effects of increasing the number of in-prompt reference examples on personalized long-form content generation (e.g., Abstract Generation) versus short-form generation (e.g., Scholarly Title Generation) tasks. For Abstract Generation, adding examples yields small but consistent gains and then plateaus: performance nudges up from about 1 example to 3–4 examples, suggesting that extra stylistic cues help open-ended expansion without overwhelming the model. In contrast, for Scholarly Title Generation, quality peaks early and then declines: 1–2 examples give the best scores, while 3–4 examples slightly hurt, likely due to prompt dilution and competing keyword signals in a short, constrained output space. These findings indicate that long-form personalized generation benefits from more reference examples, whereas short-form generation achieves optimal performance with only 1–2 examples, providing practical guidance for designing in-context learning strategies tailored to different generation tasks.

D.2 Performance Variation by Output Length

Figure 6: Abstract Generation performance across ground-truth length bins. Bin 1: 79–125, Bin 2: 126–140, Bin 3: 141–155, Bin 4: 156–170, Bin 5: 171–245. Panels show (a) OPPU, (b) PLUME, and (c) PLUME-s.

Figure  6 reports performance on the Abstract Generation task when users are grouped by the average number of ground-truth tokens. Both PLUME and PLUME-s consistently outperform OPPU across most length bins. A consistent trend appears across the OPPU, PLUME and PLUME-s models: ROUGE-1 and METEOR peak in the Bin 2 (token length 126–140) and then decline as ground-truth length increases. This trend highlights a critical characteristic in abstract generation tasks: there exists an optimal output length range for maintaining high summarization quality. When the ground-truth length exceeds a certain threshold—approximately 140 tokens—both lexical overlap (ROUGE-1) and semantic similarity (METEOR) degrade significantly, even though the models gradually increase their output length. From Table  6, while the models generate progressively longer outputs with increasing bin size, they remain shorter and more conservative than the ground-truth abstracts. Longer ground-truth abstracts contain more entities and clause structure, so partial coverage and paraphrasing are penalized more strongly, revealing capacity limits of the per user LoRA adapters together with deterministic decoding. These results suggest using length aware decoding, explicit coverage objectives for salient terms and entities, and lightweight hierarchical planning to better match long ground-truth targets.

Appendix E Case Study

To qualitatively illustrate the generation result of our approach over baseline approaches, we present a few representative examples from abstract generation and scholarly title generation tasks.

Prompt Template

You are an academic researcher. Your task is to generate an academic-style abstract that matches the author’s writing style based on the paper’s abstract. Here are reference examples (for style/tone ONLY; do not copy them): <REFERENCE_ABSTRACT> is the abstract of <REFERENCE_TITLE>. Generate a NEW abstract for the paper titled: <TARGET_TITLE>
Constraints:
• DO NOT copy any sentences from the reference abstracts. • Length:150–250 words; formal, concise, and self-contained. • Suggested structure: background/objective →\rightarrow method →\rightarrow data/pipeline →\rightarrow results/impact →\rightarrow (optional) deployment/cloud aspects. • Write only the abstract text, without headings or extra commentary.

Prompt You are user #009, and a academic researcher. Your task is to generate an academic-style abstract that matches the author’s writing style based on the paper’s abstract. Here are reference examples (for style/tone ONLY; do not copy them): “The User Centric Smart Card Ownership Model (UCOM) provides an open and dynamic smart card environment enabling cardholders to request installation/deletion of an application to which they are entitled. As in this model, smart cards are not under the control of a centralised authority; hence, it is difficult for an application provider to ascertain their trustworthiness. At present, proposed secure channel protocols for the smart card environment do not provide adequate assurance required by the UCOM. In this paper, we explore the reasons behind their failure to meet the UCOM requirements and then propose a secure and trusted channel protocol that meets them. In addition, the proposed protocol is also suitable to GlobalPlatform’s consumer-centric smart cards. A comparison of the proposed protocol with existing smart card and selected Internet protocols is provided. Then we analyse the protocol with the CasperFDR tool. Finally, we detail the implementation and the performance measurement.” is the abstract of “A Secure and Trusted Channel Protocol for the User Centric Smart Card Ownership Model.” Generate a NEW abstract for the paper titled: “Coopetitive architecture to support a dynamic and scalable NFC based mobile services architecture.” Constraints: - DO NOT copy any sentences from the reference abstracts. - Length:150–250 words; formal, concise, and self-contained. - Suggested structure: background/objective →\rightarrow method →\rightarrow data/pipeline →\rightarrow results/impact →\rightarrow (optional) deployment/cloud aspects. Write only the abstract text, without headings or extra commentary. ### Response:
Version Generated Abstract Output