tcboxmath \tl_set:Ne\tcbhighmathtcbhighmath
HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks
Abstract
Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target author using only a few reference examples while preserving the original meaning. Existing methods often struggle to achieve both high style fidelity and semantic preservation because they compress diverse references into a single static author embedding, which averages out context-dependent stylistic variation, and rely on hidden representations for style control, which entangle style with content. We propose HyperStyler, a novel architecture that decouples LAST into style selection and style realization. Stylo-navigator predicts style coordinates by jointly modeling the source context and target-author references, and Stylo-hypernet realizes them via dynamic parameter modulation instead of hidden-state injection. Our experiments on Reddit, Blog, and News datasets demonstrate that HyperStyler consistently outperforms prior methods including LLM-based approaches and generalizes robustly across domains. Notably, HyperStyler achieves superior performance with as few as 2.4% additional parameters over T5-large, while being over 1.8× faster than LLMs at inference.
1 Introduction
Even when writing the same content, individuals exhibit distinctive lexical choices, syntactic structures, and modes of expression (Stamatatos, 2009; Wang et al., 2023). Low-resource authorship style transfer (LAST) aims to rewrite a source text in the style of an arbitrary target author, preserving the original semantics given only a few reference examples Patel et al. (2024). Unlike traditional style transfer, which is often restricted to authors with massive corpora, LAST extends the scope to everyday writers with a small number of sentences. This shift offers significant practical value, enabling users to efficiently transform drafts into the nuanced voice of any desired author.
Despite its potential, high-fidelity style transfer remains a significant challenge. Initial attempts leverage in-context learning Patel et al. (2024) or inference-time control methods Khan et al. (2024); Horvitz et al. (2024a), but often yield weak stylistic transfer. More recent approaches Liu et al. (2024); Horvitz et al. (2024b) introduce unsupervised style alignment frameworks by constructing pseudo-parallel datasets. However, they still struggle to achieve strong style fidelity and semantic preservation simultaneously.
We can find the first root cause of this challenge in the stylometry literature. According to prior studies, an author’s style is not a static template but a multifaceted phenomenon that shifts with topic and register Sapkota et al. (2014); Hoover (2017); Grieve (2023). In the few-shot setting of LAST, this context-dependency introduces a fundamental problem. Since the references provided at inference time each capture the author’s style in a specific context, processing all references without explicitly identifying which is most relevant to the source text can result in a style that is either diluted into a generic average or dominated by the most salient reference. Specifically, some methods compressing references into a single static author embedding directly induce mode averaging, while other methods directly feeding all references into the context window lack any mechanism to prioritize contextually relevant references, potentially causing the model to latch onto the most stylistically prominent one.
The second cause lies in how existing methods perform style control in the hidden state space. When stylistic signals are directly injected into hidden states, they become entangled with semantic content, making it difficult to isolate style from semantic content. This is particularly problematic in LAST, where the model must handle the open-ended stylistic variation of unseen authors rather than a fixed set of style categories. Such style-content interference makes it increasingly difficult to precisely realize diverse stylistic variations while preserving the original meaning.
To address these limitations, we propose HyperStyler, a novel architecture that decouples the task into a style selection and a style realization stage via two specialized modules. First, Stylo-navigator explicitly predicts style coordinates by considering both the input context and target-author references. Second, Stylo-hypernet realizes these coordinates through dynamic parameter modulation, shifting the control mechanism to the parameter space to reduce content-style entanglement and enable fine-grained controllability. Extensive experiments across Reddit, Blog, and News domains demonstrate that HyperStyler consistently outperforms existing baselines including LLM-based approaches, and generalizes robustly across domains. Furthermore, HyperStyler is over 1.8× faster than LLMs at inference, and maintains superior performance even with only a 2.4% parameter increase over T5-large, highlighting its practical utility for time- and resource-constrained applications.
2 Related Work
2.1 Unsupervised Alignment for LAST
Due to the scarcity of parallel data, recent LAST methods commonly follow a two-stage unsupervised alignment framework (Krishna et al., 2020). The first stage trains a model to reconstruct the original text from style-neutralized paraphrases, conditioned on either a static author embedding Horvitz et al. (2024b) or a set of reference samples (Liu et al., 2024). In the second stage, the model is further aligned using filtered pseudo-parallel data. HyperStyler follows this framework but diverges by aligning toward predicted style coordinates rather than a fixed author embedding, enabling the model to account for the author’s stylistic variation.
2.2 Hypernetworks
Hypernetworks (Ha et al., 2017) generate parameters of a target model conditioned on an external signal, enabling more flexible modulation than static adapters. They have been primarily used for task- or domain-conditioned adaptation (Ivison et al., 2023; Li et al., 2024). However, their use for controlling fine-grained linguistic patterns in open-ended and few-shot settings, where the model must generalize to unseen authors and conditioning signals, remains underexplored. In this paper, we address this gap by conditioning hypernetworks on stylistic coordinates from few-shot references, enabling dynamic style control while reducing content-style entanglement.
3 HyperStyler
3.1 Overview
Given a source text and a set of references written by a target author, our goal is to generate that matches the target author’s writing style while preserving the semantics of . We explicitly decompose this process into two subtasks: (1) style selection, which infers a target style from the references and source context, and (2) stylistic realization, which rewrites in the target style without altering its meaning.
As illustrated in Figure 1, HyperStyler implements these subtasks via two modules attached to an encoder–decoder paraphraser. This architecture decouples content encoding in the encoder from stylistic realization in the decoder, aligning with established practices in style transfer Krishna et al. (2020); Lee et al. (2021); Zhao et al. (2024). First, the Stylo-navigator predicts a style coordinate from and . To ensure parameter efficiency, we reuse the backbone encoder representations as the contextual signal for without an extra encoder. Second, the Stylo-hypernet generates parameter modulations conditioned on , which are applied to the decoder as key/value prefixes in the attention layers and low-rank weight updates in the feed-forward networks (FFNs).
3.2 Stylo-navigator
We define a stylistic coordinate in the style embedding space using STYLE embedder (Wegmann et al., 2022) trained to capture content-independent stylistic representations. This reduces the influence of content semantics from the reference sentences on the control signal, encouraging to primarily reflect stylistic characteristics. Each reference sentence is mapped to a style embedding , yielding a set of reference embeddings .
The Stylo-navigator predicts a stylistic coordinate by attending reference weights conditioned on the source context via two parallel attention mechanisms. We apply self-attention over to capture inter-reference stylistic patterns that characterize the author’s uniqueness, producing . In parallel, cross-attention is applied with as queries and as keys and values, where each token-level hidden state attends to the entire reference set for fine-grained, context-dependent style selection. The token-specific results are aggregated via mean pooling to form a context-aware style query :
| (1) |
We then compute a scaled dot product between and each to obtain the contribution weight of each reference:
| (2) |
where and are layer-normalized for stability. The stylistic coordinate is obtained as the weighted sum of the reference embeddings:
| (3) |
Note that lies within the Style space, not in a newly defined space. Because the weights are conditioned on the source context, varies with the context even for the same author. As an interpolation rather than a selection, can also reach coordinates between individual references.
3.3 Stylo-hypernet
Stylo-hypernet dynamically modulates the decoder conditioned on the stylistic coordinate . Prior analyses suggest that transformer layers contribute differently to generation Langedijk et al. (2024); Alshomary et al. (2025), implying that stylistic realization may be inherently layer-dependent. Motivated by this observation and inspired by Ivison et al. (2023), we introduce learnable layer embeddings and modulate them with to construct layer-specific style signals. Concretely, we compute a style-dependent offset for each modulation target and add it to the corresponding layer embedding via a residual connection to preserve layer identity. The resulting layer-specific signals are then used to generate modulation parameters for the decoder.
Style-conditioned Layer Embeddings.
Let be a learnable embedding table associated with a modulation target (e.g., cross-attention prefix keys). Here, is the number of embeddings for target , and each row corresponds to a distinct modulation target indexed by a tuple (layer, type, position). Prefix embeddings are indexed by (layer, key/value, prefix position) and adapter embeddings are indexed only by projection type (layer, up/down).
We compute compatibility scores between and each layer embedding through a multi-head bilinear interaction:
| (4) |
where is the -th head, and are trainable projection matrices, and and denote layer-normalized vectors. Each is a style-dependent and embedding-specific signal that determines the relative contribution of the corresponding subspace of to embedding . We then form the offset by projecting with trainable matrix , scaling each head with its corresponding score, and concatenating all heads:
| (5) |
Finally, we project the offset with trainable matrix and apply layer normalization to obtain a stable style-conditioned update , which is added to the original embedding :
| (6) |
Generating Modulation Parameters.
We map the style-conditioned layer embeddings to modulation parameters using two-layer MLPs. We use dedicated generators for different modulation targets, as each target operates on distinct parameter spaces with different dimensionalities and functional roles. Here, denotes the hidden dimensionality of the underlying model.
Cross-attention prefixes modulate how the decoder references the source context during generation Li and Liang (2021). We generate key and value prefix vectors using two independent MLPs. The generated vectors are grouped into length- prefixes per layer, , and concatenated to the original keys and values along the sequence dimension, and .
Low-rank adapters target FFN layers, which have been shown to store substantial linguistic information Geva et al. (2021); Geva et al. (2022), making them well-suited for controlling surface realization such as lexical choice and syntactic patterns. For each layer , we generate low-rank down- and up-projection weights using separate MLPs, and . The resulting branch is added to the FFN output:
| (7) |
where and denote the FFN input hidden states and the activation function, respectively. In our implementation, each MLP outputs a vector of length , which is reshaped into the corresponding low-rank matrix for each layer.
3.4 Training Procedure
Due to the lack of a parallel dataset for LAST, we adopt an unsupervised training setting, and the overall training procedure consists of three stages.
Stage 1: Training the Underlying Model.
For each author, we collect an author-specific corpus consisting of a sentence set . Using a pretrained paraphraser, we generate a synthetic paraphrase for each sentence , thereby constructing synthetic pairs . To mitigate stylistic bias inherited from the pretrained paraphraser and to improve both diversity and semantic preservation, we train the underlying model with a bidirectional reconstruction objective over pairs Sjöblom et al. (2020); Ma et al. (2021). The goal of this stage is not to acquire any specific style, but to establish a semantically reliable paraphrasing backbone.
Stage 2: Training the Stylo-navigator and Stylo-hypernet.
We freeze the underlying paraphraser and integrate it with the Stylo-navigator and Stylo-hypernet to be trained simultaneously through an unsupervised reconstruction task. For each author, the reference set is embedded into the style embeddings . The Stylo-navigator is trained to identify the stylistic target within that best matches given the source context . We use the index of the target sentence as a ground-truth label and minimize the negative log-likelihood of the predicted selection probabilities :
| (8) |
To prevent the navigator’s prediction errors from propagating to the Stylo-hypernet, we use teacher-forced style conditioning, where the Stylo-hypernet is conditioned on the ground-truth style embedding (obtained by encoding with the STYLE embedder) rather than the predicted coordinate . This isolates hypernetwork optimization from navigator errors and allows the Stylo-hypernet to focus on learning precise parameter modulation. The Stylo-hypernet is optimized to maximize the reconstruction likelihood of :
| (9) |
Stage 3: Unsupervised Alignment Training.
Finally, we optimize the model for style transfer beyond reconstruction, conditioning the Stylo-hypernet on the stylistic coordinate from the Stylo-navigator. Since parallel data between source and target authors is unavailable, we construct a high-quality parallel dataset via self-distillation Zhang et al. (2019). Specifically, we generate style-transferred outputs using the Stage 2 model and filter them following the Rerank and Filtering procedure from Horvitz et al. (2024b). Unlike prior work that uses a mean-pooled reference embedding, we use predicted to evaluate style fidelity during filtering. Using the pseudo-parallel data, we jointly train the Stylo-navigator and Stylo-hypernet.
| Method | Blog | News | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | |
| STYLL(Qwen2.5-7B) | 0.811 | 0.063 | 0.428 | 0.208 | 0.812 | 0.071 | 0.408 | 0.179 | 0.739 | 0.009 | 0.479 | 0.061 |
| GPT4-turbo | 0.814 | 0.081 | 0.702 | 0.314 | 0.896 | 0.128 | 0.713 | 0.331 | 0.677 | 0.063 | 0.860 | 0.290 |
| GPT5-mini | 0.856 | 0.082 | 0.728 | 0.332 | 0.885 | 0.126 | 0.687 | 0.333 | 0.719 | 0.093 | 0.759 | 0.336 |
| GPT5.4 | 0.918 | 0.117 | 0.597 | 0.359 | 0.958 | 0.093 | 0.526 | 0.276 | 0.848 | 0.071 | 0.718 | 0.273 |
| Llama3.1-8B-Instruct | 0.756 | 0.135 | 0.587 | 0.390 | 0.875 | 0.173 | 0.539 | 0.407 | 0.819 | 0.107 | 0.493 | 0.281 |
| ParaGuideλ=200 | 0.763 | 0.053 | 0.598 | 0.235 | 0.661 | 0.069 | 0.719 | 0.301 | 0.595 | 0.038 | 0.545 | 0.187 |
| ParaGuideλ=2500 | 0.853 | 0.067 | 0.450 | 0.240 | 0.736 | 0.100 | 0.627 | 0.341 | 0.515 | 0.026 | 0.685 | 0.164 |
| StyleMC | 0.658 | 0.036 | 0.450 | 0.154 | 0.565 | 0.063 | 0.439 | 0.195 | 0.403 | 0.029 | 0.574 | 0.153 |
| ASTRAPOP | 0.578 | 0.027 | 0.728 | 0.171 | 0.997 | 0.255 | 0.014 | 0.060 | 0.840 | 0.060 | 0.170 | 0.139 |
| ASTRAPOP | 0.620 | 0.029 | 0.695 | 0.173 | 0.840 | 0.070 | 0.171 | 0.139 | 0.576 | 0.082 | 0.713 | 0.319 |
| TinyStyler | 0.897 | 0.144 | 0.352 | 0.323 | 0.791 | 0.134 | 0.603 | 0.379 | 0.605 | 0.053 | 0.582 | 0.223 |
| TinyStyler | 0.888 | 0.141 | 0.506 | 0.387 | 0.793 | 0.135 | 0.721 | 0.421 | 0.598 | 0.054 | 0.691 | 0.252 |
| TinyStyler | 0.860 | 0.122 | 0.626 | 0.399 | 0.743 | 0.129 | 0.786 | 0.434 | 0.541 | 0.057 | 0.797 | 0.278 |
| TinyStyler | 0.859 | 0.122 | 0.730 | 0.436 | 0.756 | 0.128 | 0.844 | 0.452 | 0.561 | 0.057 | 0.843 | 0.294 |
| HyperStyler | 0.806 | 0.155 | 0.475 | 0.378 | 0.676 | 0.187 | 0.684 | 0.477 | 0.581 | 0.110 | 0.615 | 0.355 |
| HyperStyler | 0.800 | 0.152 | 0.705 | 0.460 | 0.692 | 0.189 | 0.854 | 0.537 | 0.587 | 0.101 | 0.802 | 0.399 |
| HyperStyler | 0.818 | 0.152 | 0.578 | 0.418 | 0.731 | 0.183 | 0.701 | 0.489 | 0.571 | 0.098 | 0.678 | 0.370 |
| HyperStyler | 0.815 | 0.147 | 0.791 | 0.485 | 0.736 | 0.183 | 0.864 | 0.538 | 0.585 | 0.083 | 0.865 | 0.372 |
4 Experiments
4.1 Experimental Setup
Datasets.
We conduct experiments on three datasets with distinct genres: Reddit Khan et al. (2021), Blog Schler et al. (2006), and News (All-the-news). Following prior works for LAST Patel et al. (2024); Horvitz et al. (2024a); Horvitz et al. (2024b), we segment each author’s corpus into sentences and randomly sample 10 sentences per author. We filter out any samples exceeding 60 tokens and split the data by author into training, validation, and test sets with a 0.9/0.05/0.05 ratio. More details of the datasets are described in the Appendix A.
We use three evaluation sets of Reddit from Patel et al. (2024): Random, Single, and Diverse. Each split comprises 15 source and 15 target authors with 16 samples each, totaling 225 transfer directions and 3,600 transformations. We apply the same configuration to Blog and News by randomly selecting hold-out authors from the test dataset.
Evaluation Metrics.
To ensure fair comparison, we follow the evaluation protocol established in prior LAST studies Patel et al. (2024); Horvitz et al. (2024b). Away and Towards measure the degree to which the set of transferred texts moves away from the source author’s style and toward the target author’s style, respectively. These metrics are computed using a held-out UAR embedder Rivera-Soto et al. (2021) trained via contrastive learning for authorship verification. To evaluate semantic preservation, we use the Mutual Implication Score Babakov et al. (2022) as Sim. Finally, the Joint score summarizes overall performance of style transfer: , where denotes the geometric mean. (see details in C.1)
Baselines.
We compare against baselines across three categories for the LAST task.
In-context learning methods include STYLL Patel et al. (2024) with Qwen2.5, and models prompted with the instructions from Horvitz et al. (2024b), including Llama3.1, GPT4-turbo, and the reasoning enabled GPT5-mini and GPT5.4.
Inference-time control methods include ParaGuide Horvitz et al. (2024a), a diffusion-based model, and StyleMC Khan et al. (2024), which performs Metropolis-Hastings sampling guided by a future regressor.
Unsupervised alignment methods include TinyStyler Horvitz et al. (2024b), which conditions on a mean-pooled style embedding for references and utilizes self-distillation, and ASTRAPOP Liu et al. (2024), a policy optimization conditioning on the entire reference sentences. Beyond its original length-based reward, we also train ASTRAPOP using the Joint as a reward.
Implementation Details.
To ensure fair comparison, we apply two principles: (1) we unify the backbone to T5-large Raffel et al. (2020) across all trainable baselines so that performance reflects methodological rather than capacity differences, and (2) for baselines that require a style guide, we provide a mean-pooled style embedding rather than the UAR embedding used for evaluation to prevent baselines from directly optimizing the evaluation metric. We utilize off-the-shelf paraphrasing PEGASUS Zhang et al. (2020), following Horvitz et al. (2024b), and set the adapter rank to 32 and the prefix length to 5. Following TinyStyler, we apply reranking Suzgun et al. (2022) at inference time, but use the predicted as the target style instead of a mean-pooled embedding. Other details are provided in the Appendix B. Our implementation code for HyperStyler is available at https://github.com/JK-SHIN-PG/HyperStyler.
| Model | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Train: Reddit | ||||||||||||||||
| Reddit Blog | Reddit News | Blog Reddit | News Reddit | |||||||||||||
| TinyStyler | 0.768 | 0.175 | 0.653 | 0.477 | 0.716 | 0.123 | 0.621 | 0.412 | 0.672 | 0.093 | 0.797 | 0.393 | 0.428 | 0.037 | 0.843 | 0.231 |
| HyperStyler | 0.757 | 0.239 | 0.601 | 0.499 | 0.718 | 0.182 | 0.567 | 0.446 | 0.755 | 0.209 | 0.604 | 0.481 | 0.660 | 0.197 | 0.614 | 0.459 |
| Train: Blog | ||||||||||||||||
| Blog News | Blog Reddit | News Blog | Reddit Blog | |||||||||||||
| TinyStyler | 0.632 | 0.103 | 0.759 | 0.399 | 0.678 | 0.094 | 0.794 | 0.400 | 0.469 | 0.030 | 0.835 | 0.224 | 0.765 | 0.176 | 0.662 | 0.479 |
| HyperStyler | 0.630 | 0.117 | 0.720 | 0.411 | 0.618 | 0.104 | 0.718 | 0.397 | 0.596 | 0.154 | 0.718 | 0.442 | 0.762 | 0.252 | 0.639 | 0.523 |
| Train: News | ||||||||||||||||
| News Blog | News Reddit | Blog News | Reddit News | |||||||||||||
| TinyStyler | 0.470 | 0.031 | 0.837 | 0.230 | 0.426 | 0.037 | 0.846 | 0.237 | 0.633 | 0.101 | 0.754 | 0.396 | 0.713 | 0.123 | 0.613 | 0.412 |
| HyperStyler | 0.516 | 0.100 | 0.737 | 0.391 | 0.444 | 0.066 | 0.725 | 0.315 | 0.714 | 0.193 | 0.593 | 0.454 | 0.779 | 0.250 | 0.507 | 0.468 |
| Model | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Train: Reddit | ||||||||||||
| Reddit Reddit | Blog Blog | News News | ||||||||||
| TinyStyler | 0.860 | 0.122 | 0.626 | 0.399 | 0.776 | 0.111 | 0.733 | 0.388 | 0.600 | 0.017 | 0.762 | 0.123 |
| HyperStyler | 0.818 | 0.152 | 0.578 | 0.418 | 0.779 | 0.172 | 0.669 | 0.476 | 0.601 | 0.045 | 0.724 | 0.249 |
| Train: Blog | ||||||||||||
| Reddit Reddit | Blog Blog | News News | ||||||||||
| TinyStyler | 0.828 | 0.061 | 0.711 | 0.276 | 0.743 | 0.129 | 0.786 | 0.434 | 0.538 | 0.034 | 0.800 | 0.216 |
| HyperStyler | 0.787 | 0.073 | 0.641 | 0.304 | 0.731 | 0.183 | 0.701 | 0.489 | 0.556 | 0.057 | 0.746 | 0.284 |
| Train: News | ||||||||||||
| Reddit Reddit | Blog Blog | News News | ||||||||||
| TinyStyler | 0.718 | 0.029 | 0.726 | 0.182 | 0.673 | 0.046 | 0.780 | 0.243 | 0.541 | 0.057 | 0.797 | 0.278 |
| HyperStyler | 0.766 | 0.036 | 0.614 | 0.185 | 0.694 | 0.120 | 0.712 | 0.428 | 0.578 | 0.094 | 0.683 | 0.365 |
4.2 Results
Overall Performance.
Style transfer requires simultaneously achieving high style fidelity and semantic preservation, as a model biased toward preservation fails to transfer style, while one biased toward style fidelity risks distorting meaning Fu et al. (2018); Horvitz et al. (2024b). As shown in Table 1, HyperStyler strikes the best balance between the two objectives, improving Towards while maintaining competitive Sim scores, resulting in the highest Joint scores across all three domains. Further improvements are observed when reranking is applied.
Most baselines have relatively low performance on the News, as news articles are more formal and exhibit lower stylistic variability, making it more difficult to capture distinctive author-specific styles Eder et al. (2021); Wang and Riddell (2022) (Table 8). Despite this, HyperStyler outperforms all baselines, including LLMs. Additional experimental results including qualitative analysis are provided in Appendix D.
Generalization Capability.
We compare the generalization capability of HyperStyler with TinyStyler, the strongest baseline in our experiments. As shown in Table 2, TinyStyler consistently degrades when transferring from News to Reddit and Blog, where the inter-author distances to the target authors are approximately 2.04 and 1.68 larger than those in the in-domain setting, respectively (Figure 5). Given these larger distances and the high style variation in Blog and Reddit, this degradation suggests that a single mean embedding fails to provide sufficiently fine-grained style signals for such large stylistic shifts. In contrast, HyperStyler achieves consistently strong performance across most domain pairs, as its context-aware style selection and parameter modulation enable more precise stylistic adaptation across diverse domains.
We also evaluate in-domain authorship style transfer under out-of-domain training. As shown in Table 3, TinyStyler performs well when trained and evaluated within the same domain, but its performance degrades substantially when applied to other domains. Since the training domain shapes the range and granularity of styles a model learns, and each target domain may require a different level of stylistic precision, some degree of degradation under domain shift is expected. Nevertheless, HyperStyler exhibits relatively limited degradation. Notably, HyperStyler trained only on News achieves BlogBlog transfer performance close to that of TinyStyler trained directly on Blog, highlighting HyperStyler’s robust generalization capability.
4.3 Analysis and Ablation Study
Capturing Context-dependent Style Variation.
We examine whether HyperStyler navigates the style space in a context-dependent manner. Specifically, we paraphrase the original texts and have the model reconstruct them in their corresponding styles. As shown in Figure 2, TinyStyler, which is conditioned on a single static embedding, fails to faithfully reproduce the original stylistic distribution, whereas HyperStyler closely matches it. Furthermore, the predicted achieves a cosine similarity of 0.82 with the original style embedding and a Mean Reciprocal Rank (MRR) of 0.80, substantially outperforming mean pooling (cosine similarity:0.58, MRR: 0.21). These results demonstrate that the Stylo-navigator accurately identifies the context-appropriate style target. We also analyze performance under target-author style variation. Figure 3 shows that HyperStyler remains robust as variation increases, whereas the baselines degrade. This highlights that context-aware style selection enables HyperStyler to effectively handle high stylistic variation within an author’s style.
Does style selection need to be explicit?
We compare the Stylo-navigator against two alternative strategies. (1) Mean pooling: We inject the mean-pooled reference embedding into the Stylo-hypernet. (2) Implicit selection: We provide all reference embeddings and perform layer-wise style selection implicitly via cross-attention. The results in Table 4 show that both alternatives exhibit substantially lower Towards scores compared to the Stylo-navigator, demonstrating that explicit style selection conditioned on the source context is effective in improving style fidelity.
| Model | Away | Towards | Sim | Joint | |
|---|---|---|---|---|---|
| HyperStyler | 0.818 | 0.152 | 0.578 | 0.418 | |
| w/o Stylo-navigator (Mean-pooling) | 0.783 | 0.099 | 0.706 | 0.368 | |
| w/o Stylo-navigator (Implicit selection) | 0.778 | 0.114 | 0.671 | 0.384 | |
| w/o Stylo-hypernet (Global) | 0.990 | 0.006 | 0.165 | 0.016 | |
| w/o Stylo-hypernet (Layer-wise) | 0.791 | 0.121 | 0.630 | 0.394 | |
| w/o adapter in FFN | 0.800 | 0.134 | 0.608 | 0.409 | |
| w/o prefix in CrossAttn | 0.825 | 0.150 | 0.551 | 0.408 | |
| w/ prefix in SelfAttn | 0.969 | 0.016 | 0.461 | 0.076 | |
| Underlying paraphraser | 0.896 | 0.013 | 0.718 | 0.088 | |
Should style realization operate in parameter space?
We compare parameter modulation against two hidden-state injection strategies. (1) Global: We concatenate predicted to the encoder hidden states, providing the same style signal to all decoder layers. (2) Layer-wise: We generate layer-specific style embeddings and concatenate them to the encoder hidden states for each decoder layer. The global strategy fails to induce style changes, confirming that layer-wise style control is necessary. More importantly, while layer-wise injection shows some improvement, the Towards/Sim ratio of HyperStyler (0.263) is approximately 37% higher than that of layer-wise injection (0.192), indicating that parameter modulation achieves higher style fidelity for the same semantic cost compared to hidden-state injection. This demonstrates the effectiveness of parameter modulation in realizing style while preserving content.
Which decoder components should be modulated?
We compare different combinations of modulation targets. Modulating only the cross-attention (w/o Adapter in FFN) achieves high performance but falls short in style fidelity, while modulating only the FFN (w/o Prefix in CrossAttn) improves style fidelity but degrades content preservation. Modulating the self-attention layers results in broken sentence structures, interfering with the decoder’s generation process. These results suggest that jointly modulating FFN and cross-attention achieves the best balance between style fidelity and content preservation.
We further analyze the effects of varying rank and prefix length. As shown in Figure 4, increasing the rank yields only limited additional benefit, and a longer prefix does not necessarily yield further gains. Across all combinations of rank and prefix length, using both modulation components consistently outperforms variants that remove either the FFN adapter or the cross-attention prefix, further supporting the structural effectiveness of dual modulation.
4.4 Human Evaluation
We conduct a human evaluation of style transfer quality. Annotators evaluated two criteria: style fidelity (SF), where they selected whether the transferred text or the source text better matches the target author’s style, and content similarity (CS), where they rated how well the transferred text preserves the meaning of the source text. More details are provided in Appendix C.2. As shown in Table 5, HyperStyler achieves the highest SF and G.mean scores. Its SF score is statistically significantly higher than those of GPT5.4 and ParaGuide. Meanwhile, its CS score is not significantly different from ParaGuide, which achieves the highest CS score. These results further validate that HyperStyler achieves high style fidelity while preserving content comparably to the baseline, demonstrating consistency across both automated metrics and human judgment.
| Model | SF | CS | G.mean |
|---|---|---|---|
| GPT5.4 | 0.47† | 1.17 | 0.36 |
| Llama3.1 | 0.59 | 1.07†‡ | 0.41 |
| ParaGuide | 0.44†‡ | 1.20 | 0.34 |
| TinyStyler | 0.59 | 1.16 | 0.44 |
| TinyStyler | 0.51 | 1.10 | 0.38 |
| HyperStyler | 0.61 | 1.15 | 0.46 |
| HyperStyler | 0.58 | 1.19 | 0.43 |
4.5 Efficiency
Parameter Efficiency.
We evaluate the parameter efficiency of our approach by testing a constrained configuration with a reduced rank and prefix length. As summarized in Table 6, decreasing the number of parameters leads to a slight degradation in performance. Nevertheless, even with the rank and prefix length set to 1, our model still outperforms TinyStyler while adding only 2.4% of the underlying model’s parameters. This result demonstrates that HyperStyler maintains robust performance even under parameter constraints.
| Model | r | p | Away | Towards | Sim | Joint | #Params | |
|---|---|---|---|---|---|---|---|---|
| TinyStyler | - | - | 0.860 | 0.122 | 0.626 | 0.399 | 783M | - |
| HyperStyler | 1 | 1 | 0.801 | 0.141 | 0.602 | 0.413 | 802M | |
| 8 | 5 | 0.817 | 0.148 | 0.580 | 0.414 | 817M | ||
| 32 | 5 | 0.818 | 0.152 | 0.578 | 0.418 | 867M |
Inference Time and Memory.
HyperStyler achieves approximately one second per inference on a single A100 GPU, with a modest overhead over TinyStyler (Table 17). Compared to open-source LLMs, HyperStyler is over 1.8 faster and uses less than one-eighth of the VRAM, and over 2.0 faster than API-based LLMs. Notably, even with reranking applied, HyperStyler remains faster than LLM-based baselines. These results demonstrate that high style transfer performance can be achieved without compromising computational efficiency, highlighting its suitability for time- and memory-constrained practical applications.
5 Conclusion
We introduce HyperStyler, a novel architecture for LAST grounded in the stylometric view that authorship style is not static but varies with context. HyperStyler explicitly decouples the task into two stages: context-aware style selection and stylistic realization. By explicitly selecting the most contextually appropriate style from a limited set of references and realizing it through parameter-space modulation, HyperStyler effectively addresses the mode averaging and style-content entanglement problems of existing methods. Extensive experiments demonstrate that HyperStyler consistently outperforms existing baselines and generalizes robustly across diverse domains, while maintaining superior performance even with only a 2.4% increase in parameters. These results suggest that explicitly decoupling style selection and realization is a promising direction for achieving high-fidelity authorship style transfer in few-shot settings.
Limitations
While HyperStyler achieves strong performance, we acknowledge several limitations stemming from the current LAST task setting. Our study primarily focuses on short-text transformation, typically consisting of one to three sentences, following the established protocols of the LAST benchmark Patel et al. (2024). In practical applications, stylistic editing often extends beyond short texts to paragraph- or document-level inputs. However, paragraph-level authorship transfer requires dedicated solutions for defining and representing long-form stylistic elements as style conditions. Authorship style at this level encompasses compositional elements beyond sentence-level lexical and syntactic patterns, such as discourse structure, argument development, inter-sentence coherence, transitions, and narrative flow. Furthermore, extending authorship transfer to longer texts is hindered by the lack of reliable evaluation protocols for long-form style transfer, as current UAR-based evaluation metrics are primarily validated in short-text settings and may fail to capture cross-sentence stylistic coherence. These challenges highlight the need for future work on authorship transfer at the paragraph and document level, along with long-form evaluation protocols.
Another limitation is that our experiments are conducted on English corpora. Although HyperStyler is not inherently language-specific, extending LAST to multilingual or cross-lingual settings would require language-appropriate style representations capable of capturing content-independent stylistic signals across languages, as well as reliable evaluation protocols for assessing style fidelity and content preservation in multilingual settings. Addressing multilingual and cross-lingual authorship transfer remains an important direction for future work.
Ethics Considerations
Potential Misuse and Impersonation:
LAST enables effective content personalization and stylistic imitation using only a few examples. However, this technique could be exploited by malicious actors for unauthorized impersonation. High-fidelity stylistic imitation, which HyperStyler achieves, poses a significant challenge to existing AI-generated text detection methods, suggesting that a new paradigm for authorship-aware detection is required to identify sophisticated synthetic texts. We advocate for the respectful use of stylistic imitation and emphasize that the responsibility for the generated content remains with the user.
Content Risks and Potential Bias:
Our training data includes datasets from online communities such as Reddit, which inherently contain offensive language, sexual content, or unethical sentiments. In this study, we did not apply explicit pre-filtering to the training data to preserve the raw stylistic features of the source domains. Consequently, the model may inadvertently generate unethical or biased outputs. We strongly advise that robust safety filters and post-processing mechanisms must accompany any real-world deployment of the model to prevent the dissemination of harmful content.
Acknowledgments
This material is based upon work supported by the Air Force Office of Scientific Research under award number FA2386-23-1-4121, by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2024-00458720), and by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grants funded by the Korean government (MSIT) (RS-2024-00439932, SW Starlab; No.RS-2020-II201336, Artificial Intelligence graduate school support (UNIST); No.RS-2021-II212068, Artificial Intelligence Innovation Hub; RS-2025-25442824, AI Star Fellowship Program (Ulsan National Institute of Science and Technology)). The authors used a generative AI tool for linguistic refinement and grammatical editing of the manuscript.
References
- Layered insights: generalizable analysis of human authorial style by leveraging all transformer layers. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 10279–10292. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §3.3.
- A large-scale computational study of content preservation measures for text style transfer and paraphrase generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, S. Louvan, A. Madotto, and B. Madureira (Eds.), Dublin, Ireland, pp. 300–321. External Links: Link, Document Cited by: §C.1, §4.1.
- Fast and accurate deep network learning by exponential linear units (elus). In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Y. Bengio and Y. LeCun (Eds.), External Links: Link Cited by: §B.1.2.
- Acquiring a formality-informed lexical resource for style analysis. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, P. Merlo, J. Tiedemann, and R. Tsarfaty (Eds.), Online, pp. 2028–2041. External Links: Link, Document Cited by: §4.2.
- Style transfer in text: exploration and evaluation. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. External Links: Document Cited by: §4.2.
- Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 30–45. External Links: Link, Document Cited by: §3.3.
- Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 5484–5495. External Links: Link, Document Cited by: §3.3.
- Register variation explains stylometric authorship analysisRegister variation explains stylometric authorship analysis. Corpus Linguistics and Linguistic Theory 19 (1), pp. 47–77. External Links: Link, Document Cited by: §1.
- HyperNetworks. In International Conference on Learning Representations, External Links: Link Cited by: §2.2.
- STEER: unified style transfer with expert reinforcement. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 7546–7562. External Links: Link, Document Cited by: §C.2.
- Gaussian error linear units (gelus). External Links: 1606.08415, Link Cited by: §B.1.2.
- The microanalysis of style variation. Digital Scholarship in the Humanities 32 (suppl_2), pp. ii17–ii30. External Links: ISSN 2055-7671, Document, Link, https://academic.oup.com/dsh/article-pdf/32/suppl_2/ii17/21298934/fqx022.pdf Cited by: §1.
- Paraguide: guided diffusion paraphrasers for plug-and-play textual style transfer. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 18216–18224. External Links: Document Cited by: §B.1.1, §B.5, §1, §4.1, §4.1.
- TinyStyler: efficient few-shot text style transfer with authorship embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 13376–13390. External Links: Link, Document Cited by: §B.1.1, §B.2, §B.6, §1, §2.1, §3.4, §4.1, §4.1, §4.1, §4.1, §4.1, §4.2.
- HINT: hypernetwork instruction tuning for efficient zero- and few-shot generalisation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 11272–11288. External Links: Link, Document Cited by: §2.2, §3.3.
- A deep metric learning approach to account linking. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), Online, pp. 5275–5287. External Links: Link, Document Cited by: §A.1, §4.1.
- Learning to generate text in arbitrary writing styles. External Links: 2312.17242, Link Cited by: §B.3, §1, §4.1.
- Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y. He, and Y. Liu (Eds.), Online, pp. 737–762. External Links: Link, Document Cited by: §C.2, §2.1, §3.1.
- DecoderLens: layerwise interpretation of encoder-decoder transformers. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 4764–4780. External Links: Link, Document Cited by: §3.3.
- Enhancing content preservation in text style transfer using reverse attention and conditional layer normalization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 93–102. External Links: Link, Document Cited by: §3.1.
- Hypernetwork-assisted parameter-efficient fine-tuning with meta-knowledge distillation for domain knowledge disentanglement. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 1681–1695. External Links: Link, Document Cited by: §2.2.
- Prefix-tuning: optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 4582–4597. External Links: Link, Document Cited by: §3.3.
- Authorship style transfer with policy optimization. External Links: 2403.08043, Link Cited by: §B.4, §C.2, §C.2, §1, §2.1, §4.1.
- Collaborative learning of bidirectional decoders for unsupervised text style transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 9250–9266. External Links: Link, Document Cited by: §3.4.
- Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12 (2), pp. 153–157. External Links: Document Cited by: §C.2.
- Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML’10, Madison, WI, USA, pp. 807–814. External Links: ISBN 9781605589077, Link Cited by: §B.1.2.
- Low-resource authorship style transfer: can non-famous authors be imitated?. External Links: 2212.08986, Link Cited by: §A.1, §B.6, §C.1, §C.2, §1, §1, §4.1, §4.1, §4.1, §4.1, Limitations.
- Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21 (140), pp. 1–67. External Links: Link Cited by: Appendix B, §4.1.
- Learning universal authorship representations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 913–919. External Links: Link, Document Cited by: §4.1.
- Cross-topic authorship attribution: will out-of-topic data help?. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, J. Tsujii and J. Hajic (Eds.), Dublin, Ireland, pp. 1228–1237. External Links: Link Cited by: §1.
- Effects of age and gender on blogging.. In AAAI spring symposium: Computational approaches to analyzing weblogs, Vol. 6, pp. 199–205. External Links: Link Cited by: §A.2, §4.1.
- GLU variants improve transformer. External Links: 2002.05202, Link Cited by: §B.1.2.
- Paraphrase generation and evaluation on colloquial-style sentences. In Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis (Eds.), Marseille, France, pp. 1814–1822 (eng). External Links: Link, ISBN 979-10-95546-34-4 Cited by: §3.4.
- A survey of modern authorship attribution methods. Journal of the American Society for information Science and Technology 60 (3), pp. 538–556. External Links: Document Cited by: §1.
- Prompt-and-rerank: a method for zero-shot and few-shot arbitrary textual style transfer with small language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 2195–2222. External Links: Link, Document Cited by: §4.1.
- Attention is all you need. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, pp. . External Links: Link Cited by: §B.1.
- Can authorship representation learning capture stylistic features?. Transactions of the Association for Computational Linguistics 11, pp. 1416–1431. External Links: Link, Document Cited by: §1.
- CCTAA: a reproducible corpus for Chinese authorship attribution research. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, and S. Piperidis (Eds.), Marseille, France, pp. 5889–5893. External Links: Link Cited by: §4.2.
- Same author or just same topic? towards content-independent style representations. In Proceedings of the 7th Workshop on Representation Learning for NLP, S. Gella, H. He, B. P. Majumder, B. Can, E. Giunchiglia, S. Cahyawijaya, S. Min, M. Mozes, X. L. Li, I. Augenstein, A. Rogers, K. Cho, E. Grefenstette, L. Rimell, and C. Dyer (Eds.), Dublin, Ireland, pp. 249–268. External Links: Link, Document Cited by: §B.1.1, §B.3, §3.2.
- Contrastive preference optimization: pushing the boundaries of LLM performance in machine translation. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §B.4.
- PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. External Links: Link Cited by: §4.1.
- Be your own teacher: improve the performance of convolutional neural networks via self distillation. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 3712–3721. External Links: Document Cited by: §3.4.
- SC2: towards enhancing content preservation and style consistency in long text style transfer. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 9949–9960. External Links: Link, Document Cited by: §3.1.
Appendix A Data Description
A.1 Reddit
We use Million User Dataset (MUD) Khan et al. (2021), a large-scale publicly available (Apache-2.0) user text dataset collected from the social media platform Reddit. The dataset comprises over 300 million Reddit posts produced by approximately one million users over the course of one year, and includes text-based user contributions in the form of comments. Evaluation is conducted on three predefined splits from Patel et al. (2024).
- •
Diverse: Source and target authors with posts on diverse topics across 13 or more different subreddits.
- •
Random: Random source and target authors.
- •
Single: All posts belong to a popular college football subreddit.
A.2 Blog
We use the Blog Authorship Corpus Schler et al. (2006) collected from blogger.com, which consists of blog posts written by 19,320 individual bloggers. This dataset is freely available for non-commercial research purposes. This dataset is available at https://www.kaggle.com/datasets/rtatman/blog-authorship-corpus/data
A.3 News
All-the-news dataset contains news articles collected from major U.S. and English-language news outlets. This dataset is available at https://huggingface.co/datasets/rjac/all-the-news-2-1-Component-one.
Since our objective is to analyze writing characteristics at the single-author level, we apply a series of filtering steps to ensure data quality. Specifically, we remove articles with missing author information, exclude articles attributed to organizations or non-individual entities, and discard articles with multiple authors. After filtering, only articles attributed to clearly identifiable individual authors are retained.
| Dataset | #samples | #authors | #parallel pairs |
|---|---|---|---|
| 7.5M | 946K | 200K | |
| Blog | 177K | 17K | 40K |
| News | 538K | 53K | 200K |
A.4 Stylistic Distance across Datasets
Table 8 reports intra-author style variation and inter-author distance for each dataset, providing supporting statistics for the performance differences observed across domains. Style variation is measured as the mean cosine distance of each author’s style embeddings to their centroid, averaged across authors. Inter-author distance is measured as the mean pairwise cosine distance between authors in the UAR space.
| Dataset | Style variation | Inter-author distance |
|---|---|---|
| Reddit (Single) | 0.407 | 0.310 |
| Reddit (Random) | 0.414 | 0.384 |
| Reddit (Diverse) | 0.398 | 0.357 |
| Blog | 0.316 | 0.322 |
| News | 0.230 | 0.232 |
A.5 Cross-domain Inter-author Distance
Figure 5 shows the mean pairwise inter-author distances between source and target domain authors in the UAR space. Cross-domain distances are consistently larger than in-domain distances across all domain pairs. The largest differences relative to in-domain distance are observed for News to Reddit and Blog. Notably, since the Towards and Away metric is normalized by the source-target distance (Eq. 11), larger inter-author distances naturally yield lower Towards values regardless of model performance. Consequently, Towards and Joint are not directly comparable across domain pairs with different inter-author distances, and should be interpreted in terms of relative differences between models within the same domain pair.
Appendix B Implementation Details
All training experiments are conducted on two NVIDIA A100 80GB GPUs. We use a batch size of 128 for all training stages. T5-large Raffel et al. (2020) is used as the base model for HyperStyler and all trainable baselines. For each method, we select the checkpoint with the lowest validation loss. At inference time, we use sampling with a temperature of 0.8 and top- to 1.0.
B.1 HyperStyler
For all attention layers in HyperStyler, we follow the T5 backbone by adopting a pre-norm structure with layer normalization and a multi-head decomposition Vaswani et al. (2017) with . We set for simplicity. For parameter efficiency, we share the projection matrices across all embedding tables.
B.1.1 Selection of Style Embedding Space
We adopt the STYLE embedder Wegmann et al. (2022) to ensure a fair comparison with TinyStyler Horvitz et al. (2024b) and ParaGuide Horvitz et al. (2024a), which also rely on it as the style conditioning signal. This embedder is trained via contrastive learning with negative samples from the same topic and domain, encouraging the model to capture subtle stylistic signals that distinguish authors within the same topic, yielding content-independent style representations. To empirically confirm this, we apply k-means clustering with Reddit dataset. Table 21 exhibits the resulting clusters, which are clear and interpretable stylistic patterns, confirming that the STYLE embedder captures meaningful stylistic features beyond content.
B.1.2 Selection of Activation Function for FFN Adapter
We apply an activation function to the low-rank adapter used for FFN modulation (Eq. 7). To examine whether an activation function is necessary and how the choice of activation function affects performance, we compare five variants: no activation, ReLU Nair and Hinton (2010), ELU Clevert et al. (2016), GELU Hendrycks and Gimpel (2016), and GeGLU Shazeer (2020). As shown in Table 9, the variant without an activation function exhibits relatively strong style transfer, but achieves the lowest overall performance due to lower semantic preservation. In contrast, variants using activation functions generally perform better than the no-activation variant. This suggests that an adapter with an activation function is more effective for balancing the trade-off between style fidelity and semantic preservation than a simple linear low-rank transformation.
Meanwhile, the performance differences among activation functions are relatively small. Therefore, we attribute the improvement primarily to the presence of an activation function rather than to any specific choice. Since GeGLU is used in the FFN of the underlying model (google/T5-v1.1-large) and achieves competitive performance, we adopt GeGLU in our model as well.
| Activation | Away | Towards | Sim | Joint |
|---|---|---|---|---|
| w/o ACT | 0.881 | 0.165 | 0.454 | 0.386 |
| ReLU | 0.819 | 0.148 | 0.565 | 0.415 |
| ELU | 0.821 | 0.150 | 0.565 | 0.415 |
| GeLU | 0.827 | 0.151 | 0.560 | 0.411 |
| GeGLU | 0.818 | 0.152 | 0.578 | 0.418 |
| Configuration | Stage 1 | Stage 2 | Stage 3 |
|---|---|---|---|
| Learning rate | |||
| Batch size | 128 | 128 | 128 |
| Optimizer | AdamW | AdamW | AdamW |
| Weight decay | 0.01 | 0.01 | 0.01 |
| Scheduler | Constant | Cosine | Constant |
| Warm-up steps | 2000 | 2000 | 5% of max steps |
| Max steps / epochs | 200K steps | 100K steps | 3 epochs |
B.2 TinyStyler
We follow the original paper’s configuration and use the provided training code Horvitz et al. (2024b). Only for the Reddit dataset, we use the publicly available checkpoint rather than training from scratch, as we found that training from scratch in our environment yielded lower performance than originally reported.
| Configuration | Value |
|---|---|
| Pretrained Ckpt | google/t5-v1_1-large |
| Learning rate | |
| Batch size | 128 |
| Optimizer | Adam |
| Weight decay | 0.01 |
| Schedule | Constant |
| Warm-up Steps | 2000 |
| Total Steps | 150K |
B.3 StyleMC
Since the official source code for StyleMC Khan et al. (2024) is not publicly available, we implemented the method by strictly adhering to the descriptions provided in the original paper. While we made every effort to ensure a faithful reproduction, minor discrepancies may exist compared to the original implementation due to unspecified details of the algorithm or hyperparameters. For a fair comparison, we modified the baseline STYLEMC by replacing its original UAR-based implementation with the STYLE embedder Wegmann et al. (2022). An author embedding was then calculated via mean pooling, aligning it with the evaluation protocol used for other models.
| Configuration | Value |
|---|---|
| Learning rate | |
| Batch size | 128 |
| Optimizer | AdamW |
| Weight decay | 0.01 |
| Future discriminator Ckpt | facebook/opt-1.3b |
| Proposal generator Ckpt | google/t5-v1_1-large |
| Number of steps | 80Sequence length |
| 0.005 | |
| 1.0 | |
| 1.0 | |
| 0.01 |
B.4 ASTRAPOP
We conduct experiments based on the ASTRAPOP framework Liu et al. (2024) with CPO Xu et al. (2024), its best-performing variant, while modifying several components to better align it with our experimental setting. Originally, ASTRAPOP uses LLaMA-2-7B, a decoder-only model, as its backbone architecture. To ensure a fair evaluation, we replace the backbone with T5-Large. This architectural change requires decisions on how the source and reference texts are arranged in the encoder input. We also explore the JOINT reward formulation, following the filtering criterion used in TinyStyler, to examine whether it provides a more effective training signal than the original reward. To identify the configuration that performs best under our setting, we evaluate four variants combining input order and reward function:
- •
ASTRAPOP follows the original input order and reward formulation.
- •
ASTRAPOP adopts the JOINT reward formulation while preserving the original input order.
- •
ASTRAPOP places the [src] token at the beginning of the input sequence, while the original reward formulation remains unchanged.
- •
ASTRAPOPreverse, Joint combines both the reversed input order and the JOINT reward formulation.
According to Table 22, no configuration performs consistently best across datasets. We therefore report the two variants that preserve the original input order in Table 1.
In addition, to examine whether the backbone replacement puts ASTRAPOP at a disadvantage, we train ASTRAPOP with its original LLaMA-2-7B backbone and compare it with HyperStyler. As shown in Table 23, despite a more than 9 difference in model size, HyperStyler consistently achieves higher Joint scores across three datasets.
| Configuration | SFT | CPO |
|---|---|---|
| learning rate | ||
| batch size | 128 | 128 |
| Optimizer | Adam | Adam |
| # epochs | 20 | 20 |
| Max steps | 100K | 100K |
| – | 0.1 | |
| top | – | 1.0 |
| temperature | – | 0.8 |
| length penalty | – | 0.5 |
| Context Size | 512 | 512 |
| Output Size | 80 | 80 |
B.5 Paraguide
We followed the experimental setup used in the original paper Horvitz et al. (2024a).
| Configuration | Value |
|---|---|
| Pretrained Ckpt | xhan77/ssdklm |
| Learning rate | |
| Batch size | 128 |
| Optimizer | AdamW |
| Weight decay | 0.01 |
| Schedule | Constant |
| Warm-up Steps | 2000 |
| Total Steps | 150K |
| Diffusion Steps | 200 |
| Context Size | 80 |
| Output Size | 80 |
B.6 In-context Learning Methods
For GPT-based models and Llama-3.1, we use the prompt from Horvitz et al. (2024b), with default API settings for GPT-4-turbo (gpt-4-turbo-2024-04-09), GPT-5-mini (gpt-5-mini-2025-08-07), and GPT-5.4 (gpt-5.4-2026-03-05, medium). STYLL follows the original experimental setup Patel et al. (2024) using Qwen2.5-7B.
Appendix C Evaluation Details
C.1 Metric Formula Definition
We adopt the evaluation metrics proposed by Patel et al. (2024) for LAST. For any author , let denote their set of 16 posts, and denote the set of posts written by source author and style-transferred to target author . Let denote a single UAR embedding produced over a set of posts . Finally, we define scaled to the range , given by . We further define its complement as .
Away
measures how far a style-transferred text departs from the source author’s style:
| (10) |
Towards
measures how far a style-transferred text moves toward the target author’s style:
| (11) |
Sim
measures how well the transferred text preserves the meaning of the source text. The average Mutual Implication Score Babakov et al. (2022) between two sets of posts authored by and is denoted as :
| (12) |
C.2 Details on Human Evaluation
We recruit annotators from Amazon Mechanical Turk, restricting participation to workers from English-speaking countries (i.e., US, UK, Canada, Australia) with a 95% or higher approval rating. As the evaluation of style transfer is known to be difficult for humans Krishna et al. (2020); Patel et al. (2024); Hallinan et al. (2023); Liu et al. (2024), we introduce a qualification test to ensure a minimum level of annotation quality. The test consists of three items. In each item, annotators are shown five reference texts from each of two randomly selected authors and asked to identify which author wrote a held-out target text. Only annotators who correctly answer all three items are admitted to the main evaluation.
For each baseline category, we select the best-performing model for the main evaluation. The evaluation is conducted on the same 100 source-target author pairs sampled from the Reddit test set, with three annotators assigned to each model output. We exclude examples whose source texts or target-author references contain violent, sexually explicit, or profane content to minimize annotator exposure to potentially harmful or offensive material. Annotators are also informed before the task that they may encounter potentially harmful or offensive content, and only those who agree to proceed participate in the evaluation. We pay 60 cents per annotated pair, corresponding to an estimated hourly rate based on the average completion time.
For style fidelity, annotators are shown eight reference texts from the target author, along with an anonymized source text and a transferred text in randomized order. They are then asked to select which text is more likely to have been written by the target author. The final label is determined by majority vote (Krippendorff’s ), with a score of 1 assigned when the transferred text is selected and 0 otherwise. For content similarity, annotators are shown the source text and the transferred text and asked to rate their semantic similarity on a 3-point Likert scale: 0 indicates Not Similar, 1 indicates Somewhat Similar, and 2 indicates Similar. The average score across annotators is used as the final content similarity score. Detailed instructions are shown in Tables 15 and 16. Our content similarity question and rating scale are adapted from Liu et al. (2024). We use McNemar’s test McNemar (1947) for style fidelity and paired t-test for content similarity to verify whether performance differences between models are statistically significant, with .
| Instruction |
| Read the reference author’s writing samples, then decide which of the two texts is more likely written by that author based on writing style (sentence structure, word choice, tone — not topic). |
| Reference Author’s Writing Samples |
| [Writing samples are shown here] |
| Text A: [Text A is shown here] |
| Text B: [Text B is shown here] |
| Which text is more likely written by the Reference Author, based on writing style? |
| Text A |
| Text B |
| Instruction |
| Read both texts and judge how similar they are. Focus on the core content and key information conveyed, not the writing style. |
| Text A: [Text A is shown here] |
| Text B: [Text B is shown here] |
| How similar are the two texts? |
| 0 — Not Similar |
| Only small portions (less than 50%) of the passages are the same. |
| 1 — Somewhat Similar |
| Large portions (50–75%) of the passages are the same, but there are significant sections that differ or are present in only one passage. |
| 2 — Similar |
| Most of the content (75% or more) of the two passages is the same. |
Appendix D Additional Experimental Results
D.1 Analysis on the Number of References
We analyze how performance varies with the number of references . As shown in Figure 6, HyperStyler shows a consistent and substantial performance advantage over TinyStyler from . Notably, HyperStyler with only references surpasses TinyStyler’s best performance at . This suggests that HyperStyler uses the available references more effectively through context-aware style selection. These results indicate that the key factor is not simply the number of references, but how the model selects and uses stylistic evidence relevant to the source context.
| Method | Time(s) | VRAM(GiB) | |||
| In-context learning | |||||
| STYLL (Qwen2.5-7B) | 14.85 | 32.3 | |||
| GPT-4 Turbo | 2.01 | - | |||
| GPT-5 Mini | 5.14 | - | |||
| GPT-5.4 | 8.50 | - | |||
| Llama-3.1-8B-Instruct | 1.83 | 30.8 | |||
| Inference-time control | |||||
| ParaGuideλ=200 | 20.93 | 3.40 | |||
| ParaGuideλ=2500 | 20.86 | 3.40 | |||
| StyleMC | 49.51 | 9.29 | |||
| Unsupervised alignment | |||||
| ASTRAPOP | 1.72 | 3.55 | |||
| TinyStyler | 0.82 | 3.17 | |||
| TinyStyler | 1.15 | 5.05 | |||
| Proposed method | r | p | |||
| HyperStyler | 1 | 1 | 0.94 | 3.80 | |
| HyperStyler | 1 | 1 | 1.26 | 5.08 | |
| HyperStyler | 8 | 5 | 1.01 | 3.80 | |
| HyperStyler | 8 | 5 | 1.45 | 5.13 | |
| HyperStyler | 32 | 5 | 1.03 | 3.85 | |
| HyperStyler | 32 | 5 | 1.45 | 5.26 | |
D.2 Computational Cost Analysis
Table 17 reports inference latency and memory usage on the Reddit test set. Inference time is averaged over 300 instances, and VRAM is measured as peak FP32 memory usage for locally hosted models on a single NVIDIA A100 GPU. For API-based LLMs, we report wall-clock latency only, as server-side memory usage is not accessible. For reranking variants, the reported time includes both candidate generation and reranking.
| Training stage | GPUs | Time | |
|---|---|---|---|
| HyperStyler | |||
| Stage1 | A100 80G x2 | 28h | |
| Stage2 | A100 80G x2 | 16h | |
| Stage3 | A100 80G x2 | 0.5h | |
| TinyStyler | |||
| Training | A100 80G x2 | 42h | |
| Self-distillation | A100 80G x2 | 4h | |
| ASTRAPOP | |||
| SFT | A100 80G x2 | 45h | |
| CPO | A100 80G x2 | 24h | |
| Paraguide | |||
| Finetuning | A100 80G x2 | 14h | |
| StyleMC | |||
| Future regressor | A100 80G x2 | 18h | |
Table 18 reports training times on the Reddit dataset. These times characterize the practical computational cost under our experimental setup rather than provide a strictly controlled comparison of training efficiency, since methods differ in training objectives, optimization hyperparameters, and training schedules. We omit in-context learning baselines because they do not require task-specific training. HyperStyler takes 28h, 16h, and 0.5h for its three stages, respectively. Its final self-distillation stage uses a filtered set of roughly 40K instances, similar to TinyStyler’s, but requires less training time under our configuration.
D.3 Qualitative Analysis
Target-dependent style transfer.
Figure 8 visualizes t-SNE projections of style embeddings for the same source texts transferred from source author A to two target authors B and C. HyperStyler’s AB and AC outputs occupy distinct stylistic regions and are more closely aligned with the corresponding target author’s style variation. In contrast, TinyStyler and ParaGuide, which rely on static author embeddings, tend to concentrate in a particular stylistic region rather than aligning with the target-author references. Moreover, baselines that receive all reference texts as input also show weaker target-wise separation or weaker alignment with the corresponding target author’s style. These results suggest that HyperStyler’s explicit style navigation enables target-dependent style transfer, producing outputs that reflect the distinct stylistic characteristics of each target author.
Table 20 presents examples where each source text is rewritten using two different target-author reference sets. Source 1 is an argumentative reply. For Target author A, the output opens with a question that reflects the question-based style observed in the references (What is your point?), then restates the original advice in a more explicit form. For Target author B, the output stays relatively close to the source wording while reflecting the ellipsis usage observed in the references (personally….). Source 2 is a reassurance-oriented comment. Target author A makes the response warmer and more supportive (I’m glad you found it, :D), whereas Target author B keeps it concise and neutral without an emotive marker. Source 3 is an event recap. For Target author A, the output adopts a more casual, punctuation-heavy recap style (yep.., !!). For Target author B, the output becomes a more straightforward recap, consistent with the more explanatory phrasing observed in the references. These examples suggest that HyperStyler can realize target-author-specific style cues differently while preserving the source context.
| Case 1: content omission under compressive style |
|---|
| Source: Yeah, the difference between the two ranks is pretty minimal. I mentioned a few weeks back that I didn’t think we’d get a true gauge on Michigan State until the Notre Dame game … I watched that entire game … |
| Output: i didn’t think we’d get a true gauge on Michigan State until the Notre Dame game, but i definitely stand by that after the Furman game lmao |
| Case 2: over-generation under elaborative style |
| Source: Which is why I don’t respect early season rankings … AT ALL |
| Output: I’m not a big fan of early season rankings, but I’m glad someone is enjoying the game. AJ Green is a good player |
Failure mode.
We identified two representative failure modes, shown in Table 19, that arise when the information density of the source text is misaligned with the stylistic signals provided by the reference set. The first is content omission under compressive target style, which occurs when the source text is long and information-dense while the reference set reflects a short, reaction-oriented conversational style. In such cases, HyperStyler tends to compress the source text, retaining the main stance but omitting secondary propositions and supporting details. The second is over-generation under elaborative target style, which occurs when the source text is short and self-contained while the reference set reflects a more expressive and elaborative style. In such cases, the model tends to expand the output to realize the target style, introducing content that is not grounded in the source text and potentially leading to semantic drift. These cases suggest that authorship style transfer becomes particularly challenging when the amount of information that must be preserved from the source conflicts with the degree of compression or elaboration implied by the target reference style.
D.4 Analysis on Style-dependent and Layer-wise Modulation
Stylo-hypernet modulates each decoder layer via a bilinear interaction between the style coordinate and learnable layer embeddings (Eq. 4), yielding compatibility scores for each modulation target. We examine whether these scores vary across styles and layers. Using the k-means clustering results in Section B.1.1, we select the 50 sentences closest to each cluster centroid. Each sentence is fed into the Stylo-hypernet to obtain its compatibility scores , which are averaged within each cluster. The resulting scores are normalized per layer by the maximum absolute value across clusters. As shown in Figure 7, within a given layer, scores differ across style clusters. Within a given cluster, the scores also change from layer to layer. These results confirm that Stylo-hypernet assigns distinct scores across styles and layers, as intended by its design.
Appendix E Licenses and Use of Artifacts
Table 27 lists artifacts used in this work, including models for training and evaluation and software libraries, with their licenses and links to the sources. Our use of the artifacts is consistent with their licenses and intended use. In particular, artifacts released under permissive licenses are used for research in accordance with their terms, and the artifact licensed under CC BY-NC-SA 4.0 is used only in a non-commercial research context. We do not redistribute any artifact in a manner inconsistent with its original license or access conditions, and any outputs or derived materials from this work are intended only for research use.
| Target author references | HyperStyler output |
|---|---|
| Source 1: Why do you need to address him personally at all? Attack the argument, not the person. | |
| Target author A: Why not? Im not a Nasus main but Rylai procs on his ult and E.; Uhm hello what year is this? Even if hes in a straight relationship, he shouldnt do that.; Can u explain the reference?; DAE also think masc is the most opressed!!!?; TIL the internet is one person | What is your point? It is an argument, not a person. Attack the argument not the person |
| Target author B: Big difference between being a fanboi and simply not falling inline with the circlejerk….; Don’t worry /r/cars tells them everyday….; For $5,000 all you should expect is a pile of rusty parts…; Wouldn’t that be the Skrillex cut?; This thread officially proves we are no better than the YouTube comments section | Why are you even addressing him personally…. Attack the argument not the person. |
| Source 2: No prob! Easy to make the mistake. Not like Reddit is graded or anything. :) | |
| Target author A: I guess I’m building. Thank you for your help! I really appreciate it!; You guys have been really polite, I’m really surprised! Thank you for all your help so far!; Thank you for your time, looks like I’m building it! c:; Thank you for your help!; Oh, I’d rather get a new case then. […] Should I just throw it away? | I’m glad you found it. It’s easy to make a mistake, like Reddit isn’t graded! :D |
| Target author B: My two cents: If you are asking technical questions then yes definitely. […]; Hey someone photoshopped out the tea label!; I bet that the bot power is conducted through the floor.; Or poke around on the ground; That sucks. I thought dog treats were like essential kit […] | Nice. Easy to make the mistake. Like reddit isn’t graded or anything. |
| Source 3: Approx 35 people attended, which meant there was more than enough beer to go around. | |
| Target author A: and that’s a record .. i had a temporary vacation from blogging because of my preliminary exams earlier this week!!; i didnt take it badly though.. because he reasoned out that there was traffic along e. rod or something…; i was sarcastic.. coz i was reviewing.. […] luv u!!!; if there’s one way i would like to express my thoughts and ideas.. it is through writing!!; ahhh.. naaalala ko na!! | yep.. there was more than enough beer to go around, because 35 people attended!! |
| Target author B: But they’re having a party on the 17th - not massive, only 50 people - so everybody can come over and eat lots.; I’ll keep ya looped.; Well, I’m free for 10 whole weeks.; Met some new people.; Well, I’ve been on Prophet’s Inc. and Fantasy Essentials - my favourite forums - for the last few hours, chatting to random people. | It was a good night– 35 people attended and there was more than enough beer to go around. |
| Cluster | Stylistic description | Representative examples | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | First-person informal narrative reply. These instances are characterized by first-person narration, informal register, and anecdotal narrative structure. |
| ||||||
| 2 | Brief appreciative reaction with emoticons. These instances are short acknowledgment responses marked by appreciation, positive evaluation, and emoticon usage. |
| ||||||
| 6 | Terse information-seeking question. These instances are short interrogative utterances that function primarily as follow-up requests for information. |
| ||||||
| 11 | Extended clause-heavy justification reply. These instances are characterized by long, multi-clause constructions that build toward a conclusion through chained reasoning, conditional framing, or supporting elaboration. |
| ||||||
| 14 | Extended conversational reply. These instances are long, chatty responses combining informal register, colloquial markers, and loosely connected clauses. |
| ||||||
| 15 | Formulaic endorsement. These instances are short evaluative formulas used to express approval, endorsement, or visibility boosting. |
| ||||||
| 16 | Directive second-person reply. These instances express second-person guidance through imperative constructions or modal advisory forms directed at the interlocutor. |
| ||||||
| 18 | Segmented multi-line commentary. These instances are structured as short line-separated discourse units, often juxtaposing multiple evaluative or explanatory statements. |
| ||||||
| 23 | Blunt categorical assertion. These instances express direct and compact judgments or corrections in forceful declarative form. |
| ||||||
| 30 | Quote-and-correct reply structure. These instances exhibit a quote-response structure in which quoted content is followed by contradiction, correction, or practical follow-up. |
| ||||||
| 32 | Rhetorical second-person interrogative reply. These instances take interrogative form with ironic or sarcastic phrasing directed at the interlocutor, functioning as implicit challenge rather than genuine information-seeking. |
| ||||||
| 41 | Punctuation-heavy expressive reply. These instances are characterized by repeated or emphatic punctuation and expressive phrasing, often conveying heightened emotional intensity. |
| ||||||
| 45 | Emphatic congratulatory/supportive reaction. These instances convey overt positive evaluation through repeated exclamation marks, congratulatory formulas, encouragement, and emoticons. |
| ||||||
| 47 | Ellipsis-heavy hesitant reply. These instances are marked by repeated ellipses and loosely connected clauses, producing a hesitant and trailing conversational rhythm. |
| ||||||
| 48 | Extended formal expository prose. These instances are characterized by long-form prose with formal register, technical or domain-specific vocabulary, and dense information structure. |
|
| Method | Blog | News | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | |
| ASTRAPOP | 0.578 | 0.027 | 0.728 | 0.171 | 0.997 | 0.255 | 0.014 | 0.060 | 0.840 | 0.060 | 0.170 | 0.139 |
| ASTRAPOP | 0.620 | 0.029 | 0.695 | 0.173 | 0.840 | 0.070 | 0.171 | 0.139 | 0.576 | 0.082 | 0.713 | 0.319 |
| ASTRAPOP | 0.571 | 0.026 | 0.743 | 0.172 | 0.991 | 0.275 | 0.077 | 0.170 | 0.655 | 0.086 | 0.394 | 0.248 |
| ASTRAPOP | 0.596 | 0.026 | 0.725 | 0.172 | 0.988 | 0.243 | 0.125 | 0.213 | 0.950 | 0.005 | 0.202 | 0.115 |
| Method | Blog | News | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | |
| ASTRAPOP (LLaMA-2-7B) | 0.708 | 0.188 | 0.505 | 0.333 | 0.813 | 0.244 | 0.569 | 0.479 | 0.799 | 0.110 | 0.559 | 0.322 |
| HyperStyler (T5-large) | 0.818 | 0.152 | 0.578 | 0.418 | 0.731 | 0.183 | 0.701 | 0.489 | 0.571 | 0.098 | 0.678 | 0.370 |
| Method | Diverse | Random | Single | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | |
| STYLL(Qwen2.5-7B) | 0.797 | 0.069 | 0.405 | 0.200 | 0.739 | 0.074 | 0.416 | 0.240 | 0.896 | 0.047 | 0.463 | 0.184 |
| gpt-4-turbo-2024-04-09 | 0.832 | 0.073 | 0.683 | 0.296 | 0.760 | 0.087 | 0.706 | 0.332 | 0.850 | 0.083 | 0.715 | 0.313 |
| gpt-5-mini-2025-08-07 | 0.861 | 0.082 | 0.718 | 0.321 | 0.805 | 0.080 | 0.736 | 0.330 | 0.902 | 0.083 | 0.729 | 0.346 |
| gpt-5.4-2026-03-05 (medium) | 0.940 | 0.077 | 0.501 | 0.238 | 0.866 | 0.139 | 0.679 | 0.436 | 0.948 | 0.135 | 0.610 | 0.402 |
| meta-llama/Llama-3.1-8B-Instruct | 0.724 | 0.119 | 0.602 | 0.371 | 0.743 | 0.127 | 0.564 | 0.369 | 0.800 | 0.158 | 0.595 | 0.431 |
| ParaGuideλ=200 | 0.774 | 0.056 | 0.544 | 0.221 | 0.696 | 0.047 | 0.585 | 0.222 | 0.818 | 0.057 | 0.664 | 0.263 |
| ParaGuideλ=2500 | 0.859 | 0.078 | 0.381 | 0.239 | 0.801 | 0.065 | 0.456 | 0.246 | 0.900 | 0.058 | 0.512 | 0.235 |
| StyleMC | 0.603 | 0.051 | 0.462 | 0.189 | 0.625 | 0.039 | 0.453 | 0.173 | 0.746 | 0.017 | 0.435 | 0.100 |
| ASTRAPOP | 0.612 | 0.031 | 0.679 | 0.168 | 0.516 | 0.027 | 0.736 | 0.192 | 0.607 | 0.022 | 0.770 | 0.152 |
| ASTRAPOP | 0.656 | 0.035 | 0.645 | 0.168 | 0.550 | 0.029 | 0.702 | 0.205 | 0.653 | 0.022 | 0.739 | 0.145 |
| TinyStyler | 0.897 | 0.130 | 0.306 | 0.282 | 0.863 | 0.154 | 0.314 | 0.315 | 0.932 | 0.148 | 0.436 | 0.371 |
| TinyStyler | 0.883 | 0.127 | 0.462 | 0.346 | 0.854 | 0.148 | 0.465 | 0.384 | 0.927 | 0.149 | 0.592 | 0.432 |
| TinyStyler | 0.836 | 0.109 | 0.582 | 0.356 | 0.835 | 0.127 | 0.590 | 0.396 | 0.910 | 0.130 | 0.706 | 0.445 |
| TinyStyler | 0.831 | 0.111 | 0.693 | 0.393 | 0.839 | 0.125 | 0.705 | 0.434 | 0.907 | 0.131 | 0.793 | 0.480 |
| HyperStyler | 0.810 | 0.162 | 0.426 | 0.368 | 0.748 | 0.163 | 0.468 | 0.384 | 0.857 | 0.143 | 0.538 | 0.399 |
| HyperStyler | 0.803 | 0.159 | 0.650 | 0.449 | 0.746 | 0.159 | 0.709 | 0.472 | 0.849 | 0.138 | 0.765 | 0.470 |
| HyperStyler | 0.813 | 0.153 | 0.547 | 0.402 | 0.768 | 0.157 | 0.542 | 0.410 | 0.872 | 0.145 | 0.645 | 0.443 |
| HyperStyler | 0.807 | 0.149 | 0.760 | 0.467 | 0.766 | 0.148 | 0.768 | 0.480 | 0.871 | 0.145 | 0.844 | 0.508 |
| Model | Diverse | Random | Single | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | ||
| HyperStyler | 0.813 | 0.153 | 0.547 | 0.402 | 0.768 | 0.157 | 0.542 | 0.410 | 0.872 | 0.145 | 0.645 | 0.443 | |
| w/o Stylo-navigator (Mean-pooling) | 0.776 | 0.088 | 0.677 | 0.332 | 0.732 | 0.112 | 0.666 | 0.389 | 0.842 | 0.097 | 0.775 | 0.385 | |
| w/o Stylo-navigator (Implicit selection) | 0.770 | 0.102 | 0.641 | 0.348 | 0.729 | 0.124 | 0.627 | 0.392 | 0.833 | 0.116 | 0.746 | 0.412 | |
| w/o Stylo-hypernet (Global) | 0.985 | 0.009 | 0.142 | 0.015 | 0.986 | 0.008 | 0.150 | 0.033 | 0.999 | 0.000 | 0.202 | 0.002 | |
| w/o Stylo-hypernet (Layer-wise) | 0.789 | 0.119 | 0.597 | 0.374 | 0.734 | 0.123 | 0.595 | 0.388 | 0.852 | 0.121 | 0.697 | 0.421 | |
| w/o adapter in FFN | 0.794 | 0.129 | 0.579 | 0.385 | 0.747 | 0.139 | 0.565 | 0.400 | 0.859 | 0.134 | 0.681 | 0.443 | |
| w/o prefix in CrossAttn | 0.830 | 0.150 | 0.521 | 0.396 | 0.768 | 0.156 | 0.512 | 0.399 | 0.877 | 0.142 | 0.621 | 0.430 | |
| w/ prefix in SelfAttn | 0.958 | 0.022 | 0.416 | 0.094 | 0.960 | 0.016 | 0.434 | 0.081 | 0.990 | 0.010 | 0.529 | 0.071 | |
| w/o predicted in stage 3 (mean-pooling) | 0.809 | 0.143 | 0.542 | 0.399 | 0.768 | 0.157 | 0.546 | 0.407 | 0.878 | 0.133 | 0.654 | 0.421 | |
| Underlying paraphraser | 0.889 | 0.019 | 0.682 | 0.123 | 0.849 | 0.012 | 0.696 | 0.084 | 0.949 | 0.007 | 0.776 | 0.058 | |
| Rank | Prefix | Diverse | Random | Single | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | Away | Towards | Sim | Joint | ||
| 8 | - | 0.814 | 0.141 | 0.546 | 0.396 | 0.760 | 0.149 | 0.531 | 0.400 | 0.870 | 0.141 | 0.641 | 0.438 |
| 8 | 5 | 0.819 | 0.147 | 0.551 | 0.400 | 0.763 | 0.154 | 0.540 | 0.403 | 0.870 | 0.143 | 0.648 | 0.440 |
| 128 | - | 0.816 | 0.154 | 0.516 | 0.401 | 0.781 | 0.164 | 0.500 | 0.401 | 0.875 | 0.143 | 0.618 | 0.429 |
| 128 | 5 | 0.822 | 0.156 | 0.524 | 0.403 | 0.770 | 0.164 | 0.528 | 0.411 | 0.877 | 0.148 | 0.627 | 0.440 |
| - | 1 | 0.792 | 0.125 | 0.600 | 0.391 | 0.741 | 0.130 | 0.591 | 0.390 | 0.853 | 0.128 | 0.701 | 0.437 |
| 32 | 1 | 0.818 | 0.145 | 0.538 | 0.396 | 0.764 | 0.150 | 0.536 | 0.397 | 0.871 | 0.145 | 0.640 | 0.438 |
| - | 10 | 0.807 | 0.132 | 0.573 | 0.381 | 0.752 | 0.138 | 0.559 | 0.398 | 0.864 | 0.141 | 0.674 | 0.447 |
| 32 | 10 | 0.817 | 0.149 | 0.528 | 0.399 | 0.772 | 0.163 | 0.523 | 0.409 | 0.877 | 0.147 | 0.632 | 0.438 |
| Type | Artifact | License | Link |
| Model | UAR embedding | Apache-2.0 | https://huggingface.co/rrivera1849/LUAR-MUD |
| STYLE embedding | MIT | https://huggingface.co/AnnaWegmann/Style-Embedding | |
| Paraphrasing PEGASUS | Apache-2.0 | https://huggingface.co/tuner007/pegasus_paraphrase | |
| Mutual Implication Score | CC BY-NC-SA 4.0 | https://github.com/s-nlp/mutual_implication_score | |
| T5-large | Apache-2.0 | https://huggingface.co/google/t5-v1_1-large | |
| Software | HuggingFace Transformers | Apache-2.0 | https://github.com/huggingface/transformers |
| Accelerate | Apache-2.0 | https://github.com/huggingface/accelerate | |
| Scikit-learn | BSD-3-Clause | https://scikit-learn.org/stable/ | |
| NLTK | Apache-2.0 | https://www.nltk.org/ | |
| Matplotlib | PSF | https://matplotlib.org/stable/project/license.html | |
| PyTorch | License link | https://github.com/pytorch/pytorch |