arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.02772v1 [cs.CL] 02 Sep 2026
\tl_set:Ne\tcboxmath

tcboxmath \tl_set:Ne\tcbhighmathtcbhighmath

HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks

Jongkyung Shin†   Minguk Jeon   Chanwoo Park   Chiehyeon Lim† Affiliation: UNIST Affiliation: POSTECH Affiliation: POSCO Holdings Inc. {shinjk1156, rzbsys}@unist.ac.kr {cks1091, chiehyeon.lim}@postech.ac.kr
Abstract

Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target author using only a few reference examples while preserving the original meaning. Existing methods often struggle to achieve both high style fidelity and semantic preservation because they compress diverse references into a single static author embedding, which averages out context-dependent stylistic variation, and rely on hidden representations for style control, which entangle style with content. We propose HyperStyler, a novel architecture that decouples LAST into style selection and style realization. Stylo-navigator predicts style coordinates by jointly modeling the source context and target-author references, and Stylo-hypernet realizes them via dynamic parameter modulation instead of hidden-state injection. Our experiments on Reddit, Blog, and News datasets demonstrate that HyperStyler consistently outperforms prior methods including LLM-based approaches and generalizes robustly across domains. Notably, HyperStyler achieves superior performance with as few as 2.4% additional parameters over T5-large, while being over 1.8× faster than LLMs at inference.

††footnotetext: †\dagger: Corresponding authors.

1 Introduction

Even when writing the same content, individuals exhibit distinctive lexical choices, syntactic structures, and modes of expression (Stamatatos, 2009; Wang et al., 2023). Low-resource authorship style transfer (LAST) aims to rewrite a source text in the style of an arbitrary target author, preserving the original semantics given only a few reference examples Patel et al. (2024). Unlike traditional style transfer, which is often restricted to authors with massive corpora, LAST extends the scope to everyday writers with a small number of sentences. This shift offers significant practical value, enabling users to efficiently transform drafts into the nuanced voice of any desired author.

Despite its potential, high-fidelity style transfer remains a significant challenge. Initial attempts leverage in-context learning Patel et al. (2024) or inference-time control methods Khan et al. (2024); Horvitz et al. (2024a), but often yield weak stylistic transfer. More recent approaches Liu et al. (2024); Horvitz et al. (2024b) introduce unsupervised style alignment frameworks by constructing pseudo-parallel datasets. However, they still struggle to achieve strong style fidelity and semantic preservation simultaneously.

We can find the first root cause of this challenge in the stylometry literature. According to prior studies, an author’s style is not a static template but a multifaceted phenomenon that shifts with topic and register Sapkota et al. (2014); Hoover (2017); Grieve (2023). In the few-shot setting of LAST, this context-dependency introduces a fundamental problem. Since the references provided at inference time each capture the author’s style in a specific context, processing all references without explicitly identifying which is most relevant to the source text can result in a style that is either diluted into a generic average or dominated by the most salient reference. Specifically, some methods compressing references into a single static author embedding directly induce mode averaging, while other methods directly feeding all references into the context window lack any mechanism to prioritize contextually relevant references, potentially causing the model to latch onto the most stylistically prominent one.

The second cause lies in how existing methods perform style control in the hidden state space. When stylistic signals are directly injected into hidden states, they become entangled with semantic content, making it difficult to isolate style from semantic content. This is particularly problematic in LAST, where the model must handle the open-ended stylistic variation of unseen authors rather than a fixed set of style categories. Such style-content interference makes it increasingly difficult to precisely realize diverse stylistic variations while preserving the original meaning.

To address these limitations, we propose HyperStyler, a novel architecture that decouples the task into a style selection and a style realization stage via two specialized modules. First, Stylo-navigator explicitly predicts style coordinates by considering both the input context and target-author references. Second, Stylo-hypernet realizes these coordinates through dynamic parameter modulation, shifting the control mechanism to the parameter space to reduce content-style entanglement and enable fine-grained controllability. Extensive experiments across Reddit, Blog, and News domains demonstrate that HyperStyler consistently outperforms existing baselines including LLM-based approaches, and generalizes robustly across domains. Furthermore, HyperStyler is over 1.8× faster than LLMs at inference, and maintains superior performance even with only a 2.4% parameter increase over T5-large, highlighting its practical utility for time- and resource-constrained applications.

2 Related Work

2.1 Unsupervised Alignment for LAST

Due to the scarcity of parallel data, recent LAST methods commonly follow a two-stage unsupervised alignment framework (Krishna et al., 2020). The first stage trains a model to reconstruct the original text from style-neutralized paraphrases, conditioned on either a static author embedding Horvitz et al. (2024b) or a set of reference samples (Liu et al., 2024). In the second stage, the model is further aligned using filtered pseudo-parallel data. HyperStyler follows this framework but diverges by aligning toward predicted style coordinates rather than a fixed author embedding, enabling the model to account for the author’s stylistic variation.

2.2 Hypernetworks

Hypernetworks (Ha et al., 2017) generate parameters of a target model conditioned on an external signal, enabling more flexible modulation than static adapters. They have been primarily used for task- or domain-conditioned adaptation (Ivison et al., 2023; Li et al., 2024). However, their use for controlling fine-grained linguistic patterns in open-ended and few-shot settings, where the model must generalize to unseen authors and conditioning signals, remains underexplored. In this paper, we address this gap by conditioning hypernetworks on stylistic coordinates from few-shot references, enabling dynamic style control while reducing content-style entanglement.

3 HyperStyler

3.1 Overview

Figure 1: Overall architecture of HyperStyler

Given a source text xx and a set of references R={ri}i=1KR=\{r_{i}\}^{K}_{i=1} written by a target author, our goal is to generate yy that matches the target author’s writing style while preserving the semantics of xx. We explicitly decompose this process into two subtasks: (1) style selection, which infers a target style from the references and source context, and (2) stylistic realization, which rewrites xx in the target style without altering its meaning.

As illustrated in Figure 1, HyperStyler implements these subtasks via two modules attached to an encoder–decoder paraphraser. This architecture decouples content encoding in the encoder from stylistic realization in the decoder, aligning with established practices in style transfer Krishna et al. (2020); Lee et al. (2021); Zhao et al. (2024). First, the Stylo-navigator predicts a style coordinate zz from xx and RR. To ensure parameter efficiency, we reuse the backbone encoder representations Henc=Enc​(x)H_{\text{enc}}=\text{Enc}(x) as the contextual signal for xx without an extra encoder. Second, the Stylo-hypernet generates parameter modulations conditioned on zz, which are applied to the decoder as key/value prefixes in the attention layers and low-rank weight updates in the feed-forward networks (FFNs).

3.2 Stylo-navigator

We define a stylistic coordinate zz in the style embedding space using STYLE embedder (Wegmann et al., 2022) trained to capture content-independent stylistic representations. This reduces the influence of content semantics from the reference sentences on the control signal, encouraging zz to primarily reflect stylistic characteristics. Each reference sentence rir_{i} is mapped to a style embedding sis_{i}, yielding a set of reference embeddings S∈ℝK×dS\in\mathbb{R}^{K\times d}.

The Stylo-navigator predicts a stylistic coordinate zz by attending reference weights conditioned on the source context via two parallel attention mechanisms. We apply self-attention over SS to capture inter-reference stylistic patterns that characterize the author’s uniqueness, producing S~=SelfAttn⁡(S)\tilde{S}=\mathrm{SelfAttn}(S). In parallel, cross-attention is applied with HencH_{\text{enc}} as queries and SS as keys and values, where each token-level hidden state attends to the entire reference set for fine-grained, context-dependent style selection. The token-specific results are aggregated via mean pooling to form a context-aware style query q∈ℝdq\in\mathbb{R}^{d}:

q=MeanPool⁡(CrossAttn⁡(Henc,S,S)).q=\mathrm{MeanPool}\big(\mathrm{CrossAttn}(H_{\text{enc}},S,S)\big). (1)

We then compute a scaled dot product between qq and each si~∈ℝd\tilde{s_{i}}\in\mathbb{R}^{d} to obtain the contribution weight αi\alpha_{i} of each reference:

αi=exp⁡(q¯⋅si~¯/d)∑k=1Kexp⁡(q¯⋅sk~¯/d),\alpha_{i}=\frac{\exp(\bar{q}\cdot\bar{\tilde{s_{i}}}/\sqrt{d})}{\sum_{k=1}^{K}\exp(\bar{q}\cdot\bar{\tilde{s_{k}}}/\sqrt{d})}, (2)

where q¯\bar{q} and s~¯i\bar{\tilde{s}}_{i} are layer-normalized for stability. The stylistic coordinate zz is obtained as the weighted sum of the reference embeddings:

z=∑i=1Kαi​si.z=\sum_{i=1}^{K}\alpha_{i}s_{i}. (3)

Note that zz lies within the Style space, not in a newly defined space. Because the weights αi\alpha_{i} are conditioned on the source context, zz varies with the context even for the same author. As an interpolation rather than a selection, zz can also reach coordinates between individual references.

3.3 Stylo-hypernet

Stylo-hypernet dynamically modulates the decoder conditioned on the stylistic coordinate zz. Prior analyses suggest that transformer layers contribute differently to generation Langedijk et al. (2024); Alshomary et al. (2025), implying that stylistic realization may be inherently layer-dependent. Motivated by this observation and inspired by Ivison et al. (2023), we introduce learnable layer embeddings and modulate them with zz to construct layer-specific style signals. Concretely, we compute a style-dependent offset for each modulation target and add it to the corresponding layer embedding via a residual connection to preserve layer identity. The resulting layer-specific signals are then used to generate modulation parameters for the decoder.

Style-conditioned Layer Embeddings.

Let E(t)=[e1;…;eNt]∈ℝNt×deE^{(t)}=[e_{1};\dots;e_{N_{t}}]\in\mathbb{R}^{N_{t}\times d_{e}} be a learnable embedding table associated with a modulation target tt (e.g., cross-attention prefix keys). Here, NtN_{t} is the number of embeddings for target tt, and each row eje_{j} corresponds to a distinct modulation target indexed by a tuple (layer, type, position). Prefix embeddings are indexed by (layer, key/value, prefix position) and adapter embeddings are indexed only by projection type (layer, up/down).

We compute compatibility scores between zz and each layer embedding through a multi-head bilinear interaction:

bj(h)=(We​e¯j)(h)⊤​(Ws​z¯)(h),b_{j}^{(h)}=(W_{e}\bar{e}_{j})^{(h)\top}(W_{s}\bar{z})^{(h)}, (4)

where (⋅)(h)(\cdot)^{(h)} is the hh-th head, WeW_{e} and WsW_{s} are trainable projection matrices, and e¯j\bar{e}_{j} and z¯\bar{z} denote layer-normalized vectors. Each bj(h)b_{j}^{(h)} is a style-dependent and embedding-specific signal that determines the relative contribution of the corresponding subspace of zz to embedding jj. We then form the offset ojo_{j} by projecting z¯\bar{z} with trainable matrix WvW_{v}, scaling each head with its corresponding score, and concatenating all NhN_{h} heads:

oj=[bj(1)​(Wv​z¯)(1);…;bj(Nh)​(Wv​z¯)(Nh)].o_{j}=\big[b_{j}^{(1)}(W_{v}\bar{z})^{(1)};\dots;b_{j}^{(N_{h})}(W_{v}\bar{z})^{(N_{h})}\big]. (5)

Finally, we project the offset ojo_{j} with trainable matrix WoW_{o} and apply layer normalization to obtain a stable style-conditioned update Δ​ej\Delta e_{j}, which is added to the original embedding eje_{j}:

Δ​ej=LN⁡(Wo​oj),e~j=ej+Δ​ej.\Delta e_{j}=\mathrm{LN}(W_{o}o_{j}),\quad\tilde{e}_{j}=e_{j}+\Delta e_{j}. (6)
Generating Modulation Parameters.

We map the style-conditioned layer embeddings E~(t)\tilde{E}^{(t)} to modulation parameters using two-layer MLPs. We use dedicated generators for different modulation targets, as each target operates on distinct parameter spaces with different dimensionalities and functional roles. Here, dmodeld_{\text{model}} denotes the hidden dimensionality of the underlying model.

Cross-attention prefixes modulate how the decoder references the source context during generation Li and Liang (2021). We generate key and value prefix vectors using two independent MLPs. The generated vectors are grouped into length-pp prefixes per layer, PKℓ,PVℓ∈ℝp×dmodelP^{\ell}_{K},P^{\ell}_{V}\in\mathbb{R}^{p\times d_{\text{model}}}, and concatenated to the original keys and values along the sequence dimension, Kℓ′=[PKℓ;Kℓ]K^{{}^{\prime}\ell}=[P^{\ell}_{K};K^{\ell}] and Vℓ′=[PVℓ;Vℓ]V^{{}^{\prime}\ell}=[P^{\ell}_{V};V^{\ell}].

Low-rank adapters target FFN layers, which have been shown to store substantial linguistic information Geva et al. (2021); Geva et al. (2022), making them well-suited for controlling surface realization such as lexical choice and syntactic patterns. For each layer ℓ\ell, we generate low-rank down- and up-projection weights using separate MLPs, Wd​o​w​nℓ∈ℝdmodel×rW^{\ell}_{down}\in\mathbb{R}^{d_{\text{model}}\times r} and Wu​pℓ∈ℝr×dmodelW^{\ell}_{up}\in\mathbb{R}^{r\times d_{\text{model}}}. The resulting branch is added to the FFN output:

houtl=FFNl​(hinl)+σ⁡(hinl​Wd​o​w​nℓ)​Wu​pℓ,h^{l}_{\text{out}}=\mathrm{FFN}^{l}(h^{l}_{\text{in}})+\sigma(h^{l}_{\text{in}}W^{\ell}_{down})W^{\ell}_{up}, (7)

where hinh_{\text{in}} and σ⁡(⋅)\sigma(\cdot) denote the FFN input hidden states and the activation function, respectively. In our implementation, each MLP outputs a vector of length r⋅dmodelr\cdot d_{\text{model}}, which is reshaped into the corresponding low-rank matrix for each layer.

3.4 Training Procedure

Due to the lack of a parallel dataset for LAST, we adopt an unsupervised training setting, and the overall training procedure consists of three stages.

Stage 1: Training the Underlying Model.

For each author, we collect an author-specific corpus consisting of a sentence set X={xi}i=1KX=\{x_{i}\}_{i=1}^{K}. Using a pretrained paraphraser, we generate a synthetic paraphrase xi′x^{\prime}_{i} for each sentence xix_{i}, thereby constructing synthetic pairs {(xi,xi′)}i=1K\{(x_{i},x^{\prime}_{i})\}_{i=1}^{K}. To mitigate stylistic bias inherited from the pretrained paraphraser and to improve both diversity and semantic preservation, we train the underlying model with a bidirectional reconstruction objective over (xi↔xi′)(x_{i}\leftrightarrow x^{\prime}_{i}) pairs Sjöblom et al. (2020); Ma et al. (2021). The goal of this stage is not to acquire any specific style, but to establish a semantically reliable paraphrasing backbone.

Stage 2: Training the Stylo-navigator and Stylo-hypernet.

We freeze the underlying paraphraser and integrate it with the Stylo-navigator and Stylo-hypernet to be trained simultaneously through an unsupervised reconstruction task. For each author, the reference set R=XR=X is embedded into the style embeddings S∈ℝK×dS\in\mathbb{R}^{K\times d}. The Stylo-navigator is trained to identify the stylistic target within SS that best matches xix_{i} given the source context xi′x^{\prime}_{i}. We use the index ii of the target sentence as a ground-truth label and minimize the negative log-likelihood of the predicted selection probabilities α\alpha:

ℒnav=−∑logαi.\mathcal{L}_{\text{nav}}=-\sum\log\alpha_{i}. (8)

To prevent the navigator’s prediction errors from propagating to the Stylo-hypernet, we use teacher-forced style conditioning, where the Stylo-hypernet is conditioned on the ground-truth style embedding sis_{i} (obtained by encoding xix_{i} with the STYLE embedder) rather than the predicted coordinate zz. This isolates hypernetwork optimization from navigator errors and allows the Stylo-hypernet to focus on learning precise parameter modulation. The Stylo-hypernet is optimized to maximize the reconstruction likelihood of xix_{i}:

ℒhypernet=−∑logp(xi∣xi′,si).\mathcal{L}_{\text{hypernet}}=-\sum\log p(x_{i}\mid x^{\prime}_{i},s_{i}). (9)
Stage 3: Unsupervised Alignment Training.

Finally, we optimize the model for style transfer beyond reconstruction, conditioning the Stylo-hypernet on the stylistic coordinate zz from the Stylo-navigator. Since parallel data between source and target authors is unavailable, we construct a high-quality parallel dataset via self-distillation Zhang et al. (2019). Specifically, we generate style-transferred outputs using the Stage 2 model and filter them following the Rerank and Filtering procedure from Horvitz et al. (2024b). Unlike prior work that uses a mean-pooled reference embedding, we use predicted zz to evaluate style fidelity during filtering. Using the pseudo-parallel data, we jointly train the Stylo-navigator and Stylo-hypernet.

Method Reddit Blog News
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
STYLL(Qwen2.5-7B) 0.811 0.063 0.428 0.208 0.812 0.071 0.408 0.179 0.739 0.009 0.479 0.061
GPT4-turbo 0.814 0.081 0.702 0.314 0.896 0.128 0.713 0.331 0.677 0.063 0.860 0.290
GPT5-mini 0.856 0.082 0.728 0.332 0.885 0.126 0.687 0.333 0.719 0.093 0.759 0.336
GPT5.4 0.918 0.117 0.597 0.359 0.958 0.093 0.526 0.276 0.848 0.071 0.718 0.273
Llama3.1-8B-Instruct 0.756 0.135 0.587 0.390 0.875 0.173 0.539 0.407 0.819 0.107 0.493 0.281
ParaGuideλ=200 0.763 0.053 0.598 0.235 0.661 0.069 0.719 0.301 0.595 0.038 0.545 0.187
ParaGuideλ=2500 0.853 0.067 0.450 0.240 0.736 0.100 0.627 0.341 0.515 0.026 0.685 0.164
StyleMC 0.658 0.036 0.450 0.154 0.565 0.063 0.439 0.195 0.403 0.029 0.574 0.153
ASTRAPOP 0.578 0.027 0.728 0.171 0.997 0.255 0.014 0.060 0.840 0.060 0.170 0.139
ASTRAPOPJOINT{}_{\text{JOINT}} 0.620 0.029 0.695 0.173 0.840 0.070 0.171 0.139 0.576 0.082 0.713 0.319
TinyStylerREC{}_{\text{REC}} 0.897 0.144 0.352 0.323 0.791 0.134 0.603 0.379 0.605 0.053 0.582 0.223
TinyStylerREC,RERANK(5){}_{\text{REC,RERANK(5)}} 0.888 0.141 0.506 0.387 0.793 0.135 0.721 0.421 0.598 0.054 0.691 0.252
TinyStyler 0.860 0.122 0.626 0.399 0.743 0.129 0.786 0.434 0.541 0.057 0.797 0.278
TinyStylerRERANK(5){}_{\text{RERANK(5)}} 0.859 0.122 0.730 0.436 0.756 0.128 0.844 0.452 0.561 0.057 0.843 0.294
HyperStylerREC{}_{\text{REC}} 0.806 0.155 0.475 0.378 0.676 0.187 0.684 0.477 0.581 0.110 0.615 0.355
HyperStylerREC,RERANK(5){}_{\text{REC,RERANK(5)}} 0.800 0.152 0.705 0.460 0.692 0.189 0.854 0.537 0.587 0.101 0.802 0.399
HyperStyler 0.818 0.152 0.578 0.418 0.731 0.183 0.701 0.489 0.571 0.098 0.678 0.370
HyperStylerRERANK(5){}_{\text{RERANK(5)}} 0.815 0.147 0.791 0.485 0.736 0.183 0.864 0.538 0.585 0.083 0.865 0.372
Table 1: Performance comparison results on three datasets. For the Reddit dataset, we report the average performance across three splits. The highest Joint scores without and with reranking are bolded and underlined, respectively.

4 Experiments

4.1 Experimental Setup

Datasets.

We conduct experiments on three datasets with distinct genres: Reddit Khan et al. (2021), Blog Schler et al. (2006), and News (All-the-news). Following prior works for LAST Patel et al. (2024); Horvitz et al. (2024a); Horvitz et al. (2024b), we segment each author’s corpus into sentences and randomly sample 10 sentences per author. We filter out any samples exceeding 60 tokens and split the data by author into training, validation, and test sets with a 0.9/0.05/0.05 ratio. More details of the datasets are described in the Appendix A.

We use three evaluation sets of Reddit from Patel et al. (2024): Random, Single, and Diverse. Each split comprises 15 source and 15 target authors with 16 samples each, totaling 225 transfer directions and 3,600 transformations. We apply the same configuration to Blog and News by randomly selecting hold-out authors from the test dataset.

Evaluation Metrics.

To ensure fair comparison, we follow the evaluation protocol established in prior LAST studies Patel et al. (2024); Horvitz et al. (2024b). Away and Towards measure the degree to which the set of transferred texts moves away from the source author’s style and toward the target author’s style, respectively. These metrics are computed using a held-out UAR embedder Rivera-Soto et al. (2021) trained via contrastive learning for authorship verification. To evaluate semantic preservation, we use the Mutual Implication Score Babakov et al. (2022) as Sim. Finally, the Joint score summarizes overall performance of style transfer: Joint=G⁡(G⁡(Towards,Away),Sim)\textsc{Joint}=G(G(\textsc{Towards},\textsc{Away}),\textsc{Sim}), where G⁡(⋅)G(\cdot) denotes the geometric mean. (see details in C.1)

Baselines.

We compare against baselines across three categories for the LAST task.

In-context learning methods include STYLL Patel et al. (2024) with Qwen2.5, and models prompted with the instructions from Horvitz et al. (2024b), including Llama3.1, GPT4-turbo, and the reasoning enabled GPT5-mini and GPT5.4.

Inference-time control methods include ParaGuide Horvitz et al. (2024a), a diffusion-based model, and StyleMC Khan et al. (2024), which performs Metropolis-Hastings sampling guided by a future regressor.

Unsupervised alignment methods include TinyStyler Horvitz et al. (2024b), which conditions on a mean-pooled style embedding for references and utilizes self-distillation, and ASTRAPOP Liu et al. (2024), a policy optimization conditioning on the entire reference sentences. Beyond its original length-based reward, we also train ASTRAPOP using the Joint as a reward.

Implementation Details.

To ensure fair comparison, we apply two principles: (1) we unify the backbone to T5-large Raffel et al. (2020) across all trainable baselines so that performance reflects methodological rather than capacity differences, and (2) for baselines that require a style guide, we provide a mean-pooled style embedding rather than the UAR embedding used for evaluation to prevent baselines from directly optimizing the evaluation metric. We utilize off-the-shelf paraphrasing PEGASUS Zhang et al. (2020), following Horvitz et al. (2024b), and set the adapter rank to 32 and the prefix length to 5. Following TinyStyler, we apply reranking Suzgun et al. (2022) at inference time, but use the predicted zz as the target style instead of a mean-pooled embedding. Other details are provided in the Appendix B. Our implementation code for HyperStyler is available at https://github.com/JK-SHIN-PG/HyperStyler.

Model Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
Train: Reddit
Reddit →\rightarrow Blog Reddit →\rightarrow News Blog →\rightarrow Reddit News →\rightarrow Reddit
TinyStyler 0.768 0.175 0.653 0.477 0.716 0.123 0.621 0.412 0.672 0.093 0.797 0.393 0.428 0.037 0.843 0.231
HyperStyler 0.757 0.239 0.601 0.499 0.718 0.182 0.567 0.446 0.755 0.209 0.604 0.481 0.660 0.197 0.614 0.459
Train: Blog
Blog →\rightarrow News Blog →\rightarrow Reddit News →\rightarrow Blog Reddit →\rightarrow Blog
TinyStyler 0.632 0.103 0.759 0.399 0.678 0.094 0.794 0.400 0.469 0.030 0.835 0.224 0.765 0.176 0.662 0.479
HyperStyler 0.630 0.117 0.720 0.411 0.618 0.104 0.718 0.397 0.596 0.154 0.718 0.442 0.762 0.252 0.639 0.523
Train: News
News →\rightarrow Blog News →\rightarrow Reddit Blog →\rightarrow News Reddit →\rightarrow News
TinyStyler 0.470 0.031 0.837 0.230 0.426 0.037 0.846 0.237 0.633 0.101 0.754 0.396 0.713 0.123 0.613 0.412
HyperStyler 0.516 0.100 0.737 0.391 0.444 0.066 0.725 0.315 0.714 0.193 0.593 0.454 0.779 0.250 0.507 0.468
Table 2: Cross-domain authorship style transfer performance across three datasets. Source →\rightarrow Target indicates the transformation of texts from a source domain author’s style to a target domain author’s style. The highest Joint scores for each pair are in bold; values below 0.3 shaded in red. We report results without reranking.
Model Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
Train: Reddit
Reddit →\rightarrow Reddit Blog →\rightarrow Blog News →\rightarrow News
TinyStyler 0.860 0.122 0.626 0.399 0.776 0.111 0.733 0.388 0.600 0.017 0.762 0.123
HyperStyler 0.818 0.152 0.578 0.418 0.779 0.172 0.669 0.476 0.601 0.045 0.724 0.249
Train: Blog
Reddit →\rightarrow Reddit Blog →\rightarrow Blog News →\rightarrow News
TinyStyler 0.828 0.061 0.711 0.276 0.743 0.129 0.786 0.434 0.538 0.034 0.800 0.216
HyperStyler 0.787 0.073 0.641 0.304 0.731 0.183 0.701 0.489 0.556 0.057 0.746 0.284
Train: News
Reddit →\rightarrow Reddit Blog →\rightarrow Blog News →\rightarrow News
TinyStyler 0.718 0.029 0.726 0.182 0.673 0.046 0.780 0.243 0.541 0.057 0.797 0.278
HyperStyler 0.766 0.036 0.614 0.185 0.694 0.120 0.712 0.428 0.578 0.094 0.683 0.365
Table 3: In-domain authorship style transfer performance across three datasets. The highest Joint scores are in bold for each pair. Best results are shaded for each training domain: red for TinyStyler and yellow for HyperStyler.

4.2 Results

Overall Performance.

Style transfer requires simultaneously achieving high style fidelity and semantic preservation, as a model biased toward preservation fails to transfer style, while one biased toward style fidelity risks distorting meaning Fu et al. (2018); Horvitz et al. (2024b). As shown in Table 1, HyperStyler strikes the best balance between the two objectives, improving Towards while maintaining competitive Sim scores, resulting in the highest Joint scores across all three domains. Further improvements are observed when reranking is applied.

Most baselines have relatively low performance on the News, as news articles are more formal and exhibit lower stylistic variability, making it more difficult to capture distinctive author-specific styles Eder et al. (2021); Wang and Riddell (2022) (Table 8). Despite this, HyperStyler outperforms all baselines, including LLMs. Additional experimental results including qualitative analysis are provided in Appendix D.

Generalization Capability.

We compare the generalization capability of HyperStyler with TinyStyler, the strongest baseline in our experiments. As shown in Table 2, TinyStyler consistently degrades when transferring from News to Reddit and Blog, where the inter-author distances to the target authors are approximately 2.04×\times and 1.68×\times larger than those in the in-domain setting, respectively (Figure 5). Given these larger distances and the high style variation in Blog and Reddit, this degradation suggests that a single mean embedding fails to provide sufficiently fine-grained style signals for such large stylistic shifts. In contrast, HyperStyler achieves consistently strong performance across most domain pairs, as its context-aware style selection and parameter modulation enable more precise stylistic adaptation across diverse domains.

We also evaluate in-domain authorship style transfer under out-of-domain training. As shown in Table 3, TinyStyler performs well when trained and evaluated within the same domain, but its performance degrades substantially when applied to other domains. Since the training domain shapes the range and granularity of styles a model learns, and each target domain may require a different level of stylistic precision, some degree of degradation under domain shift is expected. Nevertheless, HyperStyler exhibits relatively limited degradation. Notably, HyperStyler trained only on News achieves Blog→\rightarrowBlog transfer performance close to that of TinyStyler trained directly on Blog, highlighting HyperStyler’s robust generalization capability.

4.3 Analysis and Ablation Study

Capturing Context-dependent Style Variation.

We examine whether HyperStyler navigates the style space in a context-dependent manner. Specifically, we paraphrase the original texts and have the model reconstruct them in their corresponding styles. As shown in Figure 2, TinyStyler, which is conditioned on a single static embedding, fails to faithfully reproduce the original stylistic distribution, whereas HyperStyler closely matches it. Furthermore, the predicted zz achieves a cosine similarity of 0.82 with the original style embedding and a Mean Reciprocal Rank (MRR) of 0.80, substantially outperforming mean pooling (cosine similarity:0.58, MRR: 0.21). These results demonstrate that the Stylo-navigator accurately identifies the context-appropriate style target. We also analyze performance under target-author style variation. Figure 3 shows that HyperStyler remains robust as variation increases, whereas the baselines degrade. This highlights that context-aware style selection enables HyperStyler to effectively handle high stylistic variation within an author’s style.

Figure 2: t-SNE visualization of the style embeddings for original texts, their paraphrases, and reconstructed outputs from TinyStyler and HyperStyler. Each ellipse indicates the approximate style distribution of an author.
Figure 3: Trends of Joint score with respect to the target author’s style variation on Reddit. Lines indicate fitted linear trends with 95% confidence bands.
Does style selection need to be explicit?

We compare the Stylo-navigator against two alternative strategies. (1) Mean pooling: We inject the mean-pooled reference embedding into the Stylo-hypernet. (2) Implicit selection: We provide all reference embeddings and perform layer-wise style selection implicitly via cross-attention. The results in Table 4 show that both alternatives exhibit substantially lower Towards scores compared to the Stylo-navigator, demonstrating that explicit style selection conditioned on the source context is effective in improving style fidelity.

Model Away Towards Sim Joint
HyperStyler 0.818 0.152 0.578 0.418
w/o Stylo-navigator (Mean-pooling) 0.783 0.099 0.706 0.368
w/o Stylo-navigator (Implicit selection) 0.778 0.114 0.671 0.384
w/o Stylo-hypernet (Global) 0.990 0.006 0.165 0.016
w/o Stylo-hypernet (Layer-wise) 0.791 0.121 0.630 0.394
w/o adapter in FFN 0.800 0.134 0.608 0.409
w/o prefix in CrossAttn 0.825 0.150 0.551 0.408
w/ prefix in SelfAttn 0.969 0.016 0.461 0.076
Underlying paraphraser 0.896 0.013 0.718 0.088
Table 4: Ablation study on HyperStyler. We report the averaged value across three test splits of Reddit dataset.
Should style realization operate in parameter space?

We compare parameter modulation against two hidden-state injection strategies. (1) Global: We concatenate predicted zz to the encoder hidden states, providing the same style signal to all decoder layers. (2) Layer-wise: We generate layer-specific style embeddings and concatenate them to the encoder hidden states for each decoder layer. The global strategy fails to induce style changes, confirming that layer-wise style control is necessary. More importantly, while layer-wise injection shows some improvement, the Towards/Sim ratio of HyperStyler (0.263) is approximately 37% higher than that of layer-wise injection (0.192), indicating that parameter modulation achieves higher style fidelity for the same semantic cost compared to hidden-state injection. This demonstrates the effectiveness of parameter modulation in realizing style while preserving content.

Which decoder components should be modulated?

We compare different combinations of modulation targets. Modulating only the cross-attention (w/o Adapter in FFN) achieves high performance but falls short in style fidelity, while modulating only the FFN (w/o Prefix in CrossAttn) improves style fidelity but degrades content preservation. Modulating the self-attention layers results in broken sentence structures, interfering with the decoder’s generation process. These results suggest that jointly modulating FFN and cross-attention achieves the best balance between style fidelity and content preservation.

We further analyze the effects of varying rank and prefix length. As shown in Figure 4, increasing the rank yields only limited additional benefit, and a longer prefix does not necessarily yield further gains. Across all combinations of rank and prefix length, using both modulation components consistently outperforms variants that remove either the FFN adapter or the cross-attention prefix, further supporting the structural effectiveness of dual modulation.

Figure 4: Effect of rank and prefix length with and without each modulation component, averaged across three test splits on the Reddit dataset.

4.4 Human Evaluation

We conduct a human evaluation of style transfer quality. Annotators evaluated two criteria: style fidelity (SF), where they selected whether the transferred text or the source text better matches the target author’s style, and content similarity (CS), where they rated how well the transferred text preserves the meaning of the source text. More details are provided in Appendix C.2. As shown in Table 5, HyperStyler achieves the highest SF and G.mean scores. Its SF score is statistically significantly higher than those of GPT5.4 and ParaGuide. Meanwhile, its CS score is not significantly different from ParaGuide, which achieves the highest CS score. These results further validate that HyperStyler achieves high style fidelity while preserving content comparably to the baseline, demonstrating consistency across both automated metrics and human judgment.

Model SF CS G.mean
GPT5.4 0.47† 1.17 0.36
Llama3.1 0.59 1.07†‡ 0.41
ParaGuide 0.44†‡ 1.20 0.34
TinyStyler 0.59 1.16 0.44
TinyStylerRERANK{}_{\text{RERANK}} 0.51 1.10 0.38
HyperStyler 0.61 1.15 0.46
HyperStylerRERANK{}_{\text{RERANK}} 0.58 1.19 0.43
Table 5: Human evaluation results. Bold and underlining indicate the best and second-best scores, respectively. †{\dagger} and ‡{\ddagger} denote significant differences from them, respectively. Significance is assessed at p<0.05p<0.05 using McNemar’s test for SF and paired t-tests for CS. G.mean is the geometric mean of SF and normalized CS.

4.5 Efficiency

Parameter Efficiency.

We evaluate the parameter efficiency of our approach by testing a constrained configuration with a reduced rank and prefix length. As summarized in Table 6, decreasing the number of parameters leads to a slight degradation in performance. Nevertheless, even with the rank and prefix length set to 1, our model still outperforms TinyStyler while adding only 2.4% of the underlying model’s parameters. This result demonstrates that HyperStyler maintains robust performance even under parameter constraints.

Model r p Away Towards Sim Joint #Params Δ\Delta
TinyStyler - - 0.860 0.122 0.626 0.399 783M -
HyperStyler 1 1 0.801 0.141 0.602 0.413 802M +2.4%+2.4\%
8 5 0.817 0.148 0.580 0.414 817M +4.3%+4.3\%
32 5 0.818 0.152 0.578 0.418 867M +10.7%+10.7\%
Table 6: Parameter efficiency analysis. r and p denote the adapter rank and prefix length, respectively. #Params denotes the number of parameters and Δ\Delta represents the percentage increase relative to the T5-large model.
Inference Time and Memory.

HyperStyler achieves approximately one second per inference on a single A100 GPU, with a modest overhead over TinyStyler (Table 17). Compared to open-source LLMs, HyperStyler is over 1.8×\times faster and uses less than one-eighth of the VRAM, and over 2.0×\times faster than API-based LLMs. Notably, even with reranking applied, HyperStyler remains faster than LLM-based baselines. These results demonstrate that high style transfer performance can be achieved without compromising computational efficiency, highlighting its suitability for time- and memory-constrained practical applications.

5 Conclusion

We introduce HyperStyler, a novel architecture for LAST grounded in the stylometric view that authorship style is not static but varies with context. HyperStyler explicitly decouples the task into two stages: context-aware style selection and stylistic realization. By explicitly selecting the most contextually appropriate style from a limited set of references and realizing it through parameter-space modulation, HyperStyler effectively addresses the mode averaging and style-content entanglement problems of existing methods. Extensive experiments demonstrate that HyperStyler consistently outperforms existing baselines and generalizes robustly across diverse domains, while maintaining superior performance even with only a 2.4% increase in parameters. These results suggest that explicitly decoupling style selection and realization is a promising direction for achieving high-fidelity authorship style transfer in few-shot settings.

Limitations

While HyperStyler achieves strong performance, we acknowledge several limitations stemming from the current LAST task setting. Our study primarily focuses on short-text transformation, typically consisting of one to three sentences, following the established protocols of the LAST benchmark Patel et al. (2024). In practical applications, stylistic editing often extends beyond short texts to paragraph- or document-level inputs. However, paragraph-level authorship transfer requires dedicated solutions for defining and representing long-form stylistic elements as style conditions. Authorship style at this level encompasses compositional elements beyond sentence-level lexical and syntactic patterns, such as discourse structure, argument development, inter-sentence coherence, transitions, and narrative flow. Furthermore, extending authorship transfer to longer texts is hindered by the lack of reliable evaluation protocols for long-form style transfer, as current UAR-based evaluation metrics are primarily validated in short-text settings and may fail to capture cross-sentence stylistic coherence. These challenges highlight the need for future work on authorship transfer at the paragraph and document level, along with long-form evaluation protocols.

Another limitation is that our experiments are conducted on English corpora. Although HyperStyler is not inherently language-specific, extending LAST to multilingual or cross-lingual settings would require language-appropriate style representations capable of capturing content-independent stylistic signals across languages, as well as reliable evaluation protocols for assessing style fidelity and content preservation in multilingual settings. Addressing multilingual and cross-lingual authorship transfer remains an important direction for future work.

Ethics Considerations

Potential Misuse and Impersonation:

LAST enables effective content personalization and stylistic imitation using only a few examples. However, this technique could be exploited by malicious actors for unauthorized impersonation. High-fidelity stylistic imitation, which HyperStyler achieves, poses a significant challenge to existing AI-generated text detection methods, suggesting that a new paradigm for authorship-aware detection is required to identify sophisticated synthetic texts. We advocate for the respectful use of stylistic imitation and emphasize that the responsibility for the generated content remains with the user.

Content Risks and Potential Bias:

Our training data includes datasets from online communities such as Reddit, which inherently contain offensive language, sexual content, or unethical sentiments. In this study, we did not apply explicit pre-filtering to the training data to preserve the raw stylistic features of the source domains. Consequently, the model may inadvertently generate unethical or biased outputs. We strongly advise that robust safety filters and post-processing mechanisms must accompany any real-world deployment of the model to prevent the dissemination of harmful content.

Acknowledgments

This material is based upon work supported by the Air Force Office of Scientific Research under award number FA2386-23-1-4121, by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2024-00458720), and by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grants funded by the Korean government (MSIT) (RS-2024-00439932, SW Starlab; No.RS-2020-II201336, Artificial Intelligence graduate school support (UNIST); No.RS-2021-II212068, Artificial Intelligence Innovation Hub; RS-2025-25442824, AI Star Fellowship Program (Ulsan National Institute of Science and Technology)). The authors used a generative AI tool for linguistic refinement and grammatical editing of the manuscript.

References

  • Alshomary et al. (2025) M. Alshomary, N. R. Varimalla, V. Anand, S. Muresan, and K. McKeown Layered insights: generalizable analysis of human authorial style by leveraging all transformer layers. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 10279–10292. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §3.3.
  • Babakov et al. (2022) N. Babakov, D. Dale, V. Logacheva, and A. Panchenko A large-scale computational study of content preservation measures for text style transfer and paraphrase generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, S. Louvan, A. Madotto, and B. Madureira (Eds.), Dublin, Ireland, pp. 300–321. External Links: Link, Document Cited by: §C.1, §4.1.
  • Clevert et al. (2016) D. Clevert, T. Unterthiner, and S. Hochreiter Fast and accurate deep network learning by exponential linear units (elus). In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Y. Bengio and Y. LeCun (Eds.), External Links: Link Cited by: §B.1.2.
  • Eder et al. (2021) E. Eder, U. Krieg-Holz, and U. Hahn Acquiring a formality-informed lexical resource for style analysis. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, P. Merlo, J. Tiedemann, and R. Tsarfaty (Eds.), Online, pp. 2028–2041. External Links: Link, Document Cited by: §4.2.
  • Fu et al. (2018) Z. Fu, X. Tan, N. Peng, D. Zhao, and R. Yan Style transfer in text: exploration and evaluation. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. External Links: Document Cited by: §4.2.
  • Geva et al. (2022) M. Geva, A. Caciularu, K. Wang, and Y. Goldberg Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 30–45. External Links: Link, Document Cited by: §3.3.
  • Geva et al. (2021) M. Geva, R. Schuster, J. Berant, and O. Levy Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 5484–5495. External Links: Link, Document Cited by: §3.3.
  • Grieve (2023) J. Grieve Register variation explains stylometric authorship analysisRegister variation explains stylometric authorship analysis. Corpus Linguistics and Linguistic Theory 19 (1), pp. 47–77. External Links: Link, Document Cited by: §1.
  • Ha et al. (2017) D. Ha, A. M. Dai, and Q. V. Le HyperNetworks. In International Conference on Learning Representations, External Links: Link Cited by: §2.2.
  • Hallinan et al. (2023) S. Hallinan, F. Brahman, X. Lu, J. Jung, S. Welleck, and Y. Choi STEER: unified style transfer with expert reinforcement. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 7546–7562. External Links: Link, Document Cited by: §C.2.
  • Hendrycks and Gimpel (2016) D. Hendrycks and K. Gimpel Gaussian error linear units (gelus). External Links: 1606.08415, Link Cited by: §B.1.2.
  • Hoover (2017) D. L. Hoover The microanalysis of style variation. Digital Scholarship in the Humanities 32 (suppl_2), pp. ii17–ii30. External Links: ISSN 2055-7671, Document, Link, https://academic.oup.com/dsh/article-pdf/32/suppl_2/ii17/21298934/fqx022.pdf Cited by: §1.
  • Horvitz et al. (2024a) Z. Horvitz, A. Patel, C. Callison-Burch, Z. Yu, and K. McKeown Paraguide: guided diffusion paraphrasers for plug-and-play textual style transfer. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 18216–18224. External Links: Document Cited by: §B.1.1, §B.5, §1, §4.1, §4.1.
  • Horvitz et al. (2024b) Z. Horvitz, A. Patel, K. Singh, C. Callison-Burch, K. McKeown, and Z. Yu TinyStyler: efficient few-shot text style transfer with authorship embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 13376–13390. External Links: Link, Document Cited by: §B.1.1, §B.2, §B.6, §1, §2.1, §3.4, §4.1, §4.1, §4.1, §4.1, §4.1, §4.2.
  • Ivison et al. (2023) H. Ivison, A. Bhagia, Y. Wang, H. Hajishirzi, and M. Peters HINT: hypernetwork instruction tuning for efficient zero- and few-shot generalisation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 11272–11288. External Links: Link, Document Cited by: §2.2, §3.3.
  • Khan et al. (2021) A. Khan, E. Fleming, N. Schofield, M. Bishop, and N. Andrews A deep metric learning approach to account linking. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), Online, pp. 5275–5287. External Links: Link, Document Cited by: §A.1, §4.1.
  • Khan et al. (2024) A. Khan, A. Wang, S. Hager, and N. Andrews Learning to generate text in arbitrary writing styles. External Links: 2312.17242, Link Cited by: §B.3, §1, §4.1.
  • Krishna et al. (2020) K. Krishna, J. Wieting, and M. Iyyer Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y. He, and Y. Liu (Eds.), Online, pp. 737–762. External Links: Link, Document Cited by: §C.2, §2.1, §3.1.
  • Langedijk et al. (2024) A. Langedijk, H. Mohebbi, G. Sarti, W. Zuidema, and J. Jumelet DecoderLens: layerwise interpretation of encoder-decoder transformers. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 4764–4780. External Links: Link, Document Cited by: §3.3.
  • Lee et al. (2021) D. Lee, Z. Tian, L. Xue, and N. L. Zhang Enhancing content preservation in text style transfer using reverse attention and conditional layer normalization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 93–102. External Links: Link, Document Cited by: §3.1.
  • Li et al. (2024) C. Li, L. Wang, X. Lin, S. Huang, and L. He Hypernetwork-assisted parameter-efficient fine-tuning with meta-knowledge distillation for domain knowledge disentanglement. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 1681–1695. External Links: Link, Document Cited by: §2.2.
  • Li and Liang (2021) X. L. Li and P. Liang Prefix-tuning: optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 4582–4597. External Links: Link, Document Cited by: §3.3.
  • Liu et al. (2024) S. Liu, S. Agarwal, and J. May Authorship style transfer with policy optimization. External Links: 2403.08043, Link Cited by: §B.4, §C.2, §C.2, §1, §2.1, §4.1.
  • Ma et al. (2021) Y. Ma, Y. Chen, X. Mao, and Q. Li Collaborative learning of bidirectional decoders for unsupervised text style transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 9250–9266. External Links: Link, Document Cited by: §3.4.
  • McNemar (1947) Q. McNemar Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12 (2), pp. 153–157. External Links: Document Cited by: §C.2.
  • Nair and Hinton (2010) V. Nair and G. E. Hinton Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML’10, Madison, WI, USA, pp. 807–814. External Links: ISBN 9781605589077, Link Cited by: §B.1.2.
  • Patel et al. (2024) A. Patel, N. Andrews, and C. Callison-Burch Low-resource authorship style transfer: can non-famous authors be imitated?. External Links: 2212.08986, Link Cited by: §A.1, §B.6, §C.1, §C.2, §1, §1, §4.1, §4.1, §4.1, §4.1, Limitations.
  • Raffel et al. (2020) C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21 (140), pp. 1–67. External Links: Link Cited by: Appendix B, §4.1.
  • Rivera-Soto et al. (2021) R. A. Rivera-Soto, O. E. Miano, J. Ordonez, B. Y. Chen, A. Khan, M. Bishop, and N. Andrews Learning universal authorship representations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 913–919. External Links: Link, Document Cited by: §4.1.
  • Sapkota et al. (2014) U. Sapkota, T. Solorio, M. Montes, S. Bethard, and P. Rosso Cross-topic authorship attribution: will out-of-topic data help?. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, J. Tsujii and J. Hajic (Eds.), Dublin, Ireland, pp. 1228–1237. External Links: Link Cited by: §1.
  • Schler et al. (2006) J. Schler, M. Koppel, S. Argamon, and J. W. Pennebaker Effects of age and gender on blogging.. In AAAI spring symposium: Computational approaches to analyzing weblogs, Vol. 6, pp. 199–205. External Links: Link Cited by: §A.2, §4.1.
  • Shazeer (2020) N. Shazeer GLU variants improve transformer. External Links: 2002.05202, Link Cited by: §B.1.2.
  • Sjöblom et al. (2020) E. Sjöblom, M. Creutz, and Y. Scherrer Paraphrase generation and evaluation on colloquial-style sentences. In Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis (Eds.), Marseille, France, pp. 1814–1822 (eng). External Links: Link, ISBN 979-10-95546-34-4 Cited by: §3.4.
  • Stamatatos (2009) E. Stamatatos A survey of modern authorship attribution methods. Journal of the American Society for information Science and Technology 60 (3), pp. 538–556. External Links: Document Cited by: §1.
  • Suzgun et al. (2022) M. Suzgun, L. Melas-Kyriazi, and D. Jurafsky Prompt-and-rerank: a method for zero-shot and few-shot arbitrary textual style transfer with small language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 2195–2222. External Links: Link, Document Cited by: §4.1.
  • Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, pp. . External Links: Link Cited by: §B.1.
  • Wang et al. (2023) A. Wang, C. Aggazzotti, R. Kotula, R. R. Soto, M. Bishop, and N. Andrews Can authorship representation learning capture stylistic features?. Transactions of the Association for Computational Linguistics 11, pp. 1416–1431. External Links: Link, Document Cited by: §1.
  • Wang and Riddell (2022) H. Wang and A. Riddell CCTAA: a reproducible corpus for Chinese authorship attribution research. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, and S. Piperidis (Eds.), Marseille, France, pp. 5889–5893. External Links: Link Cited by: §4.2.
  • Wegmann et al. (2022) A. Wegmann, M. Schraagen, and D. Nguyen Same author or just same topic? towards content-independent style representations. In Proceedings of the 7th Workshop on Representation Learning for NLP, S. Gella, H. He, B. P. Majumder, B. Can, E. Giunchiglia, S. Cahyawijaya, S. Min, M. Mozes, X. L. Li, I. Augenstein, A. Rogers, K. Cho, E. Grefenstette, L. Rimell, and C. Dyer (Eds.), Dublin, Ireland, pp. 249–268. External Links: Link, Document Cited by: §B.1.1, §B.3, §3.2.
  • Xu et al. (2024) H. Xu, A. Sharaf, Y. Chen, W. Tan, L. Shen, B. V. Durme, K. Murray, and Y. J. Kim Contrastive preference optimization: pushing the boundaries of LLM performance in machine translation. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §B.4.
  • Zhang et al. (2020) J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. External Links: Link Cited by: §4.1.
  • Zhang et al. (2019) L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma Be your own teacher: improve the performance of convolutional neural networks via self distillation. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 3712–3721. External Links: Document Cited by: §3.4.
  • Zhao et al. (2024) J. Zhao, Z. Guan, C. Xu, W. Zhao, and Y. Jiang SC2: towards enhancing content preservation and style consistency in long text style transfer. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 9949–9960. External Links: Link, Document Cited by: §3.1.

Appendix A Data Description

A.1 Reddit

We use Million User Dataset (MUD) Khan et al. (2021), a large-scale publicly available (Apache-2.0) user text dataset collected from the social media platform Reddit. The dataset comprises over 300 million Reddit posts produced by approximately one million users over the course of one year, and includes text-based user contributions in the form of comments. Evaluation is conducted on three predefined splits from Patel et al. (2024).

  • •

    Diverse: Source and target authors with posts on diverse topics across 13 or more different subreddits.

  • •

    Random: Random source and target authors.

  • •

    Single: All posts belong to a popular college football subreddit.

A.2 Blog

We use the Blog Authorship Corpus Schler et al. (2006) collected from blogger.com, which consists of blog posts written by 19,320 individual bloggers. This dataset is freely available for non-commercial research purposes. This dataset is available at https://www.kaggle.com/datasets/rtatman/blog-authorship-corpus/data

A.3 News

All-the-news dataset contains news articles collected from major U.S. and English-language news outlets. This dataset is available at https://huggingface.co/datasets/rjac/all-the-news-2-1-Component-one.

Since our objective is to analyze writing characteristics at the single-author level, we apply a series of filtering steps to ensure data quality. Specifically, we remove articles with missing author information, exclude articles attributed to organizations or non-individual entities, and discard articles with multiple authors. After filtering, only articles attributed to clearly identifiable individual authors are retained.

Dataset #samples #authors #parallel pairs
Reddit 7.5M 946K 200K
Blog 177K 17K 40K
News 538K 53K 200K
Table 7: Dataset statistics for training.

A.4 Stylistic Distance across Datasets

Table 8 reports intra-author style variation and inter-author distance for each dataset, providing supporting statistics for the performance differences observed across domains. Style variation is measured as the mean cosine distance of each author’s style embeddings to their centroid, averaged across authors. Inter-author distance is measured as the mean pairwise cosine distance between authors in the UAR space.

Dataset Style variation Inter-author distance
Reddit (Single) 0.407 0.310
Reddit (Random) 0.414 0.384
Reddit (Diverse) 0.398 0.357
Blog 0.316 0.322
News 0.230 0.232
Table 8: Style variation and inter-author distance across datasets. Higher values reflect greater intra-author stylistic variability and greater inter-author separation.

A.5 Cross-domain Inter-author Distance

Figure 5 shows the mean pairwise inter-author distances between source and target domain authors in the UAR space. Cross-domain distances are consistently larger than in-domain distances across all domain pairs. The largest differences relative to in-domain distance are observed for News to Reddit and Blog. Notably, since the Towards and Away metric is normalized by the source-target distance (Eq. 11), larger inter-author distances naturally yield lower Towards values regardless of model performance. Consequently, Towards and Joint are not directly comparable across domain pairs with different inter-author distances, and should be interpreted in terms of relative differences between models within the same domain pair.

Figure 5: Cross-domain inter-author distances measured in the UAR space. Bars show inter-author distance, with annotations indicating the ratio relative to the in-domain distance.

Appendix B Implementation Details

All training experiments are conducted on two NVIDIA A100 80GB GPUs. We use a batch size of 128 for all training stages. T5-large Raffel et al. (2020) is used as the base model for HyperStyler and all trainable baselines. For each method, we select the checkpoint with the lowest validation loss. At inference time, we use sampling with a temperature of 0.8 and top-pp to 1.0.

B.1 HyperStyler

For all attention layers in HyperStyler, we follow the T5 backbone by adopting a pre-norm structure with layer normalization and a multi-head decomposition Vaswani et al. (2017) with Nh=16N_{h}=16. We set d=ded=d_{e} for simplicity. For parameter efficiency, we share the projection matrices across all embedding tables.

B.1.1 Selection of Style Embedding Space

We adopt the STYLE embedder Wegmann et al. (2022) to ensure a fair comparison with TinyStyler Horvitz et al. (2024b) and ParaGuide Horvitz et al. (2024a), which also rely on it as the style conditioning signal. This embedder is trained via contrastive learning with negative samples from the same topic and domain, encouraging the model to capture subtle stylistic signals that distinguish authors within the same topic, yielding content-independent style representations. To empirically confirm this, we apply k-means clustering with Reddit dataset. Table 21 exhibits the resulting clusters, which are clear and interpretable stylistic patterns, confirming that the STYLE embedder captures meaningful stylistic features beyond content.

B.1.2 Selection of Activation Function for FFN Adapter

We apply an activation function to the low-rank adapter used for FFN modulation (Eq. 7). To examine whether an activation function is necessary and how the choice of activation function affects performance, we compare five variants: no activation, ReLU Nair and Hinton (2010), ELU Clevert et al. (2016), GELU Hendrycks and Gimpel (2016), and GeGLU Shazeer (2020). As shown in Table 9, the variant without an activation function exhibits relatively strong style transfer, but achieves the lowest overall performance due to lower semantic preservation. In contrast, variants using activation functions generally perform better than the no-activation variant. This suggests that an adapter with an activation function is more effective for balancing the trade-off between style fidelity and semantic preservation than a simple linear low-rank transformation.

Meanwhile, the performance differences among activation functions are relatively small. Therefore, we attribute the improvement primarily to the presence of an activation function rather than to any specific choice. Since GeGLU is used in the FFN of the underlying model (google/T5-v1.1-large) and achieves competitive performance, we adopt GeGLU in our model as well.

Activation Away Towards Sim Joint
w/o ACT 0.881 0.165 0.454 0.386
ReLU 0.819 0.148 0.565 0.415
ELU 0.821 0.150 0.565 0.415
GeLU 0.827 0.151 0.560 0.411
GeGLU 0.818 0.152 0.578 0.418
Table 9: Effect of activation function choice in the FFN adapter. ’w/o ACT’ denotes the setting without an activation function (no activation).
Configuration Stage 1 Stage 2 Stage 3
Learning rate 5​e−55e^{-5} 1​e−41e^{-4} 1​e−41e^{-4}
Batch size 128 128 128
Optimizer AdamW AdamW AdamW
Weight decay 0.01 0.01 0.01
Scheduler Constant Cosine Constant
Warm-up steps 2000 2000 5% of max steps
Max steps / epochs 200K steps 100K steps 3 epochs
Table 10: Training setup for HyperStyler.

B.2 TinyStyler

We follow the original paper’s configuration and use the provided training code Horvitz et al. (2024b). Only for the Reddit dataset, we use the publicly available checkpoint rather than training from scratch, as we found that training from scratch in our environment yielded lower performance than originally reported.

Configuration Value
Pretrained Ckpt google/t5-v1_1-large
Learning rate 1​e−51e^{-5}
Batch size 128
Optimizer Adam
Weight decay 0.01
Schedule Constant
Warm-up Steps 2000
Total Steps 150K
Table 11: Hyperparameters of TinyStyler.

B.3 StyleMC

Since the official source code for StyleMC Khan et al. (2024) is not publicly available, we implemented the method by strictly adhering to the descriptions provided in the original paper. While we made every effort to ensure a faithful reproduction, minor discrepancies may exist compared to the original implementation due to unspecified details of the algorithm or hyperparameters. For a fair comparison, we modified the baseline STYLEMC by replacing its original UAR-based implementation with the STYLE embedder Wegmann et al. (2022). An author embedding was then calculated via mean pooling, aligning it with the evaluation protocol used for other models.

Configuration Value
Learning rate 1​e−51e^{-5}
Batch size 128
Optimizer AdamW
Weight decay 0.01
Future discriminator Ckpt facebook/opt-1.3b
Proposal generator Ckpt google/t5-v1_1-large
Number of steps 80×\timesSequence length
αfluency\alpha_{\text{fluency}} 0.005
αstyle\alpha_{\text{style}} 1.0
αsemantic\alpha_{\text{semantic}} 1.0
αedit\alpha_{\text{edit}} 0.01
Table 12: Hyperparameters of StyleMC.

B.4 ASTRAPOP

We conduct experiments based on the ASTRAPOP framework Liu et al. (2024) with CPO Xu et al. (2024), its best-performing variant, while modifying several components to better align it with our experimental setting. Originally, ASTRAPOP uses LLaMA-2-7B, a decoder-only model, as its backbone architecture. To ensure a fair evaluation, we replace the backbone with T5-Large. This architectural change requires decisions on how the source and reference texts are arranged in the encoder input. We also explore the JOINT reward formulation, following the filtering criterion used in TinyStyler, to examine whether it provides a more effective training signal than the original reward. To identify the configuration that performs best under our setting, we evaluate four variants combining input order and reward function:

  • •

    ASTRAPOP follows the original input order and reward formulation.

  • •

    ASTRAPOPJOINT{{}_{\text{{JOINT}}}} adopts the JOINT reward formulation while preserving the original input order.

  • •

    ASTRAPOPreverse{{}_{\text{reverse}}} places the [src] token at the beginning of the input sequence, while the original reward formulation remains unchanged.

  • •

    ASTRAPOPreverse, Joint combines both the reversed input order and the JOINT reward formulation.

According to Table 22, no configuration performs consistently best across datasets. We therefore report the two variants that preserve the original input order in Table 1.

In addition, to examine whether the backbone replacement puts ASTRAPOP at a disadvantage, we train ASTRAPOP with its original LLaMA-2-7B backbone and compare it with HyperStyler. As shown in Table 23, despite a more than 9×\times difference in model size, HyperStyler consistently achieves higher Joint scores across three datasets.

Configuration SFT CPO
learning rate 5​e−55e^{-5} 1​e−51e^{-5}
batch size 128 128
Optimizer Adam Adam
# epochs 20 20
Max steps 100K 100K
β\beta – 0.1
top pp – 1.0
temperature – 0.8
length penalty α\alpha – 0.5
Context Size 512 512
Output Size 80 80
Table 13: Hyperparameters of ASTRAPOP.

B.5 Paraguide

We followed the experimental setup used in the original paper Horvitz et al. (2024a).

Configuration Value
Pretrained Ckpt xhan77/ssdklm
Learning rate 5×10−65\times 10^{-6}
Batch size 128
Optimizer AdamW
Weight decay 0.01
Schedule Constant
Warm-up Steps 2000
Total Steps 150K
Diffusion Steps 200
Context Size 80
Output Size 80
Table 14: Hyperparameters of Paraguide.

B.6 In-context Learning Methods

For GPT-based models and Llama-3.1, we use the prompt from Horvitz et al. (2024b), with default API settings for GPT-4-turbo (gpt-4-turbo-2024-04-09), GPT-5-mini (gpt-5-mini-2025-08-07), and GPT-5.4 (gpt-5.4-2026-03-05, medium). STYLL follows the original experimental setup Patel et al. (2024) using Qwen2.5-7B.

Appendix C Evaluation Details

C.1 Metric Formula Definition

We adopt the evaluation metrics proposed by Patel et al. (2024) for LAST. For any author aa, let PaP_{a} denote their set of 16 posts, and Ps→tP_{s\rightarrow t} denote the set of posts written by source author ss and style-transferred to target author tt. Let R→​(P)\vec{R}(P) denote a single UAR embedding produced over a set of posts PP. Finally, we define S⁡(𝐮,𝐯)S(\mathbf{u},\mathbf{v}) scaled to the range [0,1][0,1], given by S⁡(u→,v→)=sim⁡(u→,v→)+12S(\vec{u},\vec{v})=\frac{\mathrm{sim}(\vec{u},\vec{v})+1}{2}. We further define its complement as Sc​(u→,v→)=1−S⁡(u→,v→)S_{c}(\vec{u},\vec{v})=1-S(\vec{u},\vec{v}).

Away

measures how far a style-transferred text departs from the source author’s style:

min⁡(Sc​(R→​(Ps→t),R→​(Ps)),Sc​(R→​(Pt),R→​(Ps)))Sc​(R→​(Pt),R→​(Ps))\frac{\min\Big(S_{c}\big(\vec{R}(P_{s\to t}),\vec{R}(P_{s})\big),\,S_{c}\big(\vec{R}(P_{t}),\vec{R}(P_{s})\big)\Big)}{S_{c}\big(\vec{R}(P_{t}),\vec{R}(P_{s})\big)} (10)
Towards

measures how far a style-transferred text moves toward the target author’s style:

max⁡(S⁡(R→​(Ps→t),R→​(Pt))−S⁡(R→​(Ps),R→​(Pt)),0)Sc​(R→​(Ps),R→​(Pt))\frac{\max\Big(S\big(\vec{R}(P_{s\to t}),\vec{R}(P_{t})\big)-S\big(\vec{R}(P_{s}),\vec{R}(P_{t})\big),0\Big)}{S_{c}\big(\vec{R}(P_{s}),\vec{R}(P_{t})\big)} (11)
Sim

measures how well the transferred text preserves the meaning of the source text. The average Mutual Implication Score Babakov et al. (2022) between two sets of posts authored by aa and bb is denoted as MIS⁡(Pa,Pb)\mathrm{MIS}(P_{a},P_{b}):

max⁡(MIS⁡(Ps→t,Ps)−MIS⁡(Pt,Ps),0)1−MIS⁡(Pt,Ps)\frac{\max\Big(\mathrm{MIS}(P_{s\to t},P_{s})-\mathrm{MIS}(P_{t},P_{s}),0\Big)}{1-\mathrm{MIS}(P_{t},P_{s})} (12)

C.2 Details on Human Evaluation

We recruit annotators from Amazon Mechanical Turk, restricting participation to workers from English-speaking countries (i.e., US, UK, Canada, Australia) with a 95% or higher approval rating. As the evaluation of style transfer is known to be difficult for humans Krishna et al. (2020); Patel et al. (2024); Hallinan et al. (2023); Liu et al. (2024), we introduce a qualification test to ensure a minimum level of annotation quality. The test consists of three items. In each item, annotators are shown five reference texts from each of two randomly selected authors and asked to identify which author wrote a held-out target text. Only annotators who correctly answer all three items are admitted to the main evaluation.

For each baseline category, we select the best-performing model for the main evaluation. The evaluation is conducted on the same 100 source-target author pairs sampled from the Reddit test set, with three annotators assigned to each model output. We exclude examples whose source texts or target-author references contain violent, sexually explicit, or profane content to minimize annotator exposure to potentially harmful or offensive material. Annotators are also informed before the task that they may encounter potentially harmful or offensive content, and only those who agree to proceed participate in the evaluation. We pay 60 cents per annotated pair, corresponding to an estimated hourly rate based on the average completion time.

For style fidelity, annotators are shown eight reference texts from the target author, along with an anonymized source text and a transferred text in randomized order. They are then asked to select which text is more likely to have been written by the target author. The final label is determined by majority vote (Krippendorff’s α=0.10\alpha=0.10), with a score of 1 assigned when the transferred text is selected and 0 otherwise. For content similarity, annotators are shown the source text and the transferred text and asked to rate their semantic similarity on a 3-point Likert scale: 0 indicates Not Similar, 1 indicates Somewhat Similar, and 2 indicates Similar. The average score across annotators is used as the final content similarity score. Detailed instructions are shown in Tables 15 and 16. Our content similarity question and rating scale are adapted from Liu et al. (2024). We use McNemar’s test McNemar (1947) for style fidelity and paired t-test for content similarity to verify whether performance differences between models are statistically significant, with p<0.05p<0.05.

Instruction
Read the reference author’s writing samples, then decide which of the two texts is more likely written by that author based on writing style (sentence structure, word choice, tone — not topic).
Reference Author’s Writing Samples
[Writing samples are shown here]
Text A: [Text A is shown here]
Text B: [Text B is shown here]
Which text is more likely written by the Reference Author, based on writing style?
□\square Text A
□\square Text B
Table 15: Instruction for style fidelity evaluation.
Instruction
Read both texts and judge how similar they are. Focus on the core content and key information conveyed, not the writing style.
Text A: [Text A is shown here]
Text B: [Text B is shown here]
How similar are the two texts?
0 — Not Similar
Only small portions (less than 50%) of the passages are the same.
1 — Somewhat Similar
Large portions (50–75%) of the passages are the same, but there are significant sections that differ or are present in only one passage.
2 — Similar
Most of the content (75% or more) of the two passages is the same.
Table 16: Instruction for content similarity evaluation.

Appendix D Additional Experimental Results

D.1 Analysis on the Number of References

We analyze how performance varies with the number of references KK. As shown in Figure 6, HyperStyler shows a consistent and substantial performance advantage over TinyStyler from K=6K=6. Notably, HyperStyler with only K=9K=9 references surpasses TinyStyler’s best performance at K=16K=16. This suggests that HyperStyler uses the available references more effectively through context-aware style selection. These results indicate that the key factor is not simply the number of references, but how the model selects and uses stylistic evidence relevant to the source context.

Figure 6: Trend of Joint score across KK on the Reddit Random split, with shaded bands indicating standard error of the mean over five random samplings.
Method Time(s) VRAM(GiB)
In-context learning
STYLL (Qwen2.5-7B) 14.85 32.3
GPT-4 Turbo 2.01 -
GPT-5 Mini 5.14 -
GPT-5.4 8.50 -
Llama-3.1-8B-Instruct 1.83 30.8
Inference-time control
ParaGuideλ=200 20.93 3.40
ParaGuideλ=2500 20.86 3.40
StyleMC 49.51 9.29
Unsupervised alignment
ASTRAPOP 1.72 3.55
TinyStyler 0.82 3.17
TinyStylerRERANK(5){}_{\text{RERANK(5)}} 1.15 5.05
Proposed method r p
HyperStyler 1 1 0.94 3.80
HyperStylerRERANK(5){}_{\text{RERANK(5)}} 1 1 1.26 5.08
HyperStyler 8 5 1.01 3.80
HyperStylerRERANK(5){}_{\text{RERANK(5)}} 8 5 1.45 5.13
HyperStyler 32 5 1.03 3.85
HyperStylerRERANK(5){}_{\text{RERANK(5)}} 32 5 1.45 5.26
Table 17: Inference cost. r and p denote the adapter rank and prefix length, respectively. Time denotes the average inference time over 300 instances in the Reddit test dataset, and VRAM denotes peak memory usage.

D.2 Computational Cost Analysis

Table 17 reports inference latency and memory usage on the Reddit test set. Inference time is averaged over 300 instances, and VRAM is measured as peak FP32 memory usage for locally hosted models on a single NVIDIA A100 GPU. For API-based LLMs, we report wall-clock latency only, as server-side memory usage is not accessible. For reranking variants, the reported time includes both candidate generation and reranking.

Training stage GPUs Time
HyperStyler
Stage1 A100 80G x2 28h
Stage2 A100 80G x2 16h
Stage3 A100 80G x2 0.5h
TinyStyler
Training A100 80G x2 42h
Self-distillation A100 80G x2 4h
ASTRAPOP
SFT A100 80G x2 45h
CPO A100 80G x2 24h
Paraguide
Finetuning A100 80G x2 14h
StyleMC
Future regressor A100 80G x2 18h
Table 18: Training cost on the Reddit dataset under each method’s training configuration.

Table 18 reports training times on the Reddit dataset. These times characterize the practical computational cost under our experimental setup rather than provide a strictly controlled comparison of training efficiency, since methods differ in training objectives, optimization hyperparameters, and training schedules. We omit in-context learning baselines because they do not require task-specific training. HyperStyler takes 28h, 16h, and 0.5h for its three stages, respectively. Its final self-distillation stage uses a filtered set of roughly 40K instances, similar to TinyStyler’s, but requires less training time under our configuration.

D.3 Qualitative Analysis

Target-dependent style transfer.

Figure 8 visualizes t-SNE projections of style embeddings for the same source texts transferred from source author A to two target authors B and C. HyperStyler’s A→\rightarrowB and A→\rightarrowC outputs occupy distinct stylistic regions and are more closely aligned with the corresponding target author’s style variation. In contrast, TinyStyler and ParaGuide, which rely on static author embeddings, tend to concentrate in a particular stylistic region rather than aligning with the target-author references. Moreover, baselines that receive all reference texts as input also show weaker target-wise separation or weaker alignment with the corresponding target author’s style. These results suggest that HyperStyler’s explicit style navigation enables target-dependent style transfer, producing outputs that reflect the distinct stylistic characteristics of each target author.

Table 20 presents examples where each source text is rewritten using two different target-author reference sets. Source 1 is an argumentative reply. For Target author A, the output opens with a question that reflects the question-based style observed in the references (What is your point?), then restates the original advice in a more explicit form. For Target author B, the output stays relatively close to the source wording while reflecting the ellipsis usage observed in the references (personally….). Source 2 is a reassurance-oriented comment. Target author A makes the response warmer and more supportive (I’m glad you found it, :D), whereas Target author B keeps it concise and neutral without an emotive marker. Source 3 is an event recap. For Target author A, the output adopts a more casual, punctuation-heavy recap style (yep.., !!). For Target author B, the output becomes a more straightforward recap, consistent with the more explanatory phrasing observed in the references. These examples suggest that HyperStyler can realize target-author-specific style cues differently while preserving the source context.

Case 1: content omission under compressive style
Source: Yeah, the difference between the two ranks is pretty minimal. I mentioned a few weeks back that I didn’t think we’d get a true gauge on Michigan State until the Notre Dame game … I watched that entire game …
Output: i didn’t think we’d get a true gauge on Michigan State until the Notre Dame game, but i definitely stand by that after the Furman game lmao
Case 2: over-generation under elaborative style
Source: Which is why I don’t respect early season rankings … AT ALL
Output: I’m not a big fan of early season rankings, but I’m glad someone is enjoying the game. AJ Green is a good player
Table 19: Representative failure cases.
Failure mode.

We identified two representative failure modes, shown in Table 19, that arise when the information density of the source text is misaligned with the stylistic signals provided by the reference set. The first is content omission under compressive target style, which occurs when the source text is long and information-dense while the reference set reflects a short, reaction-oriented conversational style. In such cases, HyperStyler tends to compress the source text, retaining the main stance but omitting secondary propositions and supporting details. The second is over-generation under elaborative target style, which occurs when the source text is short and self-contained while the reference set reflects a more expressive and elaborative style. In such cases, the model tends to expand the output to realize the target style, introducing content that is not grounded in the source text and potentially leading to semantic drift. These cases suggest that authorship style transfer becomes particularly challenging when the amount of information that must be preserved from the source conflicts with the degree of compression or elaboration implied by the target reference style.

D.4 Analysis on Style-dependent and Layer-wise Modulation

Stylo-hypernet modulates each decoder layer via a bilinear interaction between the style coordinate and learnable layer embeddings (Eq. 4), yielding compatibility scores for each modulation target. We examine whether these scores vary across styles and layers. Using the k-means clustering results in Section B.1.1, we select the 50 sentences closest to each cluster centroid. Each sentence is fed into the Stylo-hypernet to obtain its compatibility scores bj(h)b_{j}^{(h)}, which are averaged within each cluster. The resulting scores are normalized per layer by the maximum absolute value across clusters. As shown in Figure 7, within a given layer, scores differ across style clusters. Within a given cluster, the scores also change from layer to layer. These results confirm that Stylo-hypernet assigns distinct scores across styles and layers, as intended by its design.

Refer to caption
Figure 7: Compatibility scores bj(1)b_{j}^{(1)} for the first head across 50 style clusters and decoder layers. Top: cross-attention prefixes (PKP_{K}, PVP_{V}, first prefix position). Bottom: FFN low-rank projections (Wd​o​w​nW_{down}, Wu​pW_{up}).

Appendix E Licenses and Use of Artifacts

Table 27 lists artifacts used in this work, including models for training and evaluation and software libraries, with their licenses and links to the sources. Our use of the artifacts is consistent with their licenses and intended use. In particular, artifacts released under permissive licenses are used for research in accordance with their terms, and the artifact licensed under CC BY-NC-SA 4.0 is used only in a non-commercial research context. We do not redistribute any artifact in a manner inconsistent with its original license or access conditions, and any outputs or derived materials from this work are intended only for research use.

Figure 8: t-SNE visualization of style-transferred outputs across methods. Each panel shows how the stylistic distribution of transferred outputs shifts from the source author texts toward two target authors’ styles (Target author B and Target author C), where ellipses indicate the approximate stylistic distribution boundary of each group.
Target author references HyperStyler output
Source 1: Why do you need to address him personally at all? Attack the argument, not the person.
Target author A: Why not? Im not a Nasus main but Rylai procs on his ult and E.; Uhm hello what year is this? Even if hes in a straight relationship, he shouldnt do that.; Can u explain the reference?; DAE also think masc is the most opressed!!!?; TIL the internet is one person What is your point? It is an argument, not a person. Attack the argument not the person
Target author B: Big difference between being a fanboi and simply not falling inline with the circlejerk….; Don’t worry /r/cars tells them everyday….; For $5,000 all you should expect is a pile of rusty parts…; Wouldn’t that be the Skrillex cut?; This thread officially proves we are no better than the YouTube comments section Why are you even addressing him personally…. Attack the argument not the person.
Source 2: No prob! Easy to make the mistake. Not like Reddit is graded or anything. :)
Target author A: I guess I’m building. Thank you for your help! I really appreciate it!; You guys have been really polite, I’m really surprised! Thank you for all your help so far!; Thank you for your time, looks like I’m building it! c:; Thank you for your help!; Oh, I’d rather get a new case then. […] Should I just throw it away? I’m glad you found it. It’s easy to make a mistake, like Reddit isn’t graded! :D
Target author B: My two cents: If you are asking technical questions then yes definitely. […]; Hey someone photoshopped out the tea label!; I bet that the bot power is conducted through the floor.; Or poke around on the ground; That sucks. I thought dog treats were like essential kit […] Nice. Easy to make the mistake. Like reddit isn’t graded or anything.
Source 3: Approx 35 people attended, which meant there was more than enough beer to go around.
Target author A: and that’s a record .. i had a temporary vacation from blogging because of my preliminary exams earlier this week!!; i didnt take it badly though.. because he reasoned out that there was traffic along e. rod or something…; i was sarcastic.. coz i was reviewing.. […] luv u!!!; if there’s one way i would like to express my thoughts and ideas.. it is through writing!!; ahhh.. naaalala ko na!! yep.. there was more than enough beer to go around, because 35 people attended!!
Target author B: But they’re having a party on the 17th - not massive, only 50 people - so everybody can come over and eat lots.; I’ll keep ya looped.; Well, I’m free for 10 whole weeks.; Met some new people.; Well, I’ve been on Prophet’s Inc. and Fantasy Essentials - my favourite forums - for the last few hours, chatting to random people. It was a good night– 35 people attended and there was more than enough beer to go around.
Table 20: Illustrative examples showing that HyperStyler generates different outputs depending on the target-author references given the same source text.
Cluster Stylistic description Representative examples
1 First-person informal narrative reply. These instances are characterized by first-person narration, informal register, and anecdotal narrative structure.
“I certainly did! […] I was 37 at the time […] haha”
“That was my sibling, haha! Siblings are cruel ;_;”
“As I was going through your list, I was thinking ‘Ha! I know all these guys!’ […]”
2 Brief appreciative reaction with emoticons. These instances are short acknowledgment responses marked by appreciation, positive evaluation, and emoticon usage.
“It did indeed work, thank you so much :D”
“I actually lost the little plastic bit on the end, but I definitely will glue it back in, thanks :)”
“I have no funny / awesome screenshots to share, but hey, maybe if I win I can take some ;D”
6 Terse information-seeking question. These instances are short interrogative utterances that function primarily as follow-up requests for information.
“What about the baby versions?”
“When did you get it?”
“What happens when you break it?”
11 Extended clause-heavy justification reply. These instances are characterized by long, multi-clause constructions that build toward a conclusion through chained reasoning, conditional framing, or supporting elaboration.
“There is virtually no chance of getting them both. If it were possible, of course
it’s a great strategy […]”
“For sure. People look more into where a player was drafted than their actual skill […]”
“The fact that the handle is his name has a lot to do with it […]”
14 Extended conversational reply. These instances are long, chatty responses combining informal register, colloquial markers, and loosely connected clauses.
“I love the song a lot haha. You should check out […]”
“Lol they were out of stock pretty much as soon as that price was up. […]”
“Yeah I was so hype when Honedge was announced […]”
15 Formulaic endorsement. These instances are short evaluative formulas used to express approval, endorsement, or visibility boosting.
“Upvote for the title!”
“This needs to be higher up in the comments!”
“Well worth the hike!”
16 Directive second-person reply. These instances express second-person guidance through imperative constructions or modal advisory forms directed at the interlocutor.
“Google a video on how to make it.”
“You should put her on the side then.”
“You should post pictures sometime of the village if you have any.”
18 Segmented multi-line commentary. These instances are structured as short line-separated discourse units, often juxtaposing multiple evaluative or explanatory statements.
“A good landing is one where no-one gets hurt. […]”
“Symbols are dangerous. I love my heritage. […]”
“Wow. […] The type of video that just leaves you speechless. […]”
23 Blunt categorical assertion. These instances express direct and compact judgments or corrections in forceful declarative form.
“It’s blue because he’s cold”
“If all that’s on your resume you’ll be fine”
“If it’s just the glass on top it’s cheap”
30 Quote-and-correct reply structure. These instances exhibit a quote-response structure in which quoted content is followed by contradiction, correction, or practical follow-up.
“> Have had our house professionally cleaned and it still smells […]
Wash the walls with diluted vinegar.”
“> and that’s a man with no engineering or mechanical educational background […]
Education is completely irrelevant here.”
“> bring em and register […] There is no legal obligation to register your firearms […]”
32 Rhetorical second-person interrogative reply. These instances take interrogative form with ironic or sarcastic phrasing directed at the interlocutor, functioning as implicit challenge rather than genuine information-seeking.
“Don’t you think that’s setting the bar a little high?”
“You sure you aren’t just hearing Hanley’s bat?”
“Are you sure your teacher isn’t Dwight Schrute?”
41 Punctuation-heavy expressive reply. These instances are characterized by repeated or emphatic punctuation and expressive phrasing, often conveying heightened emotional intensity.
“it is my top 2 as well!! fantastic story, great visuals and amazing characters!!”
“i was lucky enough to be there! […] and it was FANTASTIC!!!”
“we will make it and have a roo meet up […] BEAT THE CAPS!!”
45 Emphatic congratulatory/supportive reaction. These instances convey overt positive evaluation through repeated exclamation marks, congratulatory formulas, encouragement, and emoticons.
“Amazing!! Really beautiful gift :)”
“You can do it!! Good luck =)”
“Awww!! So sweet!! :) Congrats on the wedding!!”
47 Ellipsis-heavy hesitant reply. These instances are marked by repeated ellipses and loosely connected clauses, producing a hesitant and trailing conversational rhythm.
“that… that man has had some BAD food…”
“[…] our humanity and kindness can circumvent that…”
“this… this perfectly describes the whole datamining standpoint…”
48 Extended formal expository prose. These instances are characterized by long-form prose with formal register, technical or domain-specific vocabulary, and dense information structure.
“The events are in-game ones which cause unusual monster spawns, unusual numbers
of monsters, or give monsters new abilities.”
“Traditional IRA contributions are an above-the-line deduction that happens
before itemized or standard deductions […]”
“Finding the minimum/maximum is all about finding out the points at which
the derivative of the function is 0 […]”
Table 21: Selected clusters from a k-means clustering (k=50) in the STYLE embedding space. Sentences in each cluster group share distinct stylistic characteristics, illustrating that the STYLE embedder captures interpretable stylistic patterns independently of topic and content.
Method Reddit Blog News
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
ASTRAPOP 0.578 0.027 0.728 0.171 0.997 0.255 0.014 0.060 0.840 0.060 0.170 0.139
ASTRAPOPJOINT{}_{\text{JOINT}} 0.620 0.029 0.695 0.173 0.840 0.070 0.171 0.139 0.576 0.082 0.713 0.319
ASTRAPOPreverse{}_{\text{reverse}} 0.571 0.026 0.743 0.172 0.991 0.275 0.077 0.170 0.655 0.086 0.394 0.248
ASTRAPOPreverse,JOINT{}_{\text{reverse,JOINT}} 0.596 0.026 0.725 0.172 0.988 0.243 0.125 0.213 0.950 0.005 0.202 0.115
Table 22: Performance comparison results on the configuration of ASTRAPOP.
Method Reddit Blog News
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
ASTRAPOP (LLaMA-2-7B) 0.708 0.188 0.505 0.333 0.813 0.244 0.569 0.479 0.799 0.110 0.559 0.322
HyperStyler (T5-large) 0.818 0.152 0.578 0.418 0.731 0.183 0.701 0.489 0.571 0.098 0.678 0.370
Table 23: Comparison between ASTRAPOP with its original backbone and HyperStyler.
Method Diverse Random Single
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
STYLL(Qwen2.5-7B) 0.797 0.069 0.405 0.200 0.739 0.074 0.416 0.240 0.896 0.047 0.463 0.184
gpt-4-turbo-2024-04-09 0.832 0.073 0.683 0.296 0.760 0.087 0.706 0.332 0.850 0.083 0.715 0.313
gpt-5-mini-2025-08-07 0.861 0.082 0.718 0.321 0.805 0.080 0.736 0.330 0.902 0.083 0.729 0.346
gpt-5.4-2026-03-05 (medium) 0.940 0.077 0.501 0.238 0.866 0.139 0.679 0.436 0.948 0.135 0.610 0.402
meta-llama/Llama-3.1-8B-Instruct 0.724 0.119 0.602 0.371 0.743 0.127 0.564 0.369 0.800 0.158 0.595 0.431
ParaGuideλ=200 0.774 0.056 0.544 0.221 0.696 0.047 0.585 0.222 0.818 0.057 0.664 0.263
ParaGuideλ=2500 0.859 0.078 0.381 0.239 0.801 0.065 0.456 0.246 0.900 0.058 0.512 0.235
StyleMC 0.603 0.051 0.462 0.189 0.625 0.039 0.453 0.173 0.746 0.017 0.435 0.100
ASTRAPOP 0.612 0.031 0.679 0.168 0.516 0.027 0.736 0.192 0.607 0.022 0.770 0.152
ASTRAPOPJOINT{}_{\text{JOINT}} 0.656 0.035 0.645 0.168 0.550 0.029 0.702 0.205 0.653 0.022 0.739 0.145
TinyStylerREC{}_{\text{REC}} 0.897 0.130 0.306 0.282 0.863 0.154 0.314 0.315 0.932 0.148 0.436 0.371
TinyStylerREC,RERANK(5){}_{\text{REC,RERANK(5)}} 0.883 0.127 0.462 0.346 0.854 0.148 0.465 0.384 0.927 0.149 0.592 0.432
TinyStyler 0.836 0.109 0.582 0.356 0.835 0.127 0.590 0.396 0.910 0.130 0.706 0.445
TinyStylerRERANK(5){}_{\text{RERANK(5)}} 0.831 0.111 0.693 0.393 0.839 0.125 0.705 0.434 0.907 0.131 0.793 0.480
HyperStylerREC{}_{\text{REC}} 0.810 0.162 0.426 0.368 0.748 0.163 0.468 0.384 0.857 0.143 0.538 0.399
HyperStylerREC,RERANK(5){}_{\text{REC,RERANK(5)}} 0.803 0.159 0.650 0.449 0.746 0.159 0.709 0.472 0.849 0.138 0.765 0.470
HyperStyler 0.813 0.153 0.547 0.402 0.768 0.157 0.542 0.410 0.872 0.145 0.645 0.443
HyperStylerRERANK(5){}_{\text{RERANK(5)}} 0.807 0.149 0.760 0.467 0.766 0.148 0.768 0.480 0.871 0.145 0.844 0.508
Table 24: Performance comparison results on Reddit (Diverse / Random / Single) splits.
Model Diverse Random Single
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
HyperStyler 0.813 0.153 0.547 0.402 0.768 0.157 0.542 0.410 0.872 0.145 0.645 0.443
w/o Stylo-navigator (Mean-pooling) 0.776 0.088 0.677 0.332 0.732 0.112 0.666 0.389 0.842 0.097 0.775 0.385
w/o Stylo-navigator (Implicit selection) 0.770 0.102 0.641 0.348 0.729 0.124 0.627 0.392 0.833 0.116 0.746 0.412
w/o Stylo-hypernet (Global) 0.985 0.009 0.142 0.015 0.986 0.008 0.150 0.033 0.999 0.000 0.202 0.002
w/o Stylo-hypernet (Layer-wise) 0.789 0.119 0.597 0.374 0.734 0.123 0.595 0.388 0.852 0.121 0.697 0.421
w/o adapter in FFN 0.794 0.129 0.579 0.385 0.747 0.139 0.565 0.400 0.859 0.134 0.681 0.443
w/o prefix in CrossAttn 0.830 0.150 0.521 0.396 0.768 0.156 0.512 0.399 0.877 0.142 0.621 0.430
w/ prefix in SelfAttn 0.958 0.022 0.416 0.094 0.960 0.016 0.434 0.081 0.990 0.010 0.529 0.071
w/o predicted zz in stage 3 (mean-pooling) 0.809 0.143 0.542 0.399 0.768 0.157 0.546 0.407 0.878 0.133 0.654 0.421
Underlying paraphraser 0.889 0.019 0.682 0.123 0.849 0.012 0.696 0.084 0.949 0.007 0.776 0.058
Table 25: Ablation study results on Reddit (Diverse / Random / Single) splits.
Rank Prefix Diverse Random Single
Away Towards Sim Joint Away Towards Sim Joint Away Towards Sim Joint
8 - 0.814 0.141 0.546 0.396 0.760 0.149 0.531 0.400 0.870 0.141 0.641 0.438
8 5 0.819 0.147 0.551 0.400 0.763 0.154 0.540 0.403 0.870 0.143 0.648 0.440
128 - 0.816 0.154 0.516 0.401 0.781 0.164 0.500 0.401 0.875 0.143 0.618 0.429
128 5 0.822 0.156 0.524 0.403 0.770 0.164 0.528 0.411 0.877 0.148 0.627 0.440
- 1 0.792 0.125 0.600 0.391 0.741 0.130 0.591 0.390 0.853 0.128 0.701 0.437
32 1 0.818 0.145 0.538 0.396 0.764 0.150 0.536 0.397 0.871 0.145 0.640 0.438
- 10 0.807 0.132 0.573 0.381 0.752 0.138 0.559 0.398 0.864 0.141 0.674 0.447
32 10 0.817 0.149 0.528 0.399 0.772 0.163 0.523 0.409 0.877 0.147 0.632 0.438
Table 26: Hyperparameter study results on Reddit (Diverse / Random / Single) splits.
Type Artifact License Link
Model UAR embedding Apache-2.0 https://huggingface.co/rrivera1849/LUAR-MUD
STYLE embedding MIT https://huggingface.co/AnnaWegmann/Style-Embedding
Paraphrasing PEGASUS Apache-2.0 https://huggingface.co/tuner007/pegasus_paraphrase
Mutual Implication Score CC BY-NC-SA 4.0 https://github.com/s-nlp/mutual_implication_score
T5-large Apache-2.0 https://huggingface.co/google/t5-v1_1-large
Software HuggingFace Transformers Apache-2.0 https://github.com/huggingface/transformers
Accelerate Apache-2.0 https://github.com/huggingface/accelerate
Scikit-learn BSD-3-Clause https://scikit-learn.org/stable/
NLTK Apache-2.0 https://www.nltk.org/
Matplotlib PSF https://matplotlib.org/stable/project/license.html
PyTorch License link https://github.com/pytorch/pytorch
Table 27: Used artifacts and their licenses and links. All artifacts are used consistent with their intended use.