arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.02379v1 [cs.CL] 02 Sep 2026

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

Matteo Greco Affiliation: School of Computing and Information Systems, The University of Melbourne, Australia    Anudeex Shetty    Andrea Tagarelli Affiliation: DIMES Dept., University of Calabria, Italy    Jey Han Lau Email: matteogre01@gmail.com,{anudeex, laujh}@unimelb.edu.au, tagarelli@dimes.unical.it Affiliation: School of Computing and Information Systems, The University of Melbourne, Australia
Abstract

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 5959K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods.11 1 The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

**footnotetext: Equal contributions.

1 Introduction

Large language models (LLMs) can now generate high-quality text across multiple languages (Qin et al., 2024; Yue et al., 2025), achieving levels of fluency and coherence comparable to human-written content (Jakesch et al., 2023). These advances create new opportunities for communication, creativity, and productivity (Ahuja et al., 2023; Doddapaneni et al., 2025). Still, they also raise concerns about transparency, accountability, misuse, and the authenticity of digital content (Brundage et al., 2024; Rudolph et al., 2023; Russell et al., 2026). These risks motivate the need for robust methods to detect and attribute LLM-generated text (Huang et al., 2025; Wu et al., 2025).

Although prior research has explored LLM-generated text detection and numerous datasets have been introduced, several limitations remain as outlined in Table 1. First, most formulate detection as a binary human-versus-AI problem (He et al., 2024; Dugan et al., 2024; Thai et al., 2026), while only a limited number of works study the more challenging authorship attribution (AA) task (Uchendu et al., 2021; La Cava and Tagarelli, 2025; Shetty et al., 2026), where the objective is to identify the specific generator LLM that wrote a text. Second, existing datasets predominantly focus on short-form texts (Macko et al., 2025; Macko et al., 2023), despite the emergence of LLM-ghostwritten books (Australian Society of Authors, 2025). Even recent long-form efforts, such as Shetty et al. (2026), remain monolingual or limited to a small number of languages (Dugan et al., 2024; Tao et al., 2024). Third, methods are rarely evaluated for multiple distribution shifts, particularly cross-language transfer or generalisation to unseen generators and domains.

In this paper, we introduce MultiGhostBench, a multilingual benchmark for LLM AA that evaluates generalisation under domain, generator, and language shifts. To the best of our knowledge, MultiGhostBench is the first benchmark designed to jointly evaluate multilingual long-form LLM AA under multiple distribution shifts. To summarise, our main contributions are:

  • •

    We introduce MultiGhostBench, comprising 928928 books (average 5959K words) generated by five recent LLMs. It covers six languages and three scripts, and is designed to evaluate AA both within and across languages under domain and generator shifts.

  • •

    We conduct a comprehensive evaluation of statistical, supervised, and fingerprint-based attribution detectors, finding that no method performs consistently best across languages, data regimes, and distribution shifts.

  • •

    Our analysis reveals that Transformer-based detectors can retain generator-related information across languages, although the extent of this transfer varies across language pairs, whereas statistical and fingerprint-based approaches are more strongly affected by language.

2 Related Work

Authorship Attribution.

Research on LLM-generated text has predominantly focused on binary detection (Mitchell et al., 2023; Yang et al., 2024; Thai et al., 2026), with only a few works (Uchendu et al., 2021; La Cava and Tagarelli, 2025; Shetty et al., 2026) studying the AA task, where the goal is to identify the specific LLM that generated a text. However, these studies are predominantly based on supervised methods with limited generalisation. Shetty et al. (2026) studied long-form texts and robustness under domain and unseen-generator shifts, but their study remained restricted to English. They also developed a lightweight fingerprint-based method, trace, for the AA task.

Multilingual Authorship Attribution.

Similarly, research on multilingual LLM-generated text has primarily focused on binary detection (Wu et al., 2026; Macko and Kopál, 2026; Wang et al., 2024) rather than attribution. MULTITuDE (Macko et al., 2023) evaluates cross-language attribution performance, but its evaluation is not exhaustive, as it does not cover all languages in the dataset. Likewise, attribution experiments in M4GT-Bench (Wang et al., 2024) focus on cross-domain rather than cross-lingual generalisation. More recently, La Cava et al. (2026) provide a systematic study of multilingual attribution using MULTITuDE and MultiSocial. However, their primary setting is based on short news articles. This misses practical scenarios involving unseen generators and joint evaluation across multiple dimensions as done in this work, which we summarise in Table 1.

Dataset LLMs OOD # Lang Long? Avg. Num
# Rec. D A L Words Docs
TuringBench(2021) 2020 ✗ ✗ ✗ ✗ 11 ✗ <<200200 160160K
OpenTuring(2025) 77 ✗ ✓ ✓ ✓ 11 ✗ <<500500 497497K
GhostWriteB.(2026) 1010 ✓ ✓ ✓ ✗ 11 ✓ 5353K 325325
M4GT-Bench(2024) 66 ✗ ✓ ✗ ✗ 1212 ✗ <<500500 8787K
MULTITuDE(2023) 88 ✗ ✗ ✗ ✓ 1111 ✗ <<512512 7474K
MultiSocial(2025) 88 ✗ ✗ ✗ ✓ 2222 ✗ <<200200 472472K
MultiGhostBench 55 ✓ ✓ ✓ ✓ 66 ✓ 5959K 928928
Table 1: Overview of existing datasets for LLM-generated text attribution. ‘Rec.’ indicates whether benchmark includes recent high-capability LLMs; prior benchmarks include small-scale models (<<1010B parameters). ‘D’, ‘A’, and ‘L’ denote OOD-Domain, OOD-Author, and OOD-Language, respectively. ‘Long?’ denotes whether documents longer than 1010K words.

3 Dataset

Our proposed MultiGhostBench is a multilingual dataset of long-form texts generated by five recent LLMs (gemini-pro, gemini-flash, deepseek-v3.2, qwen3-235b, and gpt-oss; more details in Appendix Table 4) covering six languages and supporting AA under domain, author, and language shifts. A unique feature of our dataset is its language-shift setting, which enables the evaluation of whether attribution methods trained on one language can generalise to texts written in a previously unseen language.

3.1 Language Selection

MultiGhostBench covers six languages: Italian (IT), Spanish (ES), German (DE), English (EN), Chinese (ZH), and Russian (RU). These languages represent four language families: Romance (IT and ES), Germanic (DE and EN), Sino-Tibetan (ZH), and Slavic (RU). They also cover multiple writing systems, including Latin, Chinese characters, and Cyrillic scripts as well as language-specific orthographic conventions. This selection enables evaluation of AA under both closely related and more distant language and script shifts.

3.2 Multilingual Book Generation Pipeline

Human book writing is generally hierarchical: authors typically develop a book through multiple stages, such as planning, drafting, and revision Sun et al. (2022); Lee et al. (2025). We follow the generation process introduced by Shetty et al. (2026) that simulates this multi-stage writing process to create the LLM-generated books.

The book generation pipeline begins with the definition of a set of constraints, such as genres and time period. These constraints are sourced from Project Gutenberg22 2 https://www.gutenberg.org/, an online library of freely available books; see Appendix Table 3 for the list of genres, comprising both fiction and non-fiction genres. A Gutenberg book is usually multi-genre (e.g., LF (Literature & Fiction) ++ HB (History & Biographies)), therefore books in MultiGhostBench are also multi-genre. Each book is expanded iteratively by generating each subsequent segment conditioned on the outline, the previous segment, and a running summary of the narrative, i.e., updated after each generation step. This structure allows models to preserve global coherence while producing texts that exceed the length of a single generation step. We adapt the original prompt templates from Shetty et al. (2026) for the multilingual setting by translating them into each target language. The translations were verified by respective native-speaker volunteers; prompts can be found in Appendix A.1.

We provide model details for five LLMs and quantitative statistics for these LLM-generated books in Appendix Table 4 and 7, respectively. All the books were generated consistently using the same pipeline described above to minimise pipeline-specific bias. However, some LLM-specific generation bias is inevitable, which we try to reduce through our meticulous data cleaning (§3.4) and analysis (§3.5).

Lang. # Books Avg. #
Train Dev. Test Total Words
EN 9292 1010 6161 163163 64.864.8K
DE 8888 1010 5656 154154 44.444.4K
IT 9393 1010 5858 161161 50.250.2K
ES 9393 1010 6060 163163 44.944.9K
RU 7979 1010 5252 141141 83.683.6K
ZH 8282 1010 5454 146146 69.769.7K
Total 527527 6060 341341 928928 59.659.6K
Table 2: MultiGhostBench statistics: number of training, development, and test books, together with the average book length (in words) for each language. See Appendix Table 5 for detailed breakdown per LLM.

3.3 OOD Dimensions

Following Shetty et al. (2026), MultiGhostBench supports two OOD evaluation dimensions: OOD-Domain, which measures generalisation to unseen genres, and OOD-Author, which measures generalisation to unseen generators. We further introduce OOD-Language, which evaluates whether AA methods transfer to texts written in languages not observed during training.

In OOD-Domain, genres are partitioned into ID and OOD subsets for each generator LLM, ensuring that no genre appears in both partitions. Recall that these genres are sourced from the Gutenberg corpus comprising both fiction and non-fiction genres. Full statistics for genres in ID and OOD-Domain for different LLMs are provided in Appendix Table 5. In OOD-Author, we adopt a leave-one-author-out protocol: one LLM is held out at a time, methods are trained on the remaining LLMs, and evaluation is performed on books generated by the held-out model. In OOD-Language, methods are trained on one source language and evaluated on a different target language. We do not combine OOD-Language with OOD-Domain, so that the evaluation isolates the effect of language shift rather than conflating it with genre shift.

3.4 Data Cleaning

We clean the generated texts to remove generation and encoding artefacts while preserving language-specific features. Invalid control characters, markup, and non-textual symbols are removed, while formatting elements are normalised, preserving language-specific punctuation, diacritics, and non-Latin scripts. Following Shetty et al. (2026), we remove 1.51.5K tokens from both the beginning and end of each book to reduce shortcut cues from identifier markers, such as author names, titles, or publication years. As a final sanity check, we use GlotLID (Kargaran et al., 2023), an open-source language identification model, to verify that each generated book is written in its intended target language; all books in MultiGhostBench match their intended target language.

3.5 Data Exploration

Table 2 shows a summary statistics for MultiGhostBench, including the number of books per language across different splits. To provide a consistent estimate of document length across languages, we tokenise each book using a language-specific tokeniser (listed in Appendix Table 6) and convert the token counts to approximate word counts using the standard 1.5 tokens-per-word conversion.

Quantitative analysis.

The generated books are analysed using multiple quantitative textual metrics in Appendix Table 7. We compute perplexity (PPL) as a model-based measure of text predictability (Jurafsky and Martin, 2026), n-gram-based metrics for lexical overlap and diversity (Self-BLEU (S-B), self-repetition (Self-R), and nn-gram diversity (NGD)) (Zhu et al., 2018; Salkar et al., 2022; Meister et al., 2023), and compression ratio as a proxy for textual redundancy (Shaib et al., 2025). Given the multilingual nature of MultiGhostBench, token-based metrics are computed using language-dependent tokenisers (Appendix Table 6).

For comparison, we select human-written books from the Gutenberg corpus with similar properties (full-length, single-authored works from the nineteenth or twentieth century). Overall, across languages, the generated books exhibit textual properties comparable to human-written books, particularly in lexical diversity and redundancy. However, generated books show higher S-B scores and lower PPL, indicating greater inter-book lexical similarity and higher predictability for the reference model (i.e., mGPT). For Chinese and Russian, no eligible human references were available, hence they were compared with LLM-generated books in other languages. Their metric values remain within a similar range, with some variation for Russian PPL and Chinese NGD.

Human evaluation.

We conduct a small-scale human evaluation on a subset of generated books as an initial validation step. Our focus is on novels generated by gemini-pro in the LF (Literature & Fiction) genre. Following prior works (Chhun et al., 2022; Wang et al., 2025; Shetty et al., 2026), the generated books are evaluated along several dimensions that capture both local linguistic quality and broader narrative structure (more in Appendix A.3).

Appendix Table 8 summarises the annotator scores across passage-level and overall book-level dimensions. Overall, the results indicate that the generated books are generally fluent and coherent with strong passage-level scores, while narrative-level quality is less consistent across languages for subjective dimensions. Although this evaluation is limited to gemini-pro-generated literature books, the findings complement our analyses above, which show similar quality trends across LLMs. Together, these results support the overall quality of books in MultiGhostBench.

4 Authorship Attribution Methods

The selected methods span three categories: metric-based, model-based supervised, and fingerprint-based. Their multilingual adaptations are summarised below, with further detector and implementation details in Appendix B and C, respectively.

Metric-based methods.

For this category, we consider rank, entropy, and gLTR (Gehrmann et al., 2019). These methods rely on token-level probabilities produced by a reference language model. Following Macko et al. (2023), we use mGPT (Shliazhko et al., 2024) as the multilingual reference model. We adapt their extracted statistics to multi-class attribution following La Cava and Tagarelli (2025); Shetty et al. (2026).

Model-based supervised methods.

We consider n-gram (Koppel and Schler, 2004), bert-aa (Fabien et al., 2020), and detective (Guo et al., 2024), covering both non-neural and neural detectors. For n-gram, we adapt preprocessing to preserve accented and language-specific characters and use Jieba33 3 https://pypi.org/project/jieba/ segmentation for Chinese word-level features. For bert-aa, we use xlm-roberta (Conneau et al., 2020) as the multilingual encoder. For detective, we replace the original sentence-embedding model with a multilingual counterpart.

Fingerprint-based methods.

We consider trace (Shetty et al., 2026), a post-hoc fingerprint-based AA method that models transitions between token-rank or entropy statistics computed by an evaluator language model. A test text is attributed to the generator with the most similar fingerprint. We evaluate tracerank-js{}_{\texttt{rank-js}}, traceentr-js{}_{\texttt{entr-js}}, and traceentr-norm{}_{\texttt{entr-norm}}. For multilingual evaluation, we replace the original evaluator, GPT-2, with Gemma, also evaluated by Shetty et al. (2026). Previous trace evaluations were limited to English.

Refer to caption
Figure 1: Best-performing detector by language and evaluation setting under the low- and high-resource regimes. Each group includes ID, OOD-Domain, and OOD-Author. Bar height shows macro-F1F_{1}, colors identify detectors, hatches indicate evaluation settings, and split colors denote ties. The best detector varies across conditions, while more data generally improves performance.

5 Experimental Setup

Training settings.

In order to analyse how AA performance varies with respect to the amount of available training data, we follow Shetty et al. (2026) considering two training scenarios: high-resource setting where a relatively large number of labelled training books is available for each LLM (10–30 books), and low-resource setting where only a limited number of labelled training books is available for each LLM (1–5 books).

Evaluation settings.

Since MultiGhostBench includes the OOD-Author dimension, the generator of a test text may not appear among those observed during training. This corresponds to an open-set setting, where an AA method must not only distinguish among known authors, but also reject texts generated by unseen authors. To support this rejection mechanism, all methods are calibrated on the development set by selecting a confidence threshold that best balances attribution performance and rejection ability. All the thresholds can be found in Appendix Table 9.

During evaluation, the same threshold is applied across all scenarios: ID, OOD-Domain, OOD-Author, and OOD-Language. We do this because in a practical application we would not know whether a test document is in or out of the training distribution. For texts generated by known authors, predictions below the confidence threshold are counted as missed attributions. For texts generated by unseen authors, the same behaviour is instead counted as a correct rejection, since the method avoids forcing the sample into one of the known author classes. Attribution performance is evaluated using macro-F1, which gives equal weight to all generating LLMs.

6 Experimental Results

We first consider in-language evaluation (§6.1), where training and test texts are written in the same language, covering the ID, OOD-Domain, and OOD-Author settings. We then analyse cross-language evaluation (§6.2), where attribution methods are trained on one source language and evaluated on a previously unseen target language. Finally, we examine the learned representation (§6.3) space to better interpret the observed performance patterns under domain and language shifts.

6.1 In-Language Evaluation

Figure 1 reports the best-performing method for each language and evaluation setting; complete results are reported in Appendix Table 11 and 12. For completeness, the results obtained without confidence-threshold rejection are provided in Appendix Table 13 and 14.

No universal winner.

No single method consistently outperforms all others across languages, resource regimes, and evaluation settings. Nevertheless, some methods appear more frequently among the best-performing approaches. xlm-roberta is particularly competitive in the ID and OOD-Domain settings, while n-gram becomes a frequent winner in the high-resource regime, especially under OOD-Author evaluation. trace variants also obtain the best result in several language-specific configurations. In contrast, rank, entropy, and gLTR consistently achieve low performance, indicating that individual token-level statistics provide insufficient signals for reliable authorship attribution.

Refer to caption
Figure 2: Cross-language macro-F1F_{1} results for xlm-roberta in the (a) low-resource and (b) high-resource settings. Rows indicate source languages and columns target languages; the final row and column report averages excluding the diagonal. Transfer depends on both the source-target pair and the resource regime; Chinese remains the most challenging target.

More data improves performance.

In the high-resource setting, the larger amount of training data improves the best achievable performance across all ID and OOD-Domain configurations. In several cases, the best-performing detector approaches or reaches perfect macro-F1F_{1}.

Full results (Tables Appendix Table 12 and 11) show the impact of additional training data on each detector. Under OOD-Author, xlm-roberta performs worse than in the low-resource setting for Italian, Chinese, and English, whereas detective improves across all languages. n-gram benefits the most from the additional data, becoming the best-performing OOD-Author detector in five languages and the second-best in Russian. An explanation is that n-gram needs more data to estimate reliable lexical and structural patterns, whereas xlm-roberta may become overconfident on texts from unseen authors, assigning them to one of the known classes instead of rejecting them.

Robustness drops under OOD shifts.

Compared with ID evaluation, performance generally decreases under distribution shift, although the strongest high-resource detectors retain high absolute scores. The OOD-ID differences reported in Appendix Table 15 and 16 are negative in most cases, with OOD-Author generally producing larger drops than OOD-Domain.

Language differences.

The low-resource setting reveals clearer performance differences across languages. Chinese achieves the strongest results, whereas English is comparatively more challenging. Indeed, the generators produce Chinese texts of more variable quality, creating stronger generator-specific signals that attribution detectors can exploit; in contrast, English outputs’ quality may be consistently better, leading to greater stylistic overlap and weaker attribution cues. These differences become less pronounced in the high-resource regime, where most best-performing scores approach or reach perfect macro-F1F_{1}.

6.2 Cross-Language Evaluation

Figure 2 reports the cross-language macro-F1 scores of xlm-roberta under the low- and high-resource settings. Rows indicate the source language used for training, while columns indicate the target language used for evaluation. We focus the visual analysis on xlm-roberta because it achieves the strongest overall cross-language performance. detective, the other Transformer-based method, also transfers substantially better than the remaining approaches, but generally obtains lower scores than xlm-roberta, despite outperforming it in a limited number of configurations. By contrast, the metric-based, N-GRAM, and fingerprint-based methods achieve near-zero performance in most cross-language configurations, suggesting that the signals they capture are largely language-dependent and do not transfer reliably across languages. For completeness, the full results for all methods with and without thresholding are reported in Appendix Table 17 and Table 18, respectively.

Refer to caption
Figure 3: UMAP projections of xlm-roberta representations after fine-tuning on Italian. The panels show evaluation on (a) Italian in-distribution (ID), (b) Italian OOD-Domain, and (c) Spanish and (d) Chinese ID texts evaluated under OOD-Language. Generator-specific clusters remain visible in all settings, with clearer separation for Spanish than under domain shift or for Chinese.

Target language difficulty.

Chinese is the most challenging target language. Averaging across source languages, transfer to Chinese achieves a macro-F1 of 0.508 in the low-resource setting and 0.582 in the high-resource setting, the lowest average in both regimes. Russian provides an informative contrast. Although it also uses a non-Latin script, it obtains an average target score of 0.717 in the low-resource setting and 0.841 in the high-resource setting, making it the easiest target language in the latter regime. Therefore, the difficulty of transferring to Chinese cannot be explained by script differences alone. Broader typological differences, tokenisation behaviour, and language-specific distributional properties may also contribute to the observed performance gap. In particular, the distinctive NGD values observed for Chinese in the quantitative analysis may reflect distributional characteristics that make cross-language transfer more difficult, although these descriptive metrics do not directly establish the source of the performance degradation.

Source language effects.

The effectiveness of a source language varies substantially across resource regimes. In the low-resource setting, Spanish provides the strongest overall transfer, with an average macro-F1 of 0.866 across target languages, followed by Chinese with 0.823. German is the weakest source language, with an average of 0.435. This ranking changes in the high-resource setting: Italian becomes the strongest source language, with an average macro-F1 of 0.894, followed by Spanish with 0.819, whereas Russian becomes the weakest at 0.410. Chinese also drops markedly, from 0.823 to 0.629, becoming the second-weakest source language in the high-resource regime. In particular, the deterioration observed for Chinese and Russian suggests that, in these languages, training with more data may lead to overfitting to source-language-specific patterns, reducing the transferability of the learned authorship signals to other languages. The language-specific differences observed in Chinese NGD and Russian PPL may also reflect underlying distributional properties that contribute to this behaviour, although these metrics do not directly explain the transfer degradation.

Language family effects.

The clearest evidence of a language-family effect concerns Italian and Spanish. In the low-resource setting, Spanish provides the best transfer to Italian, with a macro-F1 of 0.926, while Italian is the second-best source for Spanish, reaching 0.750. With more training data, the relationship becomes fully reciprocal: Spanish-to-Italian reaches 0.857, while Italian-to-Spanish achieves 0.981, the highest score in the table. The Germanic pair (English and German) shows a weaker and less symmetric pattern. English is only the fourth-best source language for German in the low-resource setting, but becomes the strongest source in the high-resource setting, reaching a macro-F1 of 0.916. Conversely, German-to-English achieves only 0.420 in the low resource regime, but improves to 0.673 with more training data, becoming the third-best source language. Overall, language-family relatedness appears to favour transfer, particularly for the Romance pair, but it does not fully determine cross-language performance.

6.3 Representation Analysis

To better understand the behaviour of xlm-roberta and trace beyond their attribution scores, we inspect their respective representation spaces. We use UMAP McInnes and Healy (2018) to visualise how xlm-roberta embeddings and trace fingerprints change across the ID, OOD-Domain, and OOD-Language settings. The two methods offer complementary perspectives on the attribution task: xlm-roberta relies on learned representations, whereas trace operates in a statistical fingerprint space. Comparing these spaces allows us to examine how the two approaches are affected by distribution shifts.

xlm-roberta embeddings.

We focus on the representation space learned by xlm-roberta in the high-resource setting, using the model fine-tuned on Italian, which achieves the strongest average cross-language performance. We consider embeddings from Italian ID texts, Italian OOD-Domain texts, and Spanish and Chinese texts evaluated in the OOD-Language setting to examine the effects of domain and language shifts. Figure 3 illustrates the embedding space produced by xlm-roberta across the four evaluation settings. In the ID setting, texts generated by the same model form compact, well-separated clusters. Under OOD-Domain, this organisation is only partially preserved, with clusters becoming less compact and exhibiting greater overlap. This pattern is consistent with the decrease in macro-F1 from 0.956 in the ID setting to 0.676 under OOD-Domain, indicating that domain shift makes the learned representations less discriminative. When the Italian-fine-tuned model is evaluated on Spanish texts, the model-specific structure is largely preserved, with compact and well-separated clusters similar to those observed in the Italian ID setting. This is consistent with the strong Italian-to-Spanish transfer performance and suggests that the learned generator-related representations generalise effectively between closely related languages. In contrast, when the same model is evaluated on Chinese texts, a model-specific organisation remains visible, but the clusters are broader and less clearly separated than those observed for Spanish. This suggests that the learned representations retain generator-related information across languages, although their discriminative structure is increasingly affected by larger typological and script differences.

trace fingerprints.

We next analyse the entropy-based fingerprint space of trace in the high-resource setting, using Italian as the reference language. While trace remains comparatively stable in the same-language settings, its performance drops sharply under OOD-Language evaluation. Since attribution is performed by comparing distances between test and training fingerprints, their relative position provides insight into this behaviour. Figure 4 compares Italian training, ID-test, and OOD-Domain fingerprints with fingerprints from different target languages. Italian ID-test and OOD-Domain samples remain close to the training distribution, whereas fingerprints from other languages occupy clearly displaced regions. Notably, this pattern is observed not only for Chinese, which is linguistically distant from Italian, but also for Spanish, a closely related Romance language. Consequently, cross-language samples lack nearby reference fingerprints, making distance-based attribution less reliable. This suggests that the fingerprint space is strongly influenced by language-specific token transition patterns, which helps explain the poor cross-language performance of trace.

7 Conclusion

We introduced MultiGhostBench, a multilingual benchmark for LLM AA comprising 928 long-form books generated by five recent LLMs, and covering six languages from four language families and three scripts. Importantly, it is designed to evaluate generalisation under domain, generator, and language shifts. For the within-language setting (training and testing on the same language), no single method consistently outperforms the others across languages, data regimes, or the ID, OOD-Domain, and OOD-Author conditions. In the cross-language setting (OOD-Language), however, Transformer-based approaches (xlm-roberta and detective) substantially outperform other methods. Interestingly, trace and metric-based methods drop to near-zero performance. Further analysis shows that cross-language performance is influenced by source and target language, language-family proximity, and script differences. These results highlight limitations of current methods in multilingual settings. We hope MultiGhostBench will serve as a useful benchmark for developing robust attribution methods.

Limitations

Although MultiGhostBench covers six languages spanning four language families and three writing scripts, it does not capture the full diversity of multilingual LLM-generated text. In particular, it does not cover low-resource languages, languages underrepresented in LLM pre-training data, or languages with substantially different grammatical and morphological characteristics. Furthermore, we consider five recent LLMs capable of generating long-form content across the selected languages. Extending the benchmark to additional and future LLMs is an important direction for future work. Our human evaluation is also limited to a small subset of generated books. Moreover, although the generated books are generally of good quality, multilingual long-form generation remains challenging, particularly with respect to originality, self-evaluation, revision, etc. across languages. These aspects are beyond the scope of this work. Finally, our study focuses on attribution under single-author and non-adversarial settings. Extensions such as mixed authorship (i.e., a book written by multiple authors) and adversarial generalisation (such as obfuscation attacks) remain important directions for future research.

Ethical Considerations

MultiGhostBench consists of LLM-generated books for multiple languages. As with any creative text, some books may contain writing or plot elements that reflect undesirable values. However, all LLMs used in this work incorporate built-in safety guardrails, reducing the likelihood of generating such content. This is further supported by our human evaluation, where no such concerning content was identified. Overall, we aim to adhere to the ACL Code of Ethics.44 4 https://www.aclweb.org/portal/content/acl-code-ethics

Acknowledgements

This research was conducted during M. Greco’s visit at the University of Melbourne, funded by the Erasmus+ Programme. We thank the volunteer annotators from the University of Melbourne and the University of Calabria for their contribution to the human evaluation. This research was supported by The University of Melbourne’s Research Computing Services and the Petascale Campus Initiative. A. Shetty was supported by the Commonwealth through an Australian Government Research Training Program Scholarship (DOI: https://doi.org/10.82133/C42F-K220), Computing and Information Systems PhD Scholarship, and the Avashya Foundation. A. Tagarelli was partly supported by the Horizon Europe project AI-CODE, GA No. 101135437. J.H. Lau was supported by the Australian Research Council under Grant LP210200917 and DP240101006.

References

  • Ahuja et al. (2023) K. Ahuja, H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. Nambi, T. Ganu, S. Segal, M. Ahmed, K. Bali, and S. Sitaram MEGA: multilingual evaluation of generative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 4232–4267. External Links: Link, Document Cited by: §1.
  • Australian Society of Authors (2025) Australian Society of Authors Artificial intelligence. Note: [Accessed 13-02-2026] External Links: Link Cited by: §1.
  • Brundage et al. (2024) M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, H. Anderson, H. Roff, G. C. Allen, J. Steinhardt, C. Flynn, S. Ó. hÉigeartaigh, S. Beard, H. Belfield, S. Farquhar, C. Lyle, R. Crootof, O. Evans, M. Page, J. Bryson, R. Yampolskiy, and D. Amodei The malicious use of artificial intelligence: forecasting, prevention, and mitigation. External Links: 1802.07228, Link Cited by: §1.
  • Cañete et al. (2023) J. Cañete, G. Chaperon, R. Fuentes, J. Ho, H. Kang, and J. Pérez Spanish pre-trained bert model and evaluation data. External Links: 2308.02976, Link Cited by: Table 6.
  • Chan et al. (2020) B. Chan, S. Schweter, and T. Möller German’s next language model. In Proceedings of the 28th International Conference on Computational Linguistics, D. Scott, N. Bel, and C. Zong (Eds.), Barcelona, Spain (Online), pp. 6788–6796. External Links: Link, Document Cited by: Table 6.
  • Chhun et al. (2022) C. Chhun, P. Colombo, F. M. Suchanek, and C. Clavel Of human criteria and automatic metrics: a benchmark of the evaluation of story generation. In Proceedings of the 29th International Conference on Computational Linguistics, N. Calzolari, C. Huang, H. Kim, J. Pustejovsky, L. Wanner, K. Choi, P. Ryu, H. Chen, L. Donatelli, H. Ji, S. Kurohashi, P. Paggio, N. Xue, S. Kim, Y. Hahm, Z. He, T. K. Lee, E. Santus, F. Bond, and S. Na (Eds.), Gyeongju, Republic of Korea, pp. 5794–5836. External Links: Link Cited by: §A.3, §3.5.
  • Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, L. Marris, S. Petulla, C. Gaffney, A. Aharoni, N. Lintz, T. C. Pais, H. Jacobsson, I. Szpektor, N. Jiang, K. Haridasan, A. Omran, N. Saunshi, D. Bahri, G. Mishra, E. Chu, T. Boyd, B. Hekman, A. Parisi, C. Zhang, K. Kawintiranon, T. Bedrax-Weiss, O. Wang, Y. Xu, O. Purkiss, U. Mendlovic, I. Deutel, N. Nguyen, A. Langley, F. Korn, L. Rossazza, A. Ramé, S. Waghmare, H. Miller, N. Byrd, A. Sheshan, R. Hadsell, S. Bhardwaj, P. Janus, T. Rissa, D. Horgan, A. Abdagic, L. Belenki, J. Allingham, A. Singh, T. Guidroz, S. Srinivasan, H. Schmit, K. Chiafullo, A. Elisseeff, N. Jha, P. Kolhar, L. Berrada, F. Ding, X. Si, S. B. Mallick, F. Och, S. Erell, E. Ni, T. Latkar, S. Yang, P. Sirkovic, Z. Feng, R. Leland, R. Hornung, G. Wu, C. Blundell, H. Alvari, P. Huang, C. Yip, S. Deur, L. Liu, G. Surita, P. Duque, D. Damen, J. Jia, A. Guez, M. Mircea, A. Sinha, A. Magni, P. Stradomski, T. Marian, V. Galić, W. Chen, H. Husain, A. Singhal, D. Grewe, F. Aubet, S. Song, L. Blanco, L. Rechis, L. Ho, R. Munoz, K. Zheng, J. Hamrick, K. Mather, H. Taitelbaum, E. Rutherford, Y. Lei, K. Chen, A. Shukla, E. Moreira, E. Doi, B. Isik, N. Shabat, D. Rogozińska, K. Kolipaka, J. Chang, E. Vušak, S. Venkatachary, S. Noghabi, T. Bharti, Y. Jun, A. Zaks, S. Green, J. Challagundla, W. Wong, M. Mohammad, D. Hirsch, Y. Cheng, I. Naim, L. Proleev, D. Vincent, A. Singh, M. Krikun, D. Krishnan, Z. Ghahramani, A. Atias, R. Aggarwal, C. Kirov, D. Vytiniotis, C. Koh, A. Chronopoulou, P. Dogra, V. Ion, G. Tyen, J. Lee, F. Weissenberger, T. Strohman, A. Balakrishna, J. Rae, M. Velic, R. de Liedekerke, O. Elyada, W. Yuan, C. Liu, L. Shani, S. Kishchenko, B. Alessio, Y. Li, R. Song, S. Kwei, O. Jankowski, A. Pappu, Y. Namiki, Y. Ma, N. Tripuraneni, C. Cherry, M. Ikonomidis, Y. Ling, C. Ji, B. Westberg, A. Wright, D. Yu, D. Parkinson, S. Ramaswamy, J. Connor, S. H. Yeganeh, S. Grover, G. Kenwright, L. Litchev, C. Apps, A. Tomala, F. Halim, A. Castro-Ros, Z. Li, A. Boral, P. Sho, M. Yarom, E. Malmi, D. Klinghoffer, R. Lin, A. Ansell, P. K. S, S. Zhao, S. Zuo, A. Santoro, H. Cheng, S. Demmessie, Y. Liu, N. Brichtova, A. Culp, N. Braun, D. Graur, W. Ng, N. Mehta, A. Phillips, P. Sundberg, V. Godbole, F. Liu, Y. Katariya, D. Rim, M. Seyedhosseini, S. Ammirati, J. Valfridsson, M. Malihi, T. Knight, A. Toor, T. Lampe, A. Ittycheriah, L. Chiang, C. Yeung, A. Fréchette, J. Rao, H. Wang, H. Srivastava, R. Zhang, R. Rhodes, A. Brand, D. Weesner, I. Figotin, F. Gimeno, R. Fellinger, P. Marcenac, J. Leal, E. Marcus, V. Cotruta, R. Cabrera, S. Luo, D. Garrette, V. Axelrod, S. Baltateanu, D. Barker, D. Chen, H. Toma, B. Ingram, J. Riesa, C. Kulkarni, Y. Zhang, H. Liu, C. Wang, M. Polacek, W. Wu, K. Hui, A. N. Reyes, Y. Su, M. Barnes, I. Malhi, A. Siddiqui, Q. Feng, M. Damaschin, D. Pighin, A. Steiner, S. Yang, R. S. Boppana, S. Ivanov, A. Kandoor, A. Shah, A. Mujika, D. Huang, C. A. Choquette-Choo, M. Patel, T. Yu, T. Creswell, Jerry, Liu, C. Barros, Y. Razeghi, A. Roy, P. Culliton, B. Xiong, J. Pan, T. Strohmann, T. Powell, B. Seal, D. DeCarlo, P. Shyam, K. Katircioglu, X. Wang, C. Hardin, I. Odisho, J. Broder, O. Chang, A. Nair, A. Shtefan, M. O’Brien, M. Agarwal, S. Potluri, S. Goyal, A. Jhindal, S. Thakur, Y. Stuken, J. Lyon, K. Toutanova, F. Feng, A. Wu, B. Horn, A. Wang, A. Cullum, G. Taubman, D. Shrivastava, C. Shi, H. Tomlinson, R. Patel, T. Tu, A. M. Oflazer, F. Pongetti, M. Yang, A. A. Taïga, V. Perot, N. W. Pierse, F. Han, Y. Drori, I. Iturrate, A. Chakrabarti, L. Yeung, D. Dopson, Y. Chen, A. Kulshreshtha, T. Guo, P. Pham, T. Schuster, J. Chen, A. Polozov, J. Xing, H. Zhou, P. Kacham, D. Kukliansky, A. Miech, S. Yaroshenko, E. Chi, S. Douglas, H. Fei, M. Blondel, P. Myla, L. Madmoni, X. Wu, D. Keysers, K. Kjems, I. Albuquerque, L. Yu, J. D’sa, M. Plantan, V. Ionescu, J. S. Elias, A. Gupta, M. R. Vuyyuru, F. Alcober, T. Zhou, K. Ji, F. Hartmann, S. Puttagunta, H. Song, E. Amid, A. Stefanoiu, A. Lee, P. Pucciarelli, E. Wang, A. Raul, S. Petrov, I. Tian, V. Anklin, N. Nti, V. Gomes, M. Schumacher, G. Vesom, A. Panagopoulos, K. Bousmalis, D. Andor, J. Jacob, Y. Zhang, B. Rosgen, M. Kecman, M. Tung, A. Belias, N. Goodman, P. Covington, B. Wieder, N. Saxena, E. Davoodi, M. Huang, S. Maddineni, V. Roulet, F. Campbell-Ajala, P. G. Sessa, Xintian, Wu, G. Lai, P. Collins, A. Haig, V. Sakenas, X. Xu, M. Giustina, L. E. Shafey, P. Charoenpanit, S. Garg, J. Ainslie, B. Severson, M. G. Arenas, S. Pathak, S. Rajayogam, J. Feng, M. Bakker, S. Li, N. Wichers, J. Rogers, X. Geng, Y. Li, R. Jagerman, C. Jia, N. Olmert, D. Sharon, M. Mauger, S. Mariserla, H. Ma, M. Mohabey, K. Kim, A. Andreev, S. Pollom, J. Love, V. Jain, P. Agrawal, Y. Schroecker, A. Fortin, M. Warmuth, J. Liu, A. Leach, I. Blok, G. P. Girirajan, R. Aharoni, B. Uria, A. Sozanschi, D. Goldberg, L. Ionita, M. T. Ribeiro, M. Zlocha, V. Birodkar, S. Lachgar, L. Yuan, H. Choudhury, M. Ginsberg, F. Zheng, G. Dibb, E. Graves, S. Lokhande, G. Rasskin, G. Muraru, C. Quick, S. Tata, P. Sermanet, A. Chawla, I. Karo, Y. Wang, S. Zhang, O. Keller, A. Dragan, G. Su, I. Chou, X. Liu, Y. Tao, S. Prabhakara, M. Wilson, R. Liu, S. Wang, G. Evans, D. Du, A. Castaño, G. Prasad, M. E. Mahdy, S. Gerlach, M. Reid, J. Kahn, A. Zait, T. S. Pillai, T. Ulrich, G. Wang, J. Wassenberg, E. Farkash, K. Yalasangi, C. Wang, M. Bauza, S. Bucher, T. Liu, J. Yan, G. Leung, V. Sindhwani, P. Barnes, A. Singh, I. Jurin, J. Chang, N. K. Bhumihar, S. Eiger, G. Citovsky, B. Withbroe, Z. Li, S. Xue, N. D. Santo, G. Stoyanov, Y. Raimond, S. Zheng, Y. Gao, V. Listík, S. Kwasiborski, R. Saputro, A. Ozturel, G. Mallya, K. Majmundar, R. West, P. Caron, J. Wei, L. Castrejon, S. Vikram, D. Ramachandran, N. Dhawan, J. Park, S. Smoot, G. van den Driessche, Y. Blau, C. Malik, W. Liang, R. Hirsch, C. N. dos Santos, E. Weinstein, A. van den Oord, S. Lall, N. FitzGerald, Z. Jiang, X. Yang, D. Webster, A. Elqursh, A. Pope, G. Rotival, D. Raposo, W. Zhu, J. Dean, S. Alabed, D. Tran, A. Gupta, Z. Gleicher, J. Austin, E. Rosseel, M. Umekar, D. Das, Y. Sun, K. Chen, K. Misiunas, X. Zhou, Y. Di, A. Loo, J. Newlan, B. Li, V. Ramasesh, Y. Xu, A. Chen, S. Gandhe, R. Soricut, N. Gupta, S. Hu, S. El-Sayed, X. Garcia, I. Brusilovsky, P. Chen, A. Bolt, L. Huang, A. Gurney, Z. Zhang, A. Pritzel, J. Wilkiewicz, B. Seybold, B. K. Shamanna, F. Fischer, J. Dean, K. Gill, R. Mcilroy, A. Bhowmick, J. Selier, A. Yang, D. Cheng, V. Magay, J. Tan, D. Varma, C. Walder, T. Kocisky, R. Nakashima, P. Natsev, M. Kwong, I. Gog, C. Zhang, S. Dieleman, T. Jimma, A. Ryabtsev, S. Brahma, D. Steiner, D. Du, A. Žužul, M. Žanić, M. Raghavachari, W. Gierke, Z. Zheng, D. Petrova, Y. Dauphin, Y. Liu, I. Kessler, S. Hand, C. Duvarney, S. Kim, H. Lee, L. Hussenot, J. Hui, J. Smith, D. Jain, J. Xia, G. S. Tomar, K. Amiri, D. Phan, F. Fuchs, T. Weyand, N. Tomasev, A. Cordell, X. Liu, J. Mallinson, P. Joshi, A. Crawford, A. Suggala, S. Chien, N. Fernando, M. Sanchez-Vargas, D. Williams, P. Crone, X. Luo, I. Karpov, J. Shan, T. Thurk, R. Strudel, P. Voigtlaender, P. Patil, T. Dozat, A. Khodaei, S. Singla, P. Ambroszczyk, Q. Wu, Y. Chang, B. Roark, C. Hegde, T. Ding, A. Filos, Z. Wu, A. S. Pinto, S. Liu, S. Khanna, A. Pandey, S. Mcloughlin, Q. Li, S. Haves, A. Zhou, E. Buchatskaya, I. Leal, P. de Boursac, N. Akazawa, N. Anderson, T. Chen, K. Somandepalli, C. Liang, S. Goenka, S. Winkler, A. Grushetsky, Y. Ding, J. Smith, F. Ye, J. Pont-Tuset, E. Li, R. Li, T. Golany, D. Wegner, T. Jiang, O. Barak, Y. Shangguan, E. Vértes, R. Wong, J. Bornschein, A. Tudor, M. Bevilacqua, T. Schaul, A. S. Rawat, Y. Zhao, K. Axiotis, L. Meng, C. McLean, J. Lai, J. Beattie, N. Kushman, Y. Liu, B. Kutzman, F. Lang, J. Ye, P. Netrapalli, P. Mishra, M. Khan, M. Goel, R. Willoughby, D. Tian, H. Zhuang, J. Chen, Z. Tsai, T. Kementsietsidis, A. Khare, J. Keeling, K. Xu, N. Waters, F. Altché, A. Popat, B. Mittal, D. Saxton, D. E. Badawy, M. Mathieu, Z. Zheng, H. Zhou, N. Ranka, R. Shin, Q. Duan, T. Salimans, I. Mihailescu, U. Shaham, M. Chang, Y. Assael, N. Dikkala, M. Izzard, V. Cohen-Addad, C. Graves, V. Feinberg, G. Chung, D. Strouse, D. Karmon, S. Sharifzadeh, Z. Ashwood, K. Pham, J. Blanton, A. Vasiloff, J. Barber, M. Geller, A. Zhou, F. Zubach, T. Huang, L. Zhang, H. Gupta, M. Young, J. Proskurnia, R. Votel, V. Gabeur, G. Barcik, A. Tripathi, H. Yu, G. Yan, B. Changpinyo, F. Pavetić, A. Coyle, Y. Fujii, J. G. Mendez, T. Zhou, H. Rajamani, B. Hechtman, E. Cao, D. Juan, Y. Tan, V. Dalibard, Y. Du, N. Clay, K. Yao, W. Jia, D. Vijaykumar, Y. Zhou, X. Bai, W. Hung, S. Pecht, G. Todorov, N. Khadke, P. Gupta, P. Lahoti, A. Autef, K. Duddu, J. Lee-Thorp, A. Bykovsky, T. Misiunas, S. Flennerhag, S. Thangaraj, J. McGiffin, Z. Nado, M. Kunesch, A. Noever, A. Hertz, M. Liang, V. Stone, E. Palmer, S. Daruki, A. Pramanik, S. Põder, A. Kyker, M. Khan, E. Sluzhaev, M. Ritter, A. Ruderman, W. Zhou, C. Nagpal, K. Vodrahalli, G. Necula, P. Barham, E. Pavlick, J. Hartford, I. Shafran, L. Zhao, M. Mikuła, T. Eccles, H. Shimokawa, K. Garg, L. Vilnis, H. Chen, I. Shumailov, K. Lee, A. Abdelhamed, M. Xie, V. Cohen, E. Hlavnova, D. Malkin, C. Sitawarin, J. Lottes, P. Coquinot, T. Yu, S. Kumar, J. Zhang, A. Mahendru, Z. Ahmed, J. Martens, T. Chen, A. Boag, D. Peng, C. Devin, A. Klimovskiy, M. Phuong, D. Vainstein, J. Xie, B. Ramabhadran, N. Howard, X. Yu, G. Goswami, J. Cui, S. Shleifer, M. Pinto, C. Yeh, M. Yang, S. Javanmardi, D. Ethier, C. Lee, J. Orbay, S. Kotecha, C. Bromberg, P. Shaw, J. Thornton, A. G. Rosenthal, S. Gu, M. Thomas, I. Gemp, A. Ayyar, A. Ushio, A. Selvan, J. Wee, C. Liu, M. Majzoubi, W. Yu, J. Abernethy, T. Liechty, R. Pan, H. Nguyen, Qiong, Hu, S. Perrin, A. Arora, E. Pitler, W. Wang, K. Shivakumar, F. Prost, B. Limonchik, J. Wang, Y. Gao, T. Cour, S. Buch, H. Gui, M. Ivanova, P. Neubeck, K. Chan, L. Kim, H. Chen, N. Goyal, D. Chung, L. Liu, Y. Su, A. Petrushkina, J. Shen, A. Joulin, Y. Xu, S. X. Lin, Y. Kulizhskaya, C. Chelba, S. Vasudevan, E. Collins, V. Bashlovkina, T. Lu, D. Fritz, J. Park, Y. Zhou, C. Su, R. Tanburn, M. Sushkov, M. Rasquinha, J. Li, J. Prendki, Y. Li, P. LV, S. Sharma, H. Fitoussi, H. Huang, A. Dai, P. Dao, M. Burrows, H. Prior, D. Qin, G. Pundak, L. L. Sjoesund, A. Khurshudov, Z. Zhu, A. Webson, E. Kemp, T. Tan, S. Agrawal, S. Sargsyan, L. Cheng, J. Stephan, T. Kwiatkowski, D. Reid, A. Byravan, A. H. Michaely, N. Heess, L. Zhou, S. Goenka, V. Carpenter, A. Levskaya, B. Wang, R. Roberts, R. Leblond, S. Chikkerur, S. Ginzburg, M. Chang, R. Riachi, Chuqiao, Xu, Z. Borsos, M. Pliskin, J. Pawar, M. Lustman, H. Kirkwood, A. Anand, A. Chaudhary, N. Kalb, K. Milan, S. Augenstein, A. Goldie, L. Prince, K. Raman, Y. Sun, V. Xia, A. Cohen, Z. Huo, J. Camp, S. Ellis, L. Zilka, D. V. Torres, L. Patel, S. Arora, B. Chan, J. Adler, K. Ayoub, J. Liang, F. Jamil, J. Jiang, S. Baumgartner, H. Sun, Y. Karov, Y. Akulov, H. Zheng, I. Cai, C. Fantacci, J. Rubin, A. R. Acha, M. Wang, N. D’Souza, R. Sathyanarayana, S. Dai, S. Rowe, A. Simanovsky, O. Goldman, Y. Kuang, X. Pan, A. Rosenberg, T. Rojas-Esponda, P. Dutta, A. Zeng, I. Jurenka, G. Farquhar, Y. Bansal, S. Iqbal, B. Roelofs, G. Joung, P. Beak, C. Ryu, R. Poplin, Y. Wu, J. Alayrac, S. Buthpitiya, O. Ronneberger, C. Habtegebriel, W. Li, P. Cavallaro, A. Wei, G. Bensky, T. Denk, H. Ganapathy, J. Stanway, P. Joshi, F. Bertolini, J. Lo, O. Ma, Z. Charles, G. Sampemane, H. Sahni, X. Chen, H. Askham, D. Gaddy, P. Young, J. Tan, M. Eyal, A. Bražinskas, L. Zhong, Z. Wu, M. Epstein, K. Bailey, A. Hard, K. Lee, S. Goldshtein, A. Ruiz, M. Badawi, M. Lochbrunner, J. Kearns, A. Brown, F. Pardo, T. Weber, H. Yang, P. Jiang, B. Akin, Z. Fu, M. Wainwright, C. Zou, M. Gaba, P. Manzagol, W. Kan, Y. Song, K. Zainullina, R. Lin, J. Ko, S. Deshmukh, A. Jindal, J. Svensson, D. Tyam, H. Zhao, C. Kaeser-Chen, S. Baird, P. Moradi, J. Hall, Q. Guo, V. Tsang, B. Liang, F. Pereira, S. Ganesh, I. Korotkov, J. Adamek, S. Thiagarajan, V. Tran, C. Chen, C. Tar, S. Jain, I. Dasgupta, T. Bilal, D. Reitter, K. Zhao, G. Vezzani, Y. Gehman, P. Mehta, L. Beltrone, X. Dotiwalla, S. Guadarrama, Z. Abbas, S. Karp, P. Georgiev, C. Ferng, M. Brockschmidt, L. Peng, C. Hirnschall, V. Verma, Y. Bi, Y. Xiao, A. Dabush, K. Xu, P. Wallis, R. Parker, Q. Wang, Y. Xu, I. Safarli, D. Tewari, Y. Zhang, S. Kim, A. Gesmundo, M. Thomas, S. Levi, A. Chowdhury, K. Rao, P. Garst, S. Conway-Rahman, H. Ran, K. McKinney, Z. Xiao, W. Yu, R. Agrawal, A. Stjerngren, C. Ionescu, J. Chen, V. Sharma, J. Chiu, F. Liu, K. Franko, C. Sanford, X. Cai, P. Michel, S. Ganapathy, J. Labanowski, Z. Garrett, B. Vargas, S. Sun, B. Gale, T. Buschmann, G. Desjardins, N. Ghelani, P. Jain, M. Verma, C. Asawaroengchai, J. Eisenschlos, J. Harlalka, H. Kazawa, D. Metzler, J. Howland, Y. Jian, J. Ades, V. Shah, T. Gangwani, S. Lee, R. Ring, S. M. Hernandez, D. Reich, A. Sinha, A. Sathe, J. Kovac, A. Gill, A. Kannan, A. D’olimpio, M. Sevenich, J. Whang, B. Kim, K. C. Sim, J. Chen, J. Zhang, S. Lall, Y. Matias, B. Jia, A. Friesen, S. Nasso, A. Thapliyal, B. Perozzi, T. Yu, A. Shekhawat, S. Huda, P. Grabowski, E. Wang, A. Sreevatsa, H. Dib, M. Hassen, P. Schuh, V. Milutinovic, C. Welty, M. Quinn, A. Shah, B. Wang, G. Barth-Maron, J. Frye, N. Axelsson, T. Zhu, Y. Ma, I. Giannoumis, H. Sedghi, C. Ye, Y. Luan, K. Aydin, B. Chandra, V. Sampathkumar, R. Huang, V. Lavrenko, A. Eleryan, Z. Hong, S. Hansen, S. M. Carthy, B. Samanta, D. Ćevid, X. Wang, F. Li, M. Voznesensky, M. Hoffman, A. Terzis, V. Sehwag, G. Fidel, L. He, M. Cai, Y. He, A. Feng, M. Nikoltchev, S. Phatale, J. Chase, R. Lawton, M. Zhang, T. Ouyang, M. Tragut, M. H. Manshadi, A. Narayanan, J. Shen, X. Gao, T. Bolukbasi, N. Roy, X. Li, D. Golovin, L. Panait, Z. Qin, G. Han, T. Anthony, S. Kudugunta, V. Patraucean, A. Ray, X. Chen, X. Yang, T. Bhatia, P. Talluri, A. Morris, A. Ražnatović, B. Brownfield, J. An, S. Peng, P. Kane, C. Zheng, N. Duduta, J. Kessinger, J. Noraky, S. Liu, K. Rong, P. Veličković, K. Rush, A. Goldin, F. Wei, S. M. R. Garlapati, C. Pantofaru, O. Kwon, J. Ni, E. Noland, J. D. Trapani, F. Beaufays, A. G. Roy, Y. Chow, A. Turker, G. Cideron, L. Mei, J. Clark, Q. Dou, M. Bošnjak, R. Leith, Y. Du, A. Yazdanbakhsh, M. Nasr, C. Kwak, S. S. Sheth, A. Kaskasoli, A. Anand, B. Lakshminarayanan, S. Jerome, D. Bieber, C. Chu, A. Senges, T. Shen, M. Sridhar, N. Ndebele, B. Beyret, S. Mohamed, M. Chen, M. Freitag, J. Guo, L. Liu, P. Roit, H. Chen, S. Yan, T. Stone, J. Co-Reyes, J. Cole, S. Scellato, S. Azizi, H. Hashemi, A. Jin, A. Iyer, M. Valentine, A. György, A. Ahuja, D. H. Diaz, C. Lee, N. Clement, W. Kong, D. Garmon, I. Watts, K. Bhatia, K. Gupta, M. Miecnikowski, H. Vallet, A. Taly, E. Loper, S. Joshi, J. Atwood, J. Chick, M. Collier, F. Iliopoulos, R. Trostle, B. Gunel, R. Leal-Cavazos, A. M. Hrafnkelsson, M. Guzman, X. Ju, A. Forbes, J. Emond, K. Chauhan, B. Caine, L. Xiao, W. Zeng, A. Moufarek, D. Murphy, M. Meng, N. Gupta, F. Riedel, A. Das, E. Lawal, S. Narayan, T. Sosea, J. Swirhun, L. Friso, B. Neyshabur, J. Lu, S. Girgin, M. Wunder, E. Yvinec, A. Pyne, V. Carbune, S. Rijhwani, Y. Guo, T. Doshi, A. Briukhov, M. Bain, A. Hitron, X. Wang, A. Gupta, K. Chen, C. Du, W. Zhang, D. Shah, A. Akula, M. Dylla, A. Kachra, W. Kuo, T. Zou, L. Wang, L. Xu, J. Zhu, J. Snyder, S. Menon, O. Firat, I. Mordatch, Y. Yuan, N. Ponomareva, R. Blevins, L. Moore, W. Wang, P. Chen, M. Scholz, A. Dwornik, J. Lin, S. Li, D. Antognini, T. I, X. Song, M. Miller, U. Kalra, A. Raveret, O. Akerlund, F. Wu, A. Nystrom, N. Godbole, T. Liu, H. DeBalsi, J. Zhao, B. Liu, A. Caciularu, L. Lax, U. Khandelwal, V. Langston, E. Bailey, S. Lattanzi, Y. Wang, N. Kovelamudi, S. Mondal, G. Guruganesh, N. Hua, O. Roval, P. Wesołowski, R. Ingale, J. Halcrow, T. Sohn, C. Angermueller, B. Raad, E. Stickgold, E. Lu, A. Kosik, J. Xie, T. Lillicrap, A. Huang, L. L. Zhang, D. Paulus, C. Farabet, A. Wertheim, B. Wang, R. Joshi, C. Ko, Y. Wu, S. Agrawal, L. Lin, X. Sheng, P. Sung, T. Breland-King, C. Butterfield, S. Gawde, S. Singh, Q. Zhang, R. Apte, S. Shetty, A. Hutter, T. Li, E. Salesky, F. Lebron, J. Kanerva, M. Paganini, A. Nguyen, R. Vallu, J. Peter, S. Velury, D. Kao, J. Hoover, A. Bortsova, C. Bishop, S. Jakobovits, A. Agostini, A. Agarwal, C. Liu, C. Kwong, S. Tavakkol, I. Bica, A. Greve, A. GP, J. Marcus, L. Hou, T. Duerig, R. Moroshko, D. Lacey, A. Davis, J. Amelot, G. Wang, F. Kim, T. Strinopoulos, H. Wan, C. L. Lan, S. Krishnan, H. Tang, P. Humphreys, J. Bai, I. H. Shtacher, D. Machado, C. Pang, K. Burke, D. Liu, R. Aravamudhan, Y. Song, E. Hirst, A. Singh, B. Jou, L. Bai, F. Piccinno, C. K. Fu, R. Alazard, B. Meiri, D. Winter, C. Chen, M. Zhang, J. Heitkaemper, J. Lambert, J. Lee, A. Frömmgen, S. Rogulenko, P. Nair, P. Niemczyk, A. Bulyenov, B. Xu, H. Shemtov, M. Zadimoghaddam, S. Toropov, M. Wirth, H. Dai, S. Gollapudi, D. Zheng, A. Kurakin, C. Lee, K. Bullard, N. Serrano, I. Balazevic, Y. Li, J. Schalkwyk, M. Murphy, M. Zhang, K. Sequeira, R. Datta, N. Agrawal, C. Sutton, N. Attaluri, M. Chiang, W. Farhan, G. Thornton, K. Lin, T. Choma, H. Nguyen, K. Dasgupta, D. Robinson, I. Comşa, M. Riley, A. Pillai, B. Mustafa, B. Golan, A. Zandieh, J. Lespiau, B. Porter, D. Ross, S. Rajayogam, M. Agarwal, S. Venugopalan, B. Shahriari, Q. Yan, H. Xu, T. Tobin, P. Dubov, H. Shi, A. Recasens, A. Kovsharov, S. Borgeaud, L. Dery, S. Vasanth, E. Gribovskaya, L. Qiu, M. Mahdieh, W. Skut, E. Nielsen, C. Zheng, A. Yu, C. G. Bostock, S. Gupta, A. Archer, C. Rawles, E. Davies, A. Svyatkovskiy, T. Tsai, Y. Halpern, C. Reisswig, B. Wydrowski, B. Chang, J. Puigcerver, M. H. Taege, J. Li, E. Schnider, X. Li, D. Dena, Y. Xu, U. Telang, T. Shi, H. Zen, K. Kastner, Y. Ko, N. Subramaniam, A. Kumar, P. Blois, Z. Dai, J. Wieting, Y. Lu, Y. Zeldes, T. Xie, A. Hauth, A. Ţifrea, Y. Li, S. El-Husseini, D. Abolafia, H. Zhou, W. Ding, S. Ghalebikesabi, C. Guía, A. Maksai, Á. Weisz, S. Arik, N. Sukhanov, A. Świetlik, X. Jia, L. Yu, W. Wang, M. Brand, D. Bloxwich, S. Kirmani, Z. Chen, A. Go, P. Sprechmann, N. Kannen, A. Carin, P. Sandhu, I. Edkins, L. Nooteboom, J. Gupta, L. Maggiore, J. Azizi, Y. Pritch, P. Yin, M. Gupta, D. Tarlow, D. Smith, D. Ivanov, M. Babaeizadeh, A. Goel, S. Kambala, G. Chu, M. Kastelic, M. Liu, H. Soltau, A. Stone, S. Agrawal, M. Kim, K. Soparkar, S. Tadepalli, O. Bunyan, R. Soh, A. Kannan, D. Kim, B. J. Chen, A. Halumi, S. Roy, Y. Wang, O. Sercinoglu, G. Gibson, S. Bhatnagar, M. Sano, D. von Dincklage, Q. Ren, B. Mitrevski, M. Olšák, J. She, C. Doersch, Jilei, Wang, B. Liu, Q. Tan, T. Yakar, T. Warkentin, A. Ramirez, C. Lebsack, J. Dillon, R. Mathews, T. Cobley, Z. Wu, Z. Chen, J. Simon, S. Nath, T. Sainath, A. Bendebury, R. Julian, B. Mankalale, D. Ćurko, P. Zacchello, A. R. Brown, K. Sodhia, H. Howard, S. Caelles, A. Gupta, G. Evans, A. Bulanova, L. Katzen, R. Goldenberg, A. Tsitsulin, J. Stanton, B. Schillings, V. Kovalev, C. Fry, R. Shah, K. Lin, S. Upadhyay, C. Li, S. Radpour, M. Maggioni, J. Xiong, L. Haas, J. Brennan, A. Kamath, N. Savinov, A. Nagrani, T. Yacovone, R. Kappedal, K. Andriopoulos, L. Lao, Y. Li, G. Rozhdestvenskiy, K. Hashimoto, A. Audibert, S. Austin, D. Rodriguez, A. Ruoss, G. Honke, D. Karkhanis, X. Xiong, Q. Wei, J. Huang, Z. Leng, V. Premachandran, S. Bileschi, G. Evangelopoulos, T. Mensink, J. Pavagadhi, D. Teplyashin, P. Chang, L. Xue, G. Tanzer, S. Goldman, K. Patel, S. Li, J. Wiesner, I. Zheng, I. Stewart-Binks, J. Han, Z. Li, L. Luo, K. Lenc, M. Lučić, F. Xue, R. Mullins, A. Guseynov, C. Chang, I. Galatzer-Levy, A. Zhang, G. Bingham, G. Hu, A. Hartman, Y. Ma, J. Griffith, A. Irpan, C. Radebaugh, S. Yue, L. Fan, V. Ungureanu, C. Sorokin, H. Teufel, P. Li, R. Anil, D. Paparas, T. Wang, C. Lin, H. Peng, M. Shum, G. Petrovic, D. Brady, R. Nguyen, K. Macherey, Z. Li, H. Singh, M. Yenugula, M. Iinuma, X. Chen, K. Kopparapu, A. Stern, S. Dave, C. Thekkath, F. Perot, A. Kumar, F. Li, Y. Xiao, M. Bilotti, M. H. Bateni, I. Noble, L. Lee, A. Vázquez-Reina, J. Salazar, X. Yang, B. Wang, E. Gruzewska, A. Rao, S. Raghuram, Z. Xu, E. Ben-David, J. Mei, S. Dalmia, Z. Zhang, Y. Liu, G. Bansal, H. Pankov, S. Schwarcz, A. Burns, C. Chan, S. Sanghai, R. Liang, E. Liang, A. He, A. Stuart, A. Narayanan, Y. Zhu, C. Frank, B. Fatemi, A. Sabne, O. Lang, I. Bhattacharya, S. Settle, M. Wang, B. McMahan, A. Tacchetti, L. B. Soares, M. Hadian, S. Cabi, T. Chung, N. Putikhin, G. Li, J. Chen, A. Tarango, H. Michalewski, M. Kazemi, H. Masoom, H. Sheftel, R. Shivanna, A. Vadali, R. Comanescu, D. Reid, J. Moore, A. Neelakantan, M. Sander, J. Herzig, A. Rosenberg, M. Dehghani, J. Choi, M. Fink, R. Hayes, E. Ge, S. Weng, C. Ho, J. Karro, K. Krishna, L. N. Thiet, A. Skerry-Ryan, D. Eppens, M. Andreetto, N. Sarma, S. Bonacina, B. K. Ayan, M. Nawhal, Z. Shan, M. Dusenberry, S. Thakoor, S. Gubbi, D. D. Nguyen, R. Tsarfaty, S. Albanie, J. Mitrović, M. Gandhi, B. Chen, A. Epasto, G. Stephanov, Y. Jin, S. Gehman, A. Amini, J. Weber, F. Behbahani, S. Xu, M. Allamanis, X. Chen, M. Ott, C. Sha, M. Jastrzebski, H. Qi, D. Greene, X. Wu, A. Toki, D. Vlasic, J. Shapiro, R. Kotikalapudi, Z. Shen, T. Saeki, S. Xie, A. Cassirer, S. Bharadwaj, T. Kiyono, S. Bhojanapalli, E. Rosenfeld, S. Ritter, J. Mao, J. G. Oliveira, Z. Egyed, B. Bandemer, E. Parisotto, K. Kinoshita, J. Pluto, P. Maniatis, S. Li, Y. Guo, G. Ghiasi, J. Tarbouriech, S. Chatterjee, J. Jin, Katrina, Xu, J. Palomaki, S. Arnold, M. Sewak, F. Piccinini, M. Sharma, B. Albrecht, S. Purser-haskell, A. Vaswani, C. Chen, M. Wisniewski, Q. Cao, J. Aslanides, N. M. Phu, M. Sieb, L. Agubuzu, A. Zheng, D. Sohn, M. Selvi, A. Andreassen, K. Subudhi, P. Eruvbetine, O. Woodman, T. Mery, S. Krause, X. Ren, X. Ma, J. Luo, D. Chen, W. Fan, H. Griffiths, C. Schuler, A. Li, S. Zhang, J. Sarr, S. Luo, R. Patana, M. Watson, D. Naboulsi, M. Collins, S. Sidhwani, E. Hoogeboom, S. Silver, E. Caveness, X. Zhao, M. Rodriguez, M. Deines, L. Bai, P. Griffin, M. Tagliasacchi, E. Xue, S. R. Babbula, B. Pang, N. Ding, G. Shen, E. Peake, R. Crocker, S. S. Raghvendra, D. Swisher, W. Han, R. Singh, L. Wu, V. Pchelin, T. Munkhdalai, D. Alon, G. Bacon, E. Robles, J. Bulian, M. Johnson, G. Powell, F. T. Ferreira, Y. Li, F. Benzing, M. Velimirović, H. Soyer, W. Kong, Tony, Nguyên, Z. Yang, J. Liu, J. van Amersfoort, D. Gillick, B. Sun, N. Rauschmayr, K. Zhang, S. Zhan, T. Zhou, A. Frolov, C. Yang, D. Vnukov, L. Rouillard, H. Li, A. Mandhane, N. Fallen, R. Venkataraman, C. H. Hu, J. Brennan, J. Lee, J. Chang, M. Sundermeyer, Z. Pan, R. Ke, S. Tong, A. Fabrikant, W. Bono, J. Gu, R. Foley, Y. Mao, M. Delakis, D. Bhaswar, R. Frostig, N. Li, A. Zipori, C. Hope, O. Kozlova, S. Mishra, J. Djolonga, C. Schiff, M. A. Merey, E. Briakou, P. Morgan, A. Wan, A. Hassidim, R. Skerry-Ryan, K. Sengupta, M. Jasarevic, P. Kallakuri, P. Kunkle, H. Brennan, T. Lieber, H. Mansoor, J. Walker, B. Zhang, A. Xie, G. Žužić, A. Chukwuka, A. Druinsky, D. Cho, R. Yao, F. Naeem, S. Butt, E. Kim, Z. Jia, M. Jordan, A. Lelkes, M. Kurzeja, S. Wang, J. Zhao, A. Over, A. Chakladar, M. Prasetya, N. Jha, S. Ganapathy, Y. Cong, P. Shroff, C. Saroufim, S. Miryoosefi, M. Hammad, T. Nasir, W. Xi, Y. Gao, Y. Maeng, B. Hora, C. Cheng, P. Haghani, Y. Lewenberg, C. Lu, M. Matysiak, N. Raisinghani, H. Wang, L. Baugher, R. Sukthankar, M. Giang, J. Schultz, N. Fiedel, M. Chen, C. Lee, T. Dey, H. Zheng, S. Paul, C. Smith, A. Ly, Y. Wang, R. Bansal, B. Perz, S. Ricco, S. Blank, V. Keshava, D. Sharma, M. Chow, K. Lad, K. Jalan, S. Osindero, C. Swanson, J. Scott, A. Ilić, X. Li, S. R. Jonnalagadda, A. S. Soudagar, Y. Xiong, B. Batsaikhan, D. Jarrett, N. Kumar, M. Shah, M. Lawlor, A. Waters, M. Graham, R. May, S. Ramos, S. Lefdal, Z. Cankara, N. Cano, B. O’Donoghue, J. Borovik, F. Liu, J. Grimstad, M. Alnahlawi, K. Tsihlas, T. Hudson, N. Grigorev, Y. Jia, T. Huang, T. P. Igwe, S. Lebedev, X. Tang, I. Krivokon, F. Garcia, M. Tan, E. Jia, P. Stys, S. Vashishth, Y. Liang, B. Venkatraman, C. Gu, A. Kementsietsidis, C. Zhu, J. Jung, Y. Bai, M. J. Hosseini, F. Ahmed, A. Gupta, X. Yuan, S. Ashraf, S. Nigam, G. Vasudevan, P. Awasthi, A. M. Gilady, Z. Mariet, R. Eskander, H. Li, H. Hu, G. Garrido, P. Schlattner, G. Zhang, R. Saxena, P. Dević, K. Muralidharan, A. Murthy, Y. Zhou, M. Choi, A. Wongpanich, Z. Wang, P. Shah, Y. Xu, Y. Huang, S. Spencer, A. Chen, J. Cohan, J. Wang, J. Tompson, J. Wu, R. Haroun, H. Li, B. Huergo, F. Yang, T. Yin, J. Wendt, M. Bendersky, R. Chaabouni, J. Snaider, J. Ferret, A. Jindal, T. Thompson, A. Xue, W. Bishop, S. M. Phal, A. Sharma, Y. Sung, P. Radhakrishnan, M. Shomrat, R. Ingle, R. Vij, J. Gilmer, M. D. Istin, S. Sobell, Y. Lu, E. Nottage, D. Sadigh, J. Willcock, T. Zhang, S. Xu, S. Brown, K. Lee, G. Wang, Y. Zhu, Y. Tay, C. Kim, A. Gutierrez, A. Sharma, Y. Xian, S. Seo, C. Cui, E. Pochernina, C. Baetu, K. Jastrzębski, M. Ly, M. Elhawaty, D. Suh, E. Sezener, P. Wang, N. Yuen, G. Tucker, J. Cai, Z. Yang, C. Wang, A. Muzio, H. Qian, J. Yoo, D. Lockhart, K. R. McKee, M. Guo, M. Mehrotra, A. Mendonça, S. V. Mehta, S. Ben, C. Tekur, J. Mu, M. Zhu, V. Krakovna, H. Lee, A. Maschinot, S. Cevey, H. Choe, A. Bai, H. Srinivasan, D. Gasaway, N. Young, P. Siegler, D. Holtmann-Rice, V. Piratla, K. Baumli, R. Yogev, A. Hofer, H. van Hasselt, S. Grant, Y. Chervonyi, D. Silver, A. Hogue, A. Agarwal, K. Wang, P. Singh, F. Flynn, J. Lipschultz, R. David, L. Bellot, Y. Yang, L. Le, F. Graziano, K. Olszewska, K. Hui, A. Maurya, N. Parotsidis, W. Chen, T. Oguntebi, J. Kelley, A. Baddepudi, J. Mauerer, G. Shaw, A. Siegman, L. Yang, S. Shetty, S. Roy, Y. Song, W. Stokowiec, R. Burnell, O. Savant, R. Busa-Fekete, J. Miao, S. Ghosh, L. MacDermed, P. Lippe, M. Dektiarev, Z. Behrman, F. Mentzer, K. Nguyen, M. Wei, S. Verma, C. Knutsen, S. Dasari, Z. Yan, P. Mitrichev, X. Wang, V. Shejwalkar, J. Austin, S. Sunkara, N. Potti, Y. Virin, C. Wright, G. Liu, O. Riva, E. Pot, G. Kochanski, Q. Le, G. Balasubramaniam, A. Dhar, Y. Liao, A. Bloniarz, D. Shukla, E. Cole, J. Lee, S. Zhang, S. Kafle, S. Vashishtha, P. Mahmoudieh, G. Chen, R. Hoffmann, P. Srinivasan, A. D. Lago, Y. B. Shalom, Z. Wang, M. Elabd, A. Sharma, J. Oh, S. Kothawade, M. Le, M. Monteiro, S. Yang, K. Alarakyia, R. Geirhos, D. Mincu, H. Garnes, H. Kobayashi, S. Mariooryad, K. Krasowiak, Zhixin, Lai, S. Mourad, M. Wang, F. Bu, O. Aharoni, G. Chen, A. Goyal, V. Zubov, A. Bapna, E. Dabir, N. Kothari, K. Lamerigts, N. D. Cao, J. Shar, C. Yew, N. Kulkarni, D. Mahaarachchi, M. Joshi, Z. Zhu, J. Lichtarge, Y. Zhou, H. Muckenhirn, V. Selo, O. Vinyals, P. Chen, A. Brohan, V. Mehta, S. Cogan, R. Wang, T. Geri, W. Ko, W. Chen, F. Viola, K. Shivam, L. Wang, M. C. Elish, R. A. Popa, S. Pereira, J. Liu, R. Koster, D. Kim, G. Zhang, S. Ebrahimi, P. Talukdar, Y. Zheng, P. Poklukar, A. Mikhalap, D. Johnson, A. Vijayakumar, M. Omernick, M. Dibb, A. Dubey, Q. Hu, A. Suman, V. Aggarwal, I. Kornakov, F. Xia, W. Lowe, A. Kolganov, T. Xiao, V. Nikolaev, S. Hemingray, B. Li, J. Iljazi, M. Rybiński, B. Sandhu, P. Lu, T. Luong, R. Jenatton, V. Govindaraj, Hui, Li, G. Dulac-Arnold, W. Park, H. Wang, A. Modi, J. Pouget-Abadie, K. Greller, R. Gupta, R. Berry, P. Ramachandran, J. Xie, L. McCafferty, J. Wang, K. Gupta, H. Lim, B. Bratanič, A. Brock, I. Akolzin, J. Sproch, D. Karliner, D. Kim, A. Goedeckemeyer, N. Shazeer, C. Schmid, D. Calandriello, P. Bhatia, K. Choromanski, C. Montgomery, D. Dua, A. Ramalho, H. King, Y. Gao, L. Nguyen, D. Lindner, D. Pitta, O. Johnson, K. Salama, D. Ardila, M. Han, E. Farnese, S. Odoom, Z. Wang, X. Ding, N. Rink, R. Smith, H. T. Lehri, E. Cohen, N. Vats, T. He, P. Gopavarapu, A. Paszke, M. Patel, W. V. Gansbeke, L. Loher, L. Castro, M. Voitovich, T. von Glehn, N. George, S. Niklaus, Z. Eaton-Rosen, N. Rakićević, E. Jue, S. Perel, C. Zhang, Y. Bahat, A. Pouget, Z. Xing, F. Huot, A. Shenoy, T. Bos, V. Coriou, B. Richter, N. Noy, Y. Wang, S. Ontanon, S. Qin, G. Makarchuk, D. Hassabis, Z. Li, M. Sharma, K. Venkatesan, I. Kemaev, R. Daniel, S. Huang, S. Shah, O. Ponce, Warren, Chen, M. Faruqui, J. Wu, S. Andačić, S. Payrits, D. McDuff, T. Hume, Y. Cao, M. Tessler, Q. Wang, Y. Wang, I. Rendulic, E. Agustsson, M. Johnson, T. Lando, A. Howard, S. G. S. Padmanabhan, M. Daswani, A. Banino, M. Kilgore, J. Heek, Z. Ji, A. Caceres, C. Li, N. Kassner, A. Vlaskin, Z. Liu, A. Grills, Y. Hou, R. Sukkerd, G. Cheon, N. Shetty, L. Markeeva, P. Stanczyk, T. Iyer, Y. Gong, S. Gao, K. Gopalakrishnan, T. Blyth, M. Reynolds, A. Bhoopchand, M. Bilenko, D. Gharibian, V. Zayats, A. Faust, A. Singh, M. Ma, H. Jiao, S. Vijayanarasimhan, L. Aroyo, V. Yadav, S. Chakera, A. Kakarla, V. Meshram, K. Gregor, G. Botea, E. Senter, D. Jia, G. Kovacs, N. Sharma, S. Baur, K. Kang, Y. He, L. Zhuo, M. Kostelac, I. Laish, S. Peng, L. O’Bryan, D. Kasenberg, G. R. Rao, E. Leurent, B. Zhang, S. Stevens, A. Salazar, Y. Zhang, I. Lobov, J. Walker, A. Porter, M. Redshaw, H. Ke, A. Rao, A. Lee, H. Lam, M. Moffitt, J. Kim, S. Qiao, T. Koo, R. Dadashi, X. Song, M. Sundararajan, P. Xu, C. Kawamoto, Y. Zhong, C. Barbu, A. Reddy, M. Verzetti, L. Li, G. Papamakarios, H. Klimczak-Plucińska, M. Cassin, K. Kavukcuoglu, R. Swavely, A. Vaucher, J. Zhao, R. Hemsley, M. Tschannen, H. Ge, G. Menghani, Y. Yu, N. Ha, W. He, X. Wu, M. Song, R. Sterneck, S. Zinke, D. A. Calian, A. Marsden, A. C. Ruiz, M. Hessel, A. Gueta, B. Lee, B. Farris, M. Gupta, Y. Li, M. Saleh, V. Misra, K. Xiao, P. Mendolicchio, G. Buttimore, V. Krayvanova, N. Nayakanti, M. Wiethoff, Y. Pande, A. Mirhoseini, N. Lao, J. Liu, Y. Hua, A. Chen, Y. Malkov, D. Kalashnikov, S. Gupta, K. Audhkhasi, Y. Zhai, S. Kopalle, P. Jain, E. Ofek, C. Meyer, K. Baatarsukh, H. Strejček, J. Qian, J. Freedman, R. Figueira, M. Sokolik, O. Bachem, R. Lin, D. Kharrat, C. Hidey, P. Xu, D. Duan, Y. Li, M. Ersoy, R. Everett, K. Cen, R. Santamaria-Fernandez, A. Taubenfeld, I. Mackinnon, L. Deng, P. Zablotskaia, S. Viswanadha, S. Goel, D. Yates, Y. Deng, P. Choy, M. Chen, A. Sinha, A. Mossin, Y. Wang, A. Szlam, S. Hao, P. K. Rubenstein, M. Toksoz-Exley, M. Aperghis, Y. Zhong, J. Ahn, M. Isard, O. Lacombe, F. Luisier, C. Anastasiou, Y. Kalley, U. Prabhu, E. Dunleavy, S. Bijwadia, J. Mao-Jones, K. Chen, R. Pasumarthi, E. Wood, A. Dostmohamed, N. Hurley, J. Simsa, A. Parrish, M. Pajarskas, M. Harvey, O. Skopek, Y. Kochinski, J. Rey, V. Rieser, D. Zhou, S. J. Lee, T. Acharya, G. Li, J. Jiang, X. Zhang, B. Gipson, E. Mahintorabi, M. Gelmi, N. Khajehnouri, A. Yeh, K. Lee, L. Matthey, L. Baker, T. Pham, H. Fu, A. Pak, P. Gupta, C. Vasconcelos, A. Sadovsky, B. Walker, S. Hsiao, P. Zochbauer, A. Marzoca, N. Velan, J. Zeng, G. Baechler, D. Driess, D. Jain, Y. Huang, L. Tao, J. Maggs, N. Levine, J. Schneider, E. Gemzer, S. Petit, S. Han, Z. Fisher, D. Zelle, C. Biles, E. Ie, A. Fadeeva, C. Liu, J. V. Franco, A. Collister, H. Zhang, R. Wang, R. Zhao, L. Kieliger, K. Shuster, R. Zhu, B. Gong, L. Chan, R. Sun, S. Basu, R. Zimmermann, J. Hayes, A. Bapna, J. Snoek, W. Yang, P. Datta, J. A. Abdallah, K. Kilgour, L. Li, S. Mah, Y. Jun, M. Rivière, A. Karmarkar, T. Spalink, T. Huang, L. Gonzalez, D. Tran, A. Nowak, J. Palowitch, M. Chadwick, E. Talius, H. Mehta, T. Sellam, P. Fränken, M. Nicosia, K. He, A. Kini, D. Amos, S. Basu, H. Jobe, E. Shaw, Q. Xu, C. Evans, D. Ikeda, C. Yan, L. Jin, L. Wang, S. Yadav, I. Labzovsky, R. Sampath, A. Ma, C. Schumann, A. Siddhant, R. Shah, J. Youssef, R. Agarwal, N. Dabney, A. Tonioni, M. Ambar, J. Li, I. Guyon, B. Li, D. Soergel, B. Fang, G. Karadzhov, C. Udrescu, T. Trinh, V. Raunak, S. Noury, D. Guo, S. Gupta, M. Finkelstein, D. Petek, L. Liang, G. Billock, P. Sun, D. Wood, Y. Song, X. Yu, T. Matejovicova, R. Cohen, K. Andra, D. D’Ambrosio, Z. Deng, V. Nallatamby, E. Songhori, R. Dangovski, A. Lampinen, P. Botadra, A. Hillier, J. Cao, N. Baddi, A. Kuncoro, T. Yoshino, A. Bhagatwala, M. Ranzato, R. Schaeffer, T. Liu, S. Ye, O. Sarvana, J. Nham, C. Kuang, I. Gao, J. Baek, S. Mittal, A. Wahid, A. Gergely, B. Ni, J. Feldman, C. Muir, P. Lamblin, W. Macherey, E. Dyer, L. Kilpatrick, V. Campos, M. Bhutani, S. Fort, Y. Ahmad, A. Severyn, K. Chatziprimou, O. Ferludin, M. Dimarco, A. Kusupati, J. Heyward, D. Bahir, K. Villela, K. Millican, D. Marcus, S. Bahargam, C. Unlu, N. Roth, Z. Wei, S. Gopal, D. Ghoshal, E. Lee, S. Lin, J. Lees, D. Lee, A. Hosseini, C. Fan, S. Neel, M. Wu, Y. Altun, H. Cai, E. Piqueras, J. Woodward, A. Bissacco, S. Haykal, M. Bordbar, P. Sundaram, S. Hodkinson, D. Toyama, G. Polovets, A. Myers, A. Sinha, T. Levinboim, K. Krishnakumar, R. Chhaparia, T. Sholokhova, N. B. Gundavarapu, G. Jawahar, H. Qureshi, J. Hu, N. Momchev, M. Rahtz, R. Wu, A. P. S, K. Dhamdhere, M. Guo, U. Gupta, A. Eslami, M. Schain, M. Blokzijl, D. Welling, D. Orr, L. Bolelli, N. Perez-Nieves, M. Sirotenko, A. Prasad, A. Kar, B. D. B. Pigem, T. Terzi, G. Weisz, D. Ghosh, A. Mavalankar, D. Madeka, K. Daugaard, H. Adam, V. Shah, D. Berman, M. Tran, S. Baker, E. Andrejczuk, G. Chole, G. Raboshchuk, M. Mirzazadeh, T. Kagohara, S. Wu, C. Schallhart, B. Orlando, C. Wang, A. Rrustemi, H. Xiong, H. Liu, A. Vezer, N. Ramsden, S. Chang, S. Mudgal, Y. Li, N. Vieillard, Y. Hoshen, F. Ahmad, A. Slone, A. Hua, N. Potikha, M. Rossini, J. Stritar, S. Prakash, Z. Wang, X. Dong, A. Nazari, E. Nehoran, K. Tekelioglu, Y. Li, K. Badola, T. Funkhouser, Y. Li, V. Yerram, R. Ganeshan, D. Formoso, K. Langner, T. Shi, H. Li, Y. Yamamori, A. Panda, A. Saade, A. S. Scarpati, C. Breaux, C. Carey, Z. Zhou, C. Hsieh, S. Bridgers, A. Butryna, N. Gupta, V. Tulsyan, S. Woo, E. Eltyshev, W. Grathwohl, C. Parks, S. Benjamin, R. Panigrahy, S. Dodhia, D. D. Freitas, C. Sauer, W. Song, F. Alet, J. Tolins, C. Paduraru, X. Zhou, B. Albert, Z. Zhang, L. Shu, M. Bansal, S. Nguyen, A. Globerson, O. Xiao, J. Manyika, T. Hennigan, R. Rong, J. Matak, A. Bakalov, A. Sharma, D. Sinopalnikov, A. Pierson, S. Roller, G. Brown, M. Gao, T. Fukuzawa, A. Ghafouri, K. Vassigh, I. Barr, Z. Wang, A. Korsun, R. Jayaram, L. Ren, T. Zaman, S. Khan, Y. Lunts, D. Deutsch, D. Uthus, N. Katz, M. Samsikova, A. Khalifa, N. Sethi, J. Sun, L. Tang, U. Alon, X. Luo, D. Yu, A. Nayyar, B. Petrini, W. Truong, V. Hellendoorn, N. Chinaev, C. Alberti, W. Wang, J. Hu, V. Mirrokni, A. Balashankar, A. Aharon, A. Mehta, A. Iscen, J. Kready, L. Manning, A. Mohananey, Y. Chen, A. Tripathi, A. Wu, I. Petrovski, D. Hwang, M. Baeuml, S. Chandrakaladharan, Y. Liu, R. Coaguila, M. Chen, S. Ma, P. Tafti, S. Tatineni, T. Spitz, J. Ye, P. Vicol, M. Rosca, A. Puigdomènech, Z. Yahav, S. Ghemawat, H. Lin, P. Kirk, Z. Nabulsi, S. Brin, B. Bohnet, K. Caluwaerts, A. S. Veerubhotla, D. Zheng, Z. Dai, P. Petrov, Y. Xu, R. Mehran, Z. Xu, L. Zintgraf, J. Choi, S. A. Hombaiah, R. Thoppilan, S. Reddi, L. Lew, L. Li, K. Webster, K. Sawhney, L. Lamprou, S. Shakeri, M. Lunayach, J. Chen, S. Bagri, A. Salcianu, Y. Chen, Y. Donchev, C. Magister, S. Nørly, V. Rodrigues, T. Izo, H. Noga, J. Zou, T. Köppe, W. Zhou, K. Lee, X. Long, D. Eisenbud, A. Chen, C. Schenck, C. M. To, P. Zhong, E. Taropa, M. Truong, O. Levy, D. Martins, Z. Zhang, C. Semturs, K. Zhang, A. Yakubovich, P. Moreno, L. McConnaughey, D. Lu, S. Redmond, L. Weerts, Y. Bitton, T. Refice, N. Lacasse, A. Conmy, C. Tallec, J. Odell, H. Forbes-Pollard, A. Socala, J. Hoech, P. Kohli, A. Walton, R. Wang, M. Sazanovich, K. Zhu, A. Kapishnikov, R. Galt, M. Denton, B. Murdoch, C. Sikora, K. Mohamed, W. Wei, U. First, T. McConnell, L. C. Cobo, J. Qin, T. Avrahami, D. Balle, Y. Watanabe, A. Louis, A. Kraft, S. Ariafar, Y. Gu, E. Rives, C. Yoon, A. Rusu, J. Cobon-Kerr, C. Hahn, J. Luo, Yuvein, Zhu, N. Ahuja, R. Benenson, R. L. Kaufman, H. Yu, L. Hightower, J. Zhang, D. Ni, L. A. Hendricks, G. Wang, G. Yona, L. Jain, P. Barrio, S. Bhupatiraju, S. Velusamy, A. Dafoe, S. Riedel, T. Thomas, Z. Yuan, M. Bellaiche, S. Panthaplackel, K. Kloboves, S. Jauhari, C. Akbulut, T. Davchev, E. Gladchenko, D. Madras, A. Chuklin, T. Hill, Q. Yuan, M. Madhavan, L. Leonhard, D. Scandinaro, Q. Chen, N. Niu, A. Douillard, B. Damoc, Y. Onoe, F. Pedregosa, F. Bertsch, C. Leichner, J. Pagadora, J. Malmaud, S. Ponda, A. Twigg, O. Duzhyi, J. Shen, M. Wang, R. Garg, J. Chen, U. Evci, J. Lee, L. Liu, K. Kojima, M. Yamaguchi, A. Rajendran, A. Piergiovanni, V. K. Rajendran, M. Fornoni, G. Ibagon, H. Ragan, S. M. Khan, J. Blitzer, A. Bunner, G. Sun, T. Kosakai, S. Lundberg, N. Elue, K. Guu, S. Park, J. Park, A. Narayanaswamy, C. Wu, J. Mudigonda, T. Cohn, H. Mu, R. Kumar, L. Graesser, Y. Zhang, R. Killam, V. Zhuang, M. Giménez, W. A. Jishi, R. Ley-Wild, A. Zhai, K. Osawa, D. Cedillo, J. Liu, M. Upadhyay, M. Sieniek, R. Sharma, T. Paine, A. Angelova, S. Addepalli, C. Parada, K. Majumder, A. Lamp, S. Kumar, X. Deng, A. Myaskovsky, T. Sabolić, J. Dudek, S. York, F. de Chaumont Quitry, J. Nie, D. Cattle, A. Gunjan, B. Piot, W. Khawaja, S. Bang, S. Wang, S. Khodadadeh, R. R, P. Rawlani, R. Powell, K. Lee, J. Griesser, G. Oh, C. Magalhaes, Y. Li, S. Tokumine, H. N. Vogel, D. Hsu, A. BC, D. Jindal, M. Cohen, Z. Yang, J. Yuan, D. de Cesare, T. Bruguier, J. Xu, M. Roy, A. Jacovi, D. Belov, R. Arya, P. Meadowlark, S. Cohen-Ganor, W. Ye, P. Morris-Suzuki, P. Banzal, G. Song, P. Ponnuramu, F. Zhang, G. Scrivener, S. Zaiem, A. R. Rochman, K. Han, B. Ghazi, K. Lee, S. Drath, D. Suo, A. Girgis, P. Shenoy, D. Nguyen, D. Eck, S. Gupta, L. Yan, J. Carreira, A. Gulati, R. Sang, D. Mirylenka, E. Cooney, E. Chou, M. Ling, C. Fan, B. Coleman, G. Tubone, R. Kumar, J. Baldridge, F. Hernandez-Campos, A. Lazaridou, J. Besley, I. Yona, N. Bulut, Q. Wellens, A. Pierigiovanni, J. George, R. Green, P. Han, C. Tao, G. Clark, C. You, A. Abdolmaleki, J. Fu, T. Chen, A. Chaugule, A. Chandorkar, A. Rahman, W. Thompson, P. Koanantakool, M. Bernico, J. Ren, A. Vlasov, S. Vassilvitskii, M. Kula, Y. Liang, D. Kim, Y. Huang, C. Ye, D. Lepikhin, and W. Helmholz Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. External Links: 2507.06261, Link Cited by: Table 4, Table 4.
  • Conneau et al. (2020) A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 8440–8451. External Links: Link, Document Cited by: Appendix B, §4.
  • DeepSeek-AI et al. (2025) DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J. Qiu, J. Li, J. Song, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Wang, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Wang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Zhang, R. Pan, R. Wang, R. Xu, R. Zhang, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Pan, T. Wang, T. Yun, T. Pei, T. Sun, W. L. Xiao, W. Zeng, W. Zhao, W. An, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Zhang, X. Chen, X. Nie, X. Sun, X. Wang, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Song, X. Shan, X. Zhou, X. Yang, X. Li, X. Su, X. Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Y. Zhang, Y. Xu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Yu, Y. Zheng, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Tang, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Wu, Y. Ou, Y. Zhu, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Zha, Y. Xiong, Y. Ma, Y. Yan, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Huang, Z. Zhang, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Xu, Z. Wu, Z. Zhang, Z. Li, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Gao, and Z. Pan DeepSeek-v3 technical report. External Links: 2412.19437, Link Cited by: Table 4.
  • Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. External Links: 1810.04805, Link Cited by: Table 6.
  • Doddapaneni et al. (2025) S. Doddapaneni, G. Ramesh, M. Khapra, A. Kunchukuttan, and P. Kumar A primer on pretrained multilingual language models. ACM Comput. Surv. 57 (9). External Links: ISSN 0360-0300, Link, Document Cited by: §1.
  • Dugan et al. (2024) L. Dugan, A. Hwang, F. Trhlík, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, and C. Callison-Burch RAID: a shared benchmark for robust evaluation of machine-generated text detectors. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 12463–12492. External Links: Link, Document Cited by: §1.
  • Fabien et al. (2020) M. Fabien, E. Villatoro-Tello, P. Motlicek, and S. Parida BertAA : BERT fine-tuning for authorship attribution. In Proceedings of the 17th International Conference on Natural Language Processing (ICON), P. Bhattacharyya, D. M. Sharma, and R. Sangal (Eds.), Indian Institute of Technology Patna, Patna, India, pp. 127–137. External Links: Link Cited by: §4.
  • Gehrmann et al. (2019) S. Gehrmann, H. Strobelt, and A. Rush GLTR: statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, M. R. Costa-jussà and E. Alfonseca (Eds.), Florence, Italy, pp. 111–116. External Links: Link, Document Cited by: Appendix B, §4.
  • Guo et al. (2024) X. Guo, Y. He, S. Zhang, T. Zhang, W. Feng, H. Huang, and C. Ma Detective: detecting ai-generated text via multi-level contrastive learning. Advances in Neural Information Processing Systems 37, pp. 88320–88347. Cited by: Appendix B, §4.
  • He et al. (2024) X. He, X. Shen, Z. Chen, M. Backes, and Y. Zhang MGTBench: benchmarking machine-generated text detection. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS ’24, New York, NY, USA, pp. 2251–2265. External Links: ISBN 9798400706363, Link, Document Cited by: §1.
  • Huang et al. (2025) B. Huang, C. Chen, and K. Shu Authorship attribution in the era of llms: problems, methodologies, and challenges. SIGKDD Explor. Newsl. 26 (2), pp. 21–43. External Links: ISSN 1931-0145, Link, Document Cited by: §1.
  • Jakesch et al. (2023) M. Jakesch, J. T. Hancock, and M. Naaman Human heuristics for ai-generated language are flawed. Proceedings of the National Academy of Sciences 120 (11), pp. e2208839120. External Links: Document, Link, https://www.pnas.org/doi/pdf/10.1073/pnas.2208839120 Cited by: §1.
  • Jiao et al. (2025) K. Jiao, Q. Wang, L. Zhang, Z. Guo, and Z. Mao M-RangeDetector: enhancing generalization in machine-generated text detection through multi-range attention masks. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 8971–8983. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: Appendix B.
  • John and Idicula (2026) A. B. John and S. M. Idicula Authorship attribution of malayalam biblical texts using deep learning. In 2026 2nd International Conference on Advances in Intelligent Computing and Applications (AICAPS), Vol. , pp. 1–6. External Links: Document Cited by: Appendix B.
  • Jurafsky and Martin (2026) D. Jurafsky and J. H. Martin Speech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition with language models. 3rd edition. Note: Online manuscript released January 6, 2026 External Links: Link Cited by: §3.5.
  • Kargaran et al. (2023) A. H. Kargaran, A. Imani, F. Yvon, and H. Schuetze GlotLID: language identification for low-resource languages. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 6155–6218. External Links: Link, Document Cited by: §3.4.
  • Koppel and Schler (2004) M. Koppel and J. Schler Authorship verification as a one-class classification problem. In Proceedings of the twenty-first international conference on Machine learning, pp. 62. Cited by: Appendix B, §4.
  • Kuratov and Arkhipov (2019) Y. Kuratov and M. Arkhipov Adaptation of deep bidirectional multilingual transformers for russian language. External Links: 1905.07213, Link Cited by: Table 6.
  • La Cava et al. (2026) L. La Cava, D. Macko, R. Moro, I. Srba, and A. Tagarelli Authorship attribution in multilingual machine-generated texts. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 45136–45152. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §2.
  • La Cava and Tagarelli (2025) L. La Cava and A. Tagarelli OpenTuringBench: an open-model-based benchmark and framework for machine-generated text detection and attribution. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26655–26671. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Appendix B, §1, §2, Table 1, §4.
  • Lee et al. (2025) Y. Lee, S. Ka, B. Son, P. Kang, and J. Kang Navigating the path of writing: outline-guided text generation with large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track), W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), Albuquerque, New Mexico, pp. 233–250. External Links: Link, Document, ISBN 979-8-89176-194-0 Cited by: §3.2.
  • Macko et al. (2025) D. Macko, J. Kopál, R. Moro, and I. Srba MultiSocial: multilingual benchmark of machine-generated text detection of social-media texts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 727–752. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1, Table 1.
  • Macko and Kopál (2026) D. Macko and J. Kopál CEAID: benchmark of multilingual machine-generated text detection methods for Central European languages. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 29177–29190. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §2.
  • Macko et al. (2023) D. Macko, R. Moro, A. Uchendu, J. Lucas, M. Yamashita, M. Pikuliak, I. Srba, T. Le, D. Lee, J. Simko, and M. Bielikova MULTITuDE: large-scale multilingual machine-generated text detection benchmark. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 9960–9987. External Links: Link, Document Cited by: Appendix B, §1, §2, Table 1, §4.
  • Maintainers (2022) H. C. M. Maintainers Gpt2 (revision 909a290). Hugging Face. External Links: Link, Document Cited by: Table 6.
  • McInnes and Healy (2018) L. McInnes and J. Healy UMAP: uniform manifold approximation and projection for dimension reduction. pp. . External Links: Document Cited by: §6.3.
  • Meister et al. (2023) C. Meister, T. Pimentel, G. Wiher, and R. Cotterell Locally typical sampling. Transactions of the Association for Computational Linguistics 11, pp. 102–121. External Links: Link, Document Cited by: 2nd item, §3.5.
  • Mitchell et al. (2023) E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn Detectgpt: zero-shot machine-generated text detection using probability curvature. In International conference on machine learning, pp. 24950–24962. Cited by: §2.
  • Momen et al. (2025) O. Momen, M. Schaaf, and A. Mehler Filling the temporal void: recovering missing publication years in the Project Gutenberg corpus using LLMs. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 17318–17334. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: Table 3, Table 7.
  • OpenAI et al. (2025) OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V. Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, L. Gross, K. G. Guzman, J. Hallman, J. Hehir, J. Heidecke, A. Helyar, H. Hu, R. Huet, J. Huh, S. Jain, Z. Johnson, C. Koch, I. Kofman, D. Kundel, J. Kwon, V. Kyrylov, E. Y. Le, G. Leclerc, J. P. Lennon, S. Lessans, M. Lezcano-Casado, Y. Li, Z. Li, J. Lin, J. Liss, Lily, Liu, J. Liu, K. Lu, C. Lu, Z. Martinovic, L. McCallum, J. McGrath, S. McKinney, A. McLaughlin, S. Mei, S. Mostovoy, T. Mu, G. Myles, A. Neitz, A. Nichol, J. Pachocki, A. Paino, D. Palmie, A. Pantuliano, G. Parascandolo, J. Park, L. Pathak, C. Paz, L. Peran, D. Pimenov, M. Pokrass, E. Proehl, H. Qiu, G. Raila, F. Raso, H. Ren, K. Richardson, D. Robinson, B. Rotsted, H. Salman, S. Sanjeev, M. Schwarzer, D. Sculley, H. Sikchi, K. Simon, K. Singhal, Y. Song, D. Stuckey, Z. Sun, P. Tillet, S. Toizer, F. Tsimpourlas, N. Vyas, E. Wallace, X. Wang, M. Wang, O. Watkins, K. Weil, A. Wendling, K. Whinnery, C. Whitney, H. Wong, L. Yang, Y. Yang, M. Yasunaga, K. Ying, W. Zaremba, W. Zhan, C. Zhang, B. Zhang, E. Zhang, and S. Zhao Gpt-oss-120b & gpt-oss-20b model card. External Links: 2508.10925, Link Cited by: Table 4.
  • Orlando et al. (2024) R. Orlando, L. Moroni, P. Huguet Cabot, S. Conia, E. Barba, S. Orlandini, G. Fiameni, and R. Navigli Minerva LLMs: the first family of large language models trained from scratch on Italian data. In Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024), F. Dell’Orletta, A. Lenci, S. Montemagni, and R. Sprugnoli (Eds.), Pisa, Italy, pp. 707–719. External Links: Link, ISBN 979-12-210-7060-6 Cited by: Table 6.
  • Qin et al. (2024) L. Qin, Q. Chen, Y. Zhou, Z. Chen, Y. Li, L. Liao, M. Li, W. Che, and P. S. Yu Multilingual large language model: a survey of resources, taxonomy and frontiers. External Links: 2404.04925, Link Cited by: §1.
  • Rudolph et al. (2023) J. Rudolph, S. Tan, and S. Tan ChatGPT: bullshit spewer or the end of traditional assessments in higher education?. Journal of Applied Learning & Teaching 6 (1), pp. 342–363. External Links: Link Cited by: §1.
  • Russell et al. (2026) J. Russell, M. Karpinska, D. Akinode, J. Zhou, K. Thai, B. Emi, M. Spero, and M. Iyyer AI use in American newspapers is widespread, uneven, and rarely disclosed. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 14554–14580. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §1.
  • Salkar et al. (2022) N. Salkar, T. Trikalinos, B. Wallace, and A. Nenkova Self-repetition in abstractive neural summarizers. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Y. He, H. Ji, S. Li, Y. Liu, and C. Chang (Eds.), Online only, pp. 341–350. External Links: Link, Document Cited by: 3rd item, §3.5.
  • Shaib et al. (2025) C. Shaib, V. S. Govindarajan, J. Barrow, J. Sun, A. Siu, B. C. Wallace, and A. Nenkova Standardizing the measurement of text diversity: a tool and comparative analysis. In Proceedings of The 14th International Joint Conference on Natural Language Processing and The 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics: System Demonstrations, X. Liu and A. Purwarianti (Eds.), Mumbai, India, pp. 36–46. External Links: Link, Document, ISBN 979-8-89176-301-2 Cited by: 4th item, §3.5.
  • Shetty et al. (2026) A. Shetty, Q. Xu, O. Ohrimenko, and J. H. Lau Who wrote the book? detecting and attributing llm ghostwriters. External Links: 2603.28054, Link Cited by: §A.3, Appendix B, Appendix B, Appendix C, Appendix C, §1, §2, Table 1, §3.2, §3.2, §3.3, §3.4, §3.5, §4, §4, §5.
  • Shliazhko et al. (2024) O. Shliazhko, A. Fenogenova, M. Tikhonova, A. Kozlova, V. Mikhailov, and T. Shavrina MGPT: few-shot learners go multilingual. Transactions of the Association for Computational Linguistics 12, pp. 58–79. External Links: Link, Document Cited by: 5th item, Appendix B, §4.
  • Sun et al. (2022) X. Sun, Z. Sun, Y. Meng, J. Li, and C. Fan Summarize, outline, and elaborate: long-text generation via hierarchical supervision from extractive summaries. In Proceedings of the 29th International Conference on Computational Linguistics, N. Calzolari, C. Huang, H. Kim, J. Pustejovsky, L. Wanner, K. Choi, P. Ryu, H. Chen, L. Donatelli, H. Ji, S. Kurohashi, P. Paggio, N. Xue, S. Kim, Y. Hahm, Z. He, T. K. Lee, E. Santus, F. Bond, and S. Na (Eds.), Gyeongju, Republic of Korea, pp. 6392–6402. External Links: Link Cited by: §3.2.
  • Tao et al. (2024) Z. Tao, Y. Chen, D. Xi, Z. Li, and W. Xu Towards reliable detection of llm-generated texts: a comprehensive evaluation framework with cudrt. External Links: 2406.09056, Link Cited by: §1.
  • Thai et al. (2026) K. Thai, B. Emi, E. Masrour, and M. Iyyer EditLens: quantifying the extent of AI editing in text. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
  • Tyo et al. (2023) J. Tyo, B. Dhingra, and Z. C. Lipton Valla: standardizing and benchmarking authorship attribution and verification through empirical evaluation and comparative analysis. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), J. C. Park, Y. Arase, B. Hu, W. Lu, D. Wijaya, A. Purwarianti, and A. A. Krisnadhi (Eds.), Nusa Dua, Bali, pp. 649–660. External Links: Link, Document Cited by: Appendix C.
  • Uchendu et al. (2021) A. Uchendu, Z. Ma, T. Le, R. Zhang, and D. Lee TURINGBENCH: a benchmark environment for Turing test in the age of neural text generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Punta Cana, Dominican Republic, pp. 2001–2016. External Links: Link, Document Cited by: §1, §2, Table 1.
  • Wang et al. (2025) W. Wang, M. Gao, X. Hu, and X. Wan Towards a “novel” benchmark: evaluating literary fiction with large language models. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 21648–21673. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §A.3, §3.5.
  • Wang et al. (2024) Y. Wang, J. Mansurov, P. Ivanov, J. Su, A. Shelmanov, A. Tsvigun, O. Mohammed Afzal, T. Mahmoud, G. Puccetti, T. Arnold, A. Aji, N. Habash, I. Gurevych, and P. Nakov M4GT-bench: evaluation benchmark for black-box machine-generated text detection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 3964–3992. External Links: Link, Document Cited by: §2, Table 1.
  • Wu et al. (2026) J. Wu, Y. Liu, C. Zhu, H. Zhang, Z. Wu, T. Shi, Y. Du, L. Wang, W. Luo, J. Su, and D. F. Wong DetectRL-X: towards reliable multilingual and real-world LLM-generated text detection. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 38247–38294. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §2.
  • Wu et al. (2025) J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, and D. F. Wong A survey on llm-generated text detection: necessity, methods, and future directions. Computational Linguistics 51 (1), pp. 275–338. Cited by: §1.
  • Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: Table 4.
  • Yang et al. (2024) X. Yang, W. Cheng, Y. Wu, L. R. Petzold, W. Y. Wang, and H. Chen DNA-GPT: divergent n-gram analysis for training-free detection of GPT-generated text. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §2.
  • Yue et al. (2025) X. Yue, Y. Song, A. Asai, S. Kim, J. NYANDWI, S. Khanuja, A. Kantharuban, L. Sutawika, S. Ramamoorthy, and G. Neubig Pangea: a fully open multilingual multimodal llm for 39 languages. In International Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025, pp. 47758–47811. External Links: Link Cited by: §1.
  • Zhu et al. (2018) Y. Zhu, S. Lu, L. Zheng, J. Guo, W. Zhang, J. Wang, and Y. Yu Texygen: a benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, New York, NY, USA, pp. 1097–1100. External Links: ISBN 9781450356572, Link, Document Cited by: 1st item, §3.5.

Appendix

Appendix A MultiGhostBench Generation

Tag Genre
ACM Art, Culture & Media
BLP Business, Law & Politics
HB History & Biographies
JR Journals & Reports
LF Literature & Fiction
LHH Lifestyle, Health & Hobbies
SSP Social Sciences & Philosophy
ST Science & Technology
Table 3: Different genres used in this work, sourced from Project Gutenberg (Momen et al., 2025).
Model Name Context Type Checkpoint
deepseek-v3.2 (DeepSeek-AI et al., 2025) 163.8K
Open
deepseek/deepseek-v3.2
gpt-oss (OpenAI et al., 2025) 128.0K
Open
openai/gpt-oss-120b
qwen3-235b (Yang et al., 2025) 262.1K
Open
qwen/qwen3-235b-a22b-2507
gemini-flash (Comanici et al., 2025) 1.05M
Closed
google/gemini-flash
gemini-pro (Comanici et al., 2025) 1.05M
Closed
google/gemini-pro
Table 4: LLMs checkpoints used in MultiGhostBench. ‘Context’ is the total input context supported. ‘Type’ indicates whether model weights are public (
Open) or if it is proprietary LLM (
Closed).
Model Split Genres IT ZH ES DE EN RU
gemini-pro ID LHH, HB 18 8 10 18 6 11
ID HB 11 9 12 6 11 10
OOD-Domain LF 8 6 6 7 5 6
gemini-flash ID HB, ACM 30 16 30 25 31 19
OOD-Domain LHH, LF 8 5 8 7 6 6
deepseek-v3.2 ID LF 19 29 27 23 28 23
OOD-Domain HB 5 8 7 6 8 7
qwen3-235b ID SSP, LF 8 13 8 18 10 8
ID LF 9 12 11 6 10 11
OOD-Domain BLP, HB 2 1 1 2 4 2
OOD-Domain HB 1 2 2 1 3 2
OOD-Domain BLP, LHH, HB 1 4 3 4 1 2
gpt-oss ID HB, LF 24 18 22 16 24 18
OOD-Domain BLP, HB, SSP 7 5 6 5 6 5
Table 5: Detailed ID and OOD-Domain genre splits by generating model, with the number of books reported for each language. Development-set books are excluded from this breakdown.

A.1 Book Generation Prompts

For brevity, they are not listed here; they can be found in https://github.com/GrecoMT/MultiGhostBench/blob/main/prompts.py.

A.2 Quantitative Analysis

Lang. Tokeniser Reference
IT Minerva-7B Orlando et al. (2024)
DE GBERT Chan et al. (2020)
ES BETO Cañete et al. (2023)
ZH Chinese-BERT Devlin et al. (2019)
EN GPT-2 Maintainers (2022)
RU Russian-BERT Kuratov and Arkhipov (2019)
Table 6: Language-dependent tokenisers used for token-based quantitative textual metrics.
Lang Author PPL ↓\downarrow S-B ↓\downarrow Self-R ↓\downarrow NGD ↑\uparrow CR ↓\downarrow
IT gemini-pro 17.48 0.016 11.36 1.38 2.75
gemini-flash 14.99 0.017 10.12 1.58 2.86
deepseek-v3.2 18.31 0.027 10.59 1.48 3.14
qwen3-235b 18.97 0.019 11.26 1.37 2.75
gpt-oss 15.59 0.019 11.32 1.33 2.92
Human 30.07 0.004 11.08 1.484 2.56
ES gemini-pro 19.76 0.019 10.47 1.55 2.72
gemini-flash 17.00 0.014 9.43 1.65 2.85
deepseek-v3.2 20.89 0.013 11.44 1.39 2.81
qwen3-235b 20.29 0.019 9.44 1.73 2.71
gpt-oss 16.67 0.011 11.29 1.21 3.02
Human 30.01 0.004 11.04 1.39 2.73
DE gemini-pro 20.12 0.021 10.19 1.67 2.70
gemini-flash 16.52 0.020 9.51 1.73 2.88
deepseek-v3.2 21.49 0.017 11.53 1.38 2.96
qwen3-235b 23.87 0.014 10.33 1.58 2.74
gpt-oss 18.47 0.029 11.04 1.48 2.89
Human 39.62 0.004 9.33 1.76 2.68
EN gemini-pro 30.67 0.024 10.68 1.59 2.68
gemini-flash 27.81 0.016 9.92 1.65 2.82
deepseek-v3.2 31.43 0.016 11.80 1.40 2.75
qwen3-235b 28.74 0.013 11.63 1.29 2.68
gpt-oss 28.84 0.016 11.44 1.39 2.79
Human 31.36 0.001 10.25 2.14 2.65
RU gemini-pro 15.46 0.020 11.47 1.61 3.66
gemini-flash 12.16 0.022 9.58 1.90 3.98
deepseek-v3.2 15.66 0.016 10.21 1.56 3.69
qwen3-235b 16.07 0.022 11.78 1.44 3.96
gpt-oss 16.98 0.022 11.88 1.63 3.63
ZH gemini-pro 19.19 0.027 11.92 1.10 2.27
gemini-flash 14.16 0.021 10.54 1.20 2.50
deepseek-v3.2 24.55 0.019 12.27 1.08 2.29
qwen3-235b 32.17 0.010 11.60 1.09 2.33
gpt-oss 21.40 0.022 11.64 1.00 2.56
Table 7: Average quantitative textual metrics for human and LLM-generated books across languages. The human reference consists of 100 books randomly sampled per language from the Project Gutenberg corpus released by Momen et al. (2025), retaining only full-length, single-authored works originally written in the target language and published during the nineteenth or twentieth century, while excluding translations, short texts, and multi-authored works. Human values are reported only for languages for which a suitable reference sample was available.

We evaluate the generated books using several reference-free metrics capturing lexical diversity, repetition, and redundancy. Given the multilingual setting, token-based metrics are computed using language-specific tokenisers to reduce distortions caused by differences in tokenisation across languages. The analysis includes the following metrics:

  • •

    Self-BLEU (S-B): measures the lexical similarity between books generated by the same model by computing BLEU scores between each book and the remaining outputs (Zhu et al., 2018). Higher values indicate greater overlap and, consequently, lower inter-book diversity.

  • •

    N-gram Diversity (NGD): measures the variety of lexical patterns in the generated texts based on the proportion of distinct nn-grams (Meister et al., 2023). We consider nn-grams of length four, with higher values indicating greater diversity.

  • •

    Self-Repetition (Self-R): measures the extent to which nn-grams are repeated within the generated texts, capturing the tendency of a model to reuse the same lexical patterns (Salkar et al., 2022). Lower values indicate less repetition.

  • •

    Compression Ratio (CR): measures text redundancy through gzip compression (Shaib et al., 2025). More repetitive and predictable texts are generally more compressible, making this metric complementary to nn-gram-based measures.

  • •

    Perplexity (PPL): measures how predictable a text is according to a reference language model. We compute perplexity using mGPT (Shliazhko et al., 2024), providing a shared multilingual evaluation model across languages. Lower values indicate that the text is more predictable for the reference model.

A.3 Human Evaluation

Lang. Passage-level Overall
Rel. Eng. Coh. Flu. Div. Coh. Emp. Sur. Eng. Comp.
EN 4.8 4.4 4.9 4.6 4.7 4.5 3.5 4.0 4.0 3.5
DE 4.9 3.7 4.4 4.0 3.6 5.0 4.5 3.0 3.0 3.5
IT 4.8 4.1 4.6 4.2 4.8 5.0 5.0 3.5 4.0 4.0
ES 3.7 3.3 4.7 4.5 4.6 3.5 4.0 3.5 3.5 4.0
RU 4.9 2.6 3.6 3.0 2.8 3.5 1.5 2.5 1.5 2.0
ZH 4.9 4.3 4.5 4.6 4.6 3.5 4.5 2.5 3.5 3.5
Overall 4.66 3.73 4.45 4.15 4.18 4.16 3.83 3.16 3.25 3.41
Table 8: Human evaluation scores across passage-level and overall narrative quality dimensions for gemini-pro generated books in the Literature & Fiction (LF) genre.

Following Shetty et al. (2026); Chhun et al. (2022); Wang et al. (2025), we evaluate the generated books along several dimensions capturing linguistic, narrative, and structural properties at the micro-, meso-, and macro-levels:

  • •

    Relevance (Rel.): measures the degree to which the generated text follows the given instructions and remains consistent with the specified book outline.

  • •

    Engagement (Eng.): measures the extent to which the text sustains the reader’s interest and attention.

  • •

    Coherence (Coh.): evaluates how logically the narrative develops, considering the consistency of events, characters, and settings, as well as the progression from beginning to end.

  • •

    Fluency (Flu.): measures the grammatical correctness, readability, and naturalness of the language.

  • •

    Diversity (Div.): measures the degree of variation in vocabulary, sentence structure, and linguistic expression.

  • •

    Empathy (Emp.): evaluates how effectively the text conveys the characters’ emotions, motivations, and reactions.

  • •

    Surprise (Sur.): measures the degree to which the ending introduces unexpected or non-trivial narrative developments.

  • •

    Complexity (Comp.): measures the elaborateness of the narrative in terms of plot structure, character development, and overall sophistication.

Annotators.

The evaluation was conducted by volunteer annotators from our university, representing diverse demographic backgrounds. Each book was independently evaluated by two annotators, both fluent in the language of the evaluated text. Annotators received detailed instructions describing the task and evaluation criteria. Participation was voluntary and uncompensated, and all annotators were informed about the purpose of the study and their role in the evaluation.

Annotation task.

Given the length of the generated books, five passages were randomly sampled from each book while preserving the narrative flow. Annotators answered five questions per passage and five book-level questions, for a total of 3030 questions per book, using a 55-point Likert scale. Results are reported in Table 8. The evaluation provides a limited assessment of the linguistic and narrative quality of the generated books, with potential confounding factors including annotator subjectivity and the restricted coverage of models and genres.

Results.

Table 8 reports the average passage-level and overall scores. The generated books achieve strong passage-level results for relevance (4.66), coherence (4.45), fluency (4.15), and diversity (4.18), while engagement is lower (3.73). At the overall level, coherence receives the highest score (4.16), whereas the more subjective dimensions of engagement (3.25) and surprise (3.16) are weaker, indicating more predictable narratives. These results complement the automatic textual analysis (§3.5) and indicate that the generated passages are generally fluent, coherent, diverse, and aligned with the intended narrative constraints. However, the results vary across languages. In particular, Russian books obtain lower ratings across several dimensions but, importantly, maintain high relevance (4.94.9). Inter-annotator agreement remains low (avg. 0.2140.214) for a highly subjective task, varying across languages. Nevertheless, most paired ratings remain close across languages on avg. 84.484.4% by at most one point. The limited scale of the evaluation prevents broader language-specific conclusions. Overall, the results indicate adequate linguistic and narrative quality of books in MultiGhostBench.

Method Low Resource High Resource
EN DE IT ES RU ZH EN DE IT ES RU ZH
rank 0.316 0.440 0.332 0.408 0.420 0.420 0.340 0.488 0.240 0.460 0.340 0.332
entropy 0.332 0.472 0.336 0.396 0.332 0.332 0.336 0.492 0.240 0.436 0.368 0.360
gLTR 0.352 0.408 0.336 0.424 0.436 0.320 0.372 0.488 0.256 0.436 0.304 0.372
n-gram 0.577 0.433 0.514 0.667 0.325 0.470 0.757 0.937 0.946 0.964 0.658 0.874
xlm-roberta 0.610 0.660 0.520 0.490 0.610 0.600 0.660 0.860 0.490 0.670 0.670 0.740
detective 0.840 0.720 0.680 0.850 0.658 0.740 0.890 0.850 0.630 0.690 0.950 0.930
tracerank-js{}_{\texttt{rank-js}} 0.946 0.916 0.940 0.914 0.900 0.922 0.945 0.918 0.937 0.927 0.934 0.930
traceentr-js{}_{\texttt{entr-js}} 0.900 0.933 0.957 0.939 0.911 0.953 0.965 0.954 0.935 0.965 0.955 0.961
traceentr-norm{}_{\texttt{entr-norm}} -2.715 -4.460 -2.505 -2.680 -3.125 -2.710 -2.300 -2.840 -3.065 -2.265 -2.175 -2.335
Table 9: Thresholds for each method and language under the low- and high-resource settings.

A.4 Cost Analysis

Appendix Table 10 lists generation costs per LLM. In total, generating MultiGhostBench cost approximately $972972.

LLM API Cost ($/11M) Total
Input Output (in $)
gemini-pro 1.251.25 10.0010.00 760.0760.0
gemini-flash 0.300.30 2.502.50 61.661.6
deepseek-v3.2 0.200.20 0.310.31 58.158.1
qwen3-235b 0.090.09 0.550.55 45.745.7
gpt-oss 0.030.03 0.170.17 46.446.4
Total 971.8971.8
Table 10: MultiGhostBench generation cost breakdown. ‘Input’ and ‘Output’ denote API cost per 11M tokens. ‘Total’ is the cost of generating all books for a given LLM. OpenRouter API costs are as of April 2026.

Appendix B Authorship Attribution Methods (cont.)

rank.

For each text, we compute the average rank of the observed tokens.

entropy.

Similarly, we compute the average token-level entropy.

gLTR.

gLTR (Gehrmann et al., 2019) represents a text through the proportion of tokens falling into four predefined rank buckets.

For all three metric-based detectors, we follow La Cava and Tagarelli (2025); Shetty et al. (2026) and adapt the extracted statistics to multi-class attribution by training a logistic regression classifier to predict the generating LLM. Following Macko et al. (2023), we use mGPT (Shliazhko et al., 2024) as the multilingual reference language model.

n-gram.

n-gram (Koppel and Schler, 2004) uses three complementary feature representations. The char analyser preserves letters, punctuation, and spaces. The dist_char analyser replaces letters with a common placeholder, retaining structural patterns involving punctuation, numbers, symbols, and spacing. The word analyser removes punctuation and operates on word-level features. To preserve language-specific information, accented and language-specific characters are retained during preprocessing. For Chinese, word-level features are extracted after segmenting the text with Jieba.55 5 https://pypi.org/project/jieba/

bert-aa.

A multi-class classification layer is added on top of the pre-trained encoder, and the complete architecture is fine-tuned to predict the generating LLM. In our implementation, the underlying encoder is xlm-roberta (Conneau et al., 2020).

detective.

detective (Guo et al., 2024) relies on sentence embeddings learned through a multi-level contrastive objective. For AA, these sentence embeddings are used to train a multi-class classifier over the generating LLMs. The original sentence-embedding model is replaced with ZurichNLP/unsup-simcse-xlm-roberta-base,66 6 https://huggingface.co/ZurichNLP/unsup-simcse-xlm-roberta-base its multilingual counterpart. We also considered the more recent M-RangeDetector (Jiao et al., 2025) and simpler contrastive-learning approaches (John and Idicula, 2026). However, their implementations are not publicly available, and the methodological details provided were insufficient for reliable reproduction.

trace.

For each text, trace (Shetty et al., 2026) models transitions between consecutive token-rank or entropy values, producing a two-dimensional fingerprint. A reference fingerprint is constructed for each generator from its training texts, and a test text is attributed to the generator whose reference fingerprint is most similar to the fingerprint extracted from the test text. tracerank-js{}_{\texttt{rank-js}} uses rank-based transition fingerprints, whereas traceentr-js{}_{\texttt{entr-js}} and traceentr-norm{}_{\texttt{entr-norm}} use entropy-based fingerprints and compare them through Jensen-Shannon divergence and norm-based distance, respectively. We use Gemma 77 7 https://huggingface.co/google/gemma-3-1b-it as the evaluator language model.

Appendix C Experimental Details

We use OpenRouter88 8 https://openrouter.ai/ for all LLM API calls used to generate the benchmark. We adopt the same decoding parameters for all generators, setting temperature to 1.01.0, top-p to 1.01.0, top-k to 00, and maximum_output_length to 81928192 tokens. We provide a model card for each LLM used in this work in Table 4. All experiments were conducted using a single A100 GPU with CUDA 12.4 and PyTorch 2.10.0.

Since xlm-roberta cannot process complete long-form documents directly, we split each text into chunks of 512512 tokens and average the chunk-level prediction scores to obtain the final document-level prediction, following Tyo et al. (2023); Shetty et al. (2026). For detective, since it is clustering-based, we compute the proportion of top-KK results assigned to each generator across the text chunks.

All hyperparameter tuning was performed on the development set. The number of fine-tuning epochs depends on the experimental configuration. For xlm-roberta, we fine-tune the model for up to 1010 epochs and retain the checkpoint achieving the highest macro-F1 on the development set. For detective, we train for up to 5050 epochs and similarly select the checkpoint with the best development-set macro-F1. For all trace variants, we follow the hyperparameter configuration of Shetty et al. (2026). Specifically, for tracerank-js{}_{\texttt{rank-js}}, we set α=1.5\alpha=1.5 and use 100100K samples to approximate the power-law distribution. We use Gemma as the evaluator language model with a context length of 10241024 tokens. The entropy-based variants use a grid size of 5050, whereas tracerank-js{}_{\texttt{rank-js}} uses 5050 clusters.

Appendix D Additional Results

This section provides additional analyses and complete results for all evaluation settings. Figure 4 extends the qualitative analysis of trace by comparing entropy fingerprints from Italian training texts with Italian ID, Italian OOD-Domain, and Spanish and Chinese OOD-Language texts. Table 11 and Table 12 report the threshold-based in-language results for ID, OOD-Domain, and OOD-Author in the low- and high-resource settings, while Table 15 and Table 16 report the corresponding changes in macro-F1 relative to ID performance. Table 13 and Table 14 provide the ID and OOD-Domain results without threshold-based rejection. Finally, Table 17 and Table 18 report OOD-Language results across all source–target language pairs, with and without thresholding, under both resource settings.

Refer to caption
Figure 4: UMAP projections of trace entropy fingerprints using Italian as the reference language. The panels compare Italian training fingerprints with (a) Italian ID-test, (b) Italian OOD-Domain, and (c) Spanish and (d) Chinese ID fingerprints evaluated under OOD-Language. Grey points denote the Italian training references, while blue, orange, and green points denote ID, OOD-Domain, and cross-language fingerprints, respectively. Marker shapes identify the LLM generator. Italian ID-test and OOD-Domain fingerprints remain close to the training references, whereas both Spanish and Chinese fingerprints occupy clearly displaced regions, indicating that trace fingerprints are strongly language-dependent and transfer poorly under OOD-Language.
EN DE IT ES RU ZH
Method ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A
rank 0.245 0.047 0 0.114 0.191 0.080 0 0 0.105 0.080 0.066 0 0.094 0.080 0.150 0.066 0.200 0.388
entropy 0.250 0.227 0.028 0 0 0.300 0.111 0.336 0 0.120 0.107 0.176 0.286 0.116 0 0.290 0.289 0.072
gLTR 0.066 0.050 0 0.090 0.070 0 0.080 0 0 0.100 0.100 0.200 0.111 0.076 0 0.220 0.274 0
n-gram 0.447 0.335 0.500 0.639 0.634 0.309 0.493 0.277 0.060 0.581 0.600 0.974 0.605 0.517 0.040 0.960 0.792 0.843
xlm-roberta 0.700 0.531 0.485 0.620 0.593 0.548 0.842 0.426 0.317 0.886 0.654 0.280 0.661 0.630 0.429 0.882 0.872 0.696
detective 0.494 0.398 0.527 0.511 0.525 0.323 0.512 0.341 0.172 0.665 0.515 0.505 0.506 0.381 0.327 0.891 0.766 0.487
tracerank-js{}_{\texttt{rank-js}} 0.486 0.271 0.614 0.479 0.396 0.080 0.709 0.500 0.484 0.391 0.460 0.206 0.718 0.516 0.033 0.898 0.754 0.328
traceentr-js{}_{\texttt{entr-js}} 0.587 0.472 0.080 0.868 0.702 0.246 0.588 0.374 0.660 0.801 0.707 0.501 0.566 0.302 0.783 0.791 0.856 0.744
traceentr-norm{}_{\texttt{entr-norm}} 0.641 0.489 0.719 0.916 0.790 0.030 0.540 0.435 0.740 0.774 0.542 0.774 0.413 0.253 0.888 0.829 0.892 0.650
Table 11: Threshold-based low-resource results for ID, OOD-Domain, and OOD-Author evaluation in each language.
EN DE IT ES RU ZH
Method ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A ID OOD-D OOD-A
rank 0 0.142 0.200 0.057 0.125 0.317 0.095 0.046 0 0.080 0.066 0 0.166 0.093 0.114 0.114 0.118 0
entropy 0.070 0.120 0.120 0.023 0 0.200 0.180 0.129 0 0.120 0.107 0.120 0.192 0.100 0.036 0.277 0.303 0.156
gLTR 0.040 0.168 0.200 0.106 0.149 0.257 0.197 0.308 0 0.277 0.150 0.164 0.200 0.267 0 0.125 0.278 0.364
n-gram 1.000 1.000 0.980 0.803 0.659 1.000 0.896 0.725 1.000 0.869 0.633 0.992 1.000 0.966 0.733 0.971 0.959 0.946
xlm-roberta 1.000 0.924 0.334 0.949 0.775 0.660 0.956 0.676 0.203 1.000 0.933 0.441 0.977 0.960 0.474 1.000 1.000 0.653
detective 0.931 0.786 0.681 0.955 0.658 0.707 0.924 0.677 0.477 0.953 0.864 0.518 0.769 0.717 0.761 0.844 0.845 0.866
tracerank-js{}_{\texttt{rank-js}} 0.682 0.660 0.303 0.694 0.490 0.100 0.831 0.684 0.189 0.361 0.466 0.357 0.959 0.501 0.197 0.801 0.838 0.285
traceentr-js{}_{\texttt{entr-js}} 0.878 0.541 0.732 0.959 0.905 0.537 0.810 0.662 0.050 0.736 0.610 0.963 1.000 0.741 0.322 0.933 0.814 0.683
traceentr-norm{}_{\texttt{entr-norm}} 0.966 0.755 0.709 0.959 0.905 0.405 0.893 0.793 0.164 0.869 0.624 0.922 0.931 0.737 0.504 0.880 0.945 0.592
Table 12: Threshold-based high-resource results for ID, OOD-Domain, and OOD-Author evaluation in each language.
EN DE IT ES RU ZH
Method ID OOD-D ID OOD-D ID OOD-D ID OOD-D ID OOD-D ID OOD-D
rank 0.261 0.045 0.209 0.191 0.140 0.007 0.075 0.068 0.074 0.066 0.071 0.070
entropy 0.248 0.207 0.068 0.063 0.085 0.294 0.075 0.069 0.286 0.116 0.231 0.247
gLTR 0.050 0.052 0.071 0.064 0.077 0.078 0.075 0.068 0.074 0.066 0.227 0.236
n-gram 0.476 0.483 0.625 0.616 0.490 0.374 0.699 0.627 0.591 0.582 1.000 0.906
xlm-roberta 0.780 0.531 0.696 0.595 0.842 0.456 0.893 0.666 0.814 0.566 0.876 0.874
detective 0.476 0.475 0.500 0.600 0.482 0.305 0.617 0.650 0.488 0.470 0.915 0.826
tracerank-js{}_{\texttt{rank-js}} 0.440 0.294 0.479 0.376 0.694 0.579 0.395 0.420 0.718 0.509 0.926 0.787
traceentr-js{}_{\texttt{entr-js}} 0.587 0.469 0.868 0.702 0.557 0.408 0.786 0.724 0.708 0.328 0.861 0.876
traceentr-norm{}_{\texttt{entr-norm}} 0.621 0.562 0.916 0.790 0.579 0.446 0.875 0.699 0.733 0.594 0.949 0.920
Table 13: Low-resource results without threshold-based rejection for ID and OOD-Domain evaluation in each language.
EN DE IT ES RU ZH
Method ID OOD-D ID OOD-D ID OOD-D ID OOD-D ID OOD-D ID OOD-D
rank 0.070 0.078 0.068 0.060 0.070 0.055 0.075 0.068 0.074 0.082 0.082 0.082
entropy 0.129 0.080 0.068 0.063 0.180 0.150 0.075 0.167 0.228 0.169 0.300 0.300
gLTR 0.070 0.078 0.068 0.063 0.347 0.352 0.220 0.223 0.348 0.313 0.082 0.236
n-gram 1.000 1.000 1.000 1.000 1.000 0.974 1.000 0.972 1.000 1.000 1.000 1.000
xlm-roberta 1.000 1.000 0.953 0.819 0.956 0.676 1.000 0.946 1.000 0.968 1.000 1.000
detective 1.000 1.000 1.000 0.850 0.924 0.673 0.966 0.946 1.000 0.968 1.000 0.971
tracerank-js{}_{\texttt{rank-js}} 0.719 0.771 0.694 0.536 0.831 0.734 0.415 0.550 0.959 0.494 0.847 0.888
traceentr-js{}_{\texttt{entr-js}} 1.000 0.710 0.959 0.905 0.810 0.662 0.919 0.800 1.000 0.741 0.949 0.935
traceentr-norm{}_{\texttt{entr-norm}} 1.000 0.850 0.959 0.905 0.893 0.793 0.919 0.800 0.959 0.690 1.000 1.000
Table 14: High-resource results without threshold-based rejection for ID and OOD-Domain evaluation in each language.
Detector EN DE IT ES RU ZH Avg.
ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}}
rank −0.198-0.198 −0.245-0.245 0.0770.077 −0.034-0.034 0.0000.000 0.1050.105 −0.014-0.014 −0.080-0.080 −0.014-0.014 0.0560.056 0.1340.134 0.3220.322 −0.003-0.003 0.0210.021
entropy −0.023-0.023 −0.222-0.222 0.0000.000 0.3000.300 0.2250.225 −0.111-0.111 −0.013-0.013 0.0560.056 −0.170-0.170 −0.286-0.286 −0.001-0.001 −0.218-0.218 0.0030.003 −0.080-0.080
gLTR −0.016-0.016 −0.066-0.066 −0.020-0.020 −0.090-0.090 −0.080-0.080 −0.080-0.080 0.0000.000 0.1000.100 −0.035-0.035 −0.111-0.111 0.0540.054 −0.220-0.220 −0.016-0.016 −0.078-0.078
n-gram −0.112-0.112 0.0530.053 −0.005-0.005 −0.330-0.330 −0.216-0.216 −0.433-0.433 0.0190.019 0.3930.393 −0.088-0.088 −0.565-0.565 −0.168-0.168 −0.117-0.117 −0.095-0.095 −0.166-0.166
xlm-roberta −0.169-0.169 −0.215-0.215 −0.027-0.027 −0.072-0.072 −0.416-0.416 −0.525-0.525 −0.232-0.232 −0.606-0.606 −0.031-0.031 −0.232-0.232 −0.010-0.010 −0.186-0.186 −0.148-0.148 −0.306-0.306
detective −0.096-0.096 0.0330.033 0.0140.014 −0.188-0.188 −0.171-0.171 −0.340-0.340 −0.150-0.150 −0.160-0.160 −0.125-0.125 −0.179-0.179 −0.125-0.125 −0.404-0.404 −0.109-0.109 −0.206-0.206
tracerank-js{}_{\texttt{rank-js}} −0.215-0.215 0.1280.128 −0.083-0.083 −0.399-0.399 −0.209-0.209 −0.225-0.225 0.0690.069 −0.185-0.185 −0.202-0.202 −0.685-0.685 −0.144-0.144 −0.570-0.570 −0.131-0.131 −0.323-0.323
traceentr-js{}_{\texttt{entr-js}} −0.115-0.115 −0.507-0.507 −0.166-0.166 −0.622-0.622 −0.214-0.214 0.0720.072 −0.094-0.094 −0.300-0.300 −0.264-0.264 0.2170.217 0.0650.065 −0.047-0.047 −0.131-0.131 −0.198-0.198
traceentr-norm{}_{\texttt{entr-norm}} −0.152-0.152 0.0780.078 −0.126-0.126 −0.886-0.886 −0.105-0.105 0.2000.200 −0.232-0.232 0.0000.000 −0.160-0.160 0.4750.475 0.0630.063 −0.179-0.179 −0.119-0.119 −0.052-0.052
Table 15: Changes in macro-F1 under OOD evaluation relative to ID performance in the low-resource setting. For each detector and language, the table reports ΔD=F1OOD​-​D−F1ID\Delta_{\mathrm{D}}=F_{1}^{\mathrm{OOD\text{-}D}}-F_{1}^{\mathrm{ID}} and ΔA=F1OOD​-​A−F1ID\Delta_{\mathrm{A}}=F_{1}^{\mathrm{OOD\text{-}A}}-F_{1}^{\mathrm{ID}}. Negative values indicate degradation, while positive values indicate improvement. The Avg. columns report averages across languages.
Detector EN DE IT ES RU ZH Avg.
ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}} ΔD\Delta_{\mathrm{D}} ΔA\Delta_{\mathrm{A}}
rank 0.1420.142 0.2000.200 0.0680.068 0.2600.260 −0.049-0.049 −0.095-0.095 −0.014-0.014 0.0400.040 −0.073-0.073 −0.052-0.052 0.0040.004 0.0420.042 0.0130.013 0.0660.066
entropy 0.0500.050 0.1300.130 −0.023-0.023 −0.023-0.023 −0.051-0.051 −0.180-0.180 −0.013-0.013 0.0440.044 −0.079-0.079 −0.139-0.139 0.0260.026 0.0870.087 −0.015-0.015 −0.014-0.014
gLTR 0.1280.128 0.1600.160 0.0430.043 0.1510.151 0.1110.111 −0.197-0.197 −0.127-0.127 −0.113-0.113 0.0670.067 −0.200-0.200 0.1530.153 −0.079-0.079 0.0630.063 −0.046-0.046
n-gram 0.0000.000 −0.020-0.020 −0.144-0.144 0.1970.197 −0.171-0.171 0.1040.104 −0.236-0.236 0.1230.123 −0.034-0.034 −0.268-0.268 −0.012-0.012 −0.025-0.025 −0.100-0.100 0.0190.019
xlm-roberta −0.076-0.076 −0.666-0.666 −0.174-0.174 −0.289-0.289 −0.280-0.280 −0.753-0.753 −0.067-0.067 −0.559-0.559 −0.071-0.071 −0.318-0.318 0.0000.000 −0.317-0.317 −0.111-0.111 −0.484-0.484
detective −0.145-0.145 −0.250-0.250 −0.297-0.297 −0.248-0.248 −0.247-0.247 −0.447-0.447 −0.089-0.089 −0.435-0.435 −0.052-0.052 −0.008-0.008 0.0010.001 0.0220.022 −0.138-0.138 −0.228-0.228
tracerank-js{}_{\texttt{rank-js}} −0.230-0.230 −0.640-0.640 −0.182-0.182 −0.692-0.692 −0.147-0.147 −0.642-0.642 0.1050.105 −0.004-0.004 −0.200-0.200 −0.679-0.679 0.0370.037 −0.516-0.516 −0.103-0.103 −0.529-0.529
traceentr-js{}_{\texttt{entr-js}} −0.337-0.337 −0.146-0.146 −0.054-0.054 −0.422-0.422 −0.148-0.148 −0.760-0.760 −0.126-0.126 0.2270.227 −0.259-0.259 −0.037-0.037 −0.119-0.119 −0.250-0.250 −0.174-0.174 −0.231-0.231
traceentr-norm{}_{\texttt{entr-norm}} −0.211-0.211 −0.257-0.257 −0.054-0.054 −0.554-0.554 −0.100-0.100 −0.729-0.729 −0.245-0.245 0.0530.053 −0.194-0.194 −0.427-0.427 0.0650.065 −0.288-0.288 −0.123-0.123 −0.367-0.367
Table 16: Changes in macro-F1 under OOD evaluation relative to ID performance in the high-resource setting. For each detector and language, the table reports ΔD=F1OOD​-​D−F1ID\Delta_{\mathrm{D}}=F_{1}^{\mathrm{OOD\text{-}D}}-F_{1}^{\mathrm{ID}} and ΔA=F1OOD​-​A−F1ID\Delta_{\mathrm{A}}=F_{1}^{\mathrm{OOD\text{-}A}}-F_{1}^{\mathrm{ID}}. Negative values indicate degradation, while positive values indicate improvement. The Avg. columns report averages across languages.
Train Method Low Resource High Resource
EN DE IT ES RU ZH EN DE IT ES RU ZH
EN rank – 0.068 0.055 0.053 0.061 0.071 – 0 0 0 0 0
entropy – 0.068 0.055 0.053 0.061 0.071 – 0.030 0.055 0.053 0.061 0
gLTR – 0 0 0.033 0.05 0.044 – 0.103 0 0 0 0
n-gram – 0 0 0 0 0 – 0 0 0 0 0
xlm-roberta – 0.683 0.611 0.642 0.666 0.353 – 0.916 0.935 0.914 0.950 0.333
detective – 0.542 0.527 0.450 0.477 0.094 – 0.766 0.753 0.793 0.533 0.424
tracerank-js{}_{\texttt{rank-js}} – 0.112 0.404 0.295 0.203 0 – 0 0 0 0 0
traceentr-js{}_{\texttt{entr-js}} – 0 0 0 0 0 – 0 0 0 0 0
traceentr-norm{}_{\texttt{entr-norm}} – 0 0 0 0 0 – 0 0 0 0 0
DE rank 0.070 – 0.068 0.100 0.061 0.080 0.070 – 0 0 0 0
entropy 0 – 0.025 0 0.074 0.033 0 – 0.055 0.077 0.074 0.095
gLTR 0.070 – 0.213 0.175 0.236 0.092 0.072 – 0 0.023 0.057 0.133
n-gram 0.200 – 0.100 0.100 0.114 0.133 0 – 0 0 0 0
xlm-roberta 0.420 – 0.555 0.463 0.466 0.271 0.673 – 0.836 0.848 0.800 0.723
detective 0.550 – 0.650 0.602 0.600 0.100 0.766 – 0.772 0.617 0.394 0.559
tracerank-js{}_{\texttt{rank-js}} 0.057 – 0.038 0.376 0.273 0 0.150 – 0.198 0.283 0.675 0
traceentr-js{}_{\texttt{entr-js}} 0 – 0 0 0 0 0 – 0 0 0 0
traceentr-norm{}_{\texttt{entr-norm}} 0 – 0 0.160 0.08 0 0 – 0 0 0 0
IT rank 0.050 0.068 – 0.083 0 0.025 0.050 0.071 – 0.179 0.064 0.094
entropy 0.071 0.090 – 0 0.106 0.317 0.070 0.095 – 0.078 0.074 0.290
gLTR 0.050 0.090 – 0.125 0.061 0.069 0.248 0.205 – 0.302 0.061 0.260
n-gram 0.080 0.086 – 0.114 0.133 0.076 0 0 – 0 0 0
xlm-roberta 0.585 0.589 – 0.750 0.669 0.765 0.938 0.916 – 0.981 0.824 0.811
detective 0.371 0.468 – 0.471 0.206 0.304 0.895 0.749 – 0.959 0.788 0.813
tracerank-js{}_{\texttt{rank-js}} 0 0 – 0 0.08 0 0 0 – 0.189 0.05 0
traceentr-js{}_{\texttt{entr-js}} 0 0 – 0 0 0 0 0 – 0 0 0
traceentr-norm{}_{\texttt{entr-norm}} 0 0 – 0 0 0 0 0 – 0 0 0
ES rank 0.070 0.071 0.088 – 0.066 0.125 0.070 0.068 0.055 – 0.074 0.082
entropy 0.071 0.083 0.114 – 0 0.063 0.071 0.083 0.114 – 0 0.063
gLTR 0.070 0.088 0 – 0.066 0.109 0.070 0.105 0.133 – 0.114 0.100
n-gram 0 0 0 – 0 0 0 0 0 – 0 0
xlm-roberta 0.947 0.866 0.926 – 0.836 0.756 0.753 0.914 0.857 – 0.949 0.622
detective 0.300 0.266 0.483 – 0.200 0.434 0.681 0.867 0.933 – 0.847 0.681
tracerank-js{}_{\texttt{rank-js}} 0.128 0.131 0.080 – 0.146 0 0.157 0.170 0.174 – 0 0
traceentr-js{}_{\texttt{entr-js}} 0 0 0 – 0 0 0 0 0 – 0 0
traceentr-norm{}_{\texttt{entr-norm}} 0 0 0 – 0 0 0 0 0 – 0 0
RU rank 0.05 0.068 0.085 0.064 – 0.069 0.070 0.068 0.061 0.075 – 0.096
entropy 0.051 0.057 0.148 0.064 – 0.190 0.034 0.061 0.106 0.046 – 0.072
gLTR 0.063 0 0 0.076 – 0.094 0.07 0.095 0.2 0.106 – 0.229
n-gram 0 0 0 0 – 0 0 0 0 0 – 0
xlm-roberta 0.482 0.695 0.710 0.749 – 0.393 0.298 0.531 0.419 0.380 – 0.420
detective 0.063 0.283 0.114 0.250 – 0.146 0.200 0.200 0.200 0.184 – 0
tracerank-js{}_{\texttt{rank-js}} 0.125 0.413 0.305 0.214 – 0 0 0.270 0 0 – 0
traceentr-js{}_{\texttt{entr-js}} 0 0 0 0 – 0 0 0 0 0 – 0
traceentr-norm{}_{\texttt{entr-norm}} 0 0.142 0 0 – 0 0 0 0 0 – 0
ZH rank 0.070 0.068 0.055 0.053 0.061 – 0.080 0 0 0 0 –
entropy 0.070 0.196 0.082 0.048 0.074 – 0.070 0.150 0.055 0.075 0.074 –
gLTR 0.055 0.105 0.077 0.066 0.074 – 0.085 0 0 0 0.066 –
n-gram 0.080 0.068 0.077 0.075 0.061 – 0 0 0 0 0 –
xlm-roberta 0.717 0.871 0.781 0.797 0.950 – 0.315 0.711 0.719 0.715 0.683 –
detective 0.673 0.667 0.520 0.667 0.629 – 0.575 0.550 0.433 0.360 0.346 –
tracerank-js{}_{\texttt{rank-js}} 0 0.080 0 0 0.05 – 0.023 0.109 0 0 0 –
traceentr-js{}_{\texttt{entr-js}} 0 0 0 0 0 – 0 0 0 0 0 –
traceentr-norm{}_{\texttt{entr-norm}} 0 0 0 0 0 – 0 0 0 0 0 –
Table 17: Threshold-based OOD-Language results across source and target languages under low- and high-resource settings. Rows indicate the source language used for training, while columns indicate the target language used for evaluation.
Train Method Low Resource High Resource
EN DE IT ES RU ZH EN DE IT ES RU ZH
EN rank – 0.068 0.055 0.053 0.061 0.071 – 0.068 0.055 0.075 0.074 0.082
entropy – 0.068 0.055 0.053 0.061 0.071 – 0.057 0.055 0.053 0.061 0.071
gLTR – 0.138 0.077 0.064 0.074 0.059 – 0.283 0.055 0.157 0.074 0.228
n-gram – 0.068 0.16 0.064 0.156 0.059 – 0.764 0.490 0.533 0.361 0.139
xlm-roberta – 0.711 0.685 0.705 0.696 0.380 – 0.916 1 1 0.955 0.482
detective – 0.483 0.474 0.466 0.448 0.210 – 0.868 0.846 0.834 0.707 0.348
tracerank-js{}_{\texttt{rank-js}} – 0.168 0.404 0.298 0.300 0.08 – 0.130 0.298 0.172 0.342 0.137
traceentr-js{}_{\texttt{entr-js}} – 0.071 0.386 0.144 0.220 0.144 – 0.281 0.08 0.435 0.245 0.361
traceentr-norm{}_{\texttt{entr-norm}} – 0.409 0.209 0.295 0.331 0.240 – 0.076 0.077 0.4 0.119 0.244
DE rank 0.070 – 0.066 0.118 0.061 0.059 0.070 – 0.055 0.075 0.074 0.082
entropy 0.068 – 0.055 0.075 0.074 0.082 0.032 – 0.055 0.075 0.074 0.082
gLTR 0.070 – 0.176 0.175 0.236 0.085 0.070 – 0.055 0.075 0.074 0.082
n-gram 0.364 – 0.560 0.394 0.23 0.059 0.549 – 0.819 0.842 0.618 0.141
xlm-roberta 0.536 – 0.483 0.591 0.602 0.501 0.826 – 0.959 1 0.916 0.871
detective 0.633 – 0.697 0.681 0.603 0.301 0.840 – 0.878 0.962 0.488 0.668
tracerank-js{}_{\texttt{rank-js}} 0.060 – 0.060 0.384 0.258 0.109 0.155 – 0.225 0.337 0.704 0.226
traceentr-js{}_{\texttt{entr-js}} 0.271 – 0.438 0.290 0.774 0.304 0.264 – 0.375 0.266 0.544 0.355
traceentr-norm{}_{\texttt{entr-norm}} 0.225 – 0.148 0.114 0.532 0.360 0.276 – 0.366 0.391 0.488 0.264
IT rank 0.050 0.068 – 0.064 0.144 0.069 0.050 0.071 – 0.179 0.064 0.094
entropy 0.070 0.083 – 0.045 0.094 0.466 0.070 0.095 – 0.078 0.074 0.285
gLTR 0.050 0.068 – 0.064 0.074 0.059 0.248 0.172 – 0.313 0.088 0.260
n-gram 0.100 0.068 – 0.075 0.061 0.059 0.962 0.667 – 0.963 0.902 0.175
xlm-roberta 0.772 0.696 – 0.760 0.772 0.744 0.938 0.916 – 1 0.783 0.849
detective 0.431 0.432 – 0.445 0.272 0.290 0.959 0.916 – 0.941 0.871 0.832
tracerank-js{}_{\texttt{rank-js}} 0.168 0.095 – 0.146 0.211 0.070 0.370 0.124 – 0.281 0.04 0.080
traceentr-js{}_{\texttt{entr-js}} 0.163 0.418 – 0.460 0.34 0.238 0.344 0.405 – 0.286 0.394 0.144
traceentr-norm{}_{\texttt{entr-norm}} 0.173 0.243 – 0.269 0.306 0.225 0.153 0.327 – 0.426 0.340 0.059
ES rank 0.070 0.068 0.055 – 0.074 0.082 0.070 0.068 0.055 – 0.074 0.082
entropy 0.071 0.068 0.055 – 0.061 0.096 0.071 0.068 0.076 – 0 0.096
gLTR 0.070 0.068 0.055 – 0.074 0.088 0.070 0.080 0.185 – 0.133 0.096
n-gram 0.450 0.566 0.509 – 0.380 0.197 0.729 0.667 0.777 – 0.900 0.289
xlm-roberta 0.969 0.837 0.926 – 0.820 0.706 0.924 1 0.922 – 1 0.597
detective 0.574 0.499 0.461 – 0.324 0.476 0.700 0.916 1 0.324 0.818 0.700
tracerank-js{}_{\texttt{rank-js}} 0.119 0.126 0.246 – 0.382 0.008 0.171 0.163 0.254 – 0.320 0.088
traceentr-js{}_{\texttt{entr-js}} 0.275 0.512 0.317 – 0.398 0.269 0.238 0.513 0.310 – 0.381 0.226
traceentr-norm{}_{\texttt{entr-norm}} 0.393 0.558 0.474 – 0.360 0.280 0.163 0.446 0.327 – 0.136 0.229
RU rank 0.05 0.068 0.077 0.064 – 0.059 0.070 0.068 0.055 0.075 – 0.082
entropy 0.051 0.057 0.148 0.064 – 0.190 0.0347 0.059 0.148 0.046 – 0.152
gLTR 0.05 0.068 0.077 0.064 – 0.059 0.070 0.076 0.176 0.231 – 0.296
n-gram 0.070 0.068 0.055 0.053 – 0.071 0.189 0.235 0.232 0.209 – 0.197
xlm-roberta 0.590 0.805 0.716 0.819 – 0.427 0.291 0.566 0.678 0.592 – 0.344
detective 0.217 0.268 0.360 0.274 – 0.273 0.256 0.591 0.327 0.406 – 0.2
tracerank-js{}_{\texttt{rank-js}} 0.166 0.413 0.305 0.284 – 0.239 0.355 0.603 0.171 0.335 – 0.104
traceentr-js{}_{\texttt{entr-js}} 0.246 0.525 0.247 0.245 – 0.196 0.070 0.420 0.248 0.053 – 0.229
traceentr-norm{}_{\texttt{entr-norm}} 0.163 0.401 0.304 0.100 – 0.210 0.070 0.514 0.26 0.208 – 0.197
ZH rank 0.070 0.068 0.055 0.053 0.061 – 0.070 0.068 0.055 0.075 0.074 –
entropy 0.070 0.236 0.077 0.064 0.074 – 0.070 0.138 0.055 0.075 0.074 –
gLTR 0.060 0.226 0.077 0.064 0.074 – 0.070 0.206 0.156 0.146 0.135 –
n-gram 0.080 0.068 0.077 0.075 0.061 – 0.070 0.189 0.186 0.172 0.133 –
xlm-roberta 0.811 0.953 0.916 0.860 0.959 – 0.756 0.809 0.949 0.833 0.800 –
detective 0.856 0.741 0.843 0.853 0.875 – 0.351 0.561 0.544 0.621 0.712 –
tracerank-js{}_{\texttt{rank-js}} 0.080 0.205 0.077 0.142 0.156 – 0.082 0.225 0.077 0.092 0.094 –
traceentr-js{}_{\texttt{entr-js}} 0.093 0.286 0.149 0.173 0.292 – 0.235 0.223 0.149 0.225 0.209 –
traceentr-norm{}_{\texttt{entr-norm}} 0.260 0.180 0.080 0.173 0.426 – 0.264 0.172 0.077 0.247 0.240 –
Table 18: OOD-Language results across source and target languages under low- and high-resource settings without thresholding. Rows indicate the source language used for training, while columns indicate the target language used for evaluation.