KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training
Abstract
LLMs acquire vast amounts of knowledge during pre-training, but often lack the specialized knowledge needed to answer questions from niche sources such as manuals or technical documents unseen during pre-training. Continued pre-training (CPT) is widely used to inject such knowledge into model parameters. However, niche documents seldom repeat facts, making it difficult for CPT to robustly acquire such knowledge. Recent works address this by generating multiple paraphrases of the new knowledge, but paraphrasing is computationally expensive and typically requires powerful LLMs. In this work, we introduce KItCAT: Knowledge Injection via Corrupted Auto-regressive Training, a lightweight training strategy that reduces the need for paraphrasing in decoder-only LLMs. KItCAT augments standard next-token prediction by stochastically corrupting the input sequence. During training, a random subset of input tokens is replaced with other vocabulary tokens while the original next-token labels are kept unchanged. This simple intervention generates diverse training inputs from each sample, enabling large-scale data augmentation at negligible cost. We show that KItCAT consistently improves over CPT across multiple datasets and model families. Code is available at https://github.com/meghanadhpulivarthi/KItCAT.
1 Introduction
LLMs are pre-trained on massive web corpora, yet they often fail to capture specialized knowledge that is absent or underrepresented on the public web, such as information contained in proprietary manuals or technical documents. Continued Pre-Training (CPT) Ke et al. (2023); Ke et al. (2025) is a common approach for injecting such specialized knowledge into the parameters of the LLMs.
A key challenge, however, is that the niche documents seldom repeat facts or present them in varied contexts. Consequently, CPT receives limited repeated exposure to newly introduced knowledge, making it difficult for the model to robustly internalize these facts Allen-Zhu and Li (2024). In such low-diversity training settings, models can instead overfit to surface-level lexical patterns and spurious correlations that recur across epochs, rather than learning the underlying semantics.
To mitigate the limited diversity in the training documents, recent works propose to synthetically generate multiple paraphrases of the text Ovadia et al. (2024); Yang et al. (2025). While effective, these approaches rely on large proprietary LLMs Achiam et al. (2023); Team et al. (2023), which can be costly and slow when generating sizable synthetic datasets. They may also be infeasible in settings with privacy restrictions.
In this work, we ask - can we introduce effective data diversity without relying on external synthetic data generation? We observe that overfitting in low-diversity regimes is amplified by the fact that the model encounters the exact same training samples in every epoch. If the model latches onto a spurious pattern in one pass, repeated identical exposure reinforces this shortcut, degrading generalization.
Our key insight is that preventing the model from ever seeing the exact same input twice can mitigate this reinforcement effect. Motivated by this idea, we propose KItCAT - Knowledge Injection via Corrupted Auto-regressive Training, a simple yet effective strategy for injecting diversity through controlled corruption of the input. During training, KItCAT perturbs input sequences using one of four corruption schemes (Figure 1): (1) KItCAT-rand, where selected tokens are substituted with randomly sampled vocabulary items; (2) KItCAT-mask, where tokens are replaced with a special mask token (e.g., [MASK]); (3) KItCAT-SSMBA, which replaces selected tokens with contextually plausible alternatives; and (4) KItCAT-MASKER, which preferentially masks informative keywords. KItCAT-SSMBA and KItCAT-MASKER adapt advanced corruption schemes originally developed for encoder-only models to knowledge injection with decoder-only LLMs. In decoder-only models, corruption is applied solely to the conditioning input while the target labels remain unchanged. Under random replacement, the model must learn to ignore irrelevant noise, whereas in mask-based corruption, the missing information is explicitly signaled. By preventing the model from encountering identical inputs across epochs, KItCAT reduces reinforcement of spurious lexical patterns and encourages learning more semantically grounded representations.
Experiments across multiple knowledge-injection benchmarks and three model families show that KItCAT consistently improves standard CPT. Notably, KItCAT is agnostic to the input text and can be applied directly to original corpora as well as to paraphrases used to strengthen CPT. This makes it complementary to existing approaches and reduces the reliance on paraphrasing.
Overall, our contributions are:
(1) We propose Knowledge Injection via Corrupted Auto-regressive Training (KItCAT), a lightweight and generalizable data augmentation method for fine-tuning decoder-only language models by stochastically corrupting input tokens while preserving the original autoregressive training objective.
(2) We empirically demonstrate that KItCAT consistently improves standard CPT across three model families
(3) We show that KItCAT induces effective data diversity to prevent overfitting, significantly reducing the need for expensive synthetic data while remaining complementary to paraphrase-based augmentation.
2 Related Work
Approaches to injecting knowledge into LLM parameters can be broadly categorised as CPT-based methods and SFT-based methods. CPT-based methods inject knowledge using raw domain text and its paraphrases (Wu et al., 2024; Christophe et al., 2024; Ke et al., 2023; Zhao et al., 2025; Zhang et al., 2024; Wu et al., 2023; Shah et al., 2022; Xie et al., 2023). SFT-based methods use powerful LLMs to transform raw text into task-oriented formats such as QA, summarization, or reading comprehension (Ovadia et al., 2025; Bhushan et al., 2025). Both approaches benefit from high-diversity training data and, in its absence, often rely on generating large amounts of synthetic paraphrases or SFT data using powerful LLMs. In contrast, our work proposes lightweight input corruption as an alternative to expensive paraphrase generation during CPT-style knowledge injection.
Corruption can be applied to the input or to hidden representations within the model and used in both SFT and CPT training styles. Dropout (Hinton et al., 2012) corrupts hidden activations during training, while Latent Paraphrasing (Kang et al., 2024) perturbs hidden representations to facilitate knowledge injection and can be applied to both CPT and SFT.
Input-corruption techniques for training LLMs are discussed in Appendix A.2.
3 The KItCAT approach
Preliminaries:
Let be the new knowledge that we want to inject in our pretrained model . Let denote the vocabulary and be a token sequence drawn from , where .
| (1) |
Here, is a prefix of the sequence .
Knowledge Injection via Input Corruption:
The objective maximizes the likelihood of a token conditioned on its exact prefix . When seen multiple times across epochs, the model may overfit to spurious correlations between and .
To mitigate this, we propose KItCAT: Knowledge Injection via Corrupted Auto-regressive Training. Let denote a distribution over corrupted versions of the input sequence . Instead of conditioning only on the exact prefix, KItCAT trains on corrupted prefixes sampled from while preserving the target token as a likely continuation. The model is thus encouraged to predict the same target under perturbed contexts.
| (2) |
where is a prefix of the corrupted input . Intuitively, training on perturbed prefixes prevents reliance on spurious lexical patterns. Instead, it encourages the model to use information that remains present across perturbations.
To construct perturbations , we consider four stochastic in-place token modifications in the sequence : (1) KItCAT-mask, which replaces tokens with a special mask token; (2) KItCAT-rand, which replaces tokens with randomly sampled vocabulary tokens; (3) KItCAT-SSMBA, which replaces randomly selected tokens with contextually plausible alternatives; and (4) KItCAT-MASKER, which masks informative keywords identified from the training corpus. The latter two adapt the central ideas of SSMBA (Ng et al., 2020) and MASKER (Moon et al., 2021) to the conditioning inputs of decoder-only LLMs.
For every method, let denote the set of token positions selected for corruption, and let be the replacement token at each selected position . The corrupted input is therefore
For KItCAT-mask, KItCAT-rand, and KItCAT-SSMBA, we construct by independently selecting each token position with probability , the corruption probability. We resample and the replacement tokens at each epoch.
KItCAT-mask.
For each , we use the special mask token: .
KItCAT-rand.
For each , we sample a replacement uniformly from the vocabulary: .
KItCAT-SSMBA.
For each , we mask the position and sample a replacement from a frozen masked language model : , where denote with all positions in masked.
KItCAT-MASKER.
We identify keyword spans from the training corpus using TF–IDF and independently select each span, applying the corruption probability at the span level rather than the token level; contains the token positions in the selected spans and is resampled every epoch. Each selected token is replaced with the mask token: for .
For all four variants, corruption is applied only to the conditioning prefix and the target token remains unchanged.
In Section A.3, we formalize as a constrained optimization objective that enforces invariance to label-preserving corruptions, plausibly reducing overfitting.
4 Experimental Setup
Datasets:
We fine-tune an LLM via standard NTP loss to inject knowledge from four corpora – subset of PopQA Mallen et al. (2023), Companies Ovadia et al. (2025), and two Redbooks Bhushan et al. (2025). Subset of PopQA11 1 PopQA and Companies Dataset focuses on the entity centric QAs from the long tail of Wikipedia. Companies††footnotemark: contains details of 24 fictitious companies unseen during pretraining. Redbooks consists of two technical documents22 2 Book 1: Do More with Less: Automating IBM Storage FlashSystem Tasks with REST APIs, Scripting, and Ansible. Book 2: Red Hat OpenShift Container Platform on IBM Z and LinuxONE.. We evaluate the fine-tuned models using the test question-answer pairs provided with each dataset.
Baselines:
We compare with the out-of-the-box model (Instruct), standard CPT using next-token prediction (NTP), and four KItCAT corruption schemes. KItCAT-SSMBA adapts SSMBA (Ng et al., 2020) to CPT by using an encoder model to sample contextually plausible token replacements, while KItCAT-MASKER adapts MASKER (Moon et al., 2021) by preferentially masking informative keywords.
Evaluation metrics:
To test if the knowledge in the documents has been successfully injected in the model’s parameters, we evaluate the fine-tuned models using the test QAs accompanying each dataset. We use an LLM-as-a-Judge to quantify the correctness of the generated responses with respect to the gold answers. See Section A.9 for the exact prompt.
Models and Training Details:
We fine-tune Mistral-7B-Instruct-v0.3 using Huggingface’s SFTTrainer. We apply LoRA to all linear layers of the LLM and early stop based on the validation metric. Unless stated otherwise, we report results using Mistral-7B-Instruct-v0.3. To test model-family and scale generalization, we additionally evaluate Qwen3-14B and Llama-2-7B-Chat on all datasets. See Section A.5 for more training details.
| Method | RB1 | RB2 | Comp. | PopQA | Avg |
|---|---|---|---|---|---|
| Instruct | 25.6 | 11.8 | 6.8 | 10.8 | 13.8 |
| NTP | 29.7 | 14.6 | 13.6 | 12.4 | 17.6 |
| KItCAT-SSMBA | 31.5 | 17.0 | 21.2 | 22.1 | 23.0 |
| KItCAT-MASKER | 35.3 | 13.3 | 17.7 | 12.1 | 19.6 |
| KItCAT-mask | 38.3 | 18.5 | 26.8 | 19.8 | 25.9 |
| KItCAT-rand | 43.5 | 20.8 | 26.4 | 20.2 | 27.7 |
5 Experimental Results
Effectiveness of KItCAT in injecting knowledge (RQ1): To assess knowledge injection, we evaluate each adapted model on the corresponding test QAs using the LLM judge described in Section 4. Table 1 reports the results across the four corpora.
All four KItCAT variants outperform NTP on average. KItCAT-rand achieves the highest average score (27.7), a 10.0-point improvement over NTP (17.6), while KItCAT-mask also improves the average score to 26.1. The simple corruption schemes, KItCAT-mask and KItCAT-rand, together obtain the best score on three of the four datasets (both Redbooks and Companies), whereas the advanced schemes prevail only once, with KItCAT-SSMBA leading on PopQA (22.1 vs. 20.2 for KItCAT-rand). On average, simple corruption proves more effective than the more elaborate, advanced schemes. We hypothesize that perturbing random tokens in the context increases the overall diversity of the training corpus and hence, prevents overfitting to the spurious patterns thereby leading to improved performance on QA pairs from the corpus. Example outputs are provided in Section A.1.
KItCAT also generalizes to Qwen3-14B and Llama-2-7B-Chat, consistently improving over NTP across the evaluated corpora ( and ). A sensitivity analysis shows that performance is strongest at moderate corruption probabilities (–) and degrades when corruption is too weak or too strong; see Sections A.8 and .
Impact of training data size (RQ2):
Next, we note that KItCAT is complementary to other approaches for data augmentation that generate multiple paraphrases for each document present in the corpus and use them for CPT. Hence, for this experiment, we combine these two complementary methods for effective knowledge injection and observe the resultant behavior.
We consider two synthetic data generation strategies: (1) Rephrase, where an LLM Hurst et al. (2024) is prompted with multiple prompts Ovadia et al. (2025) to generate multiple paraphrases per document, and (2) Entigraph, which follows Yang et al. (2025) by extracting entities from text and prompting an LLM to describe relationships between entity pairs and triples. See Section A.9 for the exact prompts.
We control the amount of Entigraph data by randomly sampling pairs and triples until we get the desired token count. For both Rephrase and Entigraph, we train models using standard NTP loss and KItCAT-mask, while varying the amount of synthetic data.
Figure 2(a) presents the results.
We make the following observations:
(1) The gains obtained by KItCAT using only the original data (1) are higher than standard NTP on 9 Rephrase augmented data (26.84 vs. 22.35) and 5 Entigraph augmented data (26.84 vs. 24.53), demonstrating a 5–9 effective data multiplier over standard NTP.
(2) KItCAT is complementary to synthetic data generation as a data augmentation strategy. Performance consistently improves as we introduce additional paraphrased data.
In Section A.7, we present a comparison of compute costs and LLM Judge accuracy for different methods on the Companies dataset. We observe that KItCAT-mask achieves 27.6 accuracy at an estimated cost of 30 PFLOPs, matching the 27.4 accuracy of Rephrase-10 while requiring only 15.2% of its 198 PFLOPs. Combining corruption with paraphrases gives the best accuracy, confirming that KItCAT is both a cheaper partial substitute and a complementary augmentation ().
Robustness of KItCAT:
We next evaluate the robustness of KItCAT relative to the standard next-token prediction (NTP) objective by analyzing their learning dynamics. Specifically, we examine validation loss curves on QA pairs from the Companies validation set (Figure 2(b)).
We observe that the NTP objective reaches its minimum validation loss within the first seven epochs, after which it exhibits clear overfitting, as evidenced by a sharp rise in validation loss. In contrast, the two variants of KItCAT, viz. KItCAT-mask and KItCAT-rand, maintain stable validation loss throughout training.
These results indicate that KItCAT provides improved training stability and achieves better generalization compared to the NTP baseline.
6 Conclusion
We introduce KItCAT, a simple yet effective modification to autoregressive training for injecting new knowledge into LLM parameters. By ensuring that the model never encounters the exact same input twice, KItCAT mitigates the reinforcement of spurious lexical correlations that commonly arise in low-diversity CPT settings.
Experiments across four knowledge-injection benchmarks and three model families show that KItCAT consistently outperforms standard next-token prediction-based training. Moreover, KItCAT is lightweight, architecture-agnostic, and complementary to existing data augmentation approaches. Unlike synthetic data generation, it requires no external models (for paraphrasing) and incurs negligible computational overhead, making it a practical drop-in replacement for standard CPT objectives in real-world knowledge injection settings.
Limitations
A limitation of KItCAT is that, while corruption-based augmentation reduces the need for expensive synthetic paraphrase generation, it does not fully substitute for it. Existing approaches that incorporate synthesized paraphrases continue to provide complementary benefits, and the highest downstream performance is achieved when corruption and paraphrase-based augmentation are combined.
References
- Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §1.
- Physics of language models: part 3.1, knowledge storage and extraction. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §1.
- Systematic knowledge injection into large language models via diverse augmentation for domain-specific RAG. In Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 5922–5943. External Links: Link, Document Cited by: §2, §4.
- Stability and generalization. J. Mach. Learn. Res. 2, pp. 499–526. External Links: Link Cited by: §A.3.
- Masked thought: simply masking partial reasoning steps can improve mathematical reasoning learning of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 5872–5900. External Links: Link, Document Cited by: §A.2.
- An empirical survey of data augmentation for limited data learning in nlp. Transactions of the Association for Computational Linguistics 11, pp. 191–211. External Links: ISSN 2307-387X, Document, Link, https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00542/2074871/tacl_a_00542.pdf Cited by: §A.2.
- Beyond fine-tuning: unleashing the potential of continuous pretraining for clinical llms. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 10549–10561. External Links: Link, Document Cited by: §2.
- BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), pp. 4171–4186. External Links: Link, Document Cited by: §A.2.
- Train no evil: selective masking for task-guided pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y. He, and Y. Liu (Eds.), Online, pp. 6966–6974. External Links: Link, Document Cited by: §A.2.
- Improving neural networks by preventing co-adaptation of feature detectors. CoRR abs/1207.0580. External Links: Link, 1207.0580 Cited by: §2.
- Gpt-4o system card. arXiv preprint arXiv:2410.21276. Cited by: §5.
- Code-switching with word senses for pretraining in neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12889–12901. External Links: Link, Document Cited by: §A.2.
- What to hide from your students: attention-guided masked image modeling. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXX, S. Avidan, G. J. Brostow, M. Cissé, G. M. Farinella, and T. Hassner (Eds.), Lecture Notes in Computer Science, Vol. 13690, pp. 300–318. External Links: Link, Document Cited by: §A.2.
- Latent paraphrasing: perturbation on layers improves knowledge injection in language models. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Link Cited by: §2.
- Demystifying domain-adaptive post-training for financial llms. External Links: 2501.04961, Link Cited by: §1.
- Continual pre-training of language models. In Proceedings of The Eleventh International Conference on Learning Representations, Cited by: §1, §2.
- MAGNET: augmenting generative decoders with representation learning and infilling capabilities. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 27328–27346. External Links: Link Cited by: §A.2.
- BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, D. Jurafsky, J. Chai, N. Schluter, and J. R. Tetreault (Eds.), pp. 7871–7880. External Links: Link, Document Cited by: §A.2.
- MST: masked self-supervised transformer for visual representation. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, and J. W. Vaughan (Eds.), pp. 13165–13176. External Links: Link Cited by: §A.2.
- EntityBERT: entity-centric masking strategy for model pretraining for the clinical domain. In Proceedings of the 20th Workshop on Biomedical Language Processing, D. Demner-Fushman, K. B. Cohen, S. Ananiadou, and J. Tsujii (Eds.), Online, pp. 191–201. External Links: Link, Document Cited by: §A.2.
- Pre-training multilingual neural machine translation by leveraging alignment information. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, B. Webber, T. Cohn, Y. He, and Y. Liu (Eds.), pp. 2649–2663. External Links: Link, Document Cited by: §A.2.
- RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs/1907.11692. External Links: Link, 1907.11692 Cited by: §A.2.
- LLaMAX: scaling linguistic horizons of LLM by enhancing translation capabilities beyond 100 languages. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, pp. 10748–10772. External Links: Link, Document Cited by: §A.2.
- When not to trust language models: investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 9802–9822. External Links: Link, Document Cited by: §4.
- MASKER: masked keyword regularization for reliable text classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, pp. 13578–13586. External Links: Document Cited by: §A.2, §3, §4.
- SSMBA: self-supervised manifold based data augmentation for improving out-of-domain robustness. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1268–1283. External Links: Document, Link Cited by: §A.2, §3, §4.
- Knowledge-instruct: effective continual pre-training from limited data using instructions. arXiv preprint arXiv:2504.05571. Cited by: §A.10, §2, §4, §5.
- Fine-tuning or retrieval? comparing knowledge injection in LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 237–250. External Links: Link, Document Cited by: §1.
- Decoder-only llms can be masked auto-encoders. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 713–723. External Links: Link, Document Cited by: §A.2.
- Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, pp. 140:1–140:67. External Links: Link Cited by: §A.2.
- When flue meets flang: benchmarks and large pre-trained language model for financial domain. arXiv preprint arXiv:2211.00083. Cited by: §2.
- Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Cited by: §1.
- Difference-masking: choosing what to mask in continued pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), pp. 13222–13234. External Links: Link, Document Cited by: §A.2.
- PMC-llama: toward building open-source language models for medicine. Journal of the American Medical Informatics Association 31 (9), pp. 1833–1843. Cited by: §2.
- Bloomberggpt: a large language model for finance. arXiv preprint arXiv:2303.17564. Cited by: §2.
- Pixiu: a large language model, instruction data and evaluation benchmark for finance. arXiv preprint arXiv:2306.05443. Cited by: §2.
- Towards self-robust LLMs: intrinsic prompt noise resistance via coIPO. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §A.2.
- Synthetic continued pretraining. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §A.11, §1, §5.
- Chemllm: a chemical large language model. arXiv preprint arXiv:2402.06852. Cited by: §2.
- Developing chemdfm as a large language foundation model for chemistry. Cell Reports Physical Science 6 (4). Cited by: §2.
- Mask-enhanced autoregressive prediction: pay less attention to learn more. In Proceedings of the 42nd International Conference on Machine Learning, Cited by: §A.2.
Appendix A Appendix
A.1 Example Model Responses
A.2 Related Works (contd.)
Input Corruption for Training LLMs
Input corruption has long been used to learn robust representations in NLP (Chen et al., 2023). Masked Language Modeling (MLM) applies token masking to train encoder-only LMs using bidirectional context (Devlin et al., 2019; Liu et al., 2019). Encoder–decoder models such as BART (Lewis et al., 2020) and T5 (Raffel et al., 2020) further extend denoising objectives using masking, deletion, and span corruption. Several works also explore more principled corruption strategies based on lexical translations, syntactic, semantic, knowledge graphs, or attention-based signals (Lin et al., 2020; Iyer et al., 2023; Wilf et al., 2023; Gu et al., 2020; Lin et al., 2021; Li et al., 2021; Kakogeorgiou et al., 2022). Notably, SSMBA (Ng et al., 2020) generates plausible augmentations by masking and reconstructing tokens with a pretrained masked language model, while MASKER (Moon et al., 2021) masks TF-IDF-selected keywords to discourage reliance on keyword shortcuts. We adapt both strategies to the conditioning inputs of decoder-only LLMs for CPT and evaluate these adaptations in our experiments (Section 5).
More recent work has begun to study input corruption directly in decoder-only LLMs. Qiao et al. (2025); Khosla et al. (2025) use token masking to adapt decoder-only models to generate representations and infill missing text spans. Lu et al. (2024) replace randomly selected words with their multilingual translations to improve machine translation capabilities, Yang et al. (2026) introduce character-, word-, and sentence-level perturbations to improve robustness to noisy prompts, and Chen et al. (2024) mask intermediate chain-of-thought tokens during fine-tuning to encourage global reasoning. Recently, Zhuang et al. (2025) apply token masking during large-scale pre-training. Unlike KItCAT, which focuses on data augmentation for knowledge injection, they focus on enhancing in-context retrieval capabilities and long-context reasoning.
A.3 A constrained optimization perspective of our method
As discussed in Section 3, the objective can overfit to spurious correlations between prefix and target tokens , especially in low-data regimes where the same training sequences are observed repeatedly. One can mitigate this by ensuring that the predictive distribution remains stable under label-preserving perturbations of . To formalize this idea, we introduce the notation to denote the deviation in log-likelihood induced by a corrupted input context:
| (3) | ||||
where and are prefixes of and respectively. Note that this deviation can become very high if the model overfits to the prefix-target pair provided in the training data. We therefore consider the following optimization problem that minimizes the standard next-token prediction loss while bounding expected deviation.
| (4) |
Intuitively, the proposed constraint is related in spirit to Lipschitz continuity Bousquet and Elisseeff (2002), in that it bounds the sensitivity of the model’s predictions to input perturbations. However, unlike classical Lipschitz continuity defined over continuous normed spaces, our constraint is expressed as an expectation over discrete, label-preserving corruptions of the input sequence. As a result, the notion of smoothness it implies is over a discrete space, rather than gradient-based smoothness over a continuous space.
The Lagrangian of the above equation can be written as follows:
| (5) | ||||
with the constraint that the Lagrange multiplier . The Lagrange multiplier controls the trade-off between fitting the training data and enforcing stability of the model’s predictions under label-preserving corruptions. Different choices of recover familiar training objectives as special cases.
When , the constraint is ignored and the Lagrangian reduces to
which corresponds to standard autoregressive training. When , ignoring the constant offset , the Lagrangian reduces to
| (6) |
This recovers the CAT objective introduced in equation (2). Hence, KItCAT corresponds to a next-token prediction objective that is constrained to learn functions that vary smoothly under label-preserving perturbations of the input prefix. Intuitively, adding this constraint to should result in improved generalization.
A.4 Dataset Details
summarizes the training data sizes and evaluation split sizes for the CPT benchmarks.
A.5 Training Details
We fine-tune all three models, viz., Mistral-7B-Instruct-v0.3, Qwen3-14B, and Llama-2-7B-Chat, using Hugging Face’s TRL library. All models are trained on four Nvidia A100 (80GB) GPUs using bf16 precision, applying Low-Rank Adaptation (LoRA) to all linear layers. We utilize the AdamW optimizer with a warmup ratio of 0.1, training for a maximum of 64 epochs with early stopping based on validation mean token accuracy. We use LoRA rank 128 and LoRA aplha 256 throughout our experiments. For KItCAT, we treat the corruption probability as a hyperparameter and tune it on the validation set, choosing from separately for each task. For KItCAT-SSMBA, we use ModernBERT-base as the frozen masked language model to sample contextually plausible replacements.
A.6 Additional Model Results
and report the additional Qwen3-14B and Llama-2-7B-Chat experiments. All methods use the original corpus. These results show that the gains extend beyond a single model family and that more elaborate reconstruction or keyword-selection schemes do not consistently improve upon simple random corruption.
A.7 Compute Cost of Paraphrasing
reports the cost in FLOPs, separating the one-time paraphrase-generation cost from training cost. The estimates use FLOPs per generated token for a generation model with parameters and FLOPs per training token for LoRA training of an -parameter model. Training steps denote the checkpoint with maximum eval mean token accuracy. Input corruption skips generation entirely and adds only a modest increase in training cost.
A.8 Sensitivity to Corruption Probability
We sweep for both variants on Companies and Redbook 1 (). Performance has an interior optimum around – and degrades when corruption is too weak or too strong. Sensitivity is more pronounced on the data-scarce Companies corpus, where corrupting too little will result in overfitting and corrupting too much will cause confusion on what to learn.
A.9 Judge Prompts
The exact prompt used for LLM-as-a-Judge evaluation with gpt-oss-120B is shown below.
\iow_now:Ne¨\iow_now:Ne¨System:\iow_now:Ne¨You are an evaluator. Your task is to compare a Ground-truth Answer and a Prediction to decide if the Prediction correctly answers the given Question.\iow_now:Ne¨\iow_now:Ne¨Evaluation Rules:\iow_now:Ne¨\iow_now:Ne¨(1) Correctness: A correct prediction must include all essential information from the Ground-truth Answer. Extra information is allowed if it does not contradict the Ground-truth. If the Prediction states something as a possibility, treat it as a definitive statement.\iow_now:Ne¨(2) Function, Tool Names, and API Calls: If the Ground-truth Answer contains specific function names, tool names, API calls, or exact command identifiers, the Prediction must contain the same identifier(s) or clearly equivalent forms. Minor syntactic or formatting variations that do not change meaning should be treated as equivalent. For example, leading flag prefixes such as -, –, or no prefix at all when they clearly refer to the same option; underscore vs hyphen differences in identifiers when the intent is identical; surrounding punctuation or formatting differences such as backticks, quotes, parentheses, or code block notation; small whitespace differences or capitalization differences that do not change the identifier’s meaning etc.\iow_now:Ne¨However, replacements that change the actual function/tool/API name, or substitute a different command that would change the behavior are considered incorrect. Do not penalize a prediction if it contains additional function / tool / API names as long as the ones present in the Ground-Truth are covered.\iow_now:Ne¨(3) URLs: If the Ground-truth Answer contains specific URLs, the Prediction should reference the same URL or an equivalent canonical form. Minor differences that do not change the target resource (for example, presence or absence of a trailing slash, or http vs https when both resolve to the same canonical resource) should be treated as equivalent. Altering the domain, path, or query such that the resource is different is incorrect.\iow_now:Ne¨\iow_now:Ne¨Scoring Rules:\iow_now:Ne¨\iow_now:Ne¨If the Prediction is correct according to the above rules, output <score>1</score>. If the Prediction is incomplete or incorrect, output <score>0</score>.\iow_now:Ne¨\iow_now:Ne¨Output Format:\iow_now:Ne¨\iow_now:Ne¨<explanation>\iow_now:Ne¨…\iow_now:Ne¨</explanation>\iow_now:Ne¨\iow_now:Ne¨<score>\iow_now:Ne¨…\iow_now:Ne¨</score>\iow_now:Ne¨\iow_now:Ne¨First provide reasoning inside <explanation> and </explanation> tags. Then output the score as specified above within <score> and </score> tags. Do not include any extra text outside these tags.\iow_now:Ne¨\iow_now:Ne¨Human:\iow_now:Ne¨Question: {QUESTION}\iow_now:Ne¨Ground-truth Answer: {ANSWER}\iow_now:Ne¨Prediction: {ASSIST_ANSWER}
A.10 RephraseWeb Prompts for generating synthetic data
For RephraseWeb, we use the prompts proposed by Ovadia et al. (2025) to generate multiple paraphrases of the training data. They use a common system prompt, along with nine different rephrase styles to induce diversity in the generated paraphrases. Both are pasted below.
Below, we include prompts for generating paraphrases in different styles to increase diversity in the synthetic data.
A.11 EntiGraph Generation Prompts
This appendix details the prompts proposed by Yang et al. (2025) for generating synthetic data.
The pipeline consists of three stages:
(1) entity extraction from source documents,
(2) two-entity relation generation, and
(3) three-entity relation generation.
All prompts are issued to GPT-4o via the Azure OpenAI API.
A.11.1 Entity Extraction
The following system prompt is used to extract salient entities from
each source document. The model is instructed to return structured
JSON.
\iow_now:Ne¨\iow_now:Ne¨System:\iow_now:Ne¨As a knowledge analyzer, your task is to dissect and understand an\iow_now:Ne¨article provided by the user. You are required to perform the\iow_now:Ne¨following steps:\iow_now:Ne¨\iow_now:Ne¨1. Summarize the Article:\iow_now:Ne¨Provide a concise summary of the entire article, capturing the main\iow_now:Ne¨points and themes.\iow_now:Ne¨\iow_now:Ne¨2. Extract Entities:\iow_now:Ne¨Identify and list all significant "nouns" or entities mentioned within\iow_now:Ne¨the article. These entities should include but are not limited to:\iow_now:Ne¨\iow_now:Ne¨* People:\iow_now:Ne¨Any individuals mentioned in the article, using the names or\iow_now:Ne¨references provided.\iow_now:Ne¨\iow_now:Ne¨* Places:\iow_now:Ne¨Both specific locations and abstract spaces relevant to the content.\iow_now:Ne¨\iow_now:Ne¨* Objects:\iow_now:Ne¨Any concrete object that is referenced by the provided content.\iow_now:Ne¨\iow_now:Ne¨* Concepts:\iow_now:Ne¨Any significant abstract ideas or themes that are central to the\iow_now:Ne¨article’s discussion.\iow_now:Ne¨\iow_now:Ne¨Try to exhaust as many entities as possible. Your response should be\iow_now:Ne¨structured in JSON format to organize the information effectively.\iow_now:Ne¨Ensure that the summary is brief yet comprehensive, and the list of\iow_now:Ne¨entities is detailed and accurate.\iow_now:Ne¨\iow_now:Ne¨Use the following response format:\iow_now:Ne¨\iow_now:Ne¨{\iow_now:Ne¨ "summary": "<A concise summary of the article>",\iow_now:Ne¨ "entities": ["entity1", "entity2", …]\iow_now:Ne¨}The user message for entity extraction takes the following form:
\iow_now:Ne¨\iow_now:Ne¨Human:\iow_now:Ne¨### Document Content:\iow_now:Ne¨{document_content}
A.11.2 Two-Entity Relation Generation
For each pair of extracted entities , the following
system prompt instructs the model to rephrase the document content
with emphasis on each entity and analyze their interaction.
\iow_now:Ne¨\iow_now:Ne¨System:\iow_now:Ne¨You will act as a knowledge analyzer tasked with dissecting an article\iow_now:Ne¨provided by the user. Your role involves two main objectives:\iow_now:Ne¨\iow_now:Ne¨1. Rephrasing Content:\iow_now:Ne¨The user will identify two specific entities mentioned in the article.\iow_now:Ne¨You are required to rephrase the content of the article twice:\iow_now:Ne¨\iow_now:Ne¨* Once, emphasizing the first entity.\iow_now:Ne¨* Again, emphasizing the second entity.\iow_now:Ne¨\iow_now:Ne¨2. Analyzing Interactions:\iow_now:Ne¨Discuss how the two specified entities interact within the context of\iow_now:Ne¨the article.\iow_now:Ne¨\iow_now:Ne¨Your response should clearly separate the rephrased content from the\iow_now:Ne¨interaction analysis. Ensure each section includes sufficient context,\iow_now:Ne¨ideally referencing the article title to maintain clarity about the\iow_now:Ne¨discussion’s focus.\iow_now:Ne¨\iow_now:Ne¨Use the following response format:\iow_now:Ne¨\iow_now:Ne¨### Discussion of <title> in relation to <entity1>\iow_now:Ne¨<Rephrased content focusing on the first entity>\iow_now:Ne¨\iow_now:Ne¨### Discussion of <title> in relation to <entity2>\iow_now:Ne¨<Rephrased content focusing on the second entity>\iow_now:Ne¨\iow_now:Ne¨### Discussion of Interaction between <entity1> and <entity2>\iow_now:Ne¨in context of <title>\iow_now:Ne¨<Discussion on how the two entities interact within the article>The user message for two-entity relation generation takes the
following form:
\iow_now:Ne¨\iow_now:Ne¨Human:\iow_now:Ne¨### Document Content:\iow_now:Ne¨{document_content}\iow_now:Ne¨\iow_now:Ne¨### Entities:\iow_now:Ne¨- {entity1}\iow_now:Ne¨- {entity2}
A.11.3 Three-Entity Relation Generation
For each triple of extracted entities , the
following system prompt extends the two-entity setting to three
entities.
\iow_now:Ne¨\iow_now:Ne¨System:\iow_now:Ne¨You will act as a knowledge analyzer tasked with dissecting an article\iow_now:Ne¨provided by the user. Your role involves three main objectives:\iow_now:Ne¨\iow_now:Ne¨1. Rephrasing Content:\iow_now:Ne¨The user will identify three specific entities mentioned in the\iow_now:Ne¨article. You are required to rephrase the content of the article three\iow_now:Ne¨times:\iow_now:Ne¨\iow_now:Ne¨* Once, emphasizing the first entity.\iow_now:Ne¨* Again, emphasizing the second entity.\iow_now:Ne¨* Lastly, emphasizing the third entity.\iow_now:Ne¨\iow_now:Ne¨2. Analyzing Interactions:\iow_now:Ne¨Discuss how these three specified entities interact within the context\iow_now:Ne¨of the article.\iow_now:Ne¨\iow_now:Ne¨Your response should clearly separate the rephrased content from the\iow_now:Ne¨interaction analysis. Ensure each section includes sufficient context,\iow_now:Ne¨ideally referencing the article title to maintain clarity about the\iow_now:Ne¨discussion’s focus.\iow_now:Ne¨\iow_now:Ne¨Use the following response format:\iow_now:Ne¨\iow_now:Ne¨### Discussion of <title> in relation to <entity1>\iow_now:Ne¨<Rephrased content focusing on the first entity>\iow_now:Ne¨\iow_now:Ne¨### Discussion of <title> in relation to <entity2>\iow_now:Ne¨<Rephrased content focusing on the second entity>\iow_now:Ne¨\iow_now:Ne¨### Discussion of <title> in relation to <entity3>\iow_now:Ne¨<Rephrased content focusing on the third entity>\iow_now:Ne¨\iow_now:Ne¨### Discussion of Interaction between <entity1>, <entity2>, and\iow_now:Ne¨<entity3> in context of <title>\iow_now:Ne¨<Discussion on how the three entities interact within the article>The user message for three-entity relation generation takes the
following form:
\iow_now:Ne¨\iow_now:Ne¨Human:\iow_now:Ne¨### Document Content:\iow_now:Ne¨{document_content}\iow_now:Ne¨\iow_now:Ne¨### Entities:\iow_now:Ne¨- {entity1}\iow_now:Ne¨- {entity2}\iow_now:Ne¨- {entity3}