Subword Segmental BabyLMs:
Learning to Tokenise for Sample-Efficient Pretraining
Abstract
In the standard LM training pipeline, subword tokenisation is applied as a preprocessing step. Subword segmental language modelling is an alternative paradigm in which tokenisation is learned during training, allowing the model to discover subword units that optimise its training objective. In this paper, we present our submission to the 2026 BabyLM Challenge, for which we develop two new subword segmental LMs: SubSegGPT and SubSegDeBERTa. SubSegGPT is a decoder-only model that learns tokenisation during autoregressive pretraining. SubSegDeBERTa is an encoder-based model that jointly learns to generate and tokenise masked words. We train both for the Strict and Strict-small tracks. Our top submission to Strict is SubSegDeBERTa, which achieves notable gains in zero-shot evaluation. Our top submission to Strict-small is SubSegGPT, which outperforms tokenisation-based baselines. Our results show that learnable subword tokenisation can improve sample-efficiency for BabyLM pretraining. We analyse the subword learning dynamics of our models and find that tokenisation gradually converges on subword units that balance morphological alignment and fine-grained segmentation.
1 Introduction
Current language model (LM) training is developmentally implausible in many ways, one of which is its reliance on subword tokenisation. The subwords produced by standard tokenisers do not reliably align with morpheme boundaries (Batsuren et al., 2024). Moreover, the dominant paradigm of applying a fixed tokeniser during preprocessing does not mirror child language learning, as humans are not born pre-equipped with a fixed lexical vocabulary. Instead, children incrementally learn to segment speech into meaningful units during language acquisition (Jusczyk, 1999).
From a more practical, performance-driven perspective, standard NLP tokenisation poses problems for data-efficient modelling. The algorithms behind tokenisers like BPE (Sennrich et al., 2016) and ULM (Kudo, 2018) learn subword boundaries based on frequency-based objectives, with no guarantee that the resulting units are optimal for LM learnability. Given sufficient training data, neural LMs are robust to sub-optimal tokenisation. However, in small data settings, this may further compound the difficulty of learning generalisable linguistic representations.
By restricting tokenisation to preprocessing, LMs are bound to a pre-determined tokenisation scheme. Alternatively, tokenisation can be cast as a learnable component of language modelling, to be continually optimised during training. This is the motivation behind subword segmental modelling (Meyer and Buys, 2022), which unifies tokenisation and language modelling in an end-to-end trainable framework. Instead of determining subword boundaries before training, subword segmental modelling marginalises over all possible tokenisations of a sentence and allows the model to learn which subword units optimise its LM training objective.
In this paper, we present our submission11 1 Models: francois-meyer/babylm-2026-models to the 2026 BabyLM Challenge, for which we develop22 2 Code: github.com/francois-meyer/subseg-babylms two new Subword Segmental LMs: SubSegGPT and SubSegDeBERTa. SubSegGPT is a decoder-only subword segmental architecture with a GPT-2 (Radford et al., 2019) backbone. It combines subword segmental modelling with more modern, GPT-style conventions in decoder LM design and implementation. SubSegDeBERTa is a novel, encoder-based architecture and the first instantiation of subword segmental modelling for masked language modelling. It jointly learns to generate and tokenise masked words during training, conditioned on bi-directional context from a DeBERTa-v2 backbone (He et al., 2021).
We train our models for the Strict (100M words) and Strict-small (10M words) tracks and evaluate on the official evaluation pipeline, which includes zero-shot evaluations, finetuning tasks, and human likeness tests. We compare our models to equivalent tokenisation-based BabyLM baselines, based on respectively GPT-2 and DeBERTa-v2, to isolate the effect of learning tokenisation over fixed tokenisation.
In the Strict track, SubSegDeBERTa and SubSegGPT outperform tokenisation-based baselines on average across NLP tasks. SubSegDeBERTa is our strongest submission to Strict, improving average zero-shot performance by 3.16 points over GPT-2. In Strict-small, SubSegGPT outperforms baselines, but SubSegDeBERTa does not. Our two models are complementary: SubSegDeBERTa is highly competitive in the Strict track, while SubSegGPT is better suited to the more data-constrained Strict-small setting. On human likeness tasks, our models exhibit little resemblance to human language acquisition and are outperformed by tokenisation-based baselines.
Lastly, we analyse the subword learning dynamics of SubSegGPT and SubSegDeBERTa. We compare learned subword units over the course of training, tracking fertility and morpheme boundary F1. Both models undergo an initial period of rapid tokenisation change, followed by stabilised learning trajectories that converge on higher fertility (shorter subwords) and limited alignment with morpheme boundaries, which we support with a qualitative analysis of tokenised child-directed speech.
2 Background
2.1 Learning Tokenisation During Training
Without being motivated by developmental plausibility, several works have explored learning to segment language inputs during training, for tasks such as handwriting recognition (Kong et al., 2016), speech recognition (Wang et al., 2017), and translation (Kreutzer and Sokolov, 2018). These efforts equip models with the ability to discover optimal segmentation units for a given task, rather than relying on a pre-determined segmentation scheme.
Sun and Deng (2018) propose the segmental language model (SLM): an LSTM LM that computes the likelihood of a sentence by marginalising over all possible segmentations of the character sequence. Their aim is not improved performance, but rather unsupervised word segmentation for Chinese, which lacks explicit word boundaries. Kawakami et al. (2019) augment the SLM with a lexicon of high-frequency segments, improving unsupervised word segmentation. Downey et al. (2022) propose the masked SLM, a bi-directional, transformer-based SLM that outperforms recurrent SLMs on unsupervised word discovery. SLMs are trained on raw character sequences without word boundaries (whitespaces) and, as a by-product of optimising the LM training objective, they partially recover word-level segmentation.
Meyer and Buys (2022) propose the subword segmental language model (SSLM), which adapts the SLM to model subword tokenisation. SSLM assumes access to word boundaries, constrains segments to subword units (they cannot cross word boundaries), and learns subword tokenisation to optimise the LM training objective. The original SSLM (Meyer and Buys, 2022) is LSTM-based (Hochreiter and Schmidhuber, 1997). Meyer and Buys (2025) propose a transformer (Vaswani et al., 2017) SSLM, primarily as a tool to study subword learning dynamics. For both models, evaluation is limited to the Nguni languages: low-resource, agglutinative languages, for which tokenisation was hypothesised to have an outsized impact. SSLM outperformed tokenisation-based LMs on perplexity and sequence-to-sequence tasks, and performed strongly as an unsupervised morpheme segmenter.
It is unknown whether learnable subword tokenisation can improve pretraining sample-efficiency beyond the narrow linguistic scope of previous work. The BabyLM Challenge provides the ideal setting to test this. We propose two new SSLM variants, SubSegGPT and SubSegDeBERTa, which incorporate learnable subword tokenisation into modern, BabyLM-style architectures. Understanding our models requires familiarity with the subword segmental framework, so in the next subsection we provide a technical overview of SSLM.
2.2 Subword Segmental Language Modelling
SSLM (Meyer and Buys, 2022) establishes a framework for learning tokenisation during training. Any LM architecture can be adapted to the framework, removing its reliance on a fixed tokeniser. The key idea is that tokenisation is cast as a latent variable, inferred jointly with model parameters to optimise the LM objective. Adapting a model for subword segmental modelling requires augmenting its architecture with additional subnetworks, which we describe for our models in Section 3, and adopting the training algorithm of Meyer and Buys (2022), which we now summarise.
For a training sequence , SSLM still minimises the standard LM loss . However, whereas a vanilla LM computes with the chain rule over a single pre-determined subword token sequence, SSLM computes as
| (1) |
where is the set of all candidate tokenisations of : every possible way that the sequence of words can be segmented into subword units (word boundaries are enforced, so subword units cannot span across whitespaces). Each is still computed with the chain rule over the token sequence , but the overall sequence probability now incorporates multiple potential tokenisations of the training example .
Marginalising over is intractable for long sequences, so SSLM introduces two constraints:
- 1.
Subword segments cannot exceed a maximum character length , which is a hyperparameter.
- 2.
In computing the probability of a candidate tokenisation with the chain rule, each next-token probability is conditioned on the untokenised autoregressive (preceding) character-level context, so
(2) where is the character sequence in that precedes . This approximation discards tokenisation history, but enables tractable conditioning via a character-level context encoder.
Finally, to compute Equation 1 efficiently, Meyer and Buys (2022) use a dynamic programming algorithm that iteratively computes , the marginal probability of the sequence up to each character position , for (we refer the reader to Meyer and Buys (2022) for a detailed presentation of the algorithm).
3 Models
The modelling assumptions and training algorithm outlined above describe the generative model of the subword segmental framework. Parameterising this with a neural architecture requires a model capable of computing for any subword token and a mechanism for conditioning this probability on the character-level context . A vanilla LM cannot assign probabilities to arbitrary subwords, so Meyer and Buys (2022) augment the standard autoregressive architecture to enable this. Their methodology can be followed to adapt any architecture for subword segmental modelling. Doing so requires building an architecture with two components: an encoder that computes representations for a character-level context and a subword segment scorer that computes for any subword .
In this section, we present SubSegGPT and SubSegDeBERTa, two new SSLMs that adapt respectively GPT-2 (Radford et al., 2019) and DeBERTa-v2 (He et al., 2021) for learnable subword tokenisation. Their architectures are visualised in Figure 1.
3.1 SubSegGPT
SubSegGPT is a decoder-only SSLM with a GPT-2 backbone. It is a straightforward adaptation of the original, LSTM-based SSLM (Meyer and Buys, 2022) to GPT-style language modelling and makes use of the same training algorithm outlined in Section 2.2. We now describe its architecture, which closely mirrors the transformer-based SSLM of Meyer and Buys (2025), but is parameterised by the GPT-2 architecture to incorporate more recent architectural conventions and match the setup of competitive decoder-based BabyLMs.
3.1.1 Character-level history encoder
To compute we need to encode , the full character sequence preceding the subword segment . We encode with a character-level GPT-2 backbone (excluding the language modelling head), using the final-layer output embedding of the last character before to represent the sequence history and to condition next-subword probabilities .
3.1.2 Subword segment scorer
We follow previous SSLMs in computing as a mixture of two subword probabilities,
| (3) |
where is a mixture coefficient, dynamically computed for each subword with a sigmoid-activated linear projection of .
is a language modelling head that maps to a probability distribution over a fixed subword lexicon containing the most frequent subwords in the training corpus (the lexicon size is a pre-specified hyperparameter). is a small decoder subnetwork, parameterised by a 1-layer character-level LSTM,33 3 SubSegGPT and SubSegDeBERTa are transformer-based LMs, parameterised by deep transformer backbones. LSTMs appear only as 1-layer subnetworks in the LM head, conditioning or generating subwords as short character sequences. that generates subword segment one character at a time and computes the subword probability as a chain rule product of individual character probabilities. It is conditioned on the sequence history by initialising the LSTM hidden state as .
The lexicon layer and character decoder play complementary roles in next-subword prediction. directly computes probabilities of frequent subwords, such as common morphemes and words, but cannot handle infrequent, out-of-lexicon segments. By contrast, can assign probabilities to arbitrary subword segments by composing them character by character, covering rare and previously unseen subwords. The mixture gate learns to balance their contributions dynamically, based on the character-level context .
The mixture-based subword segment scorer enables SubSegGPT to compute for any candidate subword segment at any position in a training sequence. These probabilities (top of Figure 1) are passed to the dynamic programming algorithm of Section 2.2, which efficiently computes the marginal (Equation 1) via the probabilities of all candidate tokenisations (Equation 2). SubSegGPT is trained end-to-end by minimising the negative log-likelihood, jointly optimising autoregressive language modelling and subword tokenisation.
3.2 SubSegDeBERTa
We propose SubSegDeBERTa, an encoder-based, masked SSLM with a DeBERTa-v2 backbone. Meyer and Buys (2022) introduce subword segmental modelling for decoder-only LMs. Their framework is inherently autoregressive, so its extension to masked language modelling is non-trivial. In SubSegDeBERTa, we develop the first masked SSLM, incorporating learnable subword tokenisation into encoder-based pretraining. We encode bi-directional character-level context with a DeBERTa-v2 backbone, randomly mask a subset of input words, and jointly learn to generate and tokenise masked words into subword units.
3.2.1 Character-level context encoder
We randomly mask a fixed proportion of words in each training sequence. We define words as whitespace-delimited character sequences and restrict subword segments to span within word boundaries. If a word is sampled for masking, we replace its entire character sequence with a single [MASK] token (as shown in Figure 1 for the masked word “dogs”). Our bi-directional encoder is a character-level DeBERTa-v2 architecture that produces final-layer output embeddings for all characters, including [MASK] tokens.
3.2.2 Masked word scorer
In vanilla MLMs, the final-layer representation is used to predict the masked token. In SubSegDeBERTa, the masked word is not predicted as a single token. Instead, we generate masked words subword segmentally i.e. by marginalising over all possible subword tokenisations of a masked word.
Suppose we mask the word in a sentence, denoted by and consisting of characters . During training, we maximise
| (4) |
where is the set of all candidate tokenisations of , and word probabilities are conditioned on bi-directional context (all words to the left and right of the target word). The probability of each tokenisation is computed with the chain rule as
| (5) |
so this component of SubSegDeBERTa is autoregressive: masked word generation is conditioned on bi-directional context beyond the word (), but within the word it is conditioned on left-to-right context ( is the character sequence preceding subword segment within word ).
To compute the marginal of Equation 4 efficiently, we use the same dynamic programming algorithm discussed in Section 2.2 and introduce the same simplifying assumptions: we limit the length of to a pre-specified maximum number of characters and condition on untokenised character-level context, although here the context is bi-directional (Section 3.2.1). To compute the subword probability , we use the same mixture model as SubSegGPT (Section 3.1.2). However, the masked LM setting introduces a complication that requires additional subnetworks.
The GPT-2 backbone of SubSegGPT outputs contextual representations for each character in an input sequence. These representations are passed to the subword segment scorer and are used to condition next-subword probabilities at any position in the character sequence (including subword segments that start mid-word, such as “gs” in “dogs”, on the left of Figure 1). The DeBERTa backbone of SubSegDeBERTa outputs contextual representations for all characters in a sequence except masked words, whose characters are replaced by a single [MASK] token. encodes the bidirectional context , but provides only a single representation per masked word. To score subword segments at each position within , we additionally need per-position representations that encode the autoregressive within-word history .
To address this, we introduce a word context encoder to encode the context at every character in word (its role is visualised in Figure 1). This is a 1-layer character-level LSTM that processes the target word’s characters left-to-right, conditioned on in two ways: we initialise its hidden state as a learned projection of and concatenate to input character embeddings at every step. This produces a sequence of per-position context representations , that encode the bidirectional context via and the within-word history () via LSTM recurrence. This provides a contextual representation for every character position in , which is used to condition for any subword segment in , as shown for all possible subwords in the word “dogs”, on the right of Figure 1.
To compute these probabilities, we use the same mixture model as the SubSegGPT subword segment scorer (Equation 3), with passed as conditioning context, instead of . The dynamic programming algorithm of Section 2.2 computes the masked word marginal (Equation 4) via the probabilities of all candidate tokenisations (Equation 5). SubSegDeBERTa is trained end-to-end by minimising the negative log-likelihood of masked words, jointly optimising masked language modelling and subword tokenisation.
| Track | Model | BLiMP | BLiMP Sup. | EWoK | Entity | COMPS | PIQA | Avg. |
|---|---|---|---|---|---|---|---|---|
| Random chance | 50.00 | 50.00 | 50.00 | 20.00 | 50.00 | 37.50 | 42.92 | |
| Strict | GPT-2 | 74.73 | 65.00 | 54.37 | 16.91 | 55.85 | 36.62 | 50.58 |
| SubSegGPT | 78.22 | 64.78 | 51.72 | 16.12 | 53.75 | 40.15 | 50.79 | |
| DeBERTa | 73.95 | 64.20 | 52.35 | 17.46 | 53.13 | 35.30 | 49.40 | |
| SubSegDeBERTa | 78.22 | 68.73 | 55.39 | 22.28 | 57.77 | 40.02 | 53.74 | |
| Strict-Small | GPT-2 | 65.23 | 57.25 | 50.63 | 19.1 | 51.81 | 35.09 | 46.52 |
| SubSegGPT | 68.80 | 62.77 | 49.76 | 20.48 | 50.51 | 30.67 | 47.17 | |
| DeBERTa | 63.78 | 59.94 | 50.28 | 19.67 | 50.88 | 35.28 | 46.64 | |
| SubSegDeBERTa | 64.68 | 60.56 | 50.99 | 17.89 | 51.28 | 33.74 | 46.52 |
4 Experimental Setup
4.1 Pretraining
We follow the guidelines of the 2026 BabyLM Challenge (Choshen et al., 2026) to pretrain SubSegGPT and SubSegDeBERTa for Strict and Strict-small. We use the text-only datasets released by the organisers and pretrain for 10 epochs.
To test the impact of learnable tokenisation, we compare our models to fixed-tokenisation BabyLMs based on the corresponding backbone architectures and trained on the same data. For SubSegGPT, we use the GPT-2 (Radford et al., 2019) baselines released by the BabyLM Challenge organisers. For SubSegDeBERTa, we pretrain our own DeBERTa-v2 (He et al., 2021) models as baselines, matching the backbone architectural configurations of SubSegDeBERTa for comparability. Table 4 in the appendix details our baseline setup.
4.2 Hyperparameters
Our backbone architectures match the size of the BabyLM GPT-2 baseline, which corresponds to Base configurations of GPT-2 and DeBERTa. We tune pretraining hyperparameters as detailed in Appendix A. Our backbone architectures are character-based, so their embedding matrices contribute negligibly to overall parameter count. However, our subnetworks for subword segment scoring introduce additional parameters. The model sizes and hyperparameter settings of our submissions and baselines are reported in Tables 4 and 5 in the appendix.
4.3 Evaluation
We evaluate on the official 2026 BabyLM Challenge evaluation pipeline, which includes three types of tasks (full list in Appendix B). (1) Zero-shot tasks test linguistic knowledge directly from model probabilities via minimal pair evaluations. (2) Finetuning tasks test natural language understanding on (Super)GLUE (Wang et al., 2018; Wang et al., 2019) with task-specific finetuning. (3) Human-likeness tasks evaluate how well model predictions align with human psycholinguistic data (we report these results in Appendix D, as neither our models nor baselines perform well on these tasks).
4.4 Evaluating Subword Segmental LMs
The BabyLM evaluation pipeline is designed for standard LM architectures. To evaluate our models on finetuning tasks, we can apply the pipeline without change (final-layer representations are passed to classification heads). However, for zero-shot tasks, the pipeline expects per-position logits, which our models do not emit: SubSegGPT outputs a per-sentence marginal (Equation 1) and SubSegDeBERTa outputs a per-word marginal (Equation 4). We implement model-specific wrappers that transform our outputs into the quantities required for each task. We leave the official pipeline unchanged, except for one line of code (see Appendix C.1 for details). Appendix C describes the wrappers we implement to evaluate our models.
| Track | Model | BoolQ | MNLI | MRPC | MultiRC | QQP | RTE | WSC | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Majority class | 64.04 | 35.70 | 68.14 | 57.55 | 62.78 | 53.96 | 61.54 | 57.67 | |
| Strict | GPT-2 | 69.66 | 60.76 | 85.34 | 65.92 | 71.56 | 57.55 | 63.46 | 67.75 |
| SubSegGPT | 64.59 | 58.21 | 84.51 | 66.38 | 72.34 | 60.43 | 63.46 | 67.13 | |
| DeBERTa | 72.66 | 62.59 | 84.80 | 68.15 | 78.64 | 66.19 | 63.46 | 70.93 | |
| SubSegDeBERTa | 71.44 | 60.80 | 90.54 | 67.16 | 75.16 | 64.03 | 63.46 | 70.37 | |
| Strict-Small | GPT-2 | 67.71 | 49.84 | 81.37 | 65.76 | 61.67 | 56.83 | 63.46 | 63.81 |
| SubSegGPT | 64.83 | 54.65 | 85.71 | 65.64 | 71.23 | 60.43 | 63.46 | 66.56 | |
| DeBERTa | 67.52 | 47.29 | 70.59 | 67.53 | 70.17 | 56.83 | 61.54 | 63.07 | |
| SubSegDeBERTa | 68.93 | 45.78 | 81.93 | 65.80 | 62.94 | 56.83 | 65.38 | 63.94 |
5 Results
Table 1 reports zero-shot results. In the Strict track, SubSegDeBERTa achieves the highest scores across most tasks, comfortably outperforming the strongest baseline, GPT-2. SubSegGPT also outperforms both tokenisation-based baselines, suggesting that at this scale (10 epochs over 100M words) learning tokenisation during pretraining reliably improves sample-efficiency of intrinsic linguistic knowledge acquisition.
In the Strict-small track, the relative performances of our two models are reversed. SubSegGPT outperforms SubSegDeBERTa and, on average, outperforms both baselines. Performance is more mixed across individual tasks – largely because, for most tasks, all models fail to reach above-chance performance. On the only two tasks where models reliably outperform chance, SubSegGPT achieves large performance gains over its tokenisation-based equivalent, GPT-2 (+3.57 for BLiMP and +5.52 for BLiMP Supplement). SubSegDeBERTa does not outperform DeBERTa in the Strict-small track, but reaches similar performance levels on average.
Overall, these results suggest that the benefits of learnable tokenisation are scale-dependent. In the extremely low-resource setting, where even above-chance zero-shot performance is challenging, SubSegGPT already offers gains. SubSegDeBERTa requires a larger data scale for its sample-efficiency to take effect, but under those conditions it produces even more reliable gains than SubSegGPT. The relative underperformance of SubSegDeBERTa in the Strict-small track might be due to its MLM objective, which provides a sparser training signal than autoregressive modelling. SubSegDeBERTa is only trained to generate a proportion of words in each sequence, whereas SubSegGPT is trained to generate every word in the corpus.
Table 2 reports results for text classification and entailment tasks from the (Super)GLUE benchmark. With task-specific finetuning, the benefits of subword segmental pretraining diminish. In the Strict track, neither of our models consistently outperform their baselines. Among all the models we tested, DeBERTa achieves the highest average performance. At the scale of 100M words, conventional encoder-only pretraining with fixed subword tokenisation, combined with downstream finetuning, is sufficient for natural language understanding tasks. In the Strict-small track, SubSegGPT again achieves the best performance overall, as it did in zero-shot evaluation. This supports our claim that SubSegGPT provides better sample-efficiency when pretraining data is severely limited, and that this pretraining sample-efficiency transfers to improved downstream finetuning.
| SubSegGPT | SubSegDeBERTa | |
|---|---|---|
| 1M | ||
| 10M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 100M |
6 Analysing Subword Learning
In SubSegGPT and SubSegDeBERTa, we can study subword tokenisation as a learnable component of language modelling: equipped with the ability to learn tokenisation, how do subword boundaries evolve over pretraining and what are the linguistic properties of the final subword units?
6.1 Subword Learning Dynamics
We compare tokenisations of the 2022 SIGMORPHON English test set (Batsuren et al., 2022) across regular interval checkpoints (every 1M words until 10M words, every 10M words until 100M words). To extract the learned tokenisation of a sentence, we use the Viterbi algorithm to extract the highest-probability tokenisation (replacing the sum in Equations 1 and 4 with an argmax).
For each checkpoint, we quantify the properties of its subwords with three metrics. (1) Boundary flip rate measures the rate of change in tokenisation as the fraction of possible subword boundaries that change from one checkpoint to the next. (2) Fertility (Ács, 2019) is the average number of subwords per word, reflecting the granularity of subword tokenisation. (3) Morpheme boundary identification F1 measures the overlap between learned subword boundaries and ground truth morpheme boundaries, as annotated in the SIGMORPHON dataset.
Figure 2 plots the learning dynamics of SubSegGPT and SubSegDeBERTa during Strict-small pretraining. The two models exhibit similar learning trajectories. After an initial period of rapid changes, subword learning gradually stabilises and converges on a settled tokenisation scheme. This convergence is characterised by a steady increase in fertility: words are tokenised into more subword units. As a comparison, the fertility of the BabyLM baseline tokeniser (16k-vocabulary BPE) on this dataset is 1.23, so our models learn more aggressive segmentation. The subword boundaries of both models shift towards greater alignment with morphological boundaries. This alignment remains weak compared to a dedicated unsupervised morphological segmenter like Morfessor (Smit et al., 2014), which achieves 42.9% F1 on the same evaluation set, but is well above random boundary insertion at the same fertility as our models, which would achieve around 11.6% F1.
6.2 Qualitative Analysis
The subword units learned by SubSegGPT and SubSegDeBERTa reflect the tokenisation demands of sample-efficient language modelling. Our quantitative analysis suggests that this balances linguistic plausibility (partial morphological alignment) against frequency-based criteria (finer-grained segmentation). To study this trade-off qualitatively, Table 3 presents CHILDES (MacWhinney, 2000) utterances tokenised by our models, highlighting subwords corresponding to morphemes.
The examples show several instances of our models discovering morphemes as subword units, such as suffixes (“–ing”, “–s”) and compound constituents (“sand–box”). They also show instances of morphologically unsound tokenisation, some of which are model-specific: SubSegGPT performs worse as a morphological segmenter (see Figure 2) and this is reflected in examples like “playi–ng” and “jumpi–ng”. Other morphological violations can be attributed to biases introduced by our architecture design (e.g. our mixture model for subword generation) and constraints imposed by modelling assumptions. For example, subword segments cannot exceed 5 characters, so words like “kitchen” and “little” cannot be left untokenised. Removing these constraints would allow us to study truly unrestricted subword learning, but marginalising over unbounded segment lengths is computationally infeasible in our current setup.
7 Conclusion
We present two new models that learn subword tokenisation during training to optimise their pretraining objectives: SubSegGPT for autoregressive language modelling and SubSegDeBERTa for masked language modelling. Our models excel at different scales: SubSegDeBERTa is our most competitive submission in the Strict track and SubSegGPT performs strongly in the Strict-small track. By combining the SSLM framework with current best practices in sample-efficient pretraining, we show that SSLMs are effective beyond the low-resource, agglutinative languages for which they were originally proposed. More generally, our findings suggest reconsidering tokenisation conventions and identify end-to-end subword modelling as a promising future direction for BabyLM research.
8 Limitations
A disadvantage of subword segmental modelling is the additional computational complexity introduced by its training algorithm. Marginalising over several tokenisations requires more computations than using a single tokenisation, so SubSegGPT and SubSegDeBERTa have much longer training times than fixed-tokenisation BabyLMs (A100 GPU hours are listed in Table 4 in the appendix). For example, training SubSegDeBERTa required approximately 5 the A100 GPU hours of DeBERTa. Our aim is sample-efficiency (better performance with the same pretraining data budget), but this comes at the cost of compute-efficiency (more training FLOPs and longer training times).
References
- Batsuren et al. (2022) Khuyagbaatar Batsuren, Gábor Bella, Aryaman Arora, Viktor Martinovic, Kyle Gorman, Zdeněk Žabokrtský, Amarsanaa Ganbold, Šárka Dohnalová, Magda Ševčíková, Kateřina Pelegrinová, Fausto Giunchiglia, Ryan Cotterell, and Ekaterina Vylomova. 2022. The SIGMORPHON 2022 shared task on morpheme segmentation. In Proceedings of the 19th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, pages 103–116, Seattle, Washington. Association for Computational Linguistics.
- Batsuren et al. (2024) Khuyagbaatar Batsuren, Ekaterina Vylomova, Verna Dankers, Tsetsuukhei Delgerbaatar, Omri Uzan, Yuval Pinter, and Gábor Bella. 2024. Evaluating subword tokenization: Alien subword composition and oov generalization challenge. Preprint, arXiv:2404.13292.
- Chang et al. (2026) Tyler A. Chang, Catherine Arnett, et al. 2026. Global piqa: Evaluating commonsense reasoning across 100+ languages and cultures. Preprint, arXiv:2510.24081.
- Chang and Bergen (2022) Tyler A. Chang and Benjamin K. Bergen. 2022. Word acquisition in neural language models. Transactions of the Association for Computational Linguistics, 10:1–16.
- Choshen et al. (2026) Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj Sanjay Shah, Alex Warstadt, and Ethan Gotlieb Wilcox. 2026. Babylm turns 4 and goes multilingual: Call for papers for the 2026 babylm workshop. Preprint, arXiv:2602.20092.
- Clark et al. (2019) Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota. Association for Computational Linguistics.
- de Varda et al. (2024) Andrea Gregor de Varda, Marco Marelli, and Simona Amenta. 2024. Cloze probability, predictability ratings, and computational estimates for 205 English sentences, aligned with existing EEG and reading time data. Behavior Research Methods, 56:5190–5213.
- Dolan and Brockett (2005) William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
- Downey et al. (2022) C.m. Downey, Fei Xia, Gina-Anne Levow, and Shane Steinert-Threlkeld. 2022. A masked segmental language model for unsupervised natural language segmentation. In Proceedings of the 19th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, pages 39–50, Seattle, Washington. Association for Computational Linguistics.
- Giampiccolo et al. (2007) Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pages 1–9, Prague. Association for Computational Linguistics.
- He et al. (2021) Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
- Hochreiter and Schmidhuber (1997) Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.
- Ivanova et al. (2025) Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas H. Clark, Carina Kauf, Jennifer Hu, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan G. Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Joshua Tenenbaum, and Jacob Andreas. 2025. Elements of world knowledge ( EWoK ): A cognition-inspired framework for evaluating basic world knowledge in language models. Transactions of the Association for Computational Linguistics, 13:1245–1270.
- Jusczyk (1999) Peter W. Jusczyk. 1999. How infants begin to extract words from speech. Trends in Cognitive Sciences, 3(9):323–328.
- Kawakami et al. (2019) Kazuya Kawakami, Chris Dyer, and Phil Blunsom. 2019. Learning to discover, ground and use words with segmental neural language models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6429–6441, Florence, Italy. Association for Computational Linguistics.
- Khashabi et al. (2018) Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 252–262, New Orleans, Louisiana. Association for Computational Linguistics.
- Kim and Schuster (2023) Najoung Kim and Sebastian Schuster. 2023. Entity tracking in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3835–3855, Toronto, Canada. Association for Computational Linguistics.
- Kong et al. (2016) Lingpeng Kong, Chris Dyer, and Noah A. Smith. 2016. Segmental recurrent neural networks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
- Kreutzer and Sokolov (2018) Julia Kreutzer and Artem Sokolov. 2018. Learning to segment inputs for NMT favors character-level processing. In Proceedings of the 15th International Conference on Spoken Language Translation, pages 166–172, Brussels. International Conference on Spoken Language Translation.
- Kudo (2018) Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia. Association for Computational Linguistics.
- Levesque et al. (2011) Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2011. The Winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, volume 46, page 47.
- MacWhinney (2000) Brian MacWhinney. 2000. The CHILDES Project: Tools for Analyzing Talk, 3 edition. Lawrence Erlbaum Associates, Mahwah, NJ.
- Meyer and Buys (2022) Francois Meyer and Jan Buys. 2022. Subword segmental language modelling for nguni languages. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6636–6649, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
- Meyer and Buys (2025) Francois Meyer and Jan Buys. 2025. The learning dynamics of subword segmentation for morphologically diverse languages. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pages 647–661, Mumbai, India. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics.
- Misra et al. (2023) Kanishka Misra, Julia Rayz, and Allyson Ettinger. 2023. COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2928–2949, Dubrovnik, Croatia. Association for Computational Linguistics.
- Radford et al. (2019) Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
- Salazar et al. (2020) Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
- Samuel (2024) David Samuel. 2024. Berts are generative in-context learners. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA. Curran Associates Inc.
- Sennrich et al. (2016) Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
- Smit et al. (2014) Peter Smit, Sami Virpioja, Stig-Arne Grönroos, and Mikko Kurimo. 2014. Morfessor 2.0: Toolkit for statistical morphological segmentation. In Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics, pages 21–24, Gothenburg, Sweden. Association for Computational Linguistics.
- Sun and Deng (2018) Zhiqing Sun and Zhi-Hong Deng. 2018. Unsupervised neural word segmentation for Chinese via segmental language modeling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4915–4920, Brussels, Belgium. Association for Computational Linguistics.
- Vaswani et al. (2017) Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
- Wang et al. (2019) Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. SuperGLUE: a stickier benchmark for general-purpose language understanding systems. Curran Associates Inc., Red Hook, NY, USA.
- Wang et al. (2018) Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
- Wang et al. (2017) Chong Wang, Yining Wang, Po-Sen Huang, Abdelrahman Mohamed, Dengyong Zhou, and Li Deng. 2017. Sequence modeling via segmentations. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, page 3674–3683. JMLR.org.
- Warstadt et al. (2020) Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. BLiMP: A benchmark of linguistic minimal pairs for English. In Proceedings of the Society for Computation in Linguistics 2020, pages 409–410, New York, New York. Association for Computational Linguistics.
- Williams et al. (2018) Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
- Ács (2019) Judit Ács. 2019. Exploring BERT’s vocabulary.
Appendix A Hyperparameters
During development, we tuned pretraining hyperparameters by comparing performance on the Fast evaluation pipeline, which samples subsets of the zero-shot BabyLM evaluation datasets. Among standard pretraining hyperparameters, we tuned the learning rate, batch size, warmup ratio, and masking ratio (for SubSegDeBERTa). Among hyperparameters unique to subword segmental modelling, we tuned the maximum subword segment length and the subword lexicon size (the number of top-frequency character n-grams to include in the lexicon subword scorer).
We did not perform a full grid search. Instead, we started with the BabyLM baseline hyperparameters as our default setup and varied one hyperparameter at a time. This revealed that changing certain hyperparameters (batch size, warmup ratio, and maximum segment length beyond 5 characters) had little effect on performance, while others (learning rate, masking ratio, and subword lexicon size) were influential. We subsequently experimented with different learning rates (1e-3, 1e-4, 5e-4), masking ratios (0.15, 0.3, 0.4, 0.5), and subword lexicon sizes (5k, 10k, 20k, 40k), conducting most experiments in the Strict-small setting due to its faster pretraining iterations, with more limited experimentation in the Strict setting. The hyperparameters of our final submissions are reported in Tables 4 and 5.
| Hyperparameter | GPT-2 | SubSegGPT | DeBERTa-v2 | SubSegDeBERTa |
|---|---|---|---|---|
| Strict / Small | Strict / Small | Strict / Small | Strict / Small | |
| A100 training | – | 85h / 16h | 17h / 2h | 80h / 9h |
| Parameters | 98.4M | 95.8 / 103.1M | 113.2M | 140.4 / 116.9M |
| Layers | 12 | 12 | 12 | 12 |
| Hidden size | 768 | 768 | 768 | 768 |
| FF size | 3,072 | 3,072 | 3,072 | 3,072 |
| Attention heads | 12 | 12 | 12 | 12 |
| Dropout | 0.1 | 0.1 | 0.1 | 0.1 |
| Vocabulary size | 16,384 | – | 16,384 | – |
| Sequence length | 512 | 1,024 | 512 | 1,024 |
| Learning rate | 5e-5 | 5e-4 | 5e-5 | 5e-4 |
| LR scheduler | cosine | cosine | cosine | cosine |
| Warmup ratio | 0.01 | 0.1 | 0.01 | 0.1 |
| Weight decay | 0 | 0.01 | 0 | 0.01 |
| Gradient clipping | 1.0 | 1.0 | 1.0 | 1.0 |
| Masking ratio | – | – | 0.3 | 0.3 / 0.4 |
| Batch size | 16 | 16 | 16 | 16 |
| SubSegGPT | SubSegDeBERTa | |||
| Strict | Small | Strict | Small | |
| Char vocab | 742 | 416 | 744 | 418 |
| Lexicon size | 10k | 20k | 40k | 10k |
| Max segment | 5 | 5 | 5 | 5 |
| Char decoder LSTM | ||||
| – hidden size | 256 | 256 | 256 | 256 |
| – layers | 1 | 1 | 1 | 1 |
| – embedding | 128 | 128 | 128 | 128 |
| Word context encoder LSTM | ||||
| – hidden size | – | – | 768 | 768 |
| – layers | – | – | 1 | 1 |
Appendix B Evaluation Tasks
The official 2026 BabyLM Challenge evaluation pipeline contains three types of tasks for the Strict and Strict-small tracks.
- 1.
Zero-shot: BLiMP (Warstadt et al., 2020), BLiMP Supplement, EWoK (Ivanova et al., 2025), COMPS (Misra et al., 2023), Entity Tracking (Kim and Schuster, 2023), and the English subset of Global PIQA (Chang et al., 2026) (a hidden task announced shortly before the deadline).
- 2.
Finetuning: BoolQ (Clark et al., 2019), MultiRC (Khashabi et al., 2018), RTE (Giampiccolo et al., 2007), WSC (Levesque et al., 2011), MRPC (Dolan and Brockett, 2005), QQP, and MNLI (Williams et al., 2018).
- 3.
Human likeness: Reading correlations (de Varda et al., 2024) are computed based on self-paced reading times and eye tracking data. Age-of-acquisition scores (Chang and Bergen, 2022) are computed by tracking word surprisal across pretraining checkpoints and comparing learning curves to child vocabulary acquisition data.
Appendix C Evaluation Wrappers
Because our models do not output per-position logits over a fixed subword vocabulary, they cannot be evaluated directly with the official BabyLM zero-shot and human likeness evaluation pipelines. Here we describe the wrappers we implement to transform the outputs of SubSegGPT and SubSegDeBERTa (marginal probabilities) into the quantities required by each evaluation task.
C.1 SubSegGPT
Zero-shot tasks compare log-probabilities of minimal pair sentences. The evaluation pipeline computes this by summing per-position log-probabilities over the tokens in a sentence. SubSegGPT computes the log-probability of a full sentence, which we transform to per-character log-probability estimates by dividing the sentence log-probability by the number of characters in a sentence. This ensures that the per-position log-probabilities summed by the evaluation pipeline add up to the true sentence log-probability, enabling valid minimal pair comparisons. To force the evaluation pipeline to compare full sentence log-probabilities, rather than only the log-probabilities of the differing spans between minimal pairs, we disable sentence masking in the official pipeline.44 4 To evaluate SubSegGPT on zero-shot tasks, disable phrase masking by changing the following line to phrase_mask = [0 for _ in range(len(tokens))]: https://github.com/babylm-org/babylm-eval/blob/68cdd160f34826307e650c484904c274692e82ce/strict/evaluation_pipeline/sentence_zero_shot/dataset.py#L89.
Human-likeness tasks require word-level surprisal, which vanilla LMs compute by summing subword token log-probabilities. We instead derive this as
| (6) |
where both terms on the right are computed as full sequence marginals (Equation 1) using our dynamic programming algorithm.
C.2 SubSegDeBERTa
For MLMs, the BabyLM pipeline scores sentences with pseudo-log-likelihood (Salazar et al., 2020), masking each subword token in turn and summing the log-probabilities over tokens in a sentence. SubSegDeBERTa lends itself naturally to this type of evaluation, as it computes the probability of a masked word (Equation 4), which can be used for cloze-style scoring. For zero-shot tasks we mask each word in turn and sum the marginal of each word to compute the pseudo-log-likelihood of a sentence. For human-likeness tasks, we mirror the BabyLM pipeline, which estimates word surprisal in MLMs by masking the target word at the end of its context, so the model conditions only on preceding words. For reading time evaluation, we match the official evaluation pipeline by applying multi-mask ending (Samuel, 2024): appending three trailing [MASK] tokens to obtain a less restricted sentence continuation prediction.
| SubSegGPT | SubSegDeBERTa | |
|---|---|---|
| 1M | ||
| 10M | ||
| 50M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 50M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 50M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 50M | ||
| 100M | ||
| 1M | ||
| 10M | ||
| 50M | ||
| 100M |
| Track | Model | Read | AoA |
|---|---|---|---|
| Strict | GPT-2 | 6.93 | -11.58 |
| SubSegGPT | 1.56 | 0.00 | |
| DeBERTa | 4.76 | 0.00 | |
| SubSegDeBERTa | 2.47 | -20.47 | |
| Strict-small | GPT-2 | 5.63 | -12.15 |
| SubSegGPT | 2.21 | 0.00 | |
| DeBERTa | 4.17 | 0.00 | |
| SubSegDeBERTa | 2.22 | 0.00 |
Appendix D Human-Likeness Results
Neither of our models exhibits strong correlations with human psycholinguistic data, as shown in Table 7. Read scores are positive, so incorporating word-level surprisal from our models does slightly improve reading time prediction regression, but baseline surprisals lead to greater improvements than SubSegGPT and SubSegDeBERTa. AoA scores are zero or negative for all the models we tested, which shows that model word acquisition patterns exhibit no correlation with child acquisition data.