arXiv is now an independent nonprofit! Learn more
License: CC BY-SA 4.0
arXiv:2607.01502v1 [cs.CL] 01 Jul 2026

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

Jesujoba O. Alabi1, Julian Herreilers2, Badr M. Abdullah1, Dietrich Klakow1
Abstract

Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architectures in multiple languages, their effectiveness in African languages remains underexplored. In this work, we evaluate Mamba for ASR on seven South African languages. In monolingual experiments, each model is trained on 50 hours of speech per language, and we compare Mamba to a Conformer baseline of similar parameter scale. Mamba achieves similar recognition accuracy to Conformer while using fewer computational resources and training faster. We further evaluate generalization in this setting and find that both models struggle to generalize to speech that is much longer than what they were trained on. We then study multilingual ASR using Mamba models, where the baseline is pooling all languages together. On top of this, we tested three extensions: training with language-family information by adding both language and language-family embeddings as biases to the downsampled acoustic representations, and multitask learning with a CTC ASR objective and a language identification (LID) head. We find that multilingual training consistently improves performance over monolingual training. However, adding explicit language information does not improve in-domain performance but does improve cross-corpus robustness. We conducted ablation studies in low-resource multilingual settings using 5-hour and 10-hour per-language training data, where we observed gains from using language embeddings and further demonstrated that removing or altering them hurt model performance. Lastly, we analysed these embeddings and find that they do not capture linguistic similarity in a typological sense, but instead act as task-specific control vectors.

I Introduction

Automatic speech recognition (ASR) research has advanced significantly with the development of powerful sequence modeling architectures [1, 2, 3]. Attention-based models such as Conformer [4] have become strong baselines due to their ability to capture both local and global dependencies through the integration of self-attention and convolutional modules. More recently, state space model (SSM)-based architectures such as Mamba [5] have emerged as efficient alternatives, offering linear-time sequence modeling and favorable memory scaling compared to quadratic-time attention mechanisms, making them particularly attractive for long-form speech processing.

Despite these advances, research on Mamba for ASR has primarily focused on high-resource languages, a small number of low-resource languages, long-context ASR, multilingual ASR, and streaming ASR, typically in comparison with Conformer-based and other architectures [6, 7, 8]. Relatively little attention has been paid to African languages, which face a “low-resource double bind” [9]: limited annotated speech data and constrained computational resources. Although recent efforts have expanded speech resources for African languages [10, 11], it remains necessary to understand how Mamba performs under conditions of limited data and long-form speech recognition.

Long-form ASR is important in low-resource settings, where real-world speech exhibits substantial variation in utterance length, ranging from short prompted phrases to extended spontaneous speech. However, limited data availability often results in insufficient coverage of long-utterance structures during training, making it difficult for models to learn robust representations across sequence lengths. Prior work on long-form ASR in high-resource languages, including the study by [12], has shown that modeling long utterances introduces challenges in maintaining robustness over over long inputs and handling distributional shifts in sequence length. These challenges are likely to be amplified in low-resource settings. This highlights the difficulty of building robust ASR systems under limited supervision and motivates the use of multilingual training strategies, which are commonly employed to mitigate data scarcity in low-resource ASR.

To address these gaps, we conduct a systematic study of Mamba for ASR across seven South African languages. We first compare Mamba against a Conformer-based encoder baseline in monolingual settings, where each model is trained on approximately 50 hours of transcribed speech per language. Our analysis focuses on long-context recognition behavior and length generalization. We then extend the study to the multilingual setting using Mamba-based models. In this setting, we consider a pooled multilingual baseline and three extensions for incorporating language information: (i) adding language embeddings as biases to the acoustic representations, (ii) extending (i) by also incorporating language-family embeddings as biases, and (iii) a multitask learning setup with a connectionist temporal classification (CTC) ASR objective and an auxiliary language identification (LID) head. We further analyze approach (ii) under low-resource multilingual conditions (5-hour and 10-hour per-language settings), a setup commonly used to mitigate data scarcity in ASR, in order to study the effect of explicit language conditioning. Additionally, we investigate the representations learned by the language embeddings.

Our results show performance differences across languages in the monolingual setting, with Nguni languages being more challenging than Sotho languages. Mamba achieves performance comparable to Conformer models of similar parameter size, while using fewer computational resources and training faster, indicating that SSMs can match strong attention-based baselines for South African ASR. Both models show limited robustness to speech that is substantially longer than what they were trained on, highlighting challenges in length generalization. In the multilingual setting, joint training consistently improves performance over monolingual models, confirming the benefits of cross-lingual parameter sharing. However, incorporating explicit language or language-family information does not improve in-domain performance but improves cross-corpus robustness. Furthermore, in low-resource multilingual settings, explicit language embeddings provide consistent performance gains. Further analysis shows that these embeddings do not capture linguistic similarity in a typological sense, but instead act as task-specific steering vectors.

TABLE I: Dataset statistics for our experiments. Only the test splits of NCHLT and FLEURS were used.
Bucket nbl sot tsn tso ven xho zul
Swivuriso (train)
Dur (h) 50.0 50.0 50.0 50.0 50.0 50.0 50.0
Utt. 13.6k 11.6k 14.5k 11.6k 11.8k 11.4k 9.3k
Swivuriso: (dev test)
Dur (h) 14.2 25.1 25.8 25.7 11.5 25.6 25.7
Short Utt. 2266 2717 4304 3500 1195 2470 1810
Long Utt. 424 1002 965 894 460 1169 1215
NCHLT (test)
Utt. 3108 2722 2889 2905 2805 2770 2802
FLEURS (test)
Utt. - - - - 1041 854

II South African Language ASR

Language Characteristics: South Africa is a multilingual country with 11 official languages, as recognized in the Constitution, with South African Sign Language also recently granted official status. Most of these languages belong to the Bantu language family, while Afrikaans and English are Germanic Indo-European languages. The Bantu languages spoken in South Africa include the Nguni languages (isiZulu, isiXhosa, siSwati, and isiNdebele), which together with Xitsonga (from the Tswa-Ronga branch) form the broader Nguni-Tsonga grouping in many linguistic classifications. They also include the Sotho-Tswana languages (Sesotho, Sesotho sa Leboa, and Setswana), as well as Tshivenda, which belongs to the Venda branch and is often classified within a broader Sotho-Venda (or Sotho-Makua-Venda) grouping. All these languages use Latin scripts, and code-switching with English or other languages is common in informal speech.

ASR Resources: A major challenge for ASR research in African languages is resource scarcity, particularly the limited availability of large, high-quality training and evaluation datasets [13, 14]. In the case of South African languages, several resources exist, including Codeswitch Soap opera [15], NCHLT Speech corpus [16, 17], Lwazi ASR [18], Common Voice [19], FLEURS [20], Vuk’uzenzele (ViXSD) [21], and the recently released Swivuriso [10] dataset, which covers seven languages, while NCHLT Speech and Lwazi ASR cover ten. These datasets provide a foundation for ASR development.

ASR Modeling: Given the limited resources for African languages, several efforts have focused on developing ASR systems for South African languages. In particular, prior work has addressed code-switched speech involving English, including comparisons of bilingual and multilingual systems and acoustic modeling tailored to these linguistic contexts, with an emphasis on languages such as isiZulu, isiXhosa, Setswana, and Sesotho [22, 23, 24, 25]. Furthermore, multilingual speech encoders have been developed to improve representations for these languages [26, 27]. Despite this progress, there remains a need for resource-efficient ASR models, such as Mamba-based models, that maintain robust performance under limited training data, computational constraints, and varying utterance lengths.

III Experimental Setup

III-A Datasets

For our experiments and analysis, we use ASR data from three sources, which are:

Swivuriso: [10] is a large speech corpus for South African languages that covers seven languages, including isiNdebele (nbl), Sesotho (sot), Setswana (tsn), Xitsonga (tso), Tshivenda (ven), isiXhosa (xho), and isiZulu (zul). It contains scripted and unscripted speech across domains such as agriculture, healthcare, and general conversation, with over 90 speakers per language. The transcriptions include diacritics, punctuations, special symbols such as [cs] to denote code-switching, [pause] for pause, and [?].

For this study, we define short utterances as speech segments between 0 and 30 seconds, reflecting typical ASR training conditions. Segments longer than this are considered long utterances. We sample 50 hours of short utterances per language to construct a balanced dataset comprising scripted and unscripted speech from the train set. Short utterances from the dev set are used for validation, while the dev_test set is split into short and long subsets for evaluation. Table I summarizes the statistics of the data set. In the test split, short utterances outnumber long utterances by at least a factor of two in most languages, except isiZulu.

NCHLT Speech corpus: [16, 17] is another large speech corpus that covers all 11 official South African languages. It contains more than 100 hours of scripted speech and more than 200 speakers per language. We used only the test splits of the seven languages shared with Swivuriso.

FLEURS: [20] is a multilingual speech corpus that covers 102 languages, created by recording text translated by humans into the respective languages. It includes only four South African languages, of which only isiXhosa and isiZulu were used in this work. We used test splits for both languages.

III-B Dataset Preprocessing

All audio files were downsampled to 16 kHz. The transcriptions were preprocessed by normalizing them using Unicode Normalization Form C (NFC), removing all punctuation except a few. The retained punctuation includes ?!-́,.;%=+*# and numeric digits. We also removed special tokens we noticed in the data set, including [pause], [um], [cs], and [?], and lowercase all transcriptions.

III-C Model Architectures

For our experiments, we use encoder-only ASR models based on Mamba or Conformer architectures, both trained with Connectionist Temporal Classification (CTC) [28], and use character-level modeling. We chose CTC due to its simplicity, computational efficiency, and previous use in related work [12]. To ensure a fair comparison, both architectures share the same configuration, including 18 encoder layers, a hidden size of 512, and a feed-forward dimension of 2048, resulting in comparable model scales (114M parameters for Conformer and 123M parameters for Mamba).

The Conformer follows the standard architecture [4], consisting of relative positional self-attention [29], convolution modules, and feed-forward networks. In contrast, Mamba entirely replaces self-attention with a state-space sequence modeling mechanism. For the Mamba-based system, we adopt ConMamba [30], an architecture designed for speech processing that integrates a bidirectional Mamba module (BiMamba; dstate=16d_{\text{state}}=16, expand=2, dconv=4d_{\text{conv}}=4), a feed-forward network and a convolutional module to efficiently capture local and long-range dependencies.

III-D ASR Training Strategies

Our experiments consider both monolingual and multilingual training regimes. In the monolingual setting, we train a separate model for each language using language-specific data and evaluate it in the corresponding language. For the multilingual setting, we evaluate five training strategies by combining monolingual data.

  1. 1.

    Multilingual-Implicit (MI): A single model trained on the pooled data from all languages without any explicit language information.

  2. 2.

    Multilingual-Implicit Family (MIF): For each language family, one model trained on pooled data from all languages in that family, without using any explicit language-identifying information.

  3. 3.

    Multilingual Language Embedding (MLE): Similar to (1), but with the introduction of a learnable language embedding matrix E()nlang×32E^{(\ell)}\in\mathbb{R}^{n_{\text{lang}}\times 32}. For each input sequence that belongs to language \ell, the corresponding embedding is added to the acoustic representations produced by the CNN downsampling module:

    𝐡t=𝐡t+𝐞,t,\mathbf{h}^{\prime}_{t}=\mathbf{h}_{t}+\mathbf{e}_{\ell},\quad\forall t, (1)

    where 𝐡t32\mathbf{h}_{t}\in\mathbb{R}^{32} is the frame-level acoustic representation and 𝐞32\mathbf{e}_{\ell}\in\mathbb{R}^{32} is the embedding corresponding to language \ell. The resulting representations are then fed into the Mamba module. Unlike [31], we do not concatenate language embeddings to acoustic features.

  4. 4.

    Multilingual Language-Family Embedding (MLFE): We extend MLE by additionally introducing a language-family embedding matrix E(f)nfam×32E^{(f)}\in\mathbb{R}^{n_{\text{fam}}\times 32}. The acoustic representations are augmented as:

    𝐡t=𝐡t+𝐞+𝐞f,t,\mathbf{h}^{\prime}_{t}=\mathbf{h}_{t}+\mathbf{e}_{\ell}+\mathbf{e}_{f},\quad\forall t, (2)

    where 𝐞f\mathbf{e}_{f} corresponds to the embedding of the language family.

  5. 5.

    Multilingual ASR + LID (M-CTC+LID): A multitask learning setup where a shared encoder is trained jointly with a CTC-based ASR objective and an auxiliary language identification (LID) head. The total loss is:

    =CTC+λLID,\mathcal{L}=\mathcal{L}_{\text{CTC}}+\lambda\mathcal{L}_{\text{LID}}, (3)

    where λ\lambda controls the contribution of the LID objective.111λ\lambda was set at 0.1 after ablation over 0.05, 0.1, 0.2, 0.3.

III-E Training Configuration and Hyperparameters

All models were trained with three seeds for 250 (monolingual) or 100 (multilingual) epochs on NVIDIA A100 GPUs (40GB or 80GB, with 80GB used for Conformer models). Training used bfloat16 mixed-precision, a batch size of 32, and identical optimization settings following a publicly available SpeechBrain-based recipe [32].222https://github.com/mattmireles/Mamba-ASR We used the AdamW optimizer tuning the learning rate over 1e-3, 8e-4, 5e-4, 2e-4, 1e-4. All models were trained using a vocabulary of 189 characters derived from the Swivuriso combined training split. Although the target South African languages were not expected to require such a large character inventory, we observed that the dataset includes multiple characters with diacritics, as well as additional non-standard symbols, likely due to portions of the data being sourced from the web (Wikipedia). These include, among others, characters such as ß and Greek letters (e.g. σ\sigma and λ\lambda). For evaluation, we averaged the weights of the five best checkpoints per seed and report the mean performance over three seeds, with punctuation removed and Word Error Rate (WER) as the metric.

TABLE II: Performance comparison (WER %) between monolingual Conformer and ConMamba models on long and short utterances. Each model is evaluated on the same language on which it was trained on (Swivuriso). Scores are averaged over three seeds, with standard deviations reported as subscripts.
\rowcolorwhite \cellcolorwhite \cellcolorwhite \cellcolorOIblue!30Nguni \cellcolorOIblue!80!OIorange!20Tsonga \cellcolorOIgreen!35Sotho \cellcolorOIgreen!80!OIvermillion!20Venda
Setup Size nbl xho zul tso sot tsn ven Avg
Short Utterances (0-30s)
Conformer 114M 42.292.08 40.280.55 45.081.11 35.193.68 31.091.64 26.762.96 27.711.59 35.49
ConMamba 123M 40.721.19 40.181.27 44.190.99 29.921.04 29.250.10 23.300.06 22.770.14 32.91
Long Utterances (>>30s)
Conformer 114M 43.182.51 38.260.54 47.361.16 41.403.62 31.541.55 32.123.66 31.691.73 37.94
ConMamba 123M 42.111.14 38.661.46 47.091.14 35.941.00 29.700.12 28.250.12 25.980.16 35.39

IV Result and Discussion

In this section, we present the performance of Mamba for ASR across seven languages. We first compare Mamba with Conformer in a monolingual setting, followed by an evaluation of both models on long-context ASR. We then assess Mamba’s effectiveness in multilingual training by comparing five training strategies, examine its cross-dataset generalization, and finally present an ablation study of the multilingual Mamba model with language embeddings (MLE).

IV-A Monolingual ASR Performance

Table II presents the monolingual WER comparison between Conformer and Mamba models on both short and long utterances. All models were trained on short utterances and evaluated on the same language. The in-domain evaluation result shows comparable performance across languages for Conformer and ConMamba: on average, ConMamba achieves a 33.43% WER, while Conformer achieves 35.49%. Despite comparable accuracy, ConMamba, with nearly 10M additional parameters, is more efficient in training, inference, and memory usage [6]. In our experiments, Mamba required approximately 18 hours of training per language, compared to 34 hours per language for Conformer. Due to higher memory requirements, Conformer could not be trained on A100 40GB GPUs and therefore required A100 80GB GPUs, whereas Mamba could be trained on A100 40GB GPUs. Furthermore, the results indicate that the Nguni languages (nbl, xho, and zul) consistently yield the lowest performance across both models, followed by tso, a related language. In contrast, ven and tsn achieve the strongest results across both models. All monolingual models were trained under identical experimental conditions, making data imbalance an unlikely explanation. We hypothesize that this pattern may be linked to the complex agglutinative morphology of Nguni languages. Furthermore, not including English in our training could affect performance, especially in code-switching scenarios where English is often used [15]. More analysis is needed to better understand these challenges and improve performance.

IV-B Robustness to Length Variation

Evaluation on long utterances outside the training length distribution shows that WER is similar to or slightly higher than for short utterances in most languages, except xho, where it is lower. The average WER difference of 2.5%\sim 2.5\% indicates comparable length generalization across models, with similar degradation trends likely due to the CTC decoder. In  [33] it was shown that incorporating Mamba in encoder-decoder architectures helps reduce degradation for long-form ASR, and we investigate if the same holds true for encoder-only ASR.

To investigate this, we performed an ablation experiment by sampling and concatenating dev_test utterances of up to 60s, inserting 0.12 s of silence between them, to create sequences spanning progressively longer ranges (i.e. 60-90s, 90-120s, …, 210-240s). Scripted and unscripted speech from the same speaker was kept separate and concatenation was restricted to utterances from that speaker.

Refer to caption
Figure 1: Difference in WER (%) relative to baseline for Conformer and ConMamba across increasing speech length ranges.

Figure 1 shows the performance degradation of the monolingual models for nbl, tsn, ven, and zul, using inference on the original short utterances (0–30 s) as baselines. We observe that both architectures experience increased degradation as the utterance duration grows; however, ConMamba exhibits slightly less degradation overall compared to Conformer. Neither model is fully robust to longer utterances, and the degradation patterns appear similar for both architectures. Future work should evaluate these architectures on authentic long utterances to better assess their robustness as we already observe degradation when simply concatenating short utterances into ranges of at least 30s.

Thus far, we have shown that Mamba matches Conformer across all seven languages while being more computationally efficient, although both models degrade on long-form ASR. We evaluate Mamba in the multilingual setting by comparing training strategies and assessing cross-corpus generalization.

TABLE III: Performance (WER %) of multilingual Mamba models trained under different regimes on individual languages (top), evaluated on the full Swivuriso dev-test set. The bottom section reports generalization to NCHLT. Results include the originally trained monolingual baselines. Scores are averaged over three seeds; standard deviations are shown as subscripts.
\rowcolorwhite \cellcolorOIblue!30Nguni \cellcolorOIblue!80!OIorange!20Tsonga \cellcolorOIgreen!35Sotho \cellcolorOIgreen!80!OIvermillion!20Venda
Setup nbl xho zul tso sot tsn ven Avg
Swivuriso \rightarrow Swivuriso
\rowcolorLightAsh     Monolingual
Conformer 42.552.20 39.160.53 46.541.10 37.963.68 31.321.59 29.263.30 29.691.64 36.64
ConMamba 41.181.15 39.341.36 46.051.09 32.591.01 29.480.11 25.610.07 24.420.14 34.10
\rowcolorOIblue!15     Multilingual (ConMamba)
MIF 36.410.85 35.801.12 41.941.10 30.581.50 28.221.82 23.371.79 23.721.58 31.43
MI 34.760.39 34.050.20 40.330.31 27.470.44 26.410.21 21.580.28 21.690.14 29.47
MLE 34.270.07 33.710.13 39.850.11 26.850.08 26.110.06 21.280.11 21.350.08 29.06
MLFE 34.370.24 33.880.12 40.150.15 27.090.32 26.220.11 21.460.19 21.680.17 29.26
M-CTC+LID 36.280.47 35.620.34 41.770.42 29.520.56 27.450.59 22.730.33 22.730.17 30.87
Swivuriso \rightarrow NCHLT
\rowcolorLightAsh     Monolingual
Conformer 47.753.45 46.751.47 48.201.28 28.002.24 43.065.21 40.206.67 51.252.16 43.60
ConMamba 44.090.80 47.563.07 48.301.50 21.340.64 35.670.61 26.460.56 38.790.73 37.46
\rowcolorOIblue!15     Multilingual (ConMamba)
MIF 35.330.80 42.372.13 42.782.17 21.441.56 35.373.63 23.662.13 37.403.78 34.05
MI 33.220.67 39.161.22 40.091.16 19.770.50 33.441.54 21.071.00 33.681.57 31.49
MLE 30.710.59 36.731.15 36.420.38 17.440.31 30.992.06 19.431.20 29.580.59 28.76
MLFE 30.800.11 36.770.33 36.331.64 17.250.07 31.491.87 20.631.64 30.010.44 29.04
M-CTC+LID 35.011.16 40.850.82 40.783.05 20.721.06 36.012.76 22.591.85 37.020.35 33.28
TABLE IV: Generalization performance (WER %) of multilingual Mamba models trained under different training regimes on FLEURS, compared against the monolingual baseline models.
Setup \cellcolorOIblue!30xho \cellcolorOIblue!30zul Avg
Swivuriso \rightarrow FLEURS
\rowcolorLightAsh     Monolingual
Conformer 55.570.58 46.971.16 51.27
ConMamba 56.262.50 45.641.46 50.95
\rowcolorOIblue!15     Multilingual (ConMamba)
MIF 51.141.99 39.881.67 45.51
MI 48.290.33 37.490.78 42.89
MLE 47.030.10 36.590.26 41.81
MLFE 47.95.430{}_{0}.43 37.16.190{}_{0}.19 42.56
M-CTC+LID 49.980.35 39.350.28 44.66

IV-C Multilingual Mamba ASR Performance

Section IV-B presents results for multilingual models trained on short utterances from the seven Swivuriso languages and evaluated on the Swivuriso dev_test sets. We consider five multilingual training regimes described in Section III-D, including MIF, which serves as a multilingual baseline. In MIF, Mamba models are trained on the Nguni-Tsonga and Sotho-Venda language groupings, respectively. All multilingual models outperform their monolingual counterparts. The smallest improvement is achieved by MIF, with an average WER reduction of 2.68% relative to the monolingual baseline, while the fully multilingual multitask model achieves a reduction of 4%. Explicit language conditioning (MLE and MLFE) yields only marginal improvements over multilingual implicit training (MI) in terms of final WER. MLE and MLFE converge faster, achieving up to 7% lower validation error than MI during the early training epochs. MI reaches comparable validation performance after approximately 20 epochs.

At the language level, comparing MLE with the monolingual model, the Nguni languages benefit more, with nbl improving by 7%\sim 7\%, xho by 5%\sim 5\%, and zul by 6%\sim 6\%, while the Sotho languages show smaller gains of 3%\sim 3\% and 4%\sim 4\%, possibly due to their stronger monolingual performance. The higher gains for Nguni languages may be attributed to the additional training data available from related languages.

IV-D Cross-Corpus Robustness Performance

Comparing multilingual and monolingual models on cross-corpus performance, Sections IV-B and IV-B shows the results on NCHLT and FLEURS, respectively. On both datasets, we observe gaps of 8.7% and 9.14% between the monolingual ConMamba model and the best-performing multilingual model (MLE). The language-conditioned model also outperforms the other multilingual models, with improvements of 2.73% and 1.08% over the MI model on NCHLT and FLEURS, respectively. Overall, multilingual models demonstrate better cross-corpus robustness.

Thus far, we have observed that while language conditioning does not yield large in-domain improvements, it leads to stronger gains in cross-corpus generalization. Hence, we investigate this effect through an ablation study of the multilingual model with language embeddings.

TABLE V: Performance (WER, %) of multilingual Mamba models trained on 50-hour per-language data after language embedding ablations (zeroing and permutation). Evaluated on the Swivuriso dev-test set. Results are averaged over three seeds; subscripts denote standard deviations.
\rowcolorwhite \cellcolorOIblue!30Nguni \cellcolorOIblue!80!OIorange!20Tsonga \cellcolorOIgreen!35Sotho \cellcolorOIgreen!80!OIvermillion!20Venda
Setup nbl xho zul tso sot tsn ven Avg
MLE 34.270.07 33.710.13 39.850.11 26.850.08 26.110.06 21.280.11 21.350.08 29.06
Zeroed 67.4227.72 57.8322.85 57.9713.38 74.3425.08 60.1030.50 54.4435.76 69.2132.97 63.04
Permuted 129.1859.70 63.7037.44 117.9351.87 82.339.07 56.8629.80 46.8930.16 61.1133.16 79.72
TABLE VI: Performance (WER, %) of multilingual Mamba models under 5-hour and 10-hour per-language training. Evaluated on the Swivuriso dev-test set. Results are averaged over three seeds; subscripts denote standard deviations.
\rowcolorwhite \cellcolorOIblue!30Nguni \cellcolorOIblue!80!OIorange!20Tsonga \cellcolorOIgreen!35Sotho \cellcolorOIgreen!80!OIvermillion!20Venda
Setup nbl xho zul tso sot tsn ven Avg
\rowcolorLightAsh     5 hours (Swivuriso \rightarrow Swivuriso)
MI 63.710.31 63.250.30 66.630.20 62.450.59 51.940.19 48.150.38 47.130.57 57.61
MLE 61.080.24 61.480.19 65.460.34 59.800.30 50.180.37 46.320.28 45.800.64 55.73
ΔMIMLE\Delta_{\mathrm{MI-MLE}} 2.63 1.77 1.17 2.65 1.75 1.83 1.32 1.88
\rowcolorLightAsh     10 hours (Swivuriso \rightarrow Swivuriso)
MI 49.670.21 49.950.32 54.580.26 46.670.29 38.810.35 34.210.26 33.670.39 43.94
MLE 48.030.20 48.150.07 53.160.04 44.530.18 37.490.23 33.050.23 32.570.15 42.42
ΔMIMLE\Delta_{\mathrm{MI-MLE}} 1.64 1.80 1.41 2.15 1.32 1.15 1.11 1.51

IV-E Effect of Language Embedding Ablation

To assess whether the learned language embeddings actively contribute to model predictions, we perform an ablation study where the language conditioning embedding is removed at inference time by replacing it with either a zero vector or a permuted version of the original embedding (to preserve magnitude while disrupting its structure). This allows us to isolate the contribution of the language embeddings from the rest of the model, even though performance improvements on in-domain data were not observed compared to training without any language conditioning. We present the results in Table V. The results show that both ablations degrade performance across all languages, with permutation causing a greater degradation. This confirms the that the embeddings learned for the languages are informative.

IV-F Effect of Language Conditioning in Low-Resource Regimes

Given the comparable performance between MI and MLE in multilingual in-domain settings, as well as the observed robustness of MLE on out-of-domain datasets such as FLUERS and NCHLT, we hypothesize that language conditioning may improve performance in low-data scenarios. To verify this, we train MI and MLE models on 5-hour and 10-hour per-language subsets sampled from Swivuriso. Section IV-D shows the results of this synthetic low-resource regime. MLE consistently outperforms MI in most languages in both data settings, yielding at least 1% absolute WER reduction, with the average improvement being slightly higher in the 5-hour setting. This shows that language conditioning is particularly beneficial under low-resource conditions.

Refer to caption
Figure 2: Cosine Similarity of the embeddings from 50-hour per language MLE.

IV-G Analysis of Learned Language Embeddings

Given the importance of the learned language embeddings, a natural question is whether they encode linguistic information shared across languages. To investigate this, we extract the embedding matrix (for the 50-hour MLE model) and compute pairwise cosine similarities to analyze the structure of the resulting space. Figure 2 shows the resulting cosine similarity matrix, where we observe that the embedding space does not strictly reflect linguistic relatedness: linguistically similar languages do not consistently exhibit higher similarity, while some unrelated languages appear closer. This suggests that the embeddings do not encode typological similarity, but instead function as task-specific control vectors that steer the shared encoder representations to optimize the CTC objective.

V Conclusion

In this work, we systematically evaluate the Mamba architecture for ASR in seven African languages. We first compare Mamba with Conformer in monolingual settings and assess robustness to long-context speech. We then explore multilingual training strategies, including pooled data, language embeddings, hierarchical language-family embeddings, and multitask learning with an auxiliary LID objective. Our results show that Mamba achieves competitive performance while being more computationally efficient than Conformer. Multilingual training consistently outperforms monolingual models, while explicit language conditioning does not yield consistent gains in the high-resource setting yet benefits low-resource and cross-corpus settings. We further find that learned language embeddings do not reflect typological similarity but instead act as task-specific control vectors. To the best of our knowledge, no publicly available multilingual pretrained Mamba speech encoders support the studied languages, limiting direct comparison to prior work. Future work should explore larger-scale pretraining and more expressive conditioning strategies for multilingual and low-resource speech models.

VI Acknowledgments

JOA and BMA are funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – Project-ID 232722074 – SFB 1102.

VII Generative AI Use Disclosure

Generative AI tools were used exclusively for editing and polishing assistance, coding support for training and evaluation modules, code debugging, and manuscript proofreading including grammar and typographical corrections. No part of the scientific content, experimental design, analysis, citations, or conclusions was generated by AI tools.

References

  • [1] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning, Baltimore, USA, 2023.
  • [2] K. Kim, F. Wu, Y. Peng, J. Pan, P. Sridhar, K. J. Han, and S. Watanabe, “E-branchformer: Branchformer with enhanced merging for speech recognition,” in IEEE Spoken Language Technology Workshop (SLT), 2022.
  • [3] R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schlüter, and S. Watanabe, “End-to-end speech recognition: A survey,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 325–351, 2023.
  • [4] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Interspeech, Virtual, 2020.
  • [5] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in First Conference on Language Modeling, Philadelphia, USA, 2024.
  • [6] R. Zevallos, M. Cortada Garcia, S. Solito, C. Mena, A. Peiró-Lilja, and J. Hernando, “Assessing the Performance and Efficiency of Mamba ASR in Low-Resource Scenarios,” in Interspeech, Rotterdam, The Netherlands, 2025.
  • [7] T. Moriya, M. Mimura, K. Matsui, H. Sato, and K. Matsuura, “Attention-Free Dual-Mode ASR with Latency-Controlled Selective State Spaces,” in Interspeech, Rotterdam, The Netherlands, 2025.
  • [8] M. N. Ali, D. Falavigna, and A. Brutti, “Mlma: Towards multilingual asr with mamba-based architectures,” ArXiv, vol. abs/2510.18684, 2025.
  • [9] O. Ahia, J. Kreutzer, and S. Hooker, “The low-resource double bind: An empirical study of pruning for low-resource machine translation,” in Findings of the Association for Computational Linguistics: Empirical Methods in Natural Language Processing (EMNLP), Punta Cana, Dominican Republic, 2021.
  • [10] V. Marivate, K. Olaleye, S. Mundia, A. Bakainga, U. Netshifhefhe, M. Milanzie, T. H. Mogale, T. Sindane, Z. Abdulrasaq, K. Mokgosi, C. Okorie, N. Z. V. Wyk, G. Morrissey, D. Dunbar, F. Smit, T. Chidi, R. Mabuya, A. Bukula, R. Mlambo, T. Macucwa, I. Abdulmumin, , and S. Rananga, “Swivuriso: The south african next voices multilingual speech dataset,” arXiv preprint arXiv: 2512.02201, 2025.
  • [11] A. D. Diack, P. H. Nelson, K. Agbesi, A. Nakalembe, M. Mohamedkhair, V. Dube, T. Siyavora, S. Venugopalan, J. Hickey, U. Okonkwo, A. Bapna, I. Wiafe, R. D. Helegah, E. D. Atsakpo, C. Nutrokpor, F. B. P. Winful, K. K. Solaga, J.-D. Abdulai, A. O. Ekpezu, A. Niyonkuru, S. Rutunda, B. Ishimwe, M. Melese, E. Bainomugisha, J. Nakatumba‐Nabende, A. Katumba, C. Babirye, J. Mukiibi, V. Kimani, S. Kibacia, J. Maina, F. Emmah, A. I. Shekarau, I. Adamu, Y. S. Abdullahi, H. Lakougna, B. Macdonald, H. Shemtov, A. Walcott-Bryant, M. Cissé, A. Hassidim, J. Dean, and Y. Matias, “Waxal: A large-scale multilingual african language speech corpus,” 2026.
  • [12] R. Flynn and A. Ragni, “Beyond the utterance: An empirical study of very long context speech recognition,” IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 910–920, 2026.
  • [13] J. O. Alabi, M. A. Hedderich, D. I. Adelani, and D. Klakow, “Charting the landscape of African NLP: Mapping progress and shaping the road ahead,” in Conference on Empirical Methods in Natural Language Processing (EMNLP), Suzhou, China, 2025.
  • [14] S. H. Imam, B. Sani, D. K. Gete, B. Y. Ahmed, I. S. Ahmad, I. Abdulmumin, S. M. Yimam, M. Y. Bello, and S. H. Muhammad, “Automatic speech recognition for African low-resource languages: Challenges and future directions,” in 6th Workshop on African Natural Language Processing (AfricaNLP), Vienna, Austria, 2025.
  • [15] E. van der Westhuizen and T. Niesler, “A first South African corpus of multilingual code-switched soap opera speech,” in 11th International Conference on Language Resources and Evaluation (LREC), Miyazaki, Japan, 2018.
  • [16] E. Barnard, M. H. Davel, C. van Heerden, F. de Wet, and J. Badenhorst, “The NCHLT speech corpus of the South African languages,” in 4th Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU), St. Petersburg, Russia, 2014.
  • [17] J. Badenhorst and F. de Wet, “Nchlt auxiliary speech data for asr technology development in south africa,” Data in Brief, vol. 41, p. 107860, 2022.
  • [18] T. Gumede and M. Plauché, “Initial fieldwork for LWAZI: A telephone-based spoken dialog system for rural South Africa,” in 1st Workshop on Language Technologies for African Languages, Athens, Greece, 2009.
  • [19] R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in 12th Language Resources and Evaluation Conference, Marseille, France, 2020.
  • [20] A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in IEEE Spoken Language Technology Workshop (SLT), Doha, Qatar, 2022.
  • [21] J. Rajab, A. Aremu, E. A. Chimoto, D. Dunbar, G. Morrissey, F. Thior, L. Potgieter, J. Ojo, A. L. Tonja, W. N. Nekoto, P. Moiloa, J. Abbott, V. Marivate, and B. Rosman, “The esethu framework: Reimagining sustainable dataset governance and curation for low-resource languages,” in 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 2025.
  • [22] E. Yilmaz, A. Biswas, E. van der Westhuizen, F. de Wet, and T. Niesler, “Building a Unified Code-Switching ASR System for South African Languages,” in Interspeech, Hyderabad, India, 2018, pp. 1923–1927.
  • [23] A. Biswas, E. Yilmaz, E. van der Westhuizen, F. de Wet, and T. Niesler, “Code-switched automatic speech recognition in five south african languages,” Computer Speech & Language, vol. 71, p. 101262, 2022.
  • [24] A. Biswas, E. Yılmaz, F. de Wet, E. van der Westhuizen, and T. Niesler, “Semi-Supervised Acoustic Model Training for Five-Lingual Code-Switched ASR,” in Interspeech, Graz, Austria, 2019.
  • [25] A. Biswas, E. Yilmaz, F. de Wet, E. van der Westhuizen, and T. Niesler, “Semi-supervised development of ASR systems for multilingual code-switched speech in under-resourced languages,” in 12th Language Resources and Evaluation Conference (LREC), Marseille, France, 2020.
  • [26] M. Zanon Boito, V. Iyer, N. Lagos, L. Besacier, and I. Calapodescu, “mhubert-147: A compact multilingual hubert model,” in Interspeech, Kos, Greece, 2024.
  • [27] J. O. Alabi, X. Liu, D. Klakow, and J. Yamagishi, “AfriHuBERT: A self-supervised speech representation model for African languages,” in Interspeech, Rotterdam, The Netherlands, 2025.
  • [28] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in 23rd International Conference on Machine Learning (ICML), New York, USA, 2006.
  • [29] Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” in 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 2019.
  • [30] X. Jiang, Y. A. Li, A. N. Florea, C. Han, and N. Mesgarani, “Speech slytherin: Examining the performance and efficiency of mamba for speech separation, recognition, and synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025.
  • [31] J. Tian, J. Yu, C. Zhang, C. Weng, Y. Zou, and D. Yu, “Lae: Language-aware encoder for monolingual and multilingual asr,” in Interspeech, Incheon, Korea, 2022.
  • [32] M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y. Gao, R. D. Mori, and Y. Bengio, “SpeechBrain: A general-purpose speech toolkit,” 2021, arXiv:2106.04624.
  • [33] K. Miyazaki, Y. Masuyama, and M. Murata, “Exploring the Capability of Mamba in Speech Applications,” in Interspeech, Kos, Greece, 2024.