From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models
Abstract
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Index Terms:
speech emotion recognition, concept bottleneck models, large language models, interpretabilitySheffield, United Kingdom
1 Introduction
Speech emotion recognition (SER) predicts the emotion label of an utterance from the speech signal [10]. Early systems trained classifiers on a wide range of carefully selected acoustic features such as pitch, energy and spectral descriptors [18, 3]. Later, neural networks learned acoustic representations directly from the raw waveform [23], and self-supervised speech representations turned out to be better as classifier input [14]. However, besides acoustic information, speech also carries lexical content [10, 31], and subsequent systems therefore combined the speech representations with text features from the transcript [22]. More recently, SER research has adopted large audio-language models (LALMs), models that extend large language models (LLMs) with an audio encoder and predict the emotion label directly from the audio given a text instruction [25, 13]. Despite these improvements recognition performance remains low on many corpora [12, 32]. Moreover, with the direct audio input, the LALM gives no indication of which information it relies on, so its decisions cannot be examined.
Several works make use of acoustic information in textual form. EmotionThinker [26] makes its reasoning visible by keeping the audio as input and generating a reasoning text that describes prosody before giving the label. SpeechCueLLM [28] and VowelPrompt [27] make the input tangible by extracting acoustic descriptions from the audio and providing them, together with the transcript, to a text LLM. However, retraining without individual descriptions or swapping descriptions between utterances [27] does not identify which description a fixed predictor’s decision depends on; generated reasoning may not faithfully reflect how the model reached its prediction [24].
Identifying which information affects a prediction requires the ability to control the model input and knowing what such input represents. A model that uses human-readable form as input is therefore desirable. Image classification faced the same problem, and concept bottleneck models (CBMs) were introduced in response [8]. Here an image is first mapped to human-readable concepts, and the label is predicted from these alone. Hence every prediction can be traced to specific concepts and can be changed by editing them. Shin et al. [21] used such edits to examine how predictions respond to individual concepts.
We apply the CBM idea to SER: separate extractors describe each utterance by concepts that are human readable (Section 2). Then an LLM predicts the emotion label from these concepts, enabling single-concept removal with the predictor fixed to examine which class predictions change. Here concepts are the transcript, pitch, intensity and speech rate, as well as speaker age and gender. Together these describe the words spoken, intonation and speaker characteristics.
This work studies how the concepts affect emotion decisions: (1) what does each concept group contribute to recognition across corpora, before and after fine-tuning? (2) which acoustic concepts do the decisions depend on, and how does this dependence vary across corpora? The first question is addressed by comparing predictors that receive different concept groups (Section 4.1). The second by removing single acoustic concepts from the input of a predictor fine-tuned on all concepts (Section 4.2), with both analyses run on CREMA-D [2], IEMOCAP [1] and MELD [15] with three LLMs (Section 3).
2 Speech Concept Bottleneck
A concept bottleneck model operates in two cascaded stages [8]: A concept extractor maps an input to a set of concepts , and a predictor predicts the label from only,
| (1) |
Concepts are human-specified, interpretable properties of the input, whose predicted values form the input to . A concept intervention changes one concept while is held fixed and observes the change in [8, 21]. This work adopts this framework for SER: the input is an utterance, the concepts are properties of speech expressed as text, and the predictor is an LLM that processes this text to produce one emotion label from a fixed set.
The concepts used are organised into three groups, each related to how emotion is expressed in speech (see Table 1). The transcript represents content and gives the words spoken, which carry emotion information through their meaning. The acoustic concepts are pitch, intensity and speech rate, three established acoustic correlates of emotion [17, 3, 28]. In addition, the expression and perception of emotion in speech vary with the speaker’s age and gender [19, 9]. To account for these differences, the speaker concepts include the age and gender of each speaker. Each group is obtained from a separate off-the-shelf extractor that was not trained on emotion labels (see Sec. 3.1).
| Predict the emotion expressed in the utterance from the provided information. |
| Transcript: transcript |
| Pitch level: 5 levels from very low to very high. |
| Pitch variation: 5 levels from very low to very high. |
| Volume level: 5 levels from very low to very high. |
| Volume variation: 5 levels from very low to very high. |
| Speech rate: 5 levels from very slow to very fast. |
| Estimated age: 6 bands from under 20 to 60+. |
| Predicted sex: Female or Male. |
| Allowed labels: label set |
| Return exactly one label from the allowed labels. Do not provide an explanation. |
| Zero-shot | Fine-tuned | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Dataset | Model | T | A | TA | TAP | T | A | TA | TAP |
| CREMA-D | Qwen2.5 | 4.26 | 27.88 | 5.87 | 7.13 | 11.74 | 43.03 | 44.21 | 45.50 |
| Qwen2.5-Omni | 4.26 | 24.24 | 4.79 | 4.79 | 11.28 | 42.18 | 44.51 | 45.57 | |
| Llama 3.1 | 4.26 | 25.01 | 16.80 | 13.80 | 11.15 | 41.87 | 45.10 | 45.37 | |
| IEMOCAP | Qwen2.5 | 46.11 | 36.49 | 51.71 | 52.15 | 69.90 | 45.49 | 74.37 | 74.16 |
| Qwen2.5-Omni | 48.43 | 31.41 | 51.99 | 51.83 | 70.26 | 45.17 | 74.97 | 74.58 | |
| Llama 3.1 | 52.32 | 32.35 | 54.49 | 53.14 | 70.49 | 44.38 | 74.24 | 74.05 | |
| MELD | Qwen2.5 | 36.16 | 13.77 | 33.83 | 33.13 | 37.20 | 13.51 | 37.03 | 39.02 |
| Qwen2.5-Omni | 33.40 | 11.80 | 32.70 | 32.76 | 37.19 | 13.71 | 38.48 | 37.98 | |
| Llama 3.1 | 32.83 | 11.80 | 30.37 | 30.68 | 37.27 | 13.97 | 38.50 | 37.36 | |
T: transcript; A: acoustic concepts; P: speaker concepts, all given as text. Bold: highest mean per model and setting; ties share the mark. Fine-tuned: mean of three runs (seed SD 0.1–3.6 points); IEMOCAP: five-fold means. Direct-audio reference: Table 3.
3 Implementation and Experiments
3.1 Concept Extraction and Prompting
This study aims to provide insight into how the predictor makes use of the concepts, not how they are extracted. In order to extract concepts of high quality, each is obtained independently state of the art tools. The transcript is produced by Qwen3-ASR [20]. Pitch and intensity are measured with Praat [7], with the mean over the utterance giving the level and the standard deviation giving the variation. Speech rate is defined as the number of words in the transcript divided by the utterance duration [30]. Each acoustic measure is discretised into five levels with quantile thresholds fitted on the training set of each dataset and fold. The prompt is constructed with a simple description of the level (Table 1). The speaker concepts are derived from Vox-Profile [4], with the predicted age grouped into six ten-year bands from under 20 to over 60. The predicted gender agrees with the metadata for 96.3% of CREMA-D and 95.2% of IEMOCAP utterances. Age metadata exist only for CREMA-D, where the predicted band is correct for only 35.9% of utterances and within one band for 78.2%. The prompt contains the selected concept groups and allowed labels, without specifying how concept values relate to emotions.
3.2 Experimental Design
To measure what each concept group contributes, the four combinations in Table 2 are evaluated with the zero-shot LLM and with a predictor fine-tuned on each combination alone. The gain from adding a group to utterance descriptions is considered its contribution, in zero-shot setting and after fine-tuning. To identify the relation between acoustic concepts and decisions , the predictor fine-tuned on all concepts is held fixed. The same utterance is assessed with and without a specific acoustic concept , (as defined in Sec. 2 ), thus providing some indication of its relevance.
3.3 Experimental Setup
Datasets. Three corpora are used in which emotion is carried by the words and by the voice to different degrees. In CREMA-D [2], twelve fixed neutral sentences are acted in six emotions, so only the voice carries emotion. IEMOCAP [1] holds scripted and improvised dyadic sessions performed by actors, and MELD [15] multi-party television dialogue, and in both the words carry emotion as well. CREMA-D has six classes and speaker-disjoint splits, IEMOCAP four classes and session-wise five-fold cross-validation, and MELD seven classes and official splits.
Models. Three open instruction-tuned LLMs serve as the predictor: Qwen2.5-7B-Instruct [16], Qwen2.5-Omni-7B [29], which is built on Qwen2.5-7B, and Llama-3.1-8B-Instruct [5]. Qwen2.5-Omni is also given the audio directly as a reference condition, zero-shot and fine-tuned.
Training. Fine-tuning uses LoRA [6] with rank 16, and dropout 0.05 on the attention and feed-forward projections. Optimisation uses AdamW [11] with a learning rate of , 10% warm-up and linear decay, for six epochs with a batch size of 16 in BF16. Each model is fine-tuned on each combination with three seeds and on each IEMOCAP fold.
Evaluation. Performance is measured as Macro-F1 under greedy decoding, and an output that cannot be parsed as a label counts as an error. The effect of a concept removal is the change in Macro-F1 relative to the full-input score of the same run.
4 Results and Discussion
4.1 Predictive Performance
Table 2 reports Macro-F1 for each combination of concept groups, zero-shot and after fine-tuning. On CREMA-D, zero-shot models are strongly biased towards the transcript, which is detrimental. As the same sentences are spoken in every emotion, the models given the transcript alone classify every utterance as Neutral, giving 4.26 Macro-F1 for all of them. Although the acoustic concepts alone reach 27.88 for Qwen2.5, adding the transcript brings the score back to 5.87, with most predictions again Neutral, and the other two models drop in the same way.
| Zero-shot | Fine-tuned | |||
|---|---|---|---|---|
| Dataset | Concepts | Audio | Concepts | Audio |
| CREMA-D | 24.24 | 54.95 | 45.57 | 78.67 |
| IEMOCAP | 51.99 | 69.32 | 74.97 | 82.66 |
| MELD | 33.40 | 34.91 | 38.48 | 41.07 |
Audio fine-tuning also adapts the encoder and projector. Fine-tuned scores average three runs.
However, the fine-tuned models show the opposite pattern. After fine-tuning, the models decide from the acoustic concepts, and Llama 3.1 reaches 41.87 with these alone. The transcript alone still gives only 11.15. Yet adding it to the acoustic concepts now raises the score to 45.10 rather than lowering it. The gain comes from combining the two inputs, and the reversal holds for all three models.
On IEMOCAP and MELD, the same bias towards the transcript does no harm. These are conversational corpora, so the words change with the emotion, and the transcript alone already recognises it, giving 46.11 for Qwen2.5 on IEMOCAP. With the acoustic concepts alone, zero-shot Qwen2.5 scores 36.49 on IEMOCAP, and adding the transcript raises this to 51.71, where the same addition lowered the score on CREMA-D. The acoustic concepts still improve recognition on IEMOCAP, where fine-tuning with both inputs reaches 74.37 against 69.90 with the transcript, but gains on MELD are small and inconsistent across models: the same comparison gives 37.03 against 37.20 for Qwen2.5.
Adding the speaker concepts changes the fine-tuned score by less than two points on every corpus, without a consistent direction. These small gains may reflect both limited within-speaker variation and errors in the extracted speaker concepts. Even at its best, the concept input stays below direct audio, and the gap depends on the corpus. Fine-tuned Qwen2.5-Omni reaches 45.57 on CREMA-D from concepts but 78.67 from audio, while the gap shrinks to 7.69 on IEMOCAP and 2.59 on MELD (Table 3). The gap is largest where the emotion lies in the voice and smallest where it lies in the words.
4.2 Concept Interventions
Fig. 1 shows the change in Macro-F1 when one acoustic concept is removed from the input of the predictor fine-tuned on all concepts. How much a model depends on single acoustic concepts, and on which, follows the corpus. On CREMA-D, where the emotion lies in the acoustic concepts, removing any of the five lowers Qwen2.5’s Macro-F1, by 2.49 to 6.30. The largest loss comes from speech rate. In this corpus, speech rate separates Neutral from Disgust: Neutral utterances tend to be fast, with 37.0% in the fastest level and 6.5% in the slowest, whereas Disgust utterances tend to be slow, with 11.9% in the fastest level and 29.4% in the slowest. IEMOCAP, where the transcript carries more of the emotion, loses less than 0.6 for three of the five concepts. The exceptions are intensity level and variation, at 2.12 and 1.93, and intensity is again the concept that separates a class pair: 35.0% of Sad utterances have the lowest intensity level against 13.8% of Neutral. On MELD, no removal changes the score by more than 0.3. There, the emotions are spoken at nearly the same average speech rate, intensity and pitch, so the acoustic concepts separate them little: even the two classes furthest apart on speech rate, Happy and Disgust, differ by 0.7 of a level. The pattern holds for all three models, although for Qwen2.5-Omni pitch level and speech rate cost about the same on CREMA-D.
Fig. 2 removes the concept with the largest loss on CREMA-D and IEMOCAP, speech rate and intensity level, and intensity level on MELD for comparison, and follows the predictions that change. For each corpus it shows the label that changes most often, as a share of its predictions, and the label it turns into, split by the utterance’s level of the removed concept. These changes come almost entirely from the two highest or the two lowest levels of the removed concept, and few from the middle levels. Without speech rate, 48% of the Neutral predictions on CREMA-D become Disgust, all from the two fastest levels. Removing speech rate changes these Neutral predictions to Disgust, with most transitions occurring at high speech-rate levels. Without intensity level, 6% of the Sad predictions on IEMOCAP become Neutral, all from the two lowest levels. Angry and Happy predictions move to Neutral in the same way, from the loudest levels. On MELD, 7% of the Happy predictions become Neutral, mostly from the two loudest levels. Predictions also move from Neutral to Happy, despite little change in Macro-F1. The three models change the same levels and differ only in the size of the CREMA-D top level, from 32 to 40. These results reveal corpus-dependent prediction sensitivity to concept removal.
5 Conclusion
The concept bottleneck framework was adopted for SER: each utterance is represented by transcript, acoustic and speaker concepts, and an LLM predicts the emotion from these concepts alone. Experiments were conducted in zero-shot and fine-tuned settings. In the zero-shot setting, the transcript concept dominates outcomes, which is found to be detrimental on scripted corpora, where the transcript is neutral. After fine-tuning, the contributions of transcripts and acoustic concepts vary across corpora. Concept-based prediction still underperforms direct audio input, with the largest gap on CREMA-D, where the fixed transcripts carry no emotion information. Removing individual acoustic concepts from a fixed predictor reveals corpus-dependent changes in recognition performance and class predictions. The selected class transitions concentrate at particular concept levels, while on MELD predictions change despite little change in Macro-F1. These results show that aggregate scores can conceal changes in individual predictions. The concept bottleneck enables examination of dependence through controlled changes to the predictor’s inputs.
6 Compliance with Ethical Standards
This study uses existing data from the CREMA-D, IEMOCAP, and MELD datasets, with no new participant recruitment or data collection. Ethical approval was not required for this secondary analysis.
7 Acknowledgments
The authors declare no conflicts of interest. The authors used Claude Code and OpenAI Codex to assist with language editing and clarity improvements throughout the manuscript, as well as the development of experimental code. The authors reviewed and verified the resulting revisions and code and take full responsibility for the final content.
References
- [1] (2008) IEMOCAP: interactive emotional dyadic motion capture database. Lang. Resour. Eval. 42 (4), pp. 335–359. External Links: ISSN 1574-020X, 1574-0218, Document, Link Cited by: §1, §3.3.
- [2] (2014) CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset. IEEE Trans. Affect. Comput. 5 (4), pp. 377–390. External Links: ISSN 1949-3045, Document, Link Cited by: §1, §3.3.
- [3] (2016) The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing. IEEE Trans. Affect. Comput. 7 (2), pp. 190–202. External Links: ISSN 1949-3045, Document, Link Cited by: §1, §2.
- [4] (2025) Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits. arXiv. External Links: 2505.14648, Document, Link Cited by: §3.1.
- [5] (2024) The Llama 3 Herd of Models. arXiv. External Links: 2407.21783, Document, Link Cited by: §3.3.
- [6] (2022) LoRA: Low-Rank Adaptation of Large Language Models. In Proc. ICLR, External Links: Link Cited by: §3.3.
- [7] (2018) Introducing Parselmouth: A Python interface to Praat. J. Phon. 71, pp. 1–15. External Links: ISSN 0095-4470, Document, Link Cited by: §3.1.
- [8] (2020) Concept Bottleneck Models. In Proc. ICML, PMLR, Vol. 119, pp. 5338–5348. External Links: Link Cited by: §1, §2, §2.
- [9] (2018) Gender Differences in the Recognition of Vocal Emotions. Front. Psychol. 9. Note: Art. no. 882 External Links: ISSN 1664-1078, Document, Link Cited by: §2.
- [10] (2005) Toward detecting emotions in spoken dialogs. IEEE Transactions on Speech and Audio Processing 13 (2), pp. 293–303. External Links: ISSN 1063-6676, Link, Document Cited by: §1.
- [11] (2019) Decoupled Weight Decay Regularization. In Proc. ICLR, External Links: Link Cited by: §3.3.
- [12] (2024) EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark. In Interspeech 2024, pp. 1580–1584 (en). External Links: Link, Document Cited by: §1.
- [13] (2025) AA-SLLM: An Acoustically Augmented Speech Large Language Model for Speech Emotion Recognition. In Proc. Interspeech, pp. 4328–4332. External Links: Document, Link Cited by: §1.
- [14] (2021) Emotion Recognition from Speech Using wav2vec 2.0 Embeddings. In Interspeech 2021, pp. 3400–3404 (en). External Links: Link, Document Cited by: §1.
- [15] (2019) MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. In Proc. ACL, pp. 527–536. External Links: Document, Link Cited by: §1, §3.3.
- [16] (2024) Qwen2.5 Technical Report. arXiv. External Links: 2412.15115, Document, Link Cited by: §3.3.
- [17] (2003) Vocal communication of emotion: A review of research paradigms. Speech Communication 40 (1-2), pp. 227–256. External Links: ISSN 0167-6393, Document Cited by: §2.
- [18] (2009) The INTERSPEECH 2009 emotion challenge. In Interspeech 2009, pp. 312–315 (en). External Links: Link, Document Cited by: §1.
- [19] (2018) Age differences in vocal emotion perception: on the role of speaker age and listener sex. Cogn. Emot. 32 (6), pp. 1189–1204. External Links: ISSN 1464-0600, Document, Link Cited by: §2.
- [20] (2026) Qwen3-ASR Technical Report. arXiv. External Links: 2601.21337, Document, Link Cited by: §3.1.
- [21] (2023) A Closer Look at the Intervention Procedure of Concept Bottleneck Models. In Proc. ICML, PMLR, Vol. 202, pp. 31504–31520. External Links: ISSN 2640-3498, Link Cited by: §1, §2.
- [22] (2023) Using auxiliary tasks in multimodal fusion of wav2vec 2.0 and BERT for multimodal emotion recognition. In Proc. ICASSP, pp. 1–5. External Links: Link, Document Cited by: §1.
- [23] (2016) Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network. In Proc. ICASSP, pp. 5200–5204 (en). External Links: ISBN 978-1-4799-9988-0, Link, Document Cited by: §1.
- [24] (2023) Language models don't always say what they think: unfaithful explanations in chain-of-thought prompting. In Proc. NeurIPS, Vol. 36, pp. 74952–74965. External Links: Document, Link Cited by: §1.
- [25] (2024) BLSP-Emo: towards empathetic large speech-language models. In Proc. EMNLP, pp. 19186–19199. External Links: Link, Document Cited by: §1.
- [26] (2026) EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning. In Proc. ICLR, pp. 153708–153733. External Links: Link Cited by: §1.
- [27] (2026) VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation. In Proc. ICLR, pp. 20439–20460. External Links: Link Cited by: §1.
- [28] (2025) Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances. In Findings of ACL: NAACL, pp. 2202–2218. External Links: Document, Link Cited by: §1, §2.
- [29] (2025) Qwen2.5-Omni Technical Report. arXiv. External Links: 2503.20215, Document, Link Cited by: §3.3.
- [30] (2004) An acoustic study of emotions expressed in speech. In Proc. Interspeech, pp. 2193–2196. External Links: Document Cited by: §3.1.
- [31] (2018) Multimodal speech emotion recognition using audio and text. In Proc. SLT, pp. 112–118. External Links: Link, Document Cited by: §1.
- [32] (2026) VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs. arXiv (en). External Links: Link, Document, 2603.08936 Cited by: §1.