EMNLP 2026 ยท Main Conference ยท Oral Presentation
Code, the 1,200-utterance synthetic code-switching benchmark, and reproduction instructions are available at github.com/saga1214/PhraseLocalizedLCG.
Hear the difference โ language-learning content (ENโKO)
| Vanilla | Listen to the Korean phrase with the English carrier's language conditioning. | |
| Ours โ LCG | Compare the same phrase with localized Korean-language guidance. |
Global guidance flattens the embedded phrase into the carrier accent (left); LCG steers only the phrase region toward its own language (right).
Everyday situations where a carrier sentence embeds a foreign phrase. Vanilla vs. our phrase-localized LCG (ฮป=7), paper-final configuration.
The last tab embeds French and German phrases in English carriers โ the training-free control covers all five languages of the paper.
Language-learning content ENโKO
| Vanilla | |
| Ours ฮป7 |
Tour guide ENโKO
| Vanilla | |
| Ours ฮป7 |
Museum exhibition ENโJA
| Vanilla | |
| Ours ฮป7 |
Samples drawn directly from our released 1,200-utterance synthetic code-switching benchmark, comparing all three systems from the paper (unguided baseline, coupled Swap, and phrase-localized LCG).
Korean carrier sentences with embedded English phrases, drawn from the LA=1 subset: all systems synthesize the correct words, so the difference you hear is residual accent quality on the isolated segment.
Example 1 โ โpractical workshops aligned with our buying committeeโ KOโEN
| Vanilla | |
| Swap ฮป3 | |
| Ours ฮป7 |
Example 2 โ โlean household operations with zero emotional overheadโ KOโEN
| Vanilla | |
| Swap ฮป3 | |
| Ours ฮป7 |
Complete utterances with three embedded phrases each, comparing global naturalness and phrase nativeness.
English carrier ยท Japanese phrases ENโJA
| Vanilla | |
| Swap ฮป3 | |
| Ours ฮป7 |
English carrier ยท Korean phrases ENโKO
| Vanilla | |
| Swap ฮป3 | |
| Ours ฮป7 |
All audio above is synthesized from our 1,200-utterance balanced synthetic code-switching benchmark, generated with a large language model and spanning five languages across twelve directional configurations. The corpus is released with the code and documented in benchmark/.
Each utterance contains 3โ5 dense, technical or literary embedded phrases (~6 words, ~35 characters per phrase). The corpus is balanced by direction (100 utterances per direction) to enable per-direction evaluation.
@misc{lee2026phraselocalizedlanguagecontrastiveguidancetrainingfree,
title = {Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech},
author = {Che Hyun Lee and Sangkwon Park and Donghun Kang and Dongwook Lee and Youngho Cho and Heeseung Kim and Sungroh Yoon},
year = {2026},
eprint = {2609.01016},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.01016},
}