Open speech models for agents: speech recognition, live transcription and text-to-speech.
Small enough for one server, measured against the models they replace.
Models · Families · Install · Evaluate · Serving engine
| 5.69% WER FLEURS Romanian · Jackrabbit 110M Canary 1B v2: 5.95% · Whisper large-v3: 8.42% |
2.07% WER Common Voice 21 Romanian SpeD 110M: 3.47% · Parakeet 0.6B v3: 10.06% |
2,531× real time Jackrabbit 110M · 1× RTX 5090 Whisper large-v3: 102× |
CPU, Mac or GPU Amami 357M text-to-speech native runtimes, no PyTorch |
FLEURS numbers from the Open ASR Leaderboard's own runner and normalizer, every model on the same GPU. Full tables and raw outputs in surogate-speech-evals.
Surogate Speech covers the speech side of Surogate: model families for everything an agent needs to hear and say, served natively by the surogate engine. Each family does one job and ships language by language, and Romanian goes first.
| Family | Job | Released | |
|---|---|---|---|
| Jackrabbit | speech recognition, offline and streaming | Romanian: jackrabbit-110m-ro, jackrabbit-110m-ro-streaming | 116M |
| Amami | text-to-speech | Romanian: amami-357m-ro, voices Doina, Tudor, Radu | 357M |
Every family is named after a rabbit chosen for its job: the jackrabbit has the biggest ears of any hare, and the Amami rabbit is one of the few rabbits with a voice.
This repository holds the local clients, the evaluation harness behind every number on the model cards, and the frozen test sets.
For serving, use the surogate engine, which runs all three natively behind OpenAI-compatible audio routes:
surogate serve --stt surogate/jackrabbit-110m-ro
surogate serve --stt surogate/jackrabbit-110m-ro-streaming
surogate serve --tts surogate/amami-357m-ropip install "surogate-speech[asr] @ git+https://github.com/invergent-ai/surogate-speech" # speech recognition
pip install "surogate-speech[tts] @ git+https://github.com/invergent-ai/surogate-speech" # native TTS (Linux x86-64, no Torch)
pip install "surogate-speech[tts-nemo] @ git+https://github.com/invergent-ai/surogate-speech" # TTS through NeMo (CPU or GPU)
pip install "surogate-speech[eval] @ git+https://github.com/invergent-ai/surogate-speech" # evaluation harnessPython 3.10 or newer. Models download from Hugging Face on first use, at the revisions pinned in
src/surogate_speech/hub.py.
surogate-speech transcribe interviu.wav # TDT greedy
surogate-speech transcribe interviu.wav --lm # CTC + 4-gram, the best accuracy
surogate-speech listen # microphone, live partials and finals
surogate-speech listen --file interviu.wav # the same streaming path on a file
surogate-speech speak "Bună ziua! Cu ce vă pot ajuta?" --voice Doina -o doina.wav
surogate-speech speak "Bună ziua!" --voice Radu --backend nemo --device cuda -o radu.wavTalking to a running engine instead, from Python:
from surogate_speech import client
print(client.transcribe("interviu.wav"))
client.speech("Bună ziua!", voice="Tudor", out="tudor.wav")Every score runs reference and hypothesis through the language's normalizer (for Romanian,
surogate_speech.text.normalize, version surogate-ro-v1), and every result file records it,
along with the model's and the clip set's sha256.
# Speech recognition on the public Romanian test splits (FLEURS, VoxPopuli; Common Voice with --cv-repo)
surogate-speech eval asr --model-name jackrabbit-110m-ro --suite fleurs_ro,voxpopuli_ro --output asr.json
# Live-packet streaming: WER of the finals and finalization time; shard across GPUs
surogate-speech eval streaming --num-shards 8 --shard-index 0 --device cuda:0 --output shard0.json
surogate-speech eval streaming-merge --inputs shard*.json --manifest surogate-speech-data/manifest.tsv --output streaming.json
# TTS intelligibility: synthesize the 331-sentence stress set with every voice, transcribe with a judge
surogate-speech eval tts --judge nvidia/canary-1b-v2 --out tts-run/WER is corpus-level, with a 95% interval from resampling reference sentences (clips of the same
sentence move together). stress-v2 is in src/surogate_speech/sets/ with its sha256.
The code is Apache-2.0. The model weights are CC-BY-NC-4.0: free for research and other non-commercial use; for commercial use, contact Invergent. See MODEL_LICENSE.md for the base models and third-party components.