Skip to content

Repository files navigation

Surogate Speech

Surogate Speech

Open speech models for agents: speech recognition, live transcription and text-to-speech.
Small enough for one server, measured against the models they replace.

Models · Families · Install · Evaluate · Serving engine

Code: Apache 2.0 Weights: CC BY-NC 4.0 Follow on X

5.69% WER
FLEURS Romanian · Jackrabbit 110M
Canary 1B v2: 5.95% · Whisper large-v3: 8.42%
2.07% WER
Common Voice 21 Romanian
SpeD 110M: 3.47% · Parakeet 0.6B v3: 10.06%
2,531× real time
Jackrabbit 110M · 1× RTX 5090
Whisper large-v3: 102×
CPU, Mac or GPU
Amami 357M text-to-speech
native runtimes, no PyTorch

FLEURS numbers from the Open ASR Leaderboard's own runner and normalizer, every model on the same GPU. Full tables and raw outputs in surogate-speech-evals.

Surogate Speech covers the speech side of Surogate: model families for everything an agent needs to hear and say, served natively by the surogate engine. Each family does one job and ships language by language, and Romanian goes first.

Families

Family Job Released
Jackrabbit speech recognition, offline and streaming Romanian: jackrabbit-110m-ro, jackrabbit-110m-ro-streaming 116M
Amami text-to-speech Romanian: amami-357m-ro, voices Doina, Tudor, Radu 357M

Every family is named after a rabbit chosen for its job: the jackrabbit has the biggest ears of any hare, and the Amami rabbit is one of the few rabbits with a voice.

This repository holds the local clients, the evaluation harness behind every number on the model cards, and the frozen test sets.

For serving, use the surogate engine, which runs all three natively behind OpenAI-compatible audio routes:

surogate serve --stt surogate/jackrabbit-110m-ro
surogate serve --stt surogate/jackrabbit-110m-ro-streaming
surogate serve --tts surogate/amami-357m-ro

Install

pip install "surogate-speech[asr] @ git+https://github.com/invergent-ai/surogate-speech"    # speech recognition
pip install "surogate-speech[tts] @ git+https://github.com/invergent-ai/surogate-speech"    # native TTS (Linux x86-64, no Torch)
pip install "surogate-speech[tts-nemo] @ git+https://github.com/invergent-ai/surogate-speech"  # TTS through NeMo (CPU or GPU)
pip install "surogate-speech[eval] @ git+https://github.com/invergent-ai/surogate-speech"   # evaluation harness

Python 3.10 or newer. Models download from Hugging Face on first use, at the revisions pinned in src/surogate_speech/hub.py.

Use

surogate-speech transcribe interviu.wav                     # TDT greedy
surogate-speech transcribe interviu.wav --lm                # CTC + 4-gram, the best accuracy
surogate-speech listen                                      # microphone, live partials and finals
surogate-speech listen --file interviu.wav                  # the same streaming path on a file
surogate-speech speak "Bună ziua! Cu ce vă pot ajuta?" --voice Doina -o doina.wav
surogate-speech speak "Bună ziua!" --voice Radu --backend nemo --device cuda -o radu.wav

Talking to a running engine instead, from Python:

from surogate_speech import client
print(client.transcribe("interviu.wav"))
client.speech("Bună ziua!", voice="Tudor", out="tudor.wav")

Evaluate

Every score runs reference and hypothesis through the language's normalizer (for Romanian, surogate_speech.text.normalize, version surogate-ro-v1), and every result file records it, along with the model's and the clip set's sha256.

# Speech recognition on the public Romanian test splits (FLEURS, VoxPopuli; Common Voice with --cv-repo)
surogate-speech eval asr --model-name jackrabbit-110m-ro --suite fleurs_ro,voxpopuli_ro --output asr.json

# Live-packet streaming: WER of the finals and finalization time; shard across GPUs
surogate-speech eval streaming --num-shards 8 --shard-index 0 --device cuda:0 --output shard0.json
surogate-speech eval streaming-merge --inputs shard*.json --manifest surogate-speech-data/manifest.tsv --output streaming.json

# TTS intelligibility: synthesize the 331-sentence stress set with every voice, transcribe with a judge
surogate-speech eval tts --judge nvidia/canary-1b-v2 --out tts-run/

WER is corpus-level, with a 95% interval from resampling reference sentences (clips of the same sentence move together). stress-v2 is in src/surogate_speech/sets/ with its sha256.

License

The code is Apache-2.0. The model weights are CC-BY-NC-4.0: free for research and other non-commercial use; for commercial use, contact Invergent. See MODEL_LICENSE.md for the base models and third-party components.

About

Small open speech models for agents that run on your own hardware: clients, evals and test sets for Surogate Speech

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages