Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| We just open-sourced Gepard 1.0, a TTS model built for real-time conversation. It’s streaming-first: audio starts the moment text arrives, generated frame by frame instead of waiting for a full sentence. - ~555M params: Qwen3.5 0.8B backbone (14 layers) + Nemo NanoCodec (FSQ, 22.05kHz) Benchmarks (Seed-TTS-eval): we put it head-to-head against VoxCPM2, Fish-S2, OmniVoice, Qwen3-TTS, Echo-TTS, and** **Chatterbox Turbo on identical texts. Honest tradeoff: The streaming-first design costs us on speaker similarity (SIM 0.585) and WER (0.036), so it’s a strong fit where a natural realtime voice matters more than exact voice-matching. Links: vLLM serving (Cartesia compatible API) Also you can check how it works on vLLM on our website: https://www.nineninesix.ai Happy to answer questions on the architecture or the inference! [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.