r/LocalLLaMA · · 2 min read

[audio.cpp] What Does the Fox Say: 4 ASR models (Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR) in native C++/GGML, init streaming support, and 327s of audio transcribed in 2.17s.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I just pushed a new audio.cpp update with streaming support and 4 ASR models: Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR (da only). Overall 1.07x to 2.41x faster than Python.

I decided to drop Parakeet-TDT since good implementations already exist, and I find the Nemotron model more interesting.

Streaming support is also starting to land in audio.cpp.

Right now, I added initial streaming support for two models that naturally fit streaming usage: Nemotron ASR and VoxCPM2. Nemotron ASR can stream recognition results through SSE, and VoxCPM2 can receive text incrementally and return generated audio chunks. I also experimented with streaming support for Higgs Audio STT, though I still treat that path as more experimental because the model/reference behavior is less straightforward.

There is still a lot to improve. The current work is only the first step toward proper model-level and server-level streaming support. Contributions are very welcome, especially around better streaming APIs, chunk scheduling, lower TTFT, server behavior, and adding streaming paths for more models.

The headline result: Nemotron ASR transcribed a 327.6s audio file in 2.17s offline on RTX5090 with 3.18% WER, and the streaming SSE path produced the same WER with 307ms TTFT and much lower peak VRAM. VibeVoice ASR gave the lowest WER on this test, but it is much heavier. Nemotron is the most interesting result to me because the speed/VRAM tradeoff looks very strong, especially for local ASR service usage.

ASR models (No Quant):

Model Mode Dur. TTFT Wall RTF WER VRAM
Nemotron Offline 327.6s N/A 2.17s 0.0066 3.18% 8294M
Nemotron SSE 327.6s 308ms 11.62s N/A 3.18% 4382M
Nemotron SSE 1-shot 1800s N/A 53.66s 0.0298 3.45% 4167M
Higgs STT Offline 327.6s N/A 11.17s 0.0341 3.95% 12519M
Higgs STT SSE 327.6s 468ms 14.46s N/A 3.95% 6945M
VibeVoice Offline 327.6s N/A 19.24s 0.0587 0.66% 25833M
VibeVoice Offline 1800s N/A 123.72s 0.0687 1.51% 31209M

Streaming TTS:

Model Mode Memsaver TTFT Wall Audio RTF VRAM
VoxCPM2 SSE On 547ms 56.65s 314.6s N/A 7638M
VoxCPM2 SSE Off 308ms 55.75s 314.6s N/A 9091M

Repo: https://github.com/0xShug0/audio.cpp

audio.cpp is still pretty new, but the goal is becoming clearer: a ggml-based local audio framework that can handle TTS, ASR, voice cloning, long-form generation, and server-like usage without every model needing its own Python environment and custom runtime.

As always, backend/OS compatibility cannot be fully tested by one setup. If you try the Windows, Linux, CUDA, Vulkan, Metal, or CPU paths and find issues, detailed reports are very welcome.

Huge thanks to our community members for their contributions!

submitted by /u/Acceptable-Cycle4645
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA