nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| IntroductionWe're excited to introduce Nemotron-Labs-Audex-30B-A3B, a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM with 30B MoE model with 3B activated parameters. Audex-30B-A3B extends the vocabulary for discrete audio tokens used for speech and general audio outputs, as well as an audio encoder for speech and general audio inputs. Audex-30B-A3B delivers strong abilities on audio tasks (audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation) while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. Audex-30B-A3B operates in both thinking and instruct (non-thinking) modes. Quick Start
Check Model card for so much benchmarks. Additional model: [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.