Vercel — AI · · 1 min read

AI Gateway now supports streaming transcription

Mirrored from Vercel — AI for archival readability. Support the source by reading on the original site.

AI Gateway now supports streaming transcription. Previously, transcription required a complete audio file and returned the full transcript in a single response. Now you can stream audio in as it's captured and get transcript updates back as the model produces them, keeping latency low for uses like live captioning and voice input.

Streaming transcription is in beta and available through the AI SDK's streamTranscribe function with any streaming-capable transcription model.

The example below streams raw PCM audio to openai/gpt-realtime-whisper and prints each transcript delta as it arrives. The result stream also carries partial and final transcripts:

import { experimental_streamTranscribe as streamTranscribe } from 'ai';
const result = streamTranscribe({
model: 'openai/gpt-realtime-whisper',
audio: audioStream, // ReadableStream<Uint8Array | string>
inputAudioFormat: { type: 'audio/pcm', rate: 24000 },
for await (const part of result.fullStream) {
if (part.type === 'transcript-delta') {
process.stdout.write(part.delta);

The same code works across providers: swap the model string to use xai/grok-stt or any other streaming-capable transcription model.

Streaming transcription also makes it easy to add a voice mode to an agent. Stream the user's speech to a transcription model and pass the live text to your agent. The agent itself does not change: it still receives text, so this works with any text-based agent. For agents that speak back, pair it with speech generation, or use realtime voice for full two-way conversation.

For more information on streaming transcription on AI Gateway, see the documentation.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Vercel — AI