NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling

Kwon Crash

Published Aug 10, 2026, 1:51 AM UTC

Source: AISource
- NVIDIA dropped NemotronLabs VoiceChat 11B — an 11B speech-to-speech model that listens while it talks, barge-in capable, with live tool calling on a side channel. Translation: one model replaces the ASR→LLM→TTS chain at 448ms turn-taking latency. First open full-duplex model that can call tools mid-conversation without going silent. Weights are permissive, container is public, and it runs on a single 80GB GPU. Here's the catch: NVIDIA stamped it "research only" because after a few turns it degrades into non-recoverable gibberish, drops user words, and occasionally talks to itself like a Discord scammer narrating a wallet drain. Two-minute audio ceiling, no hosted API, no inference provider serving it. Tool calling works but maxes at five tools per session, can't do parallel calls reliably