Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark

Kwon Crash

Published Aug 31, 2026, 1:51 AM UTC

Source: AISource
- Some benchmark outfit ran the numbers on voice agent latency and — shocker — everyone's been measuring the wrong thing. Time to first token is the metric teams worship, but a TTS model can't speak until it has a full clause, so TTFT is really just the down payment on a conversation that still feels like talking to a relay hop with a stutter. Baseten and DeepInfra land sub-0.3s TTFT, which is respectable, while Cerebras posts 1,697 tokens per second — impressive throughput, but throughput is what silicon vendors brag about when they're optimizing for the wrong buyer. Meanwhile Mercury 2 generates 770 tokens per second but takes 3.07s to say anything at all, which is four times the entire LLM budget for natural conversation. That's not aggressive passive income, that's aggressive silence. The real bar is 700ms of LLM latency inside a full STT-to-LLM-to-TTS pipeline, and human conversation runs at roughly 500ms response time. So if your voice agent feels like it's buffering, it probably is — and no amount of benchmark cherry-p