October 2, 2026
Voice AI shares a single GPU: .wave served 56 Nemotron VoiceChat sessions
In .wave's Sep 10 benchmark, a single H100 served 56 Nemotron VoiceChat sessions without late audio.
The voice model listens to the other person while responding and keeps running even during pauses. If it misses the deadline for the next audio chunk, the listener hears a gap.
More concurrent sessions. In .wave's benchmark, the reference NVIDIA NIM implementation handled one session. With two sessions, it delivered 8–15% of audio chunks late. The WPK engine served 56 sessions on the same H100 without missing a single deadline across 84,000 audio chunks checked in three runs.
.wave moved compute scheduling onto the GPU and batched session processing. The company sets the execution order and memory layout in advance instead of deciding them during the conversation. The test used identical model weights and precision, the same audio recording in every session and two minutes of context, with latency measured on the server.
Connecting from Python. VoiceChat is available through the .wave API for voice conversations with interruptions. The guide includes a ready-to-run microphone example: install `python -m pip install websockets sounddevice`, set `DOTWAVE_API_KEY` and run the example. The connection uses WebSocket; the application's first message sends `session.start` with the `nemotron-voicechat` model, then streams PCM16 mono 24 kHz audio through `session.input_audio.append`.
As of Oct 2, the introductory rate is $0.005 per minute, or $0.30 per hour of an open session, including pauses. After signup, .wave provides a key and free credits without requiring a bank card. According to the company, these cover roughly 33 hours of VoiceChat.
The current API supports English conversations lasting up to 120 seconds with the aria voice, without custom instructions, text input or tool calls.
