.wave is a Y Combinator-backed (W25) platform for continuous inference in real-time AI. It runs stateful models that process live input throughout a session, with a Live API for hosted models and managed deployments for custom ones.
.wave: Continuous Inference for Real-Time AI
.wave is a Y Combinator-backed (W25) platform designed for continuous inference in real-time AI applications. It lets users run stateful models that process live input throughout a session, offering both a Live API for hosted models and managed deployments for custom models.
Key Features
.wave Live API: Access hosted models like NemotronLabs VoiceChat 11B (currently in public beta), with streaming STT models such as Nemotron 3.5 ASR Streaming 0.6B coming soon.
Managed Deployments: Bring your own continuous models to a managed deployment environment.
High-Performance Engine: WPK, a persistent GPU kernel, keeps recurring work on the GPU, enabling up to 56x more VoiceChat sessions per GPU and lowering GPU cost contributions by 98%.
Strict Timing Contracts: Meets recurring deadlines for each session so ASR streams stay on schedule with zero missed frames.
Use Cases
Full-Duplex Voice: Run models that need continuous, real-time voice interaction.
Streaming ASR: Process live audio streams for speech recognition without falling behind schedule.