Tavus is an AI research lab building real-time conversational video agents called PALs. Its Conversational Video Interface API combines speech recognition, turn-taking, perception, an LLM, TTS, and face rendering in one pipeline.
Tavus: Real-Time Conversational Video Agents
Tavus is a San Francisco AI research lab, founded in 2020, that builds real-time conversational video agents it calls PALs. Its Conversational Video Interface API combines speech recognition, turn-taking, perception, an LLM, TTS, and face rendering in one pipeline.
Key Features
Conversational Video Interface: A developer API for real-time face-to-face AI conversations, with bring-your-own LLM, knowledge base retrieval, memory, guardrails, and function calling.
Sparrow-2 Turn-Taking: An audio-native model that handles pauses, interruptions, backchannels, and background noise to decide when the agent should listen, wait, or speak.
Raven-1 Perception: Reads tone, prosody, facial expression, and gaze from voice, video, and screen inputs to inform responses.
Phoenix-4.5 Rendering: Real-time, lip-synced faces built from a single image or two minutes of video. Custom replica training includes a custom voice model.
Multiple Channels: PALs run over chat, voice, and video in 30+ languages, with a no-code PAL Maker builder.
Use Cases
AI SDRs: Qualified built its AI SDR, Piper, on Tavus to hold sales conversations with buyers.
Interview Practice: Final Round AI runs mock job interviews with lifelike interviewers on Tavus CVI.
Training and Intake: Role-play partners for sales and pharma rep training, and patient intake agents in healthcare.