[ STACK LIBRARY ]

Tavus

Tavus

ABOUT

Tavus is an AI research lab building real-time conversational video agents called PALs. Its Conversational Video Interface API combines speech recognition, turn-taking, perception, an LLM, TTS, and face rendering in one pipeline.

Tavus: Real-Time Conversational Video Agents Tavus is a San Francisco AI research lab, founded in 2020, that builds real-time conversational video agents it calls PALs. Its Conversational Video Interface API combines speech recognition, turn-taking, perception, an LLM, TTS, and face rendering in one pipeline. Key Features Conversational Video Interface: A developer API for real-time face-to-face AI conversations, with bring-your-own LLM, knowledge base retrieval, memory, guardrails, and function calling. Sparrow-2 Turn-Taking: An audio-native model that handles pauses, interruptions, backchannels, and background noise to decide when the agent should listen, wait, or speak. Raven-1 Perception: Reads tone, prosody, facial expression, and gaze from voice, video, and screen inputs to inform responses. Phoenix-4.5 Rendering: Real-time, lip-synced faces built from a single image or two minutes of video. Custom replica training includes a custom voice model. Multiple Channels: PALs run over chat, voice, and video in 30+ languages, with a no-code PAL Maker builder. Use Cases AI SDRs: Qualified built its AI SDR, Piper, on Tavus to hold sales conversations with buyers. Interview Practice: Final Round AI runs mock job interviews with lifelike interviewers on Tavus CVI. Training and Intake: Role-play partners for sales and pharma rep training, and patient intake agents in healthcare.

STACK COVERAGE

Text-to-Speech

DETAILS

HQ COUNTRY

San Francisco, United States

FOUNDED

2020

Is this your company? To claim this listing or update its details, contact us at sunil@uniocommunity.com

UNIOVOICE AI STACK LIBRARY
Tavus logo
Trusted by Unio

Tavus

Text-to-Speech

Tavus is an AI research lab building real-time conversational video agents called PALs. Its Conversational Video Interface API combines speech recognition, turn-taking, perception, an LLM, TTS, and face rendering in one pipeline.

Layer
Text-to-Speech
Market
Global
HQ
San Francisco, United States
Founded
2020
Listed on the Unio Voice AI Stack Library