Roark is a testing and evaluation platform for voice and chat AI teams. It simulates callers before launch and scores production calls with audio-native metrics, so you can catch problems early and keep improving your agents.
Roark: Voice AI Testing & Evals Platform
Roark is a quality assurance and evaluation platform for voice and chat AI teams. It doesn't build agents. It helps you test, monitor, and improve the ones you already have, by simulating calls before launch and scoring real production calls afterward.
Key Features
Simulation Testing: Run your agent against hundreds of simulated callers, including realistic personas and red-teaming for prompt injections. Tests cover 45 languages with native accents and background noise.
Post-Call Analysis: Score every production call on 500+ audio-native and conversational metrics, such as pronunciation, emotion, interruptions, empathy, and hallucination, measured directly from the waveform.
Self-Improving Loop: Use the prompt optimizer, ground-truth tuning, and fix verification to draft edits from failing calls, then replay the fixes against the exact callers that broke the agent.
Load and Health Tests: Run peak-volume concurrency tests and always-on health checks to catch outages before your customers do.
Enterprise Security: SOC 2 Type II compliance, HIPAA BAA availability, annual pen tests, SSO/SAML, and role-based access control.
Use Cases
Healthcare: Automatically check calls for required disclosures and identity verification.
Customer Support and Sales: Score production voice AI on pronunciation, empathy, and resolution across conversations.
Pre-Launch Validation: Validate client voice agents in staging with regression testing and CI/CD gates before they go live.
Pricing
Start free with $50 in credit, no credit card required. Detailed pricing tiers are not listed on the website.
Getting Started
Visit https://roark.ai/. Roark integrates with Vapi, Retell, LiveKit, and Pipecat, and offers Node and Python SDKs along with a REST API.