Cascaded pipeline
WHAT IT MEANS
The common way voice agents work, in three steps: speech is turned into text (STT), an AI model thinks (LLM), then text is turned into speech (TTS).
Real example
Customer speaks → text → AI reply → spoken answer.
What’s normal
Most production voice agents still use this setup, because each part can be swapped and checked.
The India angle
It lets you pick the best Indian-language STT and TTS separately, e.g. one vendor for Tamil and another for Hindi.
Questions to ask your vendor
Which STT, LLM and TTS do you use, and can we change them?
Where is most of the delay in your pipeline?
Can we bring our own models?
Related terms
Companies that do this
Other terms in Basics
Supported by







