›

Basics

›

Cascaded pipeline

Cascaded pipeline

WHAT IT MEANS

The common way voice agents work, in three steps: speech is turned into text (STT), an AI model thinks (LLM), then text is turned into speech (TTS).

Real example

Customer speaks → text → AI reply → spoken answer.

What’s normal

Most production voice agents still use this setup, because each part can be swapped and checked.

The India angle

It lets you pick the best Indian-language STT and TTS separately, e.g. one vendor for Tamil and another for Hindi.

Questions to ask your vendor

  1. Which STT, LLM and TTS do you use, and can we change them?

  2. Where is most of the delay in your pipeline?

  3. Can we bring our own models?

Was this helpful?