Transparent, Usage-Based Pricing
No black box. A voice agent is a stack — speech recognition, a language model, a voice, and a phone line. We show you the real cost of every layer, so you always know where your money goes.
Talk to us about a buildThe real cost of every layer
Most platforms bundle everything into one padded per-minute number. We don't. Below are the exact per-unit rates Yaapi meters every call against, in the providers' own units — so you can see precisely what powers your agent.
| Layer | Provider & model | Rate |
|---|---|---|
| Speech‑to‑text | Deepgram Nova‑2 | $0.0077 / min |
| Language model you choose per build | OpenAI GPT‑4o mini | $0.15 / $0.60 per 1M tokens |
| DeepSeek Chat | $0.27 / $1.10 per 1M tokens | |
| Groq Llama 3.3 70B | $0.59 / $0.79 per 1M tokens | |
| Anthropic Claude Haiku 4.5 | $1.00 / $5.00 per 1M tokens | |
| Text‑to‑speech | ElevenLabs Flash v2.5 | $0.055 / 1k chars |
| ElevenLabs Multilingual v2 | $0.11 / 1k chars | |
| Telephony | Inbound phone minutes | $0.01 / min |
Token rates show input / output. Provider rates are set by each vendor and change over time; these are the rates Yaapi currently meters every call against (as of August 2026). You pick the model that fits your quality and budget.
In practice, a typical conversation minute runs roughly $0.06–$0.10 at provider cost. Text-to-speech is usually the largest single component, so the exact figure depends on how much the agent talks and which model you run — we help you pick the stack that fits.
What Yaapi charges
The rates above are the underlying provider costs. Yaapi's own pricing is scoped to your build — every agent is custom-engineered for your business logic, not sold as a fixed tier. Tell us what you need and we'll put a number to it.