Gartner predicted that conversational AI would strip $80 billion out of contact center labor costs in 2026. That is this year. And yet almost nobody selling voice AI will tell you, in a straight line, what it actually costs to get there.
Search "voice ai pricing" and you get a wall of per-minute headline rates: 5 cents here, 12 cents there, "starting at" numbers that read like a phone plan. What none of those pages walk through is how a quoted rate turns into an actual monthly bill once you add the pieces a voice agent needs to function.
This post does that math. Not a vendor comparison table, a component breakdown, so you know what you are really paying for before you sign anything.
The sticker price is not the price
A voice AI platform's advertised rate is usually just the hosting or orchestration layer. It does not include the four things that make the call actually work:
- Telephony - the phone line connecting the call
- Speech-to-text (STT) - transcribing what the caller says
- The language model - deciding how to respond
- Text-to-speech (TTS) - generating the voice that responds
Each of those is billed separately on most developer-facing platforms, and each one has its own per-minute cost. Add them up, as the vendor numbers below show, and the headline hosting rate turns into a meaningfully higher blended rate once a call is actually live.
The three pricing models
Voice AI vendors generally charge one of three ways.
Per-minute. You pay only for talk time. This is the dominant model for developer platforms like Vapi and Retell AI, and it scales cleanly with call volume, no seat minimums.
Per-call. A flat fee per completed call regardless of length, which favors businesses with short, predictable interactions like appointment confirmations.
Subscription or flat-fee. A set number of minutes bundled into a monthly price, easier to budget but riskier if volume swings. Retell AI's pricing page documents a pay-as-you-go option with no minimum contract, which is the more forgiving structure if your call volume is unpredictable.
What actually stacks on top of the base rate
Here is what the four components cost when you check the primary sources directly.
Vapi publishes its component pricing openly: hosting starts at $0.05 a minute, Deepgram transcription runs $0.0095 to $0.0099 a minute, the OpenAI language model costs $0.0077 to $0.0452 a minute depending on which model you pick, ElevenLabs voice output runs $0.0146 to $0.0238 a minute, and Twilio telephony adds another $0.008 to $0.014 a minute. Stack the cheapest options and you land around $0.09 a minute. Stack the higher-end model and voice choices and you are closer to $0.13 to $0.15.
Retell AI's published pricing breaks out base voice infrastructure at $0.055 a minute, telephony at $0.015 a minute, the language model at $0.0016 to $0.16 a minute depending on which one you select, and text-to-speech at $0.015 to $0.040 a minute. Add-ons like PII removal cost an extra $0.01 a minute, and a branded caller ID adds $0.10 per outbound call on top of the per-minute rate.
Bland AI's pricing page lists a Start plan at $0.14 a minute plus a $0.05 a minute transfer fee for handoffs to a human, and a Build plan at $0.12 a minute plus a flat $299 a month platform fee.
ElevenLabs Agents charges $0.08 a minute for included usage and the same rate for overages, with the language model billed separately again on top of that. Exceeding your plan's concurrent call limit triggers burst pricing at a multiple of the standard rate, which is the kind of line item that does not show up until the invoice does.
Twilio, the telephony layer under many of the above, charges $0.0085 a minute for inbound calls and $0.0140 a minute for outbound, plus $1.15 a month per phone number, or $0.07 a minute if you use its own Conversation Relay AI layer.
None of these numbers are wrong on their own. They are just incomplete on their own, which is exactly why two businesses can compare "voice AI pricing" and land on wildly different real costs for what looks like the same product.
A worked example, not a quote
As an illustration, not a live quote, here is roughly what three tiers of voice agent look like once the components are stacked:
- DIY, cheapest components: hosting plus budget STT, LLM, and TTS choices lands around $0.09 to $0.12 a minute.
- Mid-tier, balanced components: a mainstream language model and a natural-sounding voice pushes that to roughly $0.15 to $0.20 a minute.
- Managed or premium: a done-for-you build with a higher-tier model, premium voice, add-ons like PII removal, and ongoing tuning typically runs $0.30 a minute or more, sometimes billed as a flat monthly package instead of metered minutes.
At 2,000 minutes a month, the difference between the cheap end and the managed end is the difference between roughly $200 and $600 a month, which is a meaningful gap for a small business deciding between assembling a voice agent themselves and buying a managed one.
Where the bill actually surprises people
Two patterns cause most of the "wait, why is this so expensive" moments.
Overage and burst pricing. Plans built around a fixed number of included minutes or a concurrent call cap charge a premium once you exceed either one. If your call volume spikes during a promotion or a seasonal rush, that is exactly when the multiplier hits.
Component creep. Every add-on, a better voice, PII redaction for compliance, a branded caller ID so calls do not get marked as spam risk, bills separately. None of these show up in the headline rate, and most businesses do not know they need them until after the agent is live.
Neither of these is a reason to avoid voice AI. Even a fully stacked component rate still lands well below what a human answering service charges per minute of talk time. It is a reason to ask a vendor for the fully-loaded rate before comparing it to anyone else's headline number, not after.
DIY component stacking versus a managed build
Assembling a voice agent from raw components gets you the lowest per-minute number on paper. It also means your team owns prompt engineering, testing across edge cases, and re-tuning every time a caller pattern changes, which is real ongoing work, not a one-time setup task.
A managed voice AI build costs more per minute but folds that work into the price: the integration with your booking or CRM system, the call flows tested against your actual business, and the tuning as your call patterns shift. For a business that wants a working agent without hiring for a skill it will only need occasionally, that trade is usually worth it. Our voice AI team builds and manages that layer directly, so the per-minute rate you see already reflects the finished product, not a starting point you still have to assemble.
If your voice agent also needs to trigger actions elsewhere, updating a record, sending a follow-up, routing a lead, that logic is closer to a broader automation build than a phone tree, and our AI automation team handles that connective layer so the voice agent is not operating in isolation from the rest of your systems.
For a closer look at how the three biggest developer platforms compare feature by feature once you are past the pricing question, our Retell AI vs Vapi vs Bland breakdown goes platform by platform. And if you are still deciding whether a full AI call center setup or a lighter answering service fits your call volume, our guide to AI call center costs covers that decision in more depth.
The rule for comparing voice AI pricing
Never compare two voice AI quotes until you know what each one includes. A hosting fee is not a per-minute rate, and a per-minute rate is not a monthly bill until telephony, the language model, and the voice are added back in. Vapi and Retell AI both publish their component pricing openly, which makes them the easiest starting point for building that comparison honestly instead of stacking two numbers that were never measuring the same thing.
If you would rather have the component math done for your actual call volume than build a spreadsheet from five different pricing pages, tell our voice AI team what your monthly call volume looks like and we will show you the real all-in rate before you commit to a platform.



