Blog
Voicetta vs. Vapi vs. Retell vs. Fonio: What Does a Voice AI Agent Actually Cost?
Voice AI pricing pages rarely tell the whole story. A platform might quote 5¢ per minute for its own infrastructure layer. Speechtotext, the LLM, texttospeech, and telephony often arrive as separate line items.
Voice AI pricing pages rarely tell the whole story. A platform might quote 5¢ per minute for its own infrastructure layer. Speech-to-text, the LLM, text-to-speech, and telephony often arrive as separate line items.
This guide compares four platforms buyers ask about most: Voicetta, Vapi, Retell AI, and Fonio. Each prices its stack differently — infrastructure, telephony, STT, LLM, and TTS. We calculated the real per-minute cost for each, not just the headline number.
The goal isn't crowning a single "cheapest" winner. Pricing shifts, and the right fit depends on your call volume and setup. The goal is showing where each cent actually goes, so you can run the math yourself.
What does each platform actually cost?
Voicetta's BYOK infrastructure runs 1.6¢/min. AI provider costs — LLM, STT, TTS — bill separately. All-in, that typically lands around 3–5¢/min.
Vapi charges 5¢/min for its hosting layer alone. AI provider costs pass through separately, the same way they do on Voicetta's BYOK tier. All-in, that usually lands around 6.5–8¢/min.
Retell AI prices infrastructure, telephony, TTS, and the LLM as four separate line items. Depending on model choice, the blended total ranges from roughly 9¢/min to over 27¢/min. That's the widest spread of the four platforms.
Fonio bundles voice usage and TTS into one rate. On the monthly plan, that's €0.10/min — about 11.7¢/min at current exchange rates. Telephony for outbound calls carries separate carrier charges on top.
None of these four numbers are directly comparable at face value. Some cover infrastructure alone; others fold in AI usage, telephony, or both. The table below breaks out each layer so you can see exactly what you're paying for before you commit.
| Platform | Infra + Telephony | AI Layer (STT/LLM/TTS) | Estimated All-In | Messaging |
|---|---|---|---|---|
| Voicetta (BYOK) | 1.6¢/min | Billed at cost by your provider | ~3–5¢/min | SMS 5¢/msg, WhatsApp 1¢/msg |
| Vapi | 5¢/min hosting | Billed at cost by your provider | ~6.5–8¢/min | SMS 0.5¢/msg |
| Retell AI | 5.5¢/min infra + 1.5¢/min telephony | 1.8–17.6¢/min (TTS + LLM) | ~9–27¢/min | Not published |
| Fonio | ~9.3–11.7¢/min bundled | TTS included; carrier fees separate | ~9.3–11.7¢/min + carrier | WhatsApp sold as separate subscription |
1. Voicetta: infrastructure priced separately from AI usage
Disclosure: Voicetta is our own product, so we're not a neutral party here — we've tried to price every platform on this list the same honest way.
Voicetta's BYOK model splits the bill into two pieces. Infrastructure — routing, retries, orchestration, and telephony — runs 1.6¢/min. AI usage bills separately, at whatever rate your chosen provider charges.
That split matters for comparison. Voicetta never marks up your AI provider bill. The 1.6¢/min covers orchestration and call handling only, making it directly comparable to another platform's infrastructure-only rate.
The cost breakdown
The 1.6¢/min splits into 1.2¢/min for voice infrastructure and 0.4¢/min for domestic telephony. That's the number Voicetta controls directly. Everything above it depends on your provider choices.
Add a typical AI stack on top — a mid-tier LLM, streaming STT, and a fast TTS model. Total BYOK cost usually lands around 3–5¢/min all-in. The exact figure depends on which providers you pick.
For teams that don't want to manage separate AI provider accounts, Voicetta also offers Managed AI. That tier starts at 12.4¢/min, with infrastructure, telephony, and AI usage bundled into one rate. A Realtime Agent tier reaches 40.6¢/min for the highest-throughput configuration.
What's included beyond the rate
The infrastructure rate covers more than phone calls. Voice, SMS, and WhatsApp all run through one execution layer. SMS is priced at 5¢ per message, and WhatsApp at 1¢ per message.
That matters because the other three platforms in this comparison price messaging very differently, if they price it at all. There's no separate per-channel product fee stacked on top at Voicetta. One rate card covers the whole conversation, not just the phone call.
AI Evaluations are included at every tier, with no extra charge. Every call and every WhatsApp or SMS thread gets graded automatically against criteria you define. That's full coverage, not a sampled audit.
Where the vendor comparison comes from
Voicetta's own rate card estimates roughly $16 in infrastructure cost at 1,000 BYOK inbound minutes. The same card estimates Vapi's comparable hosting cost at $50–60 for that volume. Retell's estimated cost lands at $55–70+ for the same 1,000 minutes.
That comparison is vendor-sourced from Voicetta's own pricing analysis. It isn't an independent audit, and competitor rates shift over time. Treat it as directional, and verify current rates before quoting them to a client.
Who it fits
Voicetta fits operators who want a finished, done-for-you system rather than a build project. The BYOK tier suits teams that already have AI provider relationships and want infrastructure priced honestly on top. It's a narrower fit for developers who want to assemble every layer themselves.
2. Vapi: infrastructure-only pricing with AI passed through at cost
Vapi prices its own hosting layer at 5¢/min. Speech-to-text, the LLM, and text-to-speech bill separately. Vapi passes through whatever your chosen AI providers charge, with no markup if you bring your own key.
Structurally, that's close to Voicetta's BYOK approach — an infrastructure fee plus pass-through AI costs. The gap shows up in the infrastructure number itself. Vapi's 5¢/min hosting rate runs roughly three times Voicetta's 1.6¢/min.
The cost breakdown
Add a typical AI stack to Vapi's 5¢/min hosting fee. All-in BYOK cost lands around 6.5–8¢/min once STT, LLM, and TTS are included. That's before any concurrency costs beyond the included tier.
Vapi includes 10 free concurrent call lines. Each additional line beyond that costs $10/month. For teams running high call volumes simultaneously, that monthly add-on can meaningfully change the total bill.
Vapi's SMS pricing is listed separately at 0.5¢ per message. That's the cheapest text rate among the platforms compared here. It's a standalone channel, though, not part of a shared cross-channel memory system.
What's included beyond the rate
Vapi's strength is developer flexibility. The platform supports composable voice pipelines and a large model marketplace. Deep API access lets technical teams build a fully custom stack from scratch.
That flexibility comes without a done-for-you deployment path. Every workflow — call flows, fallback logic, retry handling — has to be built and maintained in-house. There's no managed tier for buyers who'd rather not touch the configuration.
Who it fits
Vapi suits technical teams that want infrastructure, not a finished product. With 100,000+ developers building on the platform, it has one of the largest ecosystems in voice AI. But every deployment is a build project — the pricing reflects raw infrastructure, not a packaged operational system.
3. Retell AI: four separately priced components
Retell AI prices voice infrastructure, telephony, text-to-speech, and the LLM as four distinct line items. That component-level pricing is the most granular on this list. It's also the widest-ranging.
Voice infrastructure runs 5.5¢/min. Telephony adds another 1.5¢/min on top. TTS ranges 1.5–4¢/min depending on the provider, and LLM cost ranges 0.3–16¢/min depending on the model.
The cost breakdown
At the cheapest possible configuration, Retell's blended rate lands around 9¢/min. At the most expensive configuration — a premium LLM and TTS provider — it can exceed 27¢/min. Retell's own pricing page quotes a public range of 7–31¢/min for exactly that reason.
Monthly extras stack on top of the per-minute rate. Phone numbers run $2/month each. Concurrency beyond 20 free lines costs $8/month per additional line, and knowledge bases beyond the first 10 run $8/month each.
That component-level structure gives buyers real control. It also means two teams on the same platform can pay wildly different totals. The final number depends entirely on the model and provider choices made during setup.
What's included beyond the rate
Retell includes automated QA through its Retell Assure feature. Multi-channel support across voice, chat, email, and SMS shipped as of its January 2026 expansion. New signups get $10 in free credits to test the platform before committing.
Who it fits
Retell suits technical buyers who want granular control over each component. It works well for teams comfortable managing model selection to hit a specific target cost. The final bill depends heavily on setup choices, not on the advertised rate alone.
Buyers who skip that setup work often end up paying closer to the 27¢/min ceiling than the 9¢/min floor. Reviewing model and provider defaults before launch is worth the time. It's the difference between a competitive rate and an expensive surprise on the first invoice.
Independent user reviews on platforms like G2 are a useful sanity check before committing. They surface real deployment costs that a pricing page alone won't show. Cross-referencing a vendor's own numbers against buyer experience is good practice for any platform on this list.
4. Fonio: bundled per-minute pricing, billed in euros
Fonio bundles voice usage and text-to-speech into a single per-minute rate, rather than pricing each layer separately. The Solo plan runs €99/month for 1,000 minutes. That works out to about €0.10/min, or roughly 11.7¢/min at current exchange rates.
The annual plan drops that rate to €0.08/min, or roughly 9.3¢/min. Fonio publishes its rates in euros. US-based buyers should factor in currency conversion before comparing directly against dollar-denominated platforms.
The cost breakdown
Overage pricing tells a different story than the headline rate. Extra minutes cost €15 per 100 minutes on the monthly plan — about €0.15/min. That's noticeably higher than the €0.10/min included rate.
Any team running above its plan's minutes regularly should factor that gap in early. A property manager expecting 1,200 minutes on a 1,000-minute plan pays the higher overage rate on the last 200. Estimating volume conservatively avoids that surprise.
Telephony for outbound calls and campaigns carries separate carrier charges based on destination country. Fonio's pricing page doesn't clearly state whether inbound telephony is bundled into the per-minute rate or billed separately. That's worth confirming directly with Fonio before committing to a plan.
What's included beyond the rate
The per-minute rate includes access to 20+ voices across 25+ languages. That's useful for multilingual deployments spanning multiple markets. Call recording and a built-in scheduler are also included at every tier.
WhatsApp isn't priced per message on Fonio — it's a separate subscription product. The Solo tier starts at €79/month for 800 conversations, which works out to roughly €0.099 per conversation. That's structurally different from a flat per-message rate.
Who it fits
Fonio fits teams already operating in euro-denominated markets who want one bundled rate instead of four separate line items. The trade-off is less transparency on exactly what that bundled rate includes at the telephony layer. Buyers outside the eurozone will also need to track currency movement against their own billing.
Frequently asked questions
Why do voice AI platforms quote such different starting prices?
Some platforms quote infrastructure-only rates and pass AI provider costs through separately. Others bundle everything — telephony, STT, LLM, TTS — into one number. A 5¢/min quote and an 11¢/min quote aren't directly comparable until you know what sits inside each one.
What's the real difference between BYOK and bundled pricing?
BYOK means you supply your own AI provider keys, and the platform charges only for its infrastructure layer. Bundled pricing means the platform pays AI providers for you. That cost rolls into one number, usually with some margin built in.
Neither model is universally cheaper. BYOK tends to win at high volume, once you've already negotiated good rates with your AI providers. Bundled pricing wins on simplicity for teams that don't want to manage multiple vendor accounts.
How much does concurrency and scale add to the sticker price?
More than most buyers expect. Retell charges $8/month per concurrent call line beyond 20 free ones. Vapi charges $10/month per line beyond 10 free ones.
At high call volumes, those add-on fees compound fast. A team running 40 concurrent calls on Retell pays an extra $160/month before a single minute of usage is counted. Always ask about concurrency pricing before comparing base rates.
Does currency matter when comparing voice AI pricing?
Yes, more than it seems at first glance. Fonio prices in euros, and a €0.10/min rate converts to a slightly different dollar figure every week. Exchange rate movement alone can shift the comparison by several percent.
Always convert to your own billing currency before comparing platforms side by side. A rate that looks competitive at one exchange rate can look different a quarter later. Build in a small buffer if you're budgeting in a currency other than the one on the invoice.
Why do overage rates sometimes cost more than the included rate?
Providers often price bulk-included minutes lower than pay-as-you-go overage minutes, since the included tier is prepaid upfront. Fonio's overage rate runs about 50% higher per minute than its included rate. That pattern is worth checking on any provider's plan before committing to a specific tier.
How do messaging costs factor into total voice AI spend?
They add up faster than most buyers expect at volume. Voicetta prices SMS at 5¢ per message and WhatsApp at 1¢, both running through the same system as voice.
Vapi prices SMS separately at 0.5¢ per message, cheaper on a per-message basis. Fonio sells WhatsApp as a standalone subscription product instead of a per-message rate.
That subscription model changes the math at low conversation volumes. A team sending 200 WhatsApp conversations a month still pays Fonio's full €79 tier. Per-message pricing scales down more naturally for lighter usage.
Conclusion: compare the whole stack, not the headline number
Voicetta, Vapi, Retell, and Fonio all structure pricing differently. No single number captures the real cost of any of them. Infrastructure, telephony, STT, LLM, and TTS each carry their own rate.
That rate shifts depending on the provider and model you choose. Vapi and Voicetta both separate infrastructure from AI usage, which makes their base rates the most directly comparable pair. Retell's four-component pricing gives the most control but produces the widest possible range.
Fonio bundles more into one number but leaves telephony and currency conversion as open questions worth confirming upfront. None of these platforms price things the wrong way. They're built for different buyers with different priorities.
The only reliable comparison is your own call volume against each provider's actual rate card. The landing page number rarely survives contact with a real invoice.
Concurrency fees, overage rates, and messaging costs all move the real total. Run those numbers before you commit to a plan.
Want to see how Voicetta's BYOK pricing maps to your call volume, side by side with what you're paying today? Book a session and we'll run the numbers together — no sales script, just math.
Was this useful?
If this article helped, add Voicetta as a preferred source in Google Search. Your results can then highlight our writing.
Add as preferred source