Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news

OpenAI: Ringg Resolves 65% of Calls — Cost Impact

Ringg says its OpenAI agents resolve up to 65% of routine calls and cut selected model costs 90%. See the evidence, pricing, and buyer advice.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • Ringg says its agents resolve up to 65% of routine customer inquiries without a human and now handle more than 7 million connected calls per month.
  • Moving suitable real-time workloads from GPT-4.1 to GPT-5.6 Luna reportedly reduced model cost by about 90%; Ringg still uses GPT-4.1 for most real-time voice and chat traffic.
  • This is an OpenAI customer story, not an independent audit or a new price cut. Public materials omit token usage, call duration, telephony cost, sample definitions, and complete evaluation data.
  • Buyers should compare total cost per correctly resolved call, including speech, telephony, tools, retries, human escalation, and quality review.

Current token-cost range across Ringg's model routes

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$20.00GPT 5.6 Lunaopenai$0.2$1.20GPT 4.1openai$2.00$8.00GPT 5.6 Terraopenai$2.00$12.00GPT 5.6 Solopenai$4.00$20.00

Estimate the text-model layer of a support workflow

Assumes 75% input tokens and 25% output tokens using current per-million rates.

GPT-5.6 Luna

openai

$4.50

Input share
$1.50
Output share
$3.00

GPT-4.1

openai

$35.00

Input share
$15.00
Output share
$20.00

GPT-5.6 Terra

openai

$45.00

Input share
$15.00
Output share
$30.00

Live API rates for Ringg's reported OpenAI model stack

Model Provider Input / 1M Cached / 1M Output / 1M
GPT-5.6 Luna openai $0.2 $0.02 $1.20
GPT-4.1 openai $2.00 $0.5 $8.00
GPT-5.6 Terra openai $2.00 $0.2 $12.00
GPT-5.6 Sol openai $4.00 $0.4 $20.00

Built from pricing.json at publish time.

OpenAI has published a Ringg customer story claiming that the company’s AI agents resolve up to 65% of routine customer inquiries without human involvement. Ringg says the platform handles more than 7 million connected calls per month and its customers average a 4.8 CSAT score.

The cost result is narrower than the headline: Ringg reports that moving suitable real-time workloads from GPT-4.1 to GPT-5.6 Luna reduced model cost by approximately 90%. It still routes most real-time voice and chat traffic to GPT-4.1. This is selective model routing, not a blanket migration or a new OpenAI price cut.

What Ringg built

Ringg runs agents across voice, chat, WhatsApp, and the web. Its orchestration layer connects models to CRMs, ticketing systems, payments, scheduling tools, and internal APIs. Specialized subagents can handle qualification, verification, support, booking, and escalation while passing a conversation summary to a human when needed.

The platform combines structured filters with semantic retrieval over business records and documents. For long interactions, it creates a structured summary when context approaches roughly 80,000 tokens instead of repeatedly sending the complete history. That can control input growth, but the story does not publish the before-and-after token ledger.

How Ringg routes OpenAI models

WorkloadReported routeCost role
Most real-time voice and chat trafficGPT-4.1Production default for latency-sensitive conversations
Suitable real-time requestsGPT-5.6 LunaLower-cost route where quality, latency, and tool use pass Ringg’s tests
Post-call summaries and sentimentGPT-5.6 TerraOffline analysis; Ringg reports strong regional-language accuracy
Evaluations, prompt improvement, model-as-judgeGPT-5.6 SolHigher-capability quality-control route

Ringg says it tests historical conversations and simulated customer flows offline, introduces passing models to a small share of live traffic, and expands only after production monitoring. Its router watches regional endpoint health and latency, shifting traffic when thresholds are crossed.

Pricing impact: model cost is only one layer

The live chart, calculator, and table above use our canonical pricing data rather than freezing token rates in this article. They show why Luna can materially reduce the model layer when it passes the same acceptance gate as GPT-4.1. Current rates and context rules are maintained on the OpenAI pricing page.

The calculator is not a full voice-call quote. A deployment can also incur speech recognition, voice generation, telephony, retrieval, tool execution, storage, monitoring, retries, and human-escalation costs. Compare an independent model route through Google Gemini pricing and use the token calculator for the text portion. Our best AI for customer support guide covers broader vendor trade-offs.

Teams evaluating a separate voice stack can benchmark ElevenLabs voice agents on the same calls, languages, latency targets, and resolution rubric.

Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored ElevenLabs link at no extra cost to you. It does not affect this analysis.

What the customer results do—and do not—prove

DeploymentResult reported by OpenAI and RinggEvidence limit
Ringg overallUp to 65% of routine inquiries resolved without a humanNo common denominator, test period, or audited dataset published
Policybazaar67% of calls handled without a human; response time reduced from 8–12 minutes to under 60 secondsCase-specific workflow and measurement details are unpublished
Practo85% first-call resolution; response below three seconds; operating cost down 70%First-call resolution is not the same metric as full automation
Groww72% of selected inbound queries resolved through self-serviceScope covers stated investment-query categories, not all support

These figures are not contradictory: they describe different customers, tasks, and metrics. They also should not be averaged. OpenAI’s page does not provide raw conversations, failure rates, human-review rules, billing records, or an independent audit.

What support teams should do now

  1. Select one high-volume intent with a clear completion event and a safe human fallback.
  2. Replay the same multilingual call set through the current route and Luna, holding speech, tools, retrieval, and prompts constant.
  3. Grade correct resolution, unsafe actions, transfers, repeat contacts, latency, and customer effort—not containment alone.
  4. Record every cost layer per accepted resolution, including human review and post-call analysis.
  5. Canary the winning route on limited traffic and preserve versioned transcripts, prompts, and model snapshots.

Do not optimize for keeping callers away from people at any cost. A wrongly “resolved” insurance, healthcare, or financial request can create repeat calls, compliance exposure, and expensive remediation.

Labs coverage decision

Ringg’s customer metrics do not enter the AI Pricing Guru Labs leaderboard. GPT-5.6 Luna, Terra, and Sol already have exact-route Labs results on the fixed deterministic text suite: Luna scored 48/49, while Terra and Sol scored 49/49, all with zero endpoint errors. Those runs confirm the tracked text routes; they do not reproduce live telephony, interruptions, regional-language switching, business tools, or human handoffs.

The 65%, 90%, 97%, and customer-specific outcome claims stay outside the score. A fair replay needs a consented call corpus, pinned models and prompts, a fixed speech stack, identical backend tools and transfer rules, blinded resolution grading, latency measurement, repeat-contact follow-up, and a complete cost ledger.

Bottom line

Ringg’s case supports a practical architecture: use a tested low-cost model for routine traffic, reserve stronger models for analysis and evaluation, summarize long conversations, and escalate when the task exceeds the automation boundary.

The headline results are promising, but they remain first-party case-study claims. Run a controlled replay and buy on total cost per correctly resolved call—not model price or containment rate alone.

Sources: OpenAI’s official Ringg customer story, GPT-5.6 Luna model documentation, API pricing documentation, and Ringg’s product site. Claims, product details, and rates checked September 23, 2026 at 16:50 UTC.