This week changed how teams should budget voice and autonomous agents, even though the featured providers did not change their underlying API token rates. OpenAI launched a new full-duplex voice layer, published two GPT-6 Astra production stories, and DeepSeek V4.1 Flash posted a strong but narrowly scoped cyber-agent result.

The live tables, chart, and calculator above read current rates from our maintained dataset. They separate published prices from customer claims and private benchmark bills.

StoryWhat changedBuyer impact
GPT-Live-1OpenAI released a full-duplex voice API with a separate backend agentBudget connected session time and backend usage as two bills
GPT-6 AstraPerplexity and Cognition described end-to-end testing workflowsMeasure whether returned evidence reduces retries and review time
DeepSeek V4.1 FlashEnclave reported complete target coverage in a private cyber harnessReproduce the result with cold and warm cache accounting before routing work

GPT-Live-1 split the voice-agent bill

OpenAI’s GPT-Live-1 can listen and speak simultaneously, handle interruptions, and delegate reasoning or actions to another model. It runs through a dedicated Live endpoint rather than as a model-string swap inside an existing Realtime integration.

Its commercial consequence is architectural: elapsed voice-session usage is one cost layer, while backend tokens, tools, telephony, storage, monitoring, and human escalation form another. The voice table above shows the maintained session rates; the text table and calculator model the backend portion.

Teams should compare cost per resolved call, not the visible voice rate alone. Test interruption recovery, noisy audio, accents, tool completion, abandoned-session handling, and backend escalation before committing traffic. Our GPT-Live-1 cost guide covers the endpoint and capacity details, while the OpenAI pricing page tracks current model rates.

Astra moved from capability to workflow evidence

OpenAI says Perplexity uses GPT-6 Astra to mock dependencies and test complete application flows. Cognition says Devin uses Astra to run software and return evidence such as recordings, test reports, and screenshots.

Neither customer story changed Astra’s price or supplied a controlled cost comparison. The useful hypothesis is that a premium model can repay its token cost when it prevents retries or shortens review. Buyers should give Astra and a cheaper control the same repository state, tools, stopping rules, and acceptance checks, then compare total cost per accepted change.

The Astra workflow analysis separates customer evidence from benchmark proof. Use the token calculator for model spend, then add tool runtime, sandbox infrastructure, failed attempts, and reviewer minutes.

DeepSeek’s cyber result needs a cache label

Enclave reported that DeepSeek V4.1 Flash achieved code execution across all 11 vulnerable targets while all four fixed controls remained secure. A path audit found that six successes used the planned weakness and five found alternate routes in the private test environment.

The result is worth evaluating, but it does not establish a universal “best hacking model.” The reported bill relied overwhelmingly on cached input, and the harness has since been repaired to close unintended routes. A buyer replay needs pinned application versions, fixed controls, immutable tools, transcript review, separate intended-path grading, and both cold- and warm-cache ledgers.

DeepSeek’s API price did not change with the benchmark. The schedule above reflects its maintained peak and off-peak rates. Read the full V4.1 Flash benchmark analysis and current DeepSeek pricing before testing a route.

What buyers should do now

Treat the week’s launches and case studies as measurement prompts, not automatic migrations. For voice, record the complete session-plus-backend bill. For coding agents, require evidence that reduces independent review. For cyber evaluation, keep execution isolated and distinguish genuine target success from harness shortcuts.

For a managed DeepSeek comparison, check Novita’s current model route against the first-party API, verifying the exact model ID, cache treatment, and schedule before sending production traffic.

Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored link above at no extra cost to you. It does not affect this analysis.

Bottom line

This week’s signal was not cheaper tokens. It was better evidence about where model spend hides: behind connected time, tool use, cache assumptions, retries, and human verification.

Route on cost per accepted outcome. Keep a cheaper control, preserve complete ledgers, and promote a new model only when the whole workflow improves.

Sources: OpenAI’s official GPT-Live-1 launch, customer stories on Perplexity and Cognition, Enclave’s DeepSeek V4.1 Flash benchmark and path audit, and DeepSeek’s official pricing documentation. OpenAI claims were rechecked against the first-party source records verified in this week’s linked coverage; DeepSeek and Enclave sources returned successfully on September 17, 2026.