GPT-6 Astra Powers Devin and Perplexity — Cost Impact
OpenAI says Perplexity and Cognition use GPT-6 Astra to test end-to-end systems. See live API costs and when reduced review repays the premium.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- OpenAI says Perplexity uses GPT-6 Astra to mock services and test complete workflows, while Cognition uses it inside Devin to run software and return evidence.
- Neither case study changes Astra's API rate; the live comparison below shows its premium against lower-cost OpenAI and Anthropic routes.
- The potential saving is less supervision and rework, but OpenAI provides no controlled cost, token, or productivity comparison for either customer.
- Use Astra for difficult verification loops only when it lowers cost per accepted change; route routine work to a cheaper model.
Current API cost for self-testing agent candidates
USD per 1M tokens. Input and output rates are charted separately.
Estimate a self-testing agent workload
Assumes 75% input tokens and 25% output tokens using current per-million rates.
GPT-5.6 Sol
openai
$80.00
- Input share
- $30.00
- Output share
- $50.00
Claude Opus 5
anthropic
$100.00
- Input share
- $37.50
- Output share
- $62.50
GPT-6 Astra
openai
$200.00
- Input share
- $75.00
- Output share
- $125.00
GPT-6 Astra versus current coding-agent alternatives
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| GPT-6 Astra | openai | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol | openai | $4.00 | $0.4 | $20.00 |
| GPT-5.6 Terra | openai | $2.00 | $0.2 | $12.00 |
| GPT-5.6 Luna | openai | $0.2 | $0.02 | $1.20 |
| Claude Opus 5 | anthropic | $5.00 | $0.5 | $25.00 |
Built from pricing.json at publish time.
OpenAI has published two GPT-6 Astra customer examples: Perplexity uses it to test complete systems, while Cognition uses it inside Devin to run software and return proof that a change works.
This is production evidence, not a model launch or price change. Test whether fewer retries and reviewer minutes repay Astra’s premium; the live components above use maintained rates.
The live Perplexity page is dated September 14. We treat that as a future publication label, not the date the work occurred.
What Perplexity and Cognition are doing
Perplexity cofounder Johnny Ho says Astra can generate small test programs and realistic responses from external services such as model APIs or connectors. Those stand-ins let the company check how an application behaves from start to finish. OpenAI also says Astra writes search programs and can edit and monitor production systems. Ho says Perplexity can trust the model with full systems and check its work less frequently than with earlier generations.
Cognition is applying Astra across Devin, its CLI, and desktop products. In OpenAI’s example, Devin runs an iPhone game in a simulator, returns a recording, lists successful checks, and identifies areas it did not test. For screenshot-based bug reports, Cognition says Devin can fix the issue and return a new screenshot as evidence.
| Company | Astra workflow | Evidence returned | Claimed benefit |
|---|---|---|---|
| Perplexity | Mock dependencies and exercise an end-to-end application | Test behavior across the full workflow | Less frequent human check-ins |
| Cognition / Devin | Implement, run, and inspect software | Simulator recording, test report, screenshots | Faster review and customer response |
Pricing impact: measure verified outcomes
Astra remains the premium model in this live comparison. No rate changed with these case studies. OpenAI’s official model page also warns that long-context requests above its threshold reprice the full request; discounted Batch and Flex processing and premium Fast mode create additional routing choices. Fast mode is unavailable when Astra uses EU data residency, so European teams requiring regional processing should benchmark Standard instead.
For a self-testing agent, token cost alone is incomplete. Measure total API and tool spend divided by accepted changes, then add reviewer minutes, failed runs, reverted changes, sandbox infrastructure, and incident risk. A more expensive run can win if it replaces several attempts or produces evidence a reviewer can validate quickly.
Use the OpenAI pricing page and token calculator for the model component. Compare Anthropic pricing with identical tasks and stopping rules before standardizing on one provider.
Who benefits—and who loses
Teams with costly integration tests, UI workflows, multi-service dependencies, or difficult bug reproduction have the clearest upside. Astra supports image input, code execution, hosted shell, computer use, function calling, and other Responses tools that can connect implementation to verification.
Simple code generation, formatting, classification, and deterministic unit-test work may lose on cost. If Sol, Terra, Luna, or another provider reaches the same acceptance rate, Astra’s premium adds spend without reducing review.
Reviewers also lose if “self-testing” becomes permission to skip independent checks. Model-generated evidence can be incomplete or optimized around the model’s own assumptions. Production authorization, security review, and rollback controls remain human and system responsibilities.
What developers should do now
- Choose one production-shaped task with an objective pass condition.
- Give Astra and a cheaper control the same repository state, tools, limits, and test environment.
- Require artifacts: logs, test output, recordings, screenshots, and an explicit list of untested behavior.
- Track accepted-change rate, retries, tokens, tool calls, runtime, and reviewer minutes.
- Isolate execution, protect secrets, cap spend, and require approval for deployment or destructive actions.
- Route to Astra only when the measured verification gain improves total cost per accepted change.
For a disposable prototype environment, compare DigitalOcean development infrastructure. Keep production credentials and customer data outside agent sandboxes until controls are validated.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored link above at no extra cost to you. It does not affect this analysis.
Limits of the evidence
Both pages are OpenAI customer stories, not independent benchmarks. They provide no task counts, baseline token usage, failure distribution, reviewer-hour totals, or controlled cost comparison. Cognition’s expectation of less manual review is forward-looking, and Perplexity’s reduced-check-in claim is not quantified.
Treat the examples as workflow designs worth testing—not proof that Astra is automatically cheaper. Our earlier Astra business rollout analysis covers the broader deployment and governance controls.
Labs decision: explicit route and replay blocker
GPT-6 Astra is not inserted into the AI Pricing Guru Labs leaderboard for this story. OpenAI’s API model is live, but our current benchmark route is not authorized for Astra, and neither customer page publishes a runnable task bundle, fixed prompts, tool configuration, raw traces, token ledger, review-time baseline, or accepted-change grader.
The decision is recorded in Labs coverage notes. A future replay needs authorized Astra access plus matched repository states, identical tools and limits, artifact grading for logs, recordings, screenshots, and untested behavior, and total cost per accepted change against a cheaper control.
Bottom line
Perplexity and Cognition are using GPT-6 Astra for the step coding agents often miss: proving that their work functions in a real workflow. That can change agent economics, but only when the returned evidence reduces retries and human review.
Start with a controlled verification loop and a cheaper baseline. Keep Astra where it wins on accepted-task cost, not where a premium model merely produces a more impressive demo. The maintained OpenAI API pricing guide covers the rest of the cost stack.
Sources: OpenAI’s customer stories on Perplexity and GPT-6 Astra and Cognition, Devin, and GPT-6 Astra, plus the official GPT-6 Astra model page and API pricing. Both customer pages returned 200 in a first-party source check; the Perplexity page displayed a future September 14 publication label. Claims, capabilities, availability, and rates checked September 12, 2026 at 08:02 UTC.