Best Default AI Model? Hacker News Picks Cheap and Fast
A 63-comment Hacker News thread favors fast, economical default AI models with premium escalation. See the live API cost comparison.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- There is no Hacker News consensus on one best default: among 43 top-level replies, Anthropic models received 21 mentions, OpenAI 16, Google 11, DeepSeek 9, Z.ai/GLM 8, and xAI 6; many replies named several providers.
- The clearest workflow pattern was a fast, economical model for routine work, with premium models reserved for planning, difficult bugs, review, or failed first attempts.
- Subscription allowances, latency, writing style, harness quality, and human involvement influenced choices as much as API token prices.
- Pick a default with a replay test on your own tasks, then escalate only when a cheaper model misses a measurable acceptance check.
Default-model candidates by live token price
USD per 1M tokens. Input and output rates are charted separately.
Estimate the cost of your default route
Assumes 75% input tokens and 25% output tokens using current per-million rates.
DeepSeek V4.1 Flash
deepseek
$2.63
- Input share
- $1.13
- Output share
- $1.50
GPT-5.6 Luna
openai
$4.50
- Input share
- $1.50
- Output share
- $3.00
Gemini 3.8 Flash
$15.00
- Input share
- $5.63
- Output share
- $9.38
Claude Haiku 4.5
anthropic
$20.00
- Input share
- $7.50
- Output share
- $12.50
GLM-5.3
zai
$21.50
- Input share
- $10.50
- Output share
- $11.00
Live API rates for representative default-model candidates
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | deepseek | $0.15 | $0.0030 | $0.6 |
| GPT-5.6 Luna | openai | $0.2 | $0.02 | $1.20 |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | |
| Claude Haiku 4.5 | anthropic | $1.00 | $0.1 | $5.00 |
| GLM-5.3 | zai | $1.40 | $0.26 | $4.40 |
| Claude Sonnet 5 | anthropic | $2.00 | $0.2 | $10.00 |
| GPT-5.6 Sol | openai | $4.00 | $0.4 | $20.00 |
| Claude Opus 5 | anthropic | $5.00 | $0.5 | $25.00 |
| Claude Fable 5.1 | anthropic | $10.00 | $0.25 | $50.00 |
| GPT-6 Astra | openai | $10.00 | $1.00 | $50.00 |
Built from pricing.json at publish time.
A September 12 Hacker News thread asking developers which AI model they use by default produced a fragmented answer: there is no universal winner. The recurring strategy mattered more than the brand—use a fast, economical workhorse for most turns, then escalate selectively.
That conclusion is more useful than a popularity ranking. The thread mixes coding assistants, consumer subscriptions, direct APIs, free promotions, employer-funded access, and local models. Those choices cannot be compared on token price alone.
What Hacker News developers are using
At 17:33 UTC, the thread had 30 points and 63 comments, including 43 top-level replies. A simple mention count across those top-level replies found the following provider families. Counts overlap because many developers described multi-model workflows.
| Provider family | Top-level replies mentioning it | Common role described |
|---|---|---|
| Anthropic | 21 | Planning, complex coding, review, subscription-funded work |
| OpenAI | 16 | Fast routine work, implementation, premium escalation |
| 11 | Fast iteration, autocomplete-like coding, academic work | |
| DeepSeek | 9 | Low-cost engineering default, hosted or local use |
| Z.ai / GLM | 8 | Economical workhorse and open-weight preference |
| xAI | 6 | Specialist or second-opinion route |
| Qwen | 3 | Local and open-weight workflows |
This is an anecdotal snapshot, not a randomized poll. Model names were also inconsistent: some replies named a provider, subscription, agent harness, reasoning level, or older checkpoint rather than an exact API route.
Why cheap and fast models win the default slot
Several developers described keeping a smaller model on a short feedback loop while they remained involved. Examples included Haiku for shell questions, Gemini Flash for fast coding, Luna for routine implementation, and DeepSeek or GLM Flash for economical agent work. Premium models appeared more often as planners, reviewers, hard-bug solvers, or fallbacks.
The original poster supplied the clearest warning against a frontier default: a four-agent planning session rapidly consumed a Claude Max allowance, and a later cache miss accelerated usage again. That is one user’s experience, not an audited limit test, but it exposes the multiplier created by parallel agents and expired prompt caches.
Other replies preferred premium models from the start, arguing that fewer failed iterations can outweigh the higher unit cost. Both approaches can be rational; the deciding metric is cost per accepted task, not model prestige or price per token by itself.
What the live API comparison means
The chart, calculator, and table above normalize representative active API models mentioned in the discussion. They do not reproduce consumer-plan allowances, free promotions, employer subsidies, fast-mode premiums, reasoning settings, or local-hardware costs.
Use the token calculator to price your actual input/output mix, then compare current OpenAI pricing, Anthropic pricing, Google pricing, and DeepSeek pricing. A subscription can make a premium model feel cheaper until a session cap interrupts work; direct API billing is more measurable but transfers every token to the buyer.
OpenAI’s current direct ladder spans Luna, Terra, Sol, and Astra, while Anthropic’s runs from Haiku through Sonnet, Opus, and Fable. Google currently lists a promotional Gemini 3.8 Flash rate through December 31, and DeepSeek separates weekday peak from off-peak billing. The live components above come from the daily canonical dataset and supersede this article if an official rate changes.
Do not translate a commenter’s percentage of a Claude Max or ChatGPT allowance into an API dollar figure. Consumer subscriptions, coding-agent products, and direct APIs use different allowances, rolling windows, credits, multipliers, service tiers, and cache rules. Compare plan terms on the subscription pricing page and token billing on the provider pages.
The thread also reinforces a result from our single-model routing analysis: one evaluated default plus explicit fallbacks is easier to operate than an opaque router that changes behavior, cache affinity, and cost without clear evidence.
Choose a default model in four tests
- Select 20–50 recurring tasks with pass/fail checks: tests, accepted edits, extraction accuracy, or reviewer approval.
- Run the cheapest plausible default and one stronger control with identical prompts, tools, context, and output caps.
- Record token cost, subscription interruptions, latency, retries, tool failures, and reviewer minutes.
- Keep the cheaper default when it passes; escalate only on a failed validator, high-risk task, or demonstrated quality gap.
Teams evaluating GLM as the economical first route can compare the Z.ai coding plans with API and subscription alternatives using the same task set.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored Z.ai link at no extra cost to you. It does not affect this analysis.
Labs decision: preference evidence stays separate
Several named models already have an exact-route result or an explicit availability status in AI Pricing Guru Labs. We are not turning mentions into leaderboard points or paying to rerun a model merely because it appeared in this discussion.
The Labs coverage note records the blocker: the thread has no fixed task set, identical prompts, controlled routes, complete token and tool ledger, machine grader, or acceptance threshold. A reproducible default-model study needs a published mixed coding suite, matched routing rules, and total cost per accepted task.
Who benefits—and who loses
Developers with strong tests, clear plans, and short feedback loops benefit most from cheaper defaults because they can detect misses early. Teams with expensive failures, weak validation, or highly nuanced tasks may save money by starting with a stronger model.
Premium providers lose wasteful routine traffic when buyers adopt escalation rules. Cheap and open-weight models gain the default slot—but only if their speed, availability, and accepted-task rate hold under the buyer’s real harness.
Bottom line
Hacker News did not choose one best AI model. It surfaced a more durable answer: make the default cheap enough to use freely and capable enough to pass routine checks, then pay for premium intelligence only when the task proves it needs it.
Sources: the Hacker News Ask HN thread, HN Algolia item archive, Hacker News API item, official OpenAI API pricing, Anthropic model pricing, Google Gemini API pricing, DeepSeek models and pricing, and xAI model pricing. Thread state and official prices checked September 12, 2026 between 17:33 and 17:35 UTC. Provider mention counts are AI Pricing Guru’s text match across top-level replies and are not votes or unique-model endorsements.