This week split AI pricing into three different markets: newly launched models without complete rate cards, open weights whose infrastructure still has to be paid for, and specialist systems that claim better task economics than frontier APIs. The live table, chart, and calculator above use verified published models; they do not invent rates for Qwen3.8-Max or Shieldstral.

StoryWhat changedBuyer impact
Qwen3.8-MaxAlibaba opened API access to a 1M-context frontier model and promised weightsCapability can be tested now, but volume budgets should wait for an official rate
ShieldstralMistral released an Apache-2.0 multimodal moderation modelNo weight-license fee, but GPU, review, and operations remain part of cost
Castform retrievalA tuned Qwen3.5-4B beat Sol on one vendor-run retrieval evaluationSmall specialists may win repeated, measurable workflows after training cost is included
Grok Imagine Video 1.5xAI added text, image, and reference modes plus native 1080pCreative teams gain flexibility while paying more per generated second

Qwen3.8-Max launched before its rate card

Alibaba released Qwen3.8-Max through QwenCloud with text and image input, parallel tool use, adjustable reasoning effort, and a one-million-token context window. OpenAI- and Anthropic-compatible interfaces make evaluation straightforward, while open weights are scheduled to create a second deployment route.

The missing fact is the official token rate. Alibaba’s public pricing document still points buyers to the previous Max generation, so the Qwen baseline above is a comparison aid rather than a proxy price. Teams can test quality now, but should cap tokens and avoid approving a production budget until billing, caching, regions, and the exact model identifier are documented.

Read the full Qwen3.8-Max pricing analysis and use the Qwen3.7-Max model page as the published baseline.

Shieldstral changed moderation deployment economics

Mistral’s Shieldstral 1.0 is an Apache-2.0, 3B-class model for text, image, and combined moderation. Developers express each policy as a natural-language yes-or-no question, which may reduce retraining when rules change. Mistral says the checkpoint can fit on a single 16GB GPU.

That does not make moderation free. Mistral announced no hosted Shieldstral rate, and self-hosters still carry GPU capacity, autoscaling, observability, policy evaluation, false-positive review, and model-update costs. Run it in shadow mode and select on cost per correctly moderated item—not download price.

See the Shieldstral launch analysis and current Mistral API pricing.

Castform made the strongest cost claim

Castform and Neon reported that a reinforcement-learning post-trained Qwen3.5-4B specialist earned a higher composite reward than GPT-5.6 Sol on their agentic-retrieval evaluation with a roughly 95x measured rollout-cost advantage. The base 4B model performed poorly, making post-training—not model size alone—the key result.

The benchmark is vendor-run and does not disclose enough detail for independent reproduction. Its ratio is cost per benchmark rollout, not a universal API discount. A real break-even calculation must include training data, post-training, embeddings, database compute, idle serving capacity, evaluation, and frontier fallbacks.

Our Castform retrieval cost analysis lists the caveats. Compare verified routes on the Castform pricing page and OpenAI pricing page.

Grok Imagine Video became broader and dearer

xAI expanded Grok Imagine Video 1.5 to text-to-video, image-to-video, and reference-to-video. Text and image modes now support native 1080p; reference mode is capped at 720p and can use preset voices for selected US partners.

The new version costs more per second than its predecessor at shared resolutions, and xAI’s text-model batch discount does not apply. Video teams should test short, low-resolution candidates, promote only approved prompts to final renders, and track cost per usable clip including retries, storage, and post-production.

The Grok Imagine Video 1.5 pricing analysis covers the resolution tiers. For a separate voiceover layer, test ElevenLabs on the same approved clips and compare finished-asset cost.

Affiliate disclosure: we may earn a commission from the sponsored link above. It does not affect our analysis.

What buyers should do now

Do not fill missing rate-card cells with guesses. Put Qwen3.8-Max behind a feature flag, measure Shieldstral on labeled private data, reproduce Castform’s economics on held-out retrieval traces, and canary Grok Imagine Video 1.5 at each delivery resolution.

Use the AI token calculator for models with published token rates, then add training, tools, storage, review, retries, and infrastructure. This week’s launches reward teams that track cost per accepted result rather than comparing model headlines.

Sources: Qwen’s Qwen3.8-Max announcement, Mistral’s Shieldstral launch, Neon and Castform’s retrieval case study, and xAI’s release notes.