A how-to guide that reads like a cost playbook

On October 2, 2026, OpenAI published "A model guide for the GPT-6 family," a documentation page aimed at developers building on its newest model lineup. On its face it is a how-to guide. Read closely, it is something else: a vendor-documented statement that cost discipline, not prompt cleverness, has become the core skill of using frontier AI.

The guide's first numbered section is about production economics. It tells builders to cut context the task does not need, run independent tasks in parallel, estimate cost per successful task before deploying, and use prompt caching and compaction to manage both context and cost. Only after the money does it talk about matching models to workloads, then prompts, then long-running agent work.

For non-specialists, that ordering matters. It says the expensive part of building with AI is no longer getting a model to work at all. It is running it efficiently, at scale, without wasting tokens or human attention.

Caching: the documented economics

The guide's most concrete number is about caching. When a request reuses the same prompt prefix as an earlier request, the model can skip recomputing work it has already done. Cached input tokens, the guide says, cost "up to 95% less" than uncached input tokens, depending on the model.

OpenAI's prompt-caching documentation, retrieved separately on October 2, 2026, is consistent and adds detail: on GPT-5.6 and later models, cache writes cost 1.25 times the standard input rate, while subsequent reads cost 0.1 times that rate on most supported models and 0.05 times on GPT-6.1 Sol. That 0.05 multiple is exactly the 95% discount. The caching discount is therefore not a marketing rounding; it is a documented, model-specific price structure.

The same guide and documentation confirm a second practical capability: developers can change the model's reasoning effort mid-conversation in the API without breaking the cache, provided the request uses the documented configuration update. For a small team, this means you can run routine turns cheaply at low effort and escalate to deeper analysis only when a task needs it, without paying to reprocess the whole conversation.

One discrepancy deserves flagging. Our earlier coverage reported OpenAI caching savings of up to 90%. The new guide says up to 95%, qualified with "depending on the model." Both figures are vendor-published. The difference likely reflects different models and qualifications, since the Sol announcement separately describes cached input at 95% less than standard input pricing for that model. We do not assert which earlier figure was wrong; the dated, primary-sourced figure as of October 2, 2026 is 95% for supported models.

Matching the model and effort to the task

The guide frames model choice as an "intelligence/price tradeoff" and names three tiers: GPT-6 Astra for the hardest reasoning work, GPT-6.1 Sol for complex coding, research, and computer use, and GPT-6 Luna for focused tasks at scale such as extracting invoice fields, classifying requests, or producing structured summaries.

Reasoning effort is split into low, medium, high, and extra high or max, each tied to a task type. The guide explicitly warns against leaving extra high effort on by default, telling builders to keep it "only if the improvement justifies the added time and cost."

This is vendor guidance, but it is unusually self-limiting advice: OpenAI is telling customers that its most expensive settings are usually the wrong default. For builders on tight budgets, the practical takeaway is to benchmark task success, latency, and cost per successful task on a lower tier before paying for the top one. OpenAI's Sol announcement from September 29, 2026 supports the economics, stating Sol nearly matches Astra's intelligence on agentic coding, computer use, and professional work at one-fifth of Astra's standard token prices. Those evaluation scores are vendor-run, so treat them as claims to verify on your own workload, not settled results.

Steering, subagents, and a changed idea of prompting

For long-running work, the guide documents capabilities that were recently unusual: mid-turn steering through the Responses WebSocket API, where updates are queued and do not cancel running tools; asynchronous tool calling, where the model keeps working while a slow task such as a test suite runs; and multi-agent delegation, in which GPT-6.1 Sol assigns independent subtasks to subagents and combines their findings. Multi-agent workflows in the Responses API are, per the guide, currently in beta.

Computer use, letting the model interact directly with websites and desktop applications, is documented as available across all three models: Astra, Sol, and Luna.

One passage reads less like documentation and more like a philosophy shift. Quoting an Eric Provencher post linked from the guide, OpenAI writes: "Models have gotten much better at understanding nuance and ambiguity, so overly specific guidance can now hinder results where it previously helped."

Attribution note: this wording appears on the retrieved guide page, dated October 2, 2026, attributed there to Eric Provencher, Developer Experience at OpenAI. We could not independently retrieve the original X post at the linked address, so we treat the guide's rendering as the source of record rather than confirming the post verbatim. The substance is consistent with the guide's broader advice to replace rigid prompt recipes and blanket "always ask" rules with clear decision boundaries stating which actions the model may take independently and which require human approval.

That last point carries weight beyond cost. A vendor telling builders to explicitly define what an AI may do alone, and what needs a human sign-off, is codifying human oversight as an engineering requirement, not an afterthought.

Case studies are claims, and the advice is self-serving

The guide closes its production section with four customer examples: Harvey drafting legal documents with more context, Cognition's Devin producing test evidence engineers can review, Hex turning business questions into dashboards, and Invideo using Astra for video editing. These are case studies, not independent audits. Invideo's reported "roughly three times the success rate" on color-grading tasks and its editors creating about 50 effects in a day are self-reported vendor claims from OpenAI's pages, and we label them as such. Nothing here has been verified by a third party.

It is also fair to note the self-interest. A guide that teaches efficient model matching, caching, and tiering keeps spending inside OpenAI's ecosystem and makes its price list the frame of reference. The documentation is honest about tradeoffs, but the benchmarks and case studies are selected by the seller.

Still, the direction is real and documented. When a frontier lab publishes that the most expensive model and the highest reasoning effort are usually the wrong defaults, and that caching can cut repeated input costs by up to 95% on supported models, it lowers the barrier for small teams to build with capable AI, provided they read the pricing footnotes themselves. The skill the guide teaches, deciding what to cache, what to delegate, and what a human must approve, is one that transfers to any provider's stack.