A guide to operating a model family

OpenAI has published a model guide for GPT-6 that focuses on deploying the family rather than announcing another model. Dated 2 October, it explains how to choose among GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, set reasoning effort and manage long-running work. The central message is that capability, latency and cost should be tuned to the job instead of selecting the largest model for every request.

The guide maps Astra to the hardest reasoning work, Sol to complex coding, research and computer use, and Luna to focused high-volume tasks. It recommends starting with representative work and measuring task success, latency and cost per successful task. That last measure is important: a cheap request that needs several retries or extensive human correction can cost more than a stronger model that completes the job once.

Caching and compaction shape operating cost

OpenAI recommends prompt caching for stable instructions, tool definitions and reference material that recur across requests. Cached input can receive discounts of up to 95 per cent depending on the model. Stable content should appear before changing task details, and tool schemas should remain consistent so reuse is not broken. Teams are also directed to dashboards and diagnostics that show where cache performance falls.

For long conversations, compaction reduces the amount of context carried forward while preserving state needed to continue. Both mechanisms require testing. A compacted history can omit a critical decision, while an unexpected tool change can destroy a cache hit rate and alter costs. Applications should monitor the quality of retained state and include cache writes, long-context rates and failed reuse when estimating a complete workflow.

Reasoning is a dial, not a status symbol

The guide suggests low effort for extraction and small edits, medium effort for planning and comparison, and high effort for difficult debugging or careful review. Extra-high or maximum settings should be retained only when the improvement justifies added time and cost. GPT-6 applications can change effort during a conversation without necessarily breaking cache reuse, allowing a workflow to spend more thought only where it matters.

A production system needs policy around that dial. Letting a model raise effort without a budget can create unpredictable latency and spending. Fixing every call at low effort can reduce reliability on difficult steps. Teams can classify task stages, set ceilings and measure when higher effort changes the final result. The best setting is the lowest one that consistently meets the acceptance criteria for that stage.

Long-running agents need steering and boundaries

OpenAI highlights steering, asynchronous tools and delegation for tasks that span hours or days. Steering lets an operator update instructions during a run, while asynchronous tool calls allow unrelated work to continue as a slow process completes. GPT-6.1 Sol can delegate independent subtasks in a beta multi-agent workflow. These features reduce idle time, but they also make sequencing and responsibility more complex.

An application should identify which work can continue after a correction and which work depends on the pending result. A new instruction does not undo an action already taken, and a delegated agent can produce output based on earlier assumptions. Logs need to preserve the instruction version, tool results and hand-offs that shaped the outcome. Approval boundaries should be explicit for production changes, purchases, messages and other consequential actions.

Useful guidance still needs local evidence

The operational details deserve as much attention as headline intelligence.

Teams should also design evaluations around changes over time. A prompt and model combination that performs well at launch may behave differently after tools, reference data or product defaults change. Versioned test cases and recorded acceptance criteria make regressions visible. Human reviewers can sample both successful and failed runs, checking whether the agent reached the correct result for defensible reasons. Production monitoring should connect technical measures such as latency and cache hits with business measures such as completed tasks and corrected errors. This prevents optimisation for a cheaper call from quietly reducing the quality of the overall service.

The guide also covers computer use, recommending APIs or connected tools when they can perform a step directly and screen interaction when an application lacks an API. That hierarchy is sensible because structured tools are usually easier to validate. Computer use should receive additional checks around navigation, form submission and confirmation before irreversible actions. A visual success message is not sufficient evidence that the intended state changed correctly.

OpenAI's guide is valuable because it treats model deployment as systems engineering. Choosing a model, arranging prompts, preserving context and supervising tools all affect the final result. The recommendations should become hypotheses in a team's own evaluation suite rather than universal defaults. Representative tasks, failure analysis, cost monitoring and data controls remain the route from a capable GPT-6 demonstration to a dependable production service.