xAI has released Grok 4.6, a new model it says is designed to sustain work across longer, multi-step tasks while improving the quality of interactive and visual project work. The company positions the release as an advance on Grok 4.5 for research, knowledge work, software development and turning broad product ideas into early working applications.
The announcement is a material expansion of the Grok model line because it combines a new flagship model, immediate developer availability and explicit pricing. Grok 4.6 is available through Grok Build and Cursor, via the xAI API, and through partners including OpenRouter, Vercel and Cloudflare. xAI says it is offering double the included usage in Grok Build and Cursor for the first week of availability.
A model pitched for work that does not end after one answer
The central claim in xAI’s launch is that Grok 4.6 is better at keeping hold of complex tasks that unfold across many steps. Rather than limiting the release to a benchmark improvement, the company describes intended uses such as researching an unfamiliar topic, analysing information, working through a codebase and producing a polished work artefact or application.
That emphasis puts the model in the increasingly competitive market for agentic systems: products expected to plan, use tools, make changes and continue through a sequence of dependent actions. In practice, long-running performance depends on much more than the underlying model. The agent harness, tool permissions, context management, review controls and fallback behaviour all affect whether a multi-step workflow is reliable enough to use in production.
xAI also highlights interactive and visual work. It says Grok 4.6 produces stronger first passes for applications where a user begins with a concrete product idea, then iterates. The company’s examples describe establishing structure and visual language in an initial response, followed by refinement through feedback. These are useful capabilities for prototyping, but they are product claims rather than independent evidence that every generated interface or application will meet a team’s design, security or accessibility standards.
Benchmark comparisons require some care
xAI reports results across coding, agent and knowledge-work evaluations, including Artificial Analysis’ composite Intelligence Index, GDPVal-AA, DeepSWE, CursorBench and FrontierCode. It says Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis index, while its table shows different models leading on different tests.
These figures give prospective users a useful starting point, especially for comparing model capability within an announced release. They should not be read as a direct forecast of business outcomes. Benchmarks use controlled tasks, selected tool configurations and defined scoring rules; a production workload can introduce proprietary data, ambiguous requirements, integrations and human hand-offs that do not appear in the evaluation.
xAI says competitor figures are drawn from developers’ published system cards or benchmark leaderboards, and that the best score for each evaluation is highlighted. Organisations evaluating Grok 4.6 should test it on representative internal tasks and assess cost, latency, failure modes and review requirements alongside published scores.
Training and safety claims
The company says Grok 4.6 received a longer supplemental training run than Grok 4.5, drawing on curated model-generated reasoning and technical material, engineering data, and changes to its optimiser and training recipe. It says Grok 4.5 was used to regenerate supervised fine-tuning trajectories across reasoning efforts, agent harnesses and domains including STEM, software engineering and knowledge work, with model-based checks used to filter problematic traces.
xAI also describes reinforcement-learning tasks spanning knowledge work, coding and specialised environments such as kernel optimisation, web development and computer-aided design. These details help explain the release’s focus on extended technical work, although xAI has not used this announcement to publish a full technical report that would allow external researchers to reproduce or independently validate the training approach.
On safety, xAI says it improved and calibrated Grok 4.6’s safeguards in line with the model’s capabilities, and conducted its broadest pre-deployment testing suite so far alongside post-deployment and third-party testing. The announcement says the safety stack is intended to support legitimate work such as vulnerability patching, engineering design and AI research. It does not provide all underlying evaluation artefacts, so buyers with regulated or high-risk uses should seek the relevant security, data-processing and deployment documentation before relying on the model.
Availability and pricing put the launch in developers’ hands
Grok 4.6 can be used in Grok Build and Cursor from launch, while API access and third-party availability give development teams several routes to trial it. xAI lists pricing from US$2 per million input tokens and US$6 per million output tokens. It also lists a faster variant at twice those prices.
The published rates are only one part of a deployment calculation. Agentic workloads can make repeated calls, carry large contexts and use tools over a long trajectory, so total cost may vary substantially by task design. Teams comparing providers should measure realistic end-to-end usage and confirm the current partner pricing, rate limits, data terms and regional availability for the route they plan to use.
What to watch after launch
Grok 4.6 is a substantial release because it joins model capability claims with broad developer distribution and a clear focus on agentic work. For xAI, the more meaningful test will be whether developers see consistent gains on lengthy tasks rather than only in tightly scoped demonstrations.
For prospective users, a controlled pilot is the sensible next step: choose a bounded workflow, maintain human approval for consequential changes, compare outputs with the current model or process, and review cost and safety behaviour over multiple runs. xAI’s announcement makes Grok 4.6 available now; the practical evidence will come from how it performs under those real-world conditions.