xAI challenges GPT-5.6 with Grok 4.6, long task intelligent agents will self-check their work.

CN
1 hour ago
Grok 4.6 focuses on intelligent agents for long tasks and visualization project capabilities, achieving parity with GPT-5.6 Sol on the comprehensive intelligence index and maintaining highly competitive API pricing.

Author: xAI

Translation: Deep Tide TechFlow

Deep Tide Overview: Grok 4.6 emphasizes intelligent agents that can operate for extended periods and across multiple steps, reaching parity with GPT-5.6 Sol on the comprehensive intelligence index, which is of practical reference value for developers’ selection. xAI also provides double usage in the first week, indicating that competition among large models is shifting from one-time Q&A to agent capabilities that can actually deliver projects, which is worth tracking for practitioners interested in the intersection of AI and cryptocurrency.

Today we are releasing Grok 4.6. Grok 4.6 is built on Grok 4.5, with a special focus on long-running agents and more ambitious interaction and visual tasks. It can continuously handle complex tasks across multiple steps, whether researching topics, analyzing information, working across code repositories, or transforming ideas into polished applications or works.

Grok 4.6 has reached frontier levels on multiple agent coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is acomposite score of nine benchmarks.

Image: Source: xAI; competitive data derived from system cards or benchmark rankings publicly released by various developers.

Grok 4.6 is now available in Cursor and Grok Build. Within the first week, we are offeringdouble included usage in Grok Build and Cursor, making it easy for you to start using 4.6 immediately.

Training Grok 4.6

Grok 4.6 has undergone a longer supplementary training than Grok 4.5, utilizing curated model-generated data for reasoning and advanced technical concepts, as well as high-quality engineering data, and has improved its optimizer and training recipe. This lays a stronger foundation for subsequent SFT and RL phases.

Subsequently, we regenerated SFT trajectories using Grok 4.5, covering reasoning tasks, agent frameworks, and fields such as STEM, software engineering, and knowledge work, filtering out problematic trajectories through model-based checks. The resulting SFT checkpoints exhibit strong performance and improved behavior.

Grok 4.6 has been trained on a wide range of agent reinforcement learning tasks, including knowledge work, general coding, and specific environments targeting kernel optimization, web development, and computer-aided design.

Turning Grand Ideas into Usable Projects

We tested Grok 4.6 on projects aimed at expanding their scope and long-term working capabilities. We found that the model is particularly adept at transforming broad product ideas into usable prototypes. It can explore unfamiliar domains, build application structures, implement core interactions, and continuously optimize results through multi-round feedback.

In longer task trajectories, we are also beginning to see more self-testing and validation behaviors, with the model checking its own work before proceeding.

In visual and interactive projects, Grok 4.6's initial outputs are stronger than what we typically observe on Grok 4.5. Given a specific product idea, it can establish structures and visual languages for the application all at once. This makes it particularly useful in projects where the fastest way to achieve good results is to first create substantive content and then iterate through cycles.

Safety and Capabilities

The safety measures of Grok 4.6 have been improved and calibrated based on model capabilities.

Our safety stack is designed to maximize utility and safety in legitimate use cases, making Grok 4.6 useful and safe for tasks like vulnerability patching, accelerating engineering design cycles, and enhancing AI research.

Our safety assessment work reflects Grok 4.6's expanded capabilities, including what is our most extensive capability and safety calibration pre-deployment testing ever, as well as extensive post-deployment and third-party testing.

Benchmark Testing

Image: Comparison of scores of Grok 4.6 and competitive products on major benchmarks. Source: xAI; each benchmark's highest score is highlighted in bold, with third-party model scores taken from their own reports or publicly obtainable best results.

From the publicly available data, Grok 4.6 has matched or exceeded GPT-5.6 Sol Max onCursorBench v3.2 (69.9%), GDPVal-AA v2 (1753), AA-Briefcase (1577) andAPEX-Agents (57.5%), but still has a gap onTerminal-Bench v3.0 (26%) andDeepSWE v1.1 (65.9%); Fable 5 Max maintains a slight lead in most projects.

Getting Started with Grok 4.6

Grok 4.6 is now available in Cursor and Grok Build. It is also accessible via API and other partners, such as OpenRouter, Vercel, and Cloudflare.

Pricing starts at$2 per million input tokens and $6 per million output tokens. Additionally, there is a fast version priced at double the standard edition.

In the first week, we are offeringdouble included usage in Grok Build and Cursor, making it easy for you to start using 4.6 immediately.

Creating an API Key

Start building with Grok 4.6 immediately through the SpaceXAI API.

Read the documentation and integrate Grok 4.6 into your technology stack.

Try Grok Build for Free

Visit x.ai/build now to get started.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink