Meta Debuts AI Coding Agent Muse: Here’s How It Compares to Claude Code and Codex

CN
Decrypt
Follow
1 hour ago

Meta is the latest tech giant to ship a coding agent, racing to compete with leading AI behemoths Anthropic and OpenAI.


"We're excited to release Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, our newest model," the company wrote in an official announcement. "This marks our next step toward the frontier, with larger and much more capable models on the way."





As an agentic coding tool, Muse Code is built for software engineering across large repositories. Per Meta, it "takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results. It can coordinate multiple persistent subagents for each task, solving difficult problems faster, more accurately, and with less intervention."


The detail that stands out is the runtime. Muse Code logs every model call, tool run, approval, and edit to a local event log that acts as a single source of truth. "This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent can resume precisely where it stopped," Meta said. For long-running jobs, that's the feature that matters more than raw speed—and it's the part competitors haven't made a selling point.


It also ships with default skills. The "/plan" command turns a task into an approval-gated plan, while "/grill" stress-tests that plan until it holds up and "/goal" works toward successful completion of the objective similar to what Hermes does. Meta said it co-trained Muse Spark 1.2 with Muse Code so the core LLM and the agent work together in synergy.


The benchmarks, and the catch


Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1. Meta said it "significantly scaled up training compute on coding tasks while expanding training environment diversity, delivering improvements in code generation, complex debugging, and end-to-end developer workflows." The charts tell a clear story.


On Terminal-Bench 2.1, Muse Spark 1.2 with Muse Code scored 82.9%, behind Claude Code on Opus 5 at 86.7% but ahead of GPT-5.6 Terra on Codex (81.8%) and Grok Build (81.6%).




DeepSWE 1.1, which measures agentic coding capabilities, was closer: 59.3% for Muse versus 65.0% for Opus 5 and 64.8% for Codex. On Meta's internal coding bench, Muse hit 70.6% to Opus 5's 79.4%.


The speedup charts flip the order. Over 1,000-plus tool calls, Opus 5 posted the biggest gain versus baseline (about 74–75%), with Muse Spark 1.2 mid-pack at roughly 61–69% depending on the run. Meta's point is that the agent keeps improving as tool calls accumulate, the behavior you want from a long-horizon coder.




The most interesting demos are long-horizon and multimodal. In stress testing, Meta said Muse Code "iteratively optimized GPU kernels over 1,000+ tool calls (up to 24 hours) on Nvidia Hopper GPUs." That means it was able to improve over time.


There's also a visual-coding angle. In one demo, a user drops a fly-through video of a house into the terminal as an mp4, and Muse Code "interprets the video and produces a visually rich website with booking capabilities." Reading raw video into a working web app is the multimodal pitch Meta has been making across the Muse line.


See the launch thread:



The field is already crowded


That said, Meta is late to the fight. OpenAI's Codex already runs parallel cloud agents; DeepSeek has built its own rival to Claude Code and agentic tools like Hermes or OpenClaw are already good substitutes with more capabilities. Muse Code's edge is the crash-safe runtime and the subagent design, not benchmark supremacy.


The risk is the usual one for agentic coding: an agent that resumes after a crash and keeps calling tools for 24 hours is powerful and unpredictable. Meta is betting developers want that autonomy, and it's shipping now.


Muse Code is available for testing upon installation entering this command:

curl -fsSL https://dev.meta.ai/install.sh | bash


免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink