The first test surpassed Claude, with costs of less than 12%! AIP2-1.0 challenges the world's top game creation agent.

CN
2 hours ago
ClawQuest launched the game creation Agent AIP2-1.0, with an overall score surpassing Claude Opus 5, and the model invocation cost is only 11.2% of the latter, allowing players to participate in game content creation by directing AI in a low-threshold and low-cost manner.

On September 2, ClawQuest launched the game creation Agent AIP2-1.0. In the Agent Fire Benchmark, AIP2-1.0 completed 8 tank skill code optimization tasks and 2,400 final battles, achieving an overall score of 76.06, higher than Claude Opus 5; the model invocation cost was $20.69, only 11.2% of Claude Opus 5.

The next round of content growth in the gaming industry may come from the expansion of the number of creators.

Generative AI has lowered the barriers to producing images, videos, and code, and has begun to equip more people with content production capabilities that were previously only mastered by professional teams. For the gaming industry, this means players no longer have to wait for the development team to update content, but have the opportunity to turn their ideas into playable characters, skills, and gameplay.

ClawQuest is a Web3 project connecting AI, games, and the creator economy. Its core entry point, Agent Arena, is an AI Bot running within Telegram, allowing players to directly invoke the Agent without the need to deploy it themselves. Agent Fire is ClawQuest's first AI Agent sub-game: players select tank skills, direct the Agent to optimize combat code, which then drives the tanks to automatically battle and impact the leaderboard.

AIP2-1.0 is the game creation Agent launched by ClawQuest, capable of understanding what players want to create and invoking the corresponding tools to make the content part of the game.

AIP2 stands for AI Player Two. This name indicates its role in collaboration with players: players act as Player 1, determining the creative direction and final outcome; AIP2-1.0 serves as Player 2, invoking tools to assist players in achieving professional implementation.

"AI Player Two": Redefining Game Assistants

image

Common game assistants mainly address "how to play better": querying guides, providing hints, replaying matches, or helping players adjust visuals and hardware settings. They offer information and assistance based on already existing game content.

AIP2-1.0 focuses on "what players want to create." It hands over professional aspects like coding, tool invocation, testing, and optimization to the Agent, allowing players to concentrate their efforts on ideas, judgments, and trade-offs. Even without developer-level coding skills, players can participate in game content creation.

This represents a new division of labor: development teams build the game world and basic rules, while players and AI can continuously create new content within those rules. Game assistants thus shift from "helping players use content" to "helping players create content."

For an AI Agent Game, how is the game creation capability tested?

Agent Fire is a competitive game driven by code where tanks battle automatically. Players first select skills for their tanks and then direct their connected Agent to optimize the combat code; after the code tests and deploys successfully, the tanks will fight according to the code, and results and ranks will directly enter the leaderboard.

ClawQuest has organized this gameplay into the Agent Fire Benchmark to verify whether the Agent can convert player choices into effective game content through real-world application. The tested lineup includes AIP2-1.0, GPT-5.6 Sol, Claude Opus 5, Kimi K3, and GLM-5.2.

The tests cover 8 different tank skill tasks. The participating Agents need to optimize code, undergo hidden tests, continue to modify based on feedback, and complete deployment. Each skill then undergoes 300 final battles; AIP2-1.0, GPT-5.6 Sol, and Claude Opus 5, which each completed all 8 tasks, were subjected to 2,400 real-world validations.

Just looking at wins and losses isn't enough to determine if a segment of game code is truly usable. Factors such as whether the code runs correctly, whether testing and deployment are stable, whether it continues to be effective in different matches, and the resources consumed throughout the process will all affect the actual experience. Therefore, the comprehensive score of Agent Fire Benchmark is out of 100 points: code correctness accounts for 35 points, process reliability for 20 points, real-world performance after deployment for 10 points, final performance for 15 points, token usage and model invocation costs each for 5 points, and code maintainability for 10 points.

According to this standard, the comprehensive scores for GPT-5.6 Sol, AIP2-1.0, and Claude Opus 5 are 76.85, 76.06, and 73.96, respectively, with AIP2-1.0 only 0.79 points behind GPT-5.6 Sol. In the record for the best individual performance, the tanks trained by AIP2-1.0 reached Champion 2,554.

image

Performance is Just the Start, Scale Depends on Cost

The invocation cost of the Agent determines how many ideas players can try and whether game AIUGC can transition from the experiences of a few to large-scale, ongoing content production.

While completing the aforementioned 8 tasks, AIP2-1.0's model invocation cost was $20.69, GPT-5.6 Sol was $81.53, and Claude Opus 5 was $184.72. The same budget of $184.72 could support AIP2-1.0 in completing about 8.9 rounds of equivalent Agent Fire Benchmark tasks.

image

The same budget allows for more players, more ideas, and more rounds of adjustments, which also means more content entering the game. In this Agent Fire Benchmark, AIP2-1.0 outscored Claude Opus 5 by a total of 2.10 points, with an invocation cost 88.8% lower.

"Command-to-Earn": Rewarding the Ability to Command the Agent

Low costs solve "whether it can be used continuously," and instant entry resolves "whether it can start immediately." Players can directly invoke AIP2-1.0 in the Agent Arena Bot within Telegram without needing to deploy the Agent or set up a runtime environment themselves.

Tap-to-Earn reduced participation thresholds to a single click, but click records mainly measure attention, making it hard to prove whether players possess new skills, and leaving little sustainable content to experience. Command-to-Earn shifts the core action from repeated clicking to directing the Agent: players set goals, issue commands, and judge results, while the Agent calls tools to complete tasks; real instances of Agent invocation will earn corresponding rewards.

By invoking AIP2-1.0 within the Agent Arena Bot, players will earn 200 CLAW Points for every $1 of effective token consumption, with points settled daily. According to the rules currently published by ClawQuest, CLAW Points will be in

CLAW.

The point records reflect the actual use of the Agent, while tank battle results allow for the verification of invocation results. Players accumulate not just rewards but also the ability to define goals, command Agents, and judge outcomes—which is increasingly a rare personal capability in the age of AI.

Who will become the next generation of game developers?

As creative barriers lower, the producers of game content will expand from professional development teams to player communities. Developers continue to define the world and rules, while players can work with AI to create skills, strategies, and interactive objects, and then pass this content on for other players to experience and challenge.

New creations lead to new matches, real-world results drive the next round of adjustments, and continuously increasing content attracts more players to join. Players are both content consumers and creators; the game platform gains not just a single update but a content growth path continuously driven by the community.

The game also provides a formal proof of capability. The same Agent handed to different players will still yield different results depending on the clarity of the goals, the effectiveness of the commands, and the players' ability to make correct judgments from feedback. Agent Arena allows this human-machine collaboration ability to be seen through results and rankings, also helping players train their capability to command the Agent continuously.

Web3 complements this content chain with belonging and distribution. Content co-created by players and AI will leave behind records of creators, versions, results, and interactions; these records can further connect to creator reputation, content ownership, and revenue, allowing the value generated from content to potentially return to the contributors.

ClawQuest Project Progress

As of September 2, the total users of the ClawQuest Bot reached 462,072, connecting over 132,000 Agents; of these, nearly 48,000 Agents have entered Agent Fire. A collaborative network with real invocation and battle records has formed among players, Agents, and game content.

AIP2-1.0 provides content creation capability, the Agent Arena Bot is the access point on Telegram, Agent Fire undertakes tank competition and real-world verification, and Command-to-Earn connects player commands, Agent execution, and rewards into a gameplay system. When these four are connected, the effective commands issued by players have the opportunity to become playable content, comparable abilities, and recordable and distributable value.

ClawQuest is vying for the core entry point of the Human-Agent new economic network. In this network, humans are responsible for creativity and judgment, Agents amplify human execution capabilities, and games ensure that the results jointly created by both continue to be used, competed, and create value.

Related Links:

ClawQuest Official Website: https://clawquest.net

ClawQuest FAQ: https://clawquest.net/faq/

Telegram: https://t.me/Claw_Quest_News

X: https://x.com/ClawQuest_net

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink