A knife can save lives and also take lives.

CN
1 hour ago

A knife can save lives, but it can also take lives.

AI empowers good people, and it equally empowers bad people.

When the North Korean hacker group Lazarus systematically integrates AI into the attack chain, AI automatically scans for zero-day vulnerabilities and generates exploits in bulk; when the dark web’s WormGPT openly prices its services, allowing you to generate phishing emails and malicious scripts in bulk for a monthly fee of $60.

We must admit that bad people have turned AI into a low-cost, large-scale weapon of destruction.

Faced with AI attacks that are armed to the teeth, the only way to defeat magic is with magic.

1/ A good piece of news

The Global Cybersecurity Alliance (GCSA) has just announced that its GCSA Agent achieved a success rate of 91.3% in the CyberGym benchmark tests.

Comparing to the previous CyberGym leaderboard:

First Place: GPT-5.5-Cyber + OpenAI Agent —— 85.6%

Sixth Place: Grok 4.6 + Grok Build —— 79.7%

Seventh Place: Grok 4.5 + Grok Build —— 79.0%

The GCSA Agent is based on Grok 4.5 and 4.6. With the same underlying models, it has gained an additional 12 percentage points, surpassing both the original Grok solutions and the specifically fine-tuned OpenAI GPT-5.5-Cyber.

This indicates that in security tasks, the design of the agent framework is as crucial, if not more so, than the underlying model.

2/ What did the GCSA Agent do right?

The GCSA did not simply throw vulnerability descriptions at the large model; instead, it built a self-contained, autonomous vulnerability mining process:

Retrieve relevant code from thousands of files and millions of lines of code

Identify attack surfaces

Propose vulnerability hypotheses

Generate test inputs (PoC candidates)

Execute in a real environment

Analyze crash or execution feedback

Iterate and optimize PoC based on feedback

Verify: trigger before the patch, do not trigger after the patch

Its breakthrough lies not in the model (which is still Grok), but in the framework: how to retrieve code, propose hypotheses, parse crash logs, control the pace of iteration, and when to decisively abandon the current idea.

The model is the engine, the framework is the driving system.

The GCSA did not build a stronger engine but created a better driving system.

3/ And then?

The achievement is just the beginning; real-world security work goes far beyond this. What the GCSA truly wants is to enable AI agents to participate in the entire security lifecycle:

Autonomously discover vulnerabilities without being told where they are

Not just generate PoCs but also analyze the scope of impact

Not just report issues but also suggest and even generate fixes

Not just operate in silos but continuously monitor new risks arising from code changes

When attackers have reduced the cost of crime to nearly zero using AI, the only feasible way for the defense is to have AI patrol continuously for you, 24/7.


免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink