Humans are numb, weren't they all shouting to slow down the models a few days ago?
How come this model is being thrown out like it's free?
Yesterday Grok 4.7 faced off against MiMo v2.6, just to perform a warm-up act for these two big brothers today, right?
This early morning, OpenAI's GPT-6 Sol and GPT-6 Luna, as well as Anthropic's Claude Opus 5.5, were officially released.

I really want to ask, where is the slowdown that everyone talks about every day?
All are just players.
Moreover, I always feel this storyline is strangely familiar, so I flipped through some historical articles.

Last time, GPT-5.3 Codex clashed with Claude Opus 4.6; today, history is just repeating itself...
It's just that Claude Opus 4.6 has turned into Claude Opus 5.5.
GPT-5.3 Codex has turned into GPT-6 Sol.
The two companies are still the same companies, still those rivals who won't back down.
Let's talk one by one.
1. Claude Opus 5.5
Let's start with Claude.
To be frank, no matter how much I dislike Anthropic, no matter how I think this company is foolish, despite them sealing my account, objectively speaking, Claude is still the best and most comprehensive model in my heart.
I still remember the shock that Claude Fable 5 brought me.
This time, Claude Opus 5.5 is the first model in Anthropic's brand new Claude 5.5 series, and from my impression, it's rare for Claude to jump so many version numbers at once.
And from my brief experience (don't ask why I can still use it after being banned, it was a disposable account bought directly, and being used for about 2 hours will get me banned...), Opus 5.5 gives me a feeling of returning to Opus 4.6 in terms of communication, a very human-like quality.
Although their positioning for this thing is actually just a cost-effective model of the Claude series, it indeed reminds me of the past white moonlight.
I summarized the basic information, and it's roughly like this:
Model Name | claude-opus-5-5 |
Context | 1M Token |
Max Output | 128K Token |
Reliable Knowledge Cutoff | June 2026 |
Thinking | Adaptive Thinking, always on |
Default Effort Level | medium |
Input Price | $4 / million tokens |
Output Price | $20 / million tokens |
Cache Read | $0.2 / million tokens |
5-Minute Cache Write | $5 / million tokens |
1-Hour Cache Write | $8 / million tokens |
Compared to the previous generation Opus 5, the overall operating costs can drop directly by 40%.

At the same time, the output speed has increased by over 30% compared to Opus 5.
Currently, in some non-large tasks, Opus 5.5 can achieve performance at the level of Fable 5.1, while large and highly difficult tasks still require Fable-level models.

In coding tasks, it is basically the current SOTA.
For example, in Terminal-Bench 4.0, reflecting the ability to complete complex, multi-step tasks in a real terminal environment, Opus 5.5 refreshed the high score to 66.4%.
As long as it involves Coding, it's basically all SOTA, much like when Opus 5 first came out, it comprehensively surpassed Fable 5 in coding execution.
However, two things are somewhat different; the Terminal-Bench-Science task leans more towards scientific research, where Opus 5.5 scored 58.7%, while Astra had a higher score of 64.6%.
AutomationBench is the same; this task is actually about cross-software work, with Opus 5.5 scoring 40.0%, and GPT-6 Astra scoring 41.4%.
As for Computer Use, although GPT-6 Astra has no score, everyone knows this cannot be poor; it's definitely SOTA.
You can see the differences in the development directions of Anthropic and OpenAI; Anthropic is more inclined towards coding capabilities, while having strong aesthetics; they believe that coding is the cornerstone to AGI, whereas OpenAI focuses more on comprehensive agent capabilities, including reasoning, software operation, scientific research, and so on.
But frankly speaking, purely looking at a certain dimensional evaluation set does not truly assess a model's quality.
Just like you know, even in product development, there's a core aspect, which is your initial planning and architecture design ability; this is where Fable excels. It relies on your large parameters, on your world knowledge, on your intelligence emergence.
You say Opus 5.5 can indeed surpass Fable in certain specific executions, but to say that Opus 5.5 overall is stronger than Fable? Then I think you might as well believe that I am Qin Shi Huang.
Overall, Claude Opus 5.5 is basically the official distilled model of Fable 5.1.
It may be stronger in many vertical scenarios, with smaller parameter sizes, faster speeds, and being cheaper, but at the same time can lead to Token waste, which is actually a trap; once it's on high-difficulty tasks, some non-flagship models may endlessly think and try crazily but cannot resolve the issue, making it more costly than flagship models.

Though Opus 5.5 is indeed powerful, its output Token count per task has directly skyrocketed, so much of its intellect is gained through super long reasoning; the downside is that once it falls into a dilemma, it explodes on the spot.
So Anthropic also advises caution here.

When your results are crucial, always use the highest-level model.
Then you can take a look at the cases of Claude Opus 5.5 run by the bosses on X, they are quite stunning, especially in aesthetics and details, which have been significantly enhanced.

A comparison test by @notjazii between GPT-6 Astra and Claude Opus 5.5, showing in some scenarios, even more refined than GPT-6 Astra:
But the cost is that compared to GPT-6 Astra, it's slower and more expensive.

However, for Opus 5.5, comparing it with GPT-6 Astra is itself a compliment.
This time, Opus 5.5 has also strengthened communication.
Anthropic said one of the most common pieces of feedback they received about Opus 5 is that sometimes it writes in a roundabout way, with too much jargon and insufficient direct expression.
5.5 will place the most important information upfront, reducing odd phrasing and unnecessary explanations.
At the same time in creation, it also feels more like Claude 4.6.
So to summarize, Claude Opus 5.5 is a model that I think is fantastic.
Anthropic has moved many tasks that were only possible for flagship models down to a price point that can finally be used regularly, probably due to the immense pressure from OpenAI.
Their own cost guidelines directly recommend:
For day-to-day features, debugging, and code reviews, use Opus 5.5.
For truly critical results, or when there would be prolonged unsupervised usage, or if Opus 5.5 fails repeatedly, switch to Fable 5.1.
Fable is increasingly resembling an expert-level consultant.
Opus 5.5 feels more like a core employee.
This is Claude's new model today.
2. GPT-6 Sol and Luna
Finally, we get to something I can use normally.
GPT-6 Sol and Luna.
Because no matter how good Claude is, it's still a throwaway to me, not something I can use long term, there's no use.
So, regrettably, the mainstay is still GPT-6.
It’s particularly coincidental that when I wrote the article on GPT-6 Pro+MCP a few days ago, I even specifically mentioned:
“GPT-6 Sol is likely to be available soon. Once it's available, I might mindlessly switch to GPT-6 Sol, only switching to GPT-6 Astra for high-difficulty tasks.”
And finally, it has arrived.
Currently, it has already gone live in Codex and is directly usable.

In the past, the biggest criticism of GPT-6 Astra was that it was too expensive, and the capacity was completely insufficient.
This time, GPT-6 Sol and Luna are directly aimed at cost reduction.
Astra is still OpenAI's strongest model.
For the most difficult, most important tasks, where no compromise is desired, continue to use Astra.
However, the jobs in the real world have different scales, rhythms, and budgets.
So the significance of GPT-6 Sol and Luna is similar to that of Fable 5.1 and Claude Opus 5.5:
To bring the capabilities derived from Astra's training method down to faster and cheaper models.
Almost on the same day, both Claude and GPT-6 are discussing the same issue.
Who can get more intelligence for less money?
And that’s the Pareto frontier.
On the pricing front, the new generation is priced as follows.

Sol and Luna are 50% cheaper than the same model of GPT-5.6.

The most outrageous is this Luna; if we only look at input and output prices, it's actually cheaper than Deepseek.
Of course, we all know the real big cost comes from cache; we can see that the cache price is still about 3 times higher than Deepseek.

However, I think the most suitable scenarios for Luna are automation tasks behind various applications, just like the dozen information processing tasks that require large models behind my AIHOT, with requests numbering in the thousands daily.
In such scenarios, sometimes the cache hit rate is not particularly high; an aggregate rate of 30% is considered good, so in my view, the advantages of Luna become quite apparent.
The other basic information is as follows.

Interestingly, the effective knowledge cutoff dates are different.
Luna's is the latest, as of May 18.
Then there’s the one everyone is most concerned about, GPT-6 Sol. Frankly speaking, in terms of capabilities, it's somewhat below my expectations, but in terms of cost reduction, it's better than I anticipated.

We can see that it only slightly outperforms GPT-5.6 Sol in AA, but Opus 5.5 has dramatically taken the lead.
Although AA's benchmark is sometimes criticized by many for lacking precision in many details, the overall direction won't be significantly off.
Thus, GPT-6 Sol is more based on new model architecture, aiming to maintain a slight edge over GPT-5.6 Sol and significantly reduce costs.
For example, this AutomationBench, which measures the capability mentioned when discussing Claude, is one where Opus 5.5 performs strongly, measuring Agents' execution across 47 tools for real business processes in sales, marketing, operations, customer service, finance, HR, etc.
GPT-6 Sol xhigh:
33.2%.
Average cost per task:
$0.27.
GPT-6 Astra low is at 30.3%, but task costs are 3.9 times that of Sol.

Fable 5.1 plus the Opus 5 fallback is at 31.4%, costing at least 8.9 times what Sol costs.
Note that this comparison is still against Claude Opus 5, as Opus 5.5 was released almost simultaneously with GPT-6 Sol; thus, OpenAI's chart couldn't fit it in time, and even if it did, it wouldn't look great for OpenAI.
However, this chart clearly elucidates Sol's positioning.
The strength of GPT lies in its token efficiency, and when it comes to saving money, it truly does manage to save.

Moreover, compared to before, there's an update I think is fantastic—error rate in facts. I keep saying that GPT is almost becoming my fact-checking assistant; it really has very few hallucinations, and it's exceptionally strong in this regard. I used to think 5.6 was pretty good, but they have significantly optimized it this time, bringing it nearly up to the level of GPT-6 Astra. Here, we can also note the reasoning levels: if you are using light or medium reasoning, you may often make factual errors. This is why I have always recommended that everyone use the high setting for their daily needs, and if you're using Luna, you must have the highest setting.

I also tested it myself; there is a noticeable downgrade in aesthetics and some details.
For instance, manipulating Blender to build a motorcycle shows that the degree of detail completion is significantly lacking.

This is GPT-6 Astra's Temple of Heaven.

And this is GPT-6 Sol's Temple of Heaven; many details are problematic, the door literally has bugs.

However, the cost has dropped by about 70%, so overall, it can still be accepted.
While I was writing, GPT-6 Sol and GPT-6 Luna had already started pushing in ChatGPT Work and Codex.

But it's a bit awkward that OpenAI said these models are not yet available in Chat, meaning that in chat mode, they are still providing GPT-5.6 Sol.
This part I'm a bit puzzled about; since your costs have already significantly decreased, why not also switch to a cost-lower model in chat?
Final Thoughts
I know that everyone, upon reaching here, will definitely want to ask one question.
So how should I choose?
There are just too many models now, with new models being released every day; it’s really exhausting.
I can only describe my choice as best as possible.
1. If you can subscribe to Claude without being banned.
Logically, I truly dislike Anthropic. However, from the user's experience perspective, if you can subscribe to Claude without being banned and use it long-term, I still recommend you subscribe to Claude.
For complex planning and tasks, use Claude Fable 5.1 for planning, and Opus 5.5 for execution; this might currently be the best combination.
2. If you have been banned from Claude but can subscribe to overseas models.
Then just mindlessly subscribe to ChatGPT; honestly speaking, in terms of model capabilities, it's still slightly below Claude, be it in creation, cognitive insights, or coding.
But it has almost the best C-end experience, unlimited quota in chat mode, and the most user-friendly Codex client, and GPT-6 in this world is still T0 level strength.
For complex planning and tasks, use GPT-6 Astra or the advanced GPT-6 Pro, and utilize GPT-6 Astra high or GPT-5.6 Sol xhigh for execution; GPT-6 Luna can be used for large-scale automated tasks, which will work very well.
3. If you can only subscribe to domestic models.
Frankly speaking, the companies Qwen, Kimi, GLM, MiMo, and DeepSeek do not have a distinctly superior differentiator; use whichever you are accustomed to. The only issue is that the model tiering is not as strong as the latter two; Kimi K3 and Qwen 3.8 Max are more suited for planning and schemes, while GLM-5.3 and MiMo v2.6 Pro, DeepSeek V4.1 Flash are more suited for execution.
Just remember to always subscribe for just one month at a time; not too long, at most buy a quarterly card, and definitely do not buy an annual card.
OpenAI and Anthropic are still too mature, and the current environment is not quite the same as before.
In the past, these two were responsible for delivering cutting-edge intelligence, in other words, focusing on high-end.
Now, domestic models aim to do mid-range and cost-effective.
But now, their supply chains seem more mature, just like Apple, starting to cover the whole domain; I want all your mid-range, even your ultra-low-end.
However, for all users, this is a good thing.
Because now the intelligence appears to be becoming increasingly affordable.
In the past, for equivalent AI performance, its cost nearly dropped by 47% every quarter; in a few months to half a year, it may really become like coal, water, and electricity, an asset everyone can afford.
At that time, it may truly be a great flourishing era.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。