Author: Jose Antonio Lanz
Translated by: Shenchao TechFlow
Senchao Introduction: On September 3rd, OpenAI launched GPT-6 Astra, and early testers treated it as an open stress test: some recreated the entire Manhattan in Unreal Engine in a week, some created a browser shooting game with online servers in a day, and others had it write a Bach choral piece achieving historically best results. However, when it comes to writing articles, the same model performs even worse than its predecessor, with writing benchmarks dropping by about 80 Elo. In summary: it excels at tasks with objective right and wrong, but falters on tasks based on taste, and all this comes at a price 2.5 times that of the previous generation.
Price and Positioning: 2.5 Times That of the Previous Generation
On September 3rd, OpenAI released GPT-6 Astra, with an input price of **10** per million tokens and an output price of **50 per million tokens (one token is approximately three-quarters of a word, also the billing unit for various AI companies), which is 2.5 times that of its predecessor model. OpenAI President Greg Brockman declared at the launch that AGI has arrived.
Astra's top selling point is "computer control": the model no longer throws a list of operations at you but directly takes over the mouse and keyboard on a real desktop. In the OSWorld 2.0 test (which measures the proportion of agents autonomously completing daily desktop tasks), OpenAI's reported score is 72.6%, with a single task taking about 40 minutes; by comparison, Sol's score is 65.7% but takes 75 minutes.
It is also the first model in OpenAI's history deemed to have reached the "critical threshold" for cybersecurity, meaning it can independently discover unknown software vulnerabilities and construct usable attack code without humans pointing out a hole first.
Visual Understanding: Building Manhattan Street by Street
It turns out Astra is ridiculously strong in visual understanding and spatial perception.
Investor and former HyperWrite CEO Matt Shumer had Astra immerse itself in Unreal Engine (the game engine behind Fortnite) for a week, and as a result, it generated a replica of Manhattan. He released a flying preview and said the model was "refined street by street until each one was perfect."
Even more ridiculous followed. Shumer made Astra build a survival world, with each character running a copy of the model, and then let it run all night. The next day he heard someone talking in the living room, thinking there was a thief, and when he went to check, he found "they had already started chatting."
Max Weinbach fed the model photos of Apple Park, requesting a reconstruction using free 3D modeling software Blender, and his feedback was, "it looks absurdly good." Tom Krcha provided just a photo of a house and received an editable indoor geometry that can run at 60 frames per second, with appliances and toys included. Pietro Schirano compressed the process into one action: placing a pin on a map and requesting the surrounding area's 3D, saying "it just did it for you."
Developer SuSu scaled this idea to a city level: Astra used Three.js (a JavaScript library for rendering 3D in regular browsers without downloads) to reconstruct Hangzhou and its surrounding towns within 24 minutes, including West Lake, Leifeng Pagoda, tea hills, and wetlands.
Programming: Games People Really Play
Games are currently the hottest use case and also Astra's most brilliant field.
Former Apple UX/UI designer and AI developer Anshu Chimala created a 3D game in just 45 minutes, consuming very little usage quota. He described Astra as a "turbocharged AGI machine in the realm of 3D games." His method is more valuable than the bragging itself: connecting the model to Blender, having it generate concept art of its desired style, and then commanding it to iterate constantly until the in-game screenshot matches the reference image at 60 frames. Astra modeled every asset and also generated the textures.
Former developers at Coinbase and Eleven Labs, Rishi Prasad created "Astral War" in a day: a browser shooting game with authoritative online servers, 12-person rooms, controller support, and voice chat. He described it as a "huge leap in visual fidelity" compared to what he made a month ago with Claude Opus 5.
Some even skipped the design phase. An AI developer with the pseudonym Daniel showed Astra a mobile game advertisement, asking it to create a playable browser version of what was in the ad, and in under 30 minutes it was "close enough."
Computer Control and Illustration: Drawing with the Mouse
Japanese illustrator Taiyaki Sun conducted the most straightforward validation of computer control in this batch of tests: rather than asking for a picture, she handed Astra a hand-drawn line art file and let it color the drawing like a human colorist in Clip Studio Paint using the mouse.
Astra created layers, zoomed in on the image, selected brushes, and filled in colors. According to this artist (translated from a Japanese post), she merely watched from the side throughout the process. This session ran on a $100 Pro package at maximum computational power, consuming 21% of the quota.
Other users also shared entertaining videos: Astra, through computer control (taking over your computer screen rather than using MCP servers or API keys), was able to fully replicate their photos in drawing software.
Music: Bach Test
GPT-6 Astra has relatively good taste for an LLM.
Auggie, who operates the "Augmented Fifth" newsletter, has an informal benchmark: using a fixed prompt, asking the model to write a four-part choral piece in the style of Bach, in G minor, 3/4 time, formatted in LilyPond text that can be compiled into sheet music, then scoring it according to music college harmonic rules.
Astra achieved the best score ever on this test: no voice guidance errors (meaning no melody line collided with conflicts prohibited by Bach's rules), and in the harmony, there appeared a Neapolitan sixth, a half-diminished chord typically found in Mozart and Beethoven's works. Auggie noted this as "the first model to write passing tones on this benchmark."
OpenAI's own table also points to the same conclusion. In OpenScore String Quartets (measuring the accuracy of the model reading and transcribing classical scores), Astra reached 0.84, while Sol only scored 0.19.
Doctor and veteran AI tester Derya Unutmaz had the model create a fully playable virtual piano, incorporating all six of Bach's "Brandenburg Concertos." He wrote that this "crazy model finished the whole task in about 11 minutes."
It is important to emphasize that GPT-6 Astra is an LLM, not an audio/music model. Its understanding of music likely comes from sheet music and text data, rather than from real associations of sound and music in the training data, so these results are quite impressive for a text model, but would not be considered remarkable if produced by specialized AIs like Suno.
Writing: The Areas Where It Fails
Oh my, how much people miss GPT-4o.
Like usual, OpenAI's models excel in programming but falter in writing, at least when not heavily prompted, given context, and guided. To be fair, writing is neither OpenAI's strong suit nor its primary focus.
Louis-François Bouchard ran an internal benchmark, scoring the model's writing based on Elo (a rating system originating from chess that ranks based on wins and losses in pairwise matchups) within his team’s editorial tone. Astra ranked 11th with a score of 1995; its predecessor ranked 6th with 2156. Each of Astra’s scripts costs about $0.26, approximately 1.8 times that of Sol.
Bouchard described the results as "surprisingly disappointing" and expressed that he had not expected this at all.
Giuseppe Paleologo, author of a bestselling guide on quantitative portfolio management, had Astra generate novel ideas for optimal portfolio diversification but received "obviously exaggerated hodgepodge," wrapped in text that could easily be identified as machine-written. His conclusion: "True creativity is still a long way off."
Mia AI Lab shares a similar view, acknowledging that Astra might be the strongest on some tasks, but calling it "boring," lacking personality, and advising to avoid it for any creative work.
Ingar Haaland ran the cleanest version of the test: having Astra write four paragraphs in his own style, close enough that Pangram (a tool that identifies AI-written text by comparing it against hundreds of thousands of human and machine samples) couldn't detect it. The result: "Pangram was not fooled."
In other words, this model lacks creativity, and its outputs can easily be identified as AI-generated, not due to any watermark but because of the way it writes and expresses itself.
Independent assessments align with these complaints. Artificial Analysis recorded a drop of about 80 Elo on the GDPval-AA v2 (adapted from OpenAI’s proprietary dataset covering economically valuable tasks across 44 professions), with customer service and long context reasoning also showing minor regressions.
It is not all one-sided. Silas Alberti from Cognition told OpenAI that Astra's writing made Devin's test report clearer; Every's contributing author Katie Parrott had Astra draft a preliminary version of her evaluation, and even her CEO did not realize it wasn't written by her after reading it.
The disparity between these two halves seems to be key: Astra excels at tasks with verifiable standard answers, a plausible chord, a renderable grid, a form that can be submitted; whereas in areas where standards are based on taste, it becomes mediocre.
How Much Does It Cost to Explore?
Astra is being pushed to ChatGPT Plus, Pro, Business, Enterprise users, as well as APIs, Microsoft Azure, and AWS Bedrock, with enterprise access defaulted off and requiring manual activation by administrators. Its advanced cybersecurity features are still locked behind OpenAI's Daybreak program; this decision has proven wise within 48 hours, as Reuters reported that OpenAI's agents were already trading illicit tactics on a German website.
The prediction market previously gave Astra a 72% probability of being released by September 30, but it arrived on the 3rd.
In the Artificial Analysis Intelligence Index (a third-party comprehensive assessment of reasoning, knowledge, and programming), Astra scored 61.2, GPT-5.6 Sol scored 60.9, and Anthropic's Claude Fable 5.1 scored 65.7, while Astra's pricing is 2.5 times that of Sol.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。