CyrilXBT
CyrilXBT|Sep 10, 2026 06:47
i spent 20 hours trying to figure out why gpt-6 astra and claude fable 5.1 have the exact same price tag and completely different bills. both launched two days apart. both charge $10 input, $50 output per million tokens. identical on paper. not identical once you actually run them. astra's real shift is computer use. it operates software directly, fills forms, updates crm records, runs research, builds a site and qa-checks its own frontend. 72.6% on osworld 2.0 in about 40 minutes per task versus sol's 65.7% in 75. and a genuinely striking safety number, confirmed straight from openai's own page: on a new eval built specifically after the hugging face incident, sol went beyond its authorized scope 48% of the time without production safeguards. astra did it 0% of the time. fable 5.1's real edge is stamina and coding. 52.6% on terminal-bench-science versus opus 5's 29% and sol's 22.4%. 55.8% on terminal-bench 4.0. three effort levels, low and medium often matching fable 5's output at a fraction of the cost. here's what took me most of those 20 hours to actually track down. cache reads. fable 5.1 charges $0.25 per million. astra charges $1, four times more. agentic work re-reads the same context dozens of times in a single task, so cache reads dwarf fresh input on any real workload. that line decides your actual bill, not the sticker price. astra also adds a surcharge past 272k input tokens, jumping to $20/$75 for the whole request. worth being honest about the benchmark comparisons circulating this week too. neither lab tested against the other directly. even the one shared benchmark, osworld 2.0, gets reported under three different scoring methods between the two announcements. any table putting them head to head is comparing different measurements, not the same test. gemini 3.8 flash isn't even trying to compete on intelligence here, $0.75 input, $3.75 output, a 13x gap under both. most daily tasks, classifying a message, summarizing a call, extracting fields from an invoice, don't need a frontier model at all, and running them on one anyway means paying frontier prices for capability you never use. the actual takeaway after 20 hours. astra wins on computer use, speed, and staying inside task boundaries. fable 5.1 wins on coding, research stamina, and the cache-read line that decides real agentic bills. flash wins on daily-task cost by an order of magnitude. the right move isn't picking one winner, it's splitting your workload across tiers and building your system so the model underneath is replaceable, because two frontier labs just shipped within 48 hours of each other, and that alone tells you single-vendor lock-in is a real risk, not a theoretical one.(CyrilXBT)
Share To

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads