I spent last week going through GPT-6 Astra’s benchmark results line by line. The number that got my attention was $26,000, roughly what it cost to run one version of a test that Astra also aced for a lot less money, under a different set of rules. If you’re the person at your firm who gets asked “should we use this thing,” that’s worth knowing before you answer.
OpenAI released GPT-6 Astra last week, and it scored high across a wide range of benchmarks. One of the loudest headlines was a 99.9% on ARC-AGI-3, a test built by an organization called ARC that keeps raising the bar as models get better at clearing it. This is their third version. Models had been scoring below 50% on it right up until Astra. A 99.9% looked like a breakthrough, and a lot of LinkedIn posts treated it like one.
But there’s a catch. That 99.9% came from OpenAI’s own custom harness. A “harness” is the setup wrapped around a model during a test, and it can change what the model is allowed to do. In this case, the custom harness let Astra remember things across the challenge that the standard harness didn’t allow. Give a model memory it didn’t have before, and it solves things a lot faster. Under the standard rules, Astra scored 64%. Same model, same test, a 36-point swing depending on which version you were shown.
That’s not the only place cost shows up, either. ARC published what it actually cost to hit that score, which is rare. Most benchmarks never report this. Getting Astra to that result ran into the tens of thousands of dollars, while human reviewers solving the same category of problem were getting paid a couple hundred bucks. Yes, the AI got there. It also cost about a hundred times more than paying a person. We’re getting closer to a world where the question isn’t “can AI do this,” it’s “is it worth it.”
That one benchmark isn’t the whole picture, though. On the broadest scorecards, the Intelligence Index and the Coding Agent Index (worth bookmarking if you’re not already following them), Astra is actually behind Claude Fable 5.1, which released the same week. Not by a lot, but it’s behind. Where Astra does stand out is cost: it’s cheaper to run than Fable 5.1, just scoring a bit lower for it. Picture the classic cost-versus-performance curve: Fable 5.1 costs more and does more, Astra costs less and does a little less, and Astra ends up owning the “best value at this price point” corner of that chart.
Model ML’s composite testing, which grades models specifically on finance and accounting work, tells a similar story. Astra comes out ahead in most categories, but usually by a small step, not a leap. On PowerPoint creation, it’s basically tied with Fable 5.1 but noticeably cheaper to run. On analytical finance, it actually costs more for a small improvement. And Gemini is the cheapest of the three by a wide margin, trailing the other two by only a little on accuracy. If getting the answer exactly right matters most, Astra looks like the pick. If cost is the constraint, Gemini is the clear call, for barely any drop in performance.
Here’s what I want you to walk away with:
- A benchmark score is a condition, not a constant. Ask what setup produced it before you trust it.
- Cost and capability don’t necessarily move together. If someone’s telling you how capable a model is, ask them what it costs, too.
- Composite indices, the Intelligence Index and Model ML’s composite, tell you more than any single benchmark headline.
- Before you adopt anything, ask which configuration produced the score and what it costs per task.
- AI adoption is a break-even question, not a leaderboard contest.
Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: https://everydaycpe.com/courses/technical-information-technology/gpt-6-astra-capability-vs-cost/


Leave a Reply