A few weeks ago, Kimi K3 came out and lit up my feed with benchmark charts showing it beating models three times its price. What almost nobody mentioned in the retweets was that the score was for front-end development. On most other tasks, it wasn’t close. It got reported as “Kimi is just better,” when really it was “Kimi is better at this one thing.”
That’s basically the whole story behind a new benchmark Model ML just published for finance work, and I think it’s worth ten minutes of your attention or you can just read this instead.
What Model ML actually tested
Model ML builds AI tools for investment banks, private equity firms, and consultants. They ran more than a dozen AI models, closed-source frontier models alongside the strongest open-weight ones, against hundreds of real finance tasks, split into five categories: PowerPoint creation, Excel model building, analytical finance, multi-step financial workflows, and document retrieval. Not trivia questions. Actual work, scored against rubrics built by finance professionals.
On a single task, the gap between the best- and worst-scoring frontier model was over 30 percentage points. Pick one model and run everything through it, and that gap means you’re guaranteed to be leaving performance on the table somewhere.
The “jagged frontier” problem
A while back I was at a course at Wharton where Ethan Mollick, who’s well known in the AI world, described what he calls the “jagged frontier.” AI models don’t improve along a clean, even line. They get very, very good in some spots and stay weak in others, a jagged shape instead of a smooth one. Model ML’s results are a clean illustration of that: across their five categories, no single model won all five. The highest scorer and the best-value pick were almost never the same model.
Cost is part of the model, not separate from it
Model ML’s line is that cost is a property of the whole kitchen, not just the chef. An API call is the chef, but the bill also covers every retry, every tool call, and every bit of context that gets re-sent along the way. To make the comparison fair, they held that “kitchen” (their own agent harness) fixed and only swapped which model was doing the cooking.
That’s what let them compute a real dollar figure per task, not just a token count and it’s why open-weight models showed up so well. On multi-document retrieval, GLM 5.2 scored within a few points of category-leader Gemini 3.6 Flash at a small fraction of the cost. And on single-filing extraction, a cheaper model reached roughly 93% of the top model’s accuracy for about 2% of the cost.
What this means for you
If you’re evaluating an AI vendor’s tools, or advising a client who is, a benchmark score by itself doesn’t tell you much. A vendor saying their tool “scored 90% on finance tasks” could mean almost anything. Ask which task category that’s from, because pulling a number out of a filing and building a leveraged buyout model are genuinely different skills.
The other question worth asking is how the vendor routes work. A lot of AI products right now are, functionally, a wrapper around one frontier model’s API. If a platform can explain that it sends different task types to different models based on cost and accuracy, that’s a real signal of engineering, not just packaging. If they can’t answer that question, or won’t, that tells you something too.
Key Takeaways
- No single AI model wins across finance tasks — the highest scorer and the best-value pick were almost never the same model.
- Accuracy gaps between models can swing as much as 30 points depending on the specific task.
- Cost is a property of the whole system — the harness, the retries, the re-sent context — not just the underlying model.
- Open-weight models trail the frontier but cost a fraction as much, and they’re increasingly competitive on high-volume work like retrieval.
- Ask vendors for cost-per-task data and how they route work across models — not just a headline benchmark score.
Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: https://everydaycpe.com/courses/technical-finance/the-model-ml-composite/


Leave a Reply