I spent three days this week reading Wall Street Journal headlines about Chinese AI models, and I’ll be honest — some of it felt like fear mongering. Tech CEOs sounding alarms. Government officials weighing crackdowns. Stock prices swinging on a single model release. So let’s set the headlines aside for a second and talk about what’s actually going on.
The trigger for all this was a model called Kimi K3, released by Moonshot AI out of China. It landed in third place globally for intelligence — right behind OpenAI’s newest model and Anthropic’s Fable model. That’s notable, because open-source models out of China have historically lagged the frontier by about six months. Kimi K3 closing that gap is the first time in a while an open-source model has gotten this close.
You’ve probably also seen headlines saying Kimi “ranked number one.” Here’s the honest version: it ranked #1 in one specific category — frontend design — on one benchmarking site, LMArena. That’s a real data point, not a small one. But it’s not the same as being crowned the best model overall, and it’s worth knowing the difference before you repeat that claim to anyone.
This Isn’t the First Time
DeepSeek did almost exactly this back in January of 2025 — came out of nowhere, cheap, capable, open-source, and rattled markets the same way. The pattern isn’t new. What’s actually new is the trend underneath it. There’s a chart I like from Theory Ventures that tracks open-weight models against closed-weight models on Chatbot Arena going back to 2023. In 2023, the gap between the two was huge. It roughly halved by 2024, halved again by 2025, and now in 2026 we’re at the point where they’re trading the lead back and forth. The frontier models are still generally ahead — but the distance has gotten a lot smaller than most people realize.
There’s one more wrinkle worth a footnote: Anthropic and OpenAI have both accused some of these Chinese labs of distillation — training a new model using an existing model’s outputs, instead of building from scratch. It’s a real cost-saver, since you skip a lot of the expensive work of generating original training data. Nothing has come of these accusations so far, and I’m not here to litigate who’s right. I just think it’s a good reminder that a benchmark number is rarely the whole story.
The Part That Actually Matters: Model Tiering
Here’s the part of this story that barely made the headlines, and it’s the part I think is actually useful. If you’re a typical business user, your experience of AI is probably one chat window — ChatGPT, Copilot, whatever your firm gave you. But on the development side, where people are more hands-on, we’re seeing more and more teams build what’s basically a tiering system into how they use these models.
The idea: use a frontier model — something like Fable or GPT-5.6 Sol — to do the high-level planning. Then hand the actual grunt work, the repetitive stuff, the high-volume stuff, off to a cheaper model. Think of it like staffing an engagement. The frontier model is your senior manager — making the judgment calls, setting the plan, reviewing the output. The cheaper model is your staff associate — doing the volume of the actual work. I do this myself: on a complex project, I’ll do the planning with Fable, then switch down to Sonnet to actually get the work done. Better plan, cheaper execution, more bang for the buck.
And here’s the thing about “cheap”: cheap doesn’t mean bad anymore. These open-source models — whatever you think about how they got built — are genuinely capable, and the floor for what a cheap model can do has gotten very high. If a tool or a vendor is cheaper than it used to be, that’s often a sign they’ve shifted more work onto a tiered, cost-efficient model for the simpler tasks — not that quality dropped.
Key Takeaways
- Kimi K3 ranked #3 globally for intelligence and #1 in one specific category (frontend design) on one site — know the difference before repeating either claim.
- The performance gap between open-source and closed-source models has been shrinking fast: roughly half the size every year since 2023.
- Model tiering — a frontier model for planning, a cheaper model for execution — is becoming the standard architecture, not an edge case.
- A vendor or tool getting cheaper doesn’t automatically mean lower quality. Ask what’s handled by which model, and why.
- Where’s the line between the simpler work and the more complex work in your own AI use — and what happens when something falls in between?
Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: Open vs. Closed Source: Model Tiering


Leave a Reply