Kimi K3, Moonshot AI's new frontier model, is official: announced as "Open Frontier Intelligence" on July 16, 2026, it is available on Kimi.com, Kimi Work, Kimi Code and the API — after a wave-based rollout that started with the Kimi CLI. The official sheet: 2,800 billion parameters, a one-million-token context window, native multimodal, and open weights promised for July 27. The price list, meanwhile, says it all: $3 per million input tokens, $15 for output, $0.30 on cache hits. Translation: the budget champion now charges frontier prices. That is the real news — not the model's size. Here is what the announcement really says, what the published benchmarks are worth, and why, even if K3 is excellent, it will change almost nothing for your business — for a reason that has nothing to do with China.
Key takeaways
- ▸K3 is official since July 16, 2026: 2,800 billion parameters (MoE), 1-million-token context, native multimodal. Available on Kimi.com, Kimi Work, Kimi Code and the API; open weights announced for July 27, 2026. Benchmark card published (14 measurements).
- ▸The price marks a turning point: $3 / $15 per million tokens ($0.30 on cache hits), with a one-million-token context window. That is Claude Sonnet 5 territory — and three to four times the price of its predecessor K2.6 ($0.95 / $4.00). Moonshot is no longer selling volume: it is selling performance.
- ▸Moonshot's full card (14 benchmarks) puts K3 ahead of Claude Opus 4.8 on all 14 lines (including 1,668 vs 1,600 Elo on GDPval-AA v2) and first on 5 of them — but behind Claude Fable 5 on knowledge work and most coding lines. Vendor figures, not yet cross-checked independently.
- ▸The measured baseline, however, is solid. Kimi K2.6 ranks 4th worldwide on Artificial Analysis' Intelligence Index (54 points versus 57 for the leaders): the world's best open-weight model even before K3.
- ▸Our conclusion, which contradicts the headline: however excellent K3 turns out to be, it will change nothing for you. An SMB's problem has never been picking the right model. It is being able to switch.
In this article
1. What is confirmed, what remains unclear
Confirmed. The official announcement has landed: Moonshot presents K3 as its "Open Frontier Intelligence" — 2,800 billion parameters (mixture-of-experts architecture), a one-million-token context window, native multimodality (the model understands images without a separate module), and an always-on reasoning mode. The rollout, started in waves on July 16 through the Kimi CLI, now covers Kimi.com, Kimi Work, Kimi Code and the API. Open weights are announced for July 27, 2026. The price list is live — we break it down below.
Called in advance. We wrote it in the first version of this article: the date did not leak through an engineer, but through the store. On the Kimi Open Platform, a credit top-up promotion named "K3 launch" runs from July 15 to August 11, 2026 (10 to 30% bonus). You do not name a commercial operation after a model you are not about to serve — those credits have to be honored. The marketing department had spilled the beans before the comms department. It was right, give or take a day.
Settled since. The "2,500 to 3,000 billion parameters" rumor — which we refused to endorse as long as it rested on a single anonymous source — is closed: 2,800 billion, officially. And our architectural bet is confirmed: we wrote that linear attention was "the most likely path" to serve a one-million-token context (the team had admitted in a December 2025 AMA that it was "too costly to serve" back then). That is exactly what Kimi Delta Attention does — K3's hybrid linear attention, up to 6.3x faster decoding in million-token contexts — backed by Attention Residuals (~25% higher training efficiency at less than 2% additional cost, per Moonshot).
Still unclear. Only two things: the weights themselves (announced for July 27, not yet downloadable — exact license to be checked at release), and any independent measurement of performance. The 14 published benchmarks remain vendor figures.
Kimi K3 — official spec sheet
| Maker | Moonshot AI (Beijing) |
| Official announcement | July 16, 2026 ("Open Frontier Intelligence") |
| Parameters | 2,800 billion (mixture-of-experts) |
| Context | 1,000,000 tokens |
| Modalities | Text + vision (native multimodal) |
| Architecture | Kimi Delta Attention (hybrid linear attention, up to 6.3x faster decoding at 1M context) + Attention Residuals |
| API pricing | $3 / M input tokens · $15 / M output · $0.30 / M on cache hits |
| Availability | Kimi.com, Kimi Work, Kimi Code, API (platform.kimi.ai) |
| Open weights | Announced for July 27, 2026 |
| Stated focus | Long-horizon agentic coding, self-evolving workflows |
Source: Moonshot AI's official announcement and tech blog (07/16-17/2026).
2. Pricing: the end of the Kimi discount
Here is the table that should catch the eye of any executive paying an API bill — prices per million tokens:
| Model | Input | Output |
|---|---|---|
| Kimi K3 · 1M context | $3 ($0.30 on cache hit) | $15 |
| Kimi K2.6 / K2.7-Code | $0.95 | $4.00 |
| Claude Sonnet 5 | $3 | $15 |
| Claude Opus 4.8 | $5 | $25 |
| GPT-5.6 Sol | $5 | $30 |
| Claude Fable 5 | $10 | $50 |
Three readings of this table, from the most obvious to the most interesting.
One. Moonshot is changing business models. At $3 / $15, K3 costs three times K2.6's input price, almost four times on output — and lines up exactly with Claude Sonnet 5. The message is clear: Kimi no longer sells itself as the low-cost alternative, but as a frontier model priced accordingly. Which is consistent with the money invested: $500 million raised in January — earmarked, according to press reports, for "K3 development and compute capacity expansion" —, then $2 billion in May at a $20 billion valuation. A model subsidized at a loss does not pay that back.
Two. The real price is the cache hit. $0.30 per million tokens on already-seen input is ten times less than the nominal rate. And that is precisely how agents consume: re-reading the same context (instructions, knowledge base, history) a hundred times, adding only a few new tokens each turn. Combined with the one-million-token window, this pricing describes a model designed for long agentic workflows — not for chat.
Three. The trade-off narrows, but does not flip. Against Opus 4.8 ($5 / $25) and Fable 5 ($10 / $50), K3 remains 40 to 70% cheaper on output — the line item that explodes in agentic usage. But the days of switching to Kimi blindly to divide your bill by six are over. From now on, you measure: at comparable quality, the price gap justifies the test; at lower quality, K2.6 at $0.95 remains the market's best value for volume work.
3. How does it stack up against Opus 4.8?
The trade press headlines that K3 is "expected to close the gap" with Anthropic's Opus 4.8. Early user reports point the same way — some even put it ahead on long coding tasks. Our position: plausible, and not yet demonstrated. No independent benchmark has been published at the time of writing. We have seen too many models shine on social media for three days before falling back in the measurements to treat impressions as data.
The full benchmark card is now out. After three early scores relayed on launch night, Moonshot published its complete card: 14 benchmarks, "all maxed out on thinking effort" — every model measured at its maximum reasoning effort. It also corrects the early figures: 1,668 Elo on GDPval-AA v2 (not 1,687), 1,548 on AA-Briefcase (not 1,527), and GPT-5.6 Sol drops to 90.4% on BrowseComp (not 92.2) — which hands that line's first place to K3. The caveat stands: these are vendor-published figures, not yet cross-checked by an independent measurement.
| Benchmark | Kimi K3 | Claude Fable 5 | Claude Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|
| Knowledge work (GDPval-AA v2, Elo) | 1,668 | 1,760 | 1,600 | 1,748 |
| Knowledge work (AA-Briefcase, Elo) | 1,548 | 1,583 | 1,354 | 1,495 |
| Business tasks (JobBench, %) | 52.9 | 57.4 | 48.4 | 46.5 |
| Spreadsheets (SpreadsheetBench 2, %) | 34.8 | 34.7 | 31.6 | 32.4 |
| Agentic browsing (BrowseComp, %) | 91.2 | 88.0 | 84.3 | 90.4 |
| Coding (FrontierSWE, %) | 81.2 | 86.6 | 66.7 | 71.3 |
| Terminal (Terminal Bench 2.1, %) | 88.3 | 84.6 | 84.6 | 88.8 |
| Long-run coding (SWE Marathon, %) | 42.0 | 35.0 | 40.0 | 39.0 |
Excerpt from Moonshot's card (retrieved 07/17/2026) — 8 of 14 benchmarks, all models at maximum thinking effort. Vendor figures, pending independent verification.
If these scores hold up, they describe exactly a Fable / Sol-class model: ahead of Claude Opus 4.8 on all 14 published lines, first on 5 of them (agentic browsing, spreadsheets, automation, Program Bench, long-run coding), behind Claude Fable 5 on knowledge work and most coding lines — the first-place tally reads 7 for Fable 5, 5 for K3, 2 for GPT-5.6 Sol. In other words: the level the price announces. To be confirmed — we will update this section as soon as independent measurements land.
What is measured independently, on the other hand, is the starting point. Kimi K2.6 ranks 4th worldwide on Artificial Analysis' Intelligence Index, at 54 points versus 57 for the three closed leaders at the time of measurement, in April 2026 (Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.4 — since superseded by Opus 4.8, GPT-5.6 and others; the index has not been re-run yet). The world's best open-weight model, three points from the frontier. If K3 only gained those three points, the press headline would be factually fulfilled — and the asking price justified.
One corrective, still, to avoid howling with the wolves. The benchmarks Moonshot puts forward are in-house benchmarks ("Kimi Code Bench", "Kimi Claw"). On its own results table, K2.7-Code lost 11 cells out of 12 against GPT-5.5 and Claude Opus 4.8. The gap has narrowed; it had not disappeared. Anyone telling you today that K3 "has overtaken" the closed frontier is over-reading: even by Moonshot's own card, Claude Fable 5 keeps 7 first places out of 14 — including, an irony that lends the card credibility, Moonshot's own in-house benchmark (Kimi Code Bench 2.0: Fable 5 at 76.9, K3 at 72.9). Ahead of Opus 4.8, however, is what all 14 published lines say — pending independent validation. We are testing K3 through the Kimi CLI and will publish ours right here.
4. Why it will change nothing for you
Here is the part other articles will not write, because it hurts the click.
K3 being very good will change nothing about your situation. Neither to make you switch to Moonshot, nor to keep you away from it. And it is not a geopolitical question: it is an architecture question.
Remember June 13, 2026. That day, invoking national security, the US government ordered Anthropic to suspend access to Fable 5 and Mythos 5 for all foreign nationals. Unable to verify its users' nationality in real time, Anthropic did the only thing possible: cut both models for everyone, everywhere. Two weeks of downtime. Restored on July 1st.
Ask yourself the only question that matters: if your automations had been running on that model, what would have happened on the morning of June 13?
A pipeline with the provider hard-coded would have stopped. A pipeline where the model is a parameter, backed by a fallback model, would have switched over and kept running without anyone picking up the phone. The difference is not decided on the day of the outage: it is decided on day one, and it costs one line of configuration. That is not genius, it is hygiene — but it is exactly what separates an automation from a gamble.
That is why K3's arrival leaves us cold: if your pipeline is well built, testing Kimi costs you one configuration line and an afternoon of measurement. If it is badly built, no model in the world will save you — you will just be captive to another provider, in another country. This is exactly the reversibility principle we apply in our AI automation projects.
Our position.
The event is not that a Chinese model is closing in on the Americans — nor even that it now charges American prices. It is that the model layer has become a substitutable, politically exposed commodity, whatever its flag. The skill that holds value is no longer picking the right model. It is building processes that survive changing it.
5. Three things nobody else will tell you
Self-hosting is out of reach for an SMB
"Open weights" does not mean "on your premises". Kimi K2.5's weights weigh around 595 gigabytes. You need multi-GPU data-center hardware. With K3's official 2,800 billion parameters, it will be worse. Moonshot does not publish smaller versions (8, 32, 70 billion), despite community demand. In practice, open-weight sovereignty is consumed through a European hosting provider, not in your server room.
Calling Moonshot's API means sending your data to China
No more and no less serious than sending it to the United States — but no less either. It is a transfer that must be documented and framed under GDPR. Not an ideological debate, a contract clause. And remember the deadline coming up: the AI Act's transparency obligation (Article 50) applies from August 2, 2026 — your users must know they are interacting with an AI, whatever model sits behind it.
The best model is almost never the point
In the pipelines we deliver, the gains come from data quality, step decomposition and error recovery — not from the three index points separating 1st from 4th. A significantly cheaper model that does the job is a better engineering choice than a slightly stronger model that ruins the margin. And since K3's pricing, that trade-off also applies inside the Kimi catalog: for sorting, extraction or summarization, K2.6 at $0.95 remains unbeatable.
6. FAQ
Kimi K3 is Moonshot AI's (Beijing) frontier AI model, announced on July 16, 2026: 2,800 billion parameters in a mixture-of-experts architecture, a one-million-token context window, native multimodality, and a stated focus on long-horizon agentic coding. Its weights go open on July 27, 2026, which should make it the world's largest open-weight model.
Yes. Officially announced on July 16, 2026, Kimi K3 is available on Kimi.com, Kimi Work, Kimi Code and through the API (platform.kimi.ai), after a wave-based rollout that started with the Kimi CLI. Open weights are announced for July 27, 2026.
$3 per million input tokens, $15 for output, and $0.30 on cache hits (already-seen input), with a one-million-token context window. That is Claude Sonnet 5's price point, and three to four times the price of its predecessor Kimi K2.6 ($0.95 / $4.00), which remains available.
According to Moonshot's full card, yes — across all 14 published benchmarks: 1,668 vs 1,600 Elo on GDPval-AA v2, 1,548 vs 1,354 on AA-Briefcase, 91.2% vs 84.3% on BrowseComp. K3 still trails Claude Fable 5 on knowledge work and most coding lines (7 first places out of 14 for Fable 5, vs 5 for K3 and 2 for GPT-5.6 Sol) — and these are vendor figures, not yet confirmed by an independent measurement.
Almost: Kimi K3 officially has 2,800 billion parameters (mixture-of-experts architecture), per Moonshot's July 16, 2026 announcement — so the initial rumor (2,500 to 3,000 billion) was within range. For scale: the entire K2 line ran on 1,000 billion parameters. K3 adds a one-million-token context and native multimodality.
Yes: Moonshot has announced the weights for July 27, 2026. The exact license remains to be checked at release; the entire K2 series shipped under a modified MIT license, weights included — the only constraint being to display the Kimi brand above 100 million monthly active users or $20 million in monthly revenue, a threshold irrelevant for an SMB.
Wrong question. The goal is not to change providers, it is to be able to — and to arbitrate task by task, based on cost, quality and availability. A properly designed architecture lets you test Kimi K3 on a real task in an afternoon, and automatically fails over to a backup model the day one becomes unavailable.
Sources
- Official Kimi K3 announcement and tech blog "Open Frontier Intelligence" — Moonshot AI (07/16/2026): 2,800B parameters, 1M context, native multimodal, Kimi Delta Attention (6.3x decoding), Attention Residuals, open weights 07/27/2026
- Official Moonshot AI channels — K3 teaser (07/15/2026),
kimi-cli/kimi-coderepositories (07/14/2026), K3 price list on the Kimi platform (retrieved 07/16/2026) - "Moonshot's upcoming Kimi 3 is expected to close the gap with Anthropic's Opus 4.8" — TechCrunch (07/16/2026)
- Full Kimi K3 benchmark card — Moonshot AI, retrieved 07/17/2026: 14 benchmarks, all models at maximum thinking effort, incl. GDPval-AA v2 (1,668 Elo), AA-Briefcase (1,548 Elo), BrowseComp (91.2%). Corrects the early figures relayed on 07/16 (1,687 / 1,527; GPT-5.6 Sol 92.2 → 90.4 on BrowseComp); vendor figures, not cross-checked
- Moonshot founders' AMA — official Kimi platform blog (12/03/2025)
- Kimi K2.6, best open-weight model (Intelligence Index: 54, 4th worldwide) — Artificial Analysis (04/20/2026)
- Kimi K2.7-Code model card — HuggingFace (06/12/2026): in-house benchmarks and modified MIT license
- 2,500-3,000 billion parameter rumor — AIBase (04/29 and 07/02/2026); settled by the official announcement: 2,800 billion
- "Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI" — The Decoder (07/17/2026)
- "K3 launch" promotion (07/15 → 08/11/2026), Kimi Open Platform — relayed via screenshot (@kimmonismus, 07/14/2026)
- $500M raise earmarked for K3 — The Decoder (01/01/2026); $2B at a $20B valuation — TechCrunch (05/07/2026)
- Suspension of Fable 5 and Mythos 5 for foreign nationals, by order of the US government — Anthropic and Al Jazeera (06/13/2026)
- Competitor model pricing — Anthropic documentation; OpenAI, GPT-5.6 launch (07/09/2026)
The benchmarks cited are published by Moonshot and not yet confirmed by independent measurements.8) are not confirmed by independent measurements. The bets expressed in this article are JAIKIN's opinions, flagged as such. Article updated continuously during the K3 rollout.