What are the strengths of Kimi K3? China's AI model enters top tier

Editor︰Ivy Cin

The recently launched Chinese AI model, Kimi K3 of Moonshot AI, made an immediate splash on major global AI benchmark leaderboards upon release, with certain performance metrics even outperforming GPT and Claude. This signifies that China's overall AI capabilities have officially entered the top tier globally.

So, what makes Kimi K3 so impressive? What are its key strengths, and why has it sent fresh shockwaves through the tech community?

Kimi K3: The world's largest open-source large language model

Many people are familiar with Kimi, which was developed by a Beijing-based tech company called "Moonshot AI" (月之暗面).

Its four core founding team members are all graduates of the Department of Computer Science at Tsinghua University. Kimi was first launched in October 2023, and over the past three years it has gathered a considerable user base. The newly released K3 is an entirely new version.

Regarding Kimi K3, the point most frequently highlighted in the news is its scale: its parameter scale reaches up to 2.8 trillion, making it currently the world's largest open-source large language model (LLM) in terms of parameter scale.

We often hear "open-source" and "closed-source", but what do they specifically mean?

"Closed-source" means that the source code and working mechanisms of an AI model are not publicly disclosed, only providing access services through an Application Programming Interface (API) or web page.

Currently, major models from the United States, such as ChatGPT, Gemini, and Claude, all adopt a closed-source approach.

"Open-source" means that the source code and related parameters of a model are publicly disclosed, and anyone can download, view, and modify them. Currently, major AI models from China, such as DeepSeek and Kimi, all adopt an open-source approach.

By continuing to opt for an open-source approach, Kimi K3 effectively makes the "blueprints" of this super-brain freely accessible to scientists and entrepreneurs worldwide. This, in turn, helps lower the barriers for industries of all kinds to adopt AI.

At the 2026 World Artificial Intelligence Conference held in Shanghai, the Kimi K3 exhibition area attracted a large number of users to come and experience it
At the 2026 World Artificial Intelligence Conference held in Shanghai, the Kimi K3 exhibition area attracted a large number of users to come and experience it. (Image Source: VCG)

Kimi K3: World-class "IQ" level

How should we understand the previously mentioned parameter scale of 2.8 trillion?

For human brain, the more neurons and the more complex the network, the smarter it is and the better its memory. For an AI model, more parameters mean it can accommodate more code and knowledge during the training process.

2.8 trillion parameters means the model can fit more knowledge and patterns into its "brain", understanding more, thinking deeper, and answering more accurately.

It also handles the analysis of both text and images, enabling it to tackle more complex and difficult tasks, with outstanding performance in areas such as programming, visual understanding, knowledge work, and long-range task processing.

This brings us to Kimi K3's second advantage: its "IQ" level has reached a world-class standard.

On the global AI comprehensive intelligence leaderboard Artificial Analysis, Kimi K3 is ranked 3rd globally overall, behind only the closed-source top American models Claude Fable 5 and GPT-5.6 Sol.

And on the anonymous front-end programming evaluation leaderboard recognised by developers worldwide, Arena, Kimi K3 reached the top spot within 24 hours of its release, winning first place in 6 out of 7 specific programming tracks.

Even Tesla founder Elon Musk publicly stated "impressive" on social media.

On the global AI evaluation leaderboard Arena, Kimi K3 reached the top spot within 24 hours of its release
On the global AI evaluation leaderboard Arena, Kimi K3 reached the top spot within 24 hours of its release, surpassing American AI giants Claude Fable 5 and GPT-5.6 Sol. (Web Image)

Kimi K3: 1 million token context window with a strong memory

The third advantage of Kimi K3 is its one-million-token context window, which is officially referred to as "native long context capability," and it supports visual understanding, meaning it can process text and images simultaneously.

The context window can be understood as the AI's "short-term memory".

An ordinary AI might forget what was said earlier in a conversation; however, Kimi K3 has a stronger memory and can be "fed" more data in one go, such as a complete set of instruction manuals, hundreds of design drawings, or legal contracts of several hundred thousand words, all of which it can digest at once, making it more suitable for processing large documents and long-form content.

Kimi K3: Running larger models with less computing power

It is also worth mentioning that Kimi K3's performance rivals the world's top level, but for the same volume of calls, the API call cost is only about half that of similar overseas closed-source models, significantly reducing training and inference costs.

The reason it can achieve this is that instead of solely relying on brute computing power, Kimi K3 has developed a proprietary technology named the KDA mixed linear attention architecture, carrying out original innovation in the underlying architecture.

This significantly reduces the model's operational cache consumption, increases inference efficiency by 2.5 times with the same amount of computing power, achieving the goal of "running a larger model with less computing power."

Kimi K3, the AI model of China's Moonshot AI
As soon as Kimi K3 was launched, it garnered global attention, with Wall Street describing it as another "DeepSeek moment" for Chinese AI. (Image Source: VCG)

Li Xiaodong (李曉東), director of the Internet Governance Research Centre at Tsinghua University, noted that this proves Chinese enterprises can build LLMs that rival the world's best through self-developed algorithmic architectures, without relying on closed-source code or exclusive compute monopolies.

For a long time, the high-end model space was dominated by overseas closed-source products; by claiming top spot in core coding benchmarks, Kimi K3 has shattered that status quo.

Speaking of which, it is believed that many people are eager to give it a try. However, it should be noted that due to a surge in users in a short period, Kimi K3 is approaching its capacity limit, and to ensure system stability, subscriptions for new consumer-end (C-end) users have been suspended.

It is reported that Kimi K3 is making every effort to expand its computing power. Once the new computing power is in place, subscription slots will be gradually opened up, allowing more people to experience this more intelligent and powerful AI model together.

Read more:

Understanding the "raising lobsters" OpenClaw craze

Generating "director-level" videos with a single sentence, how powerful is China's AI Seedance 2.0?

Series of Articles: Can China become the "world's token factory"?

Related Tags
AI

How does computing power flow like water and electricity?|Token Factory Ⅴ

How can China's token exporting be realised?|Token Factory Ⅳ

What is computing-electricity synergy and why is it important to AI?|Token Factory Ⅲ

The limit of AI is computing power? It's all about energy|Token Factory Ⅱ

Tech guardians|How to protect Chinese white dolphin, the "National Treasure of the Sea" ?

WeChat