This website uses cookies

Read our Privacy policy and Terms of use for more information.

When this article was first published, Securities Times, a Chinese national financial newspaper, reported that local regulators had granted approvals to 14 large language models (LLMs) for public use. That first batch of approvals included 40 AI models in total.

Among the 14 LLMs greenlit for commercial deployment were products from well-known companies such as smartphone maker Xiaomi and the startup 01.AI started by Kai-Fu Lee.

The full list was not disclosed, so we focused on five prominent Chinese model families. We have kept that original snapshot below and added a 2026 update to each entry.

The Original Five Chinese LLM Families—Updated for 2026

  1. CPM (2020; 2.6B): China’s first large-scale pre-trained model created by the Beijing Academy of Artificial Intelligence (BAAI) and Tsinghua University.

    2026 update: CPM is now best understood as a historical starting point. The broader OpenBMB ecosystem has moved toward the smaller MiniCPM family, including multimodal and on-device models such as MiniCPM-o 4.5 and MiniCPM-V 4.5.

  2. ERNIE 3.0 Titan (2021; 260B): A big brother of Baidu’s ERNIE 3.0 foundation model having 10B parameters. It was designed to explore the performance of scaling up the original model. At the time of its creation, it was the largest Chinese dense pre-training model.

    2026 update: Baidu’s open family has advanced to ERNIE 4.5. The official release includes text and vision-language models, mixture-of-experts variants with 3B or 47B activated parameters, and a 0.3B dense model; the weights and ERNIEKit tooling are available under Apache 2.0.

  3. Qwen 1.5 family (2023; 1.8B, 7B, 14B): The Qwen models distinguish themselves through extensive training on diverse datasets, enabling superior performance across a variety of tasks including language understanding, coding, and mathematics, as well as offering specialized capabilities in vision-language tasks.

    2026 update: Qwen has progressed to Qwen3, a much broader family of dense and mixture-of-experts models. Qwen3 introduced thinking and non-thinking modes, stronger reasoning and tool use, and official support for more than 100 languages and dialects.

  4. Yi family (2023; 6B, 34B): These models, developed by 01.AI, represent a pioneering force in bilingual large language models, demonstrating exceptional strength in language understanding and commonsense reasoning. Particularly notable for their performance on both English and Chinese benchmarks, they secured top rankings on leaderboards such as AlpacaEval and the Hugging Face Open LLM Leaderboard at the time.

    2026 update: The latest major open update in 01.AI’s official Yi repository remains Yi 1.5, released in 2024 with improvements in coding, mathematics, reasoning, and instruction following. Yi remains useful as a bilingual open model family, although its public release cadence has been quieter than DeepSeek’s or Qwen’s.

  5. Baichuan 2 (2023; 7B, 13B): By providing both Base and Chat models in various sizes, including an efficient 4-bit quantized version, Baichuan 2 facilitated a wide range of research and commercial applications, lowering the barrier for innovation and enabling more developers to incorporate advanced AI capabilities into their projects.

    2026 update: Baichuan’s open work has expanded beyond the original general-purpose text models. Its official repositories now include Baichuan-Omni-1.5 for text, image, audio, and video understanding, plus the medical-focused Baichuan-M3-235B released in 2026.

Chinese LLMs in 2026: What's New

China's leading model families now compete on reasoning, coding, long context, agentic tool use, multimodality, and efficient open-weight deployment—not only on parameter count. Developers can now choose among several mature ecosystems rather than a single national frontrunner.

DeepSeek's official repositories currently document the V3/V3.2 line and the R1 reasoning family; an official DeepSeek-R2 release is not listed as of August 2026. Qwen3 adds dense and mixture-of-experts checkpoints with thinking and non-thinking modes, while Moonshot AI's Kimi K2 focuses on coding, tool use, and agentic workflows.

  1. DeepSeek V3/V3.2 and R1: DeepSeek-V3 uses a mixture-of-experts architecture with 671B total parameters and 37B activated per token, while R1 is the company's open reasoning line. V3.2 adds newer experimental work, but references to “R2” should be treated as speculation until DeepSeek publishes an official model or repository. Official DeepSeek repositories

  2. Qwen3: Alibaba's Qwen3 family includes dense and mixture-of-experts models in multiple sizes. Its signature feature is the ability to switch between thinking mode for harder reasoning tasks and non-thinking mode for faster general responses, with official support for more than 100 languages and dialects. Official repository

  3. Kimi K2: Moonshot AI's open-weight Kimi K2 is a 1T-parameter mixture-of-experts model with 32B activated parameters. It is optimized for coding, tool calling, and agentic tasks, and the official deployment guide supports engines including vLLM, SGLang, KTransformers, and TensorRT-LLM. Official repository

We also did a deep dive into Chinese models: Kimi K2 vs DeepSeek-R1 vs Qwen3 vs GLM-4.5

FAQ

What are the leading Chinese LLMs in 2026?

The most visible open-weight families include DeepSeek V3/V3.2 and R1, Alibaba's Qwen3, and Moonshot AI's Kimi K2. The best choice depends on whether the priority is reasoning, multilingual work, coding, tool use, deployment cost, or local control.

Is Qwen3 open source?

Qwen3 model weights and code are publicly available through Alibaba's official repositories and model pages. Developers should still check the license attached to the exact checkpoint they plan to use, because terms can vary across releases and derivatives.

What is DeepSeek-R2, and has it been released?

DeepSeek-R2 is a widely discussed name for a possible successor to DeepSeek-R1, but DeepSeek's official repositories do not list an R2 release as of August 2026. For factual comparisons, use the officially published R1 and V3/V3.2 materials rather than rumored R2 specifications.

What is Kimi K2 best used for?

Kimi K2 is designed for agentic intelligence, including coding, tool calling, and multi-step workflows. Its open weights and support across major serving engines also make it relevant to teams evaluating self-hosted agent infrastructure.

Can Chinese LLMs be deployed locally?

Many can, because DeepSeek, Qwen, and Kimi publish open weights or official deployment guidance. Small Qwen3 checkpoints are accessible on modest hardware, while the largest mixture-of-experts models usually require quantization, multiple accelerators, or a hosted inference provider.

If you’ve found this article valuable, subscribe for free to our newsletter.

We post helpful lists and bite-sized explanations daily on our X (Twitter). Let’s connect!

Reply

Avatar

or to participate

Keep Reading

View more
caret-right