Skip to content
Daily Edition · AI industry recordEdition of Sunday, September 6, 2026
Live desk ●
LaunchNews Report2 min readUpdatedByAI Tools Daily

Kimi K2.5 Released: Trillion-Parameter MoE Trained on 15T Mixed Visual-Text Tokens

Moonshot AI released and open-sourced Kimi K2.5 on January 27, 2026, adding large-scale visual data atop the K2 trillion-parameter base.

ShareXLinkedIn

Moonshot AI officially released Kimi K2.5 on January 27, 2026: built on the trillion-parameter MoE architecture and trained on 15 trillion mixed visual and text tokens, with clear gains over K2 in multimodal understanding and agentic tasks.

Since K2 debuted in July 2025 as the world's largest open model, Moonshot has kept a six-month cadence — K2.6 followed in April 2026. China's open-source trio — DeepSeek, Qwen and Kimi — now rotate steadily at the front of the open-model pack.

For developers, K2.5 remains open and commercially usable, deployable via the Kimi open platform and mainstream inference frameworks.

Vision in Pretraining: Open Source Closes the Multimodal Gap

What matters most about K2.5 is not the parameter count — the trillion-parameter MoE base carries over from K2 (see our K2 coverage) — but the training recipe: 15 trillion mixed visual and text tokens. China's open-source trio had excelled at text reasoning and coding, with multimodal understanding the main gap versus closed flagships. Mixing vision into pretraining rather than bolting on an adapter means agents can natively see screens, read charts and drive interfaces — the prerequisite for browser-operation and computer-use agent tasks.

The six-month cadence is its own signal: from K2 (July 2025) to K2.5 (January 2026) to K2.6 (April 2026), Moonshot pulled open-flagship refresh rates level with closed vendors, while DeepSeek and Qwen iterated in parallel (see our R1 and Qwen3 coverage) — the trio's rotation keeps compressing the old assumption that open source trails closed by half a year.

Our Take

K2.5 marks the open-source arms race pivoting from parameter counts to modality mix and agent practicality: scale is no longer news; the modality ratio of training data is where differentiation lives. For anyone choosing a stack: agent workflows that need visual understanding can now seriously evaluate open models rather than merely tolerate them. For the industry: the trio's six-month rhythm means every closed vendor's lead window keeps shrinking, and the inference price war will only intensify (see our DeepSeek R2 coverage).

This article aggregates official announcements and public reporting; original sources are linked below.

AI Tools Daily is a bilingual newsroom covering AI tool launches, product updates and industry trends. Editorial standards · Report a correction

Tools in this story

All stories in this section · Launch