Qwen 3 Released: MoE Architecture Sets a New Bar for Open Source
Alibaba releases the Qwen 3 series; the flagship Qwen3-235B-A22B rivals top models like DeepSeek-R1 and o1 on major benchmarks.
Alibaba Cloud officially released the Qwen 3 series on April 29, 2025. The flagship Qwen3-235B-A22B uses a mixture-of-experts (MoE) architecture with 235 billion total and 22 billion active parameters, rivaling top models like DeepSeek-R1 and o1 on coding, math and general-capability benchmarks.
Qwen 3 open-sourced eight models at once — 2 MoE and 6 dense models from 0.6B to 235B — covering deployment from on-device to cloud. It pioneered a hybrid reasoning mode that switches seamlessly between thinking and fast responses, with native support for 119 languages and dialects.
The Full-Size Matrix: An Ecosystem Play, Not a Single Shot
Eight models in one drop, from 0.6B edge to the 235B flagship — that is Qwen's core difference from single-model rivals: whether you run on a phone, one GPU or a cluster, there is a size inside the same stack. This family lock-in is exactly how Llama built its ecosystem, now replicated at a faster cadence. On Hugging Face's derivative-model boards, the Qwen line has long led downloads and fine-tunes.
Hybrid Reasoning: Thinking as a Switch
The pioneering hybrid mode lets one model toggle between deep thinking and fast response — turning the reasoning paradigm o1 opened (see our coverage) into a controllable option rather than a separate product line. It beat GPT-5's automatic routing to the idea by three months, converging on the same thesis: reasoning ends up a built-in capability of every model, not a category.
The Cost Card: Deployment at a Quarter of the Price
Officially, deployment cost is roughly 1/4 that of comparable models — and with Apache 2.0 commercial use, that hits the core demand of enterprise self-hosting (see our local LLM deployment review). Together with DeepSeek's cheap APIs and Kimi's trillion-parameter open weights, the Chinese relay pushed open-source competition from capability parity to price-performance dominance (see our coverage).
Our Take
Every model ships Apache 2.0 and commercially usable; analysts read Qwen 3 as China's open models entering the global first tier. The deeper lesson is ecosystem-position logic: benchmark scores get eclipsed by the next release, but once developers build toolchains and train derivatives on your full-size matrix, migration costs compound — Qwen is not winning a launch, it is occupying a position.
This article aggregates official announcements and public reporting; original sources are linked below.
Source:通义千问官方博客
AI Tools Daily is a bilingual newsroom covering AI tool launches, product updates and industry trends. Editorial standards · Report a correction