YYaaa News

Qwen-Image-3.0 released | 4.5k tokens in, 9-grid infographics out, 10px text rendered accurately

TL;DR

QwenTeam released Qwen-Image-3.0 on July 21, pushing input length to 4.5k tokens (from 1k). It can produce 9-grid infographics, newspapers, exam papers, storyboards, and multi-layer UI mocks in a single generation. Detail-level rendering handles 10px text, academic formulas, hair strands and skin pores. Supports 12 languages, 100+ art styles, and can pull in live web info.

QwenTeam on July 21 released Qwen-Image-3.0 — repositioning image generation from "pretty" to "actually usable." The headline change: up to 4.5k token input (Qwen-Image-2.0 was 1k), enabling single-shot generation of high-information-density outputs — 9-grid infographics, newspapers, exam papers, storyboards, multi-layer nested UI mocks.

Detail-level rendering is the real gain. Qwen-Image-3.0 can accurately render 10 px small text, academic formulas, and paper annotations; at the micro-texture layer, hair strands, skin pores and restoration of traditional paintings all reproduce. It's the first Chinese-native image model to sell "pore-level detail" as headline capability — a lane previously held by Midjourney v6 and FLUX.1.

Supports 12 languages, 100+ art styles, and UI-mock simulation — the last one means Qwen-Image-3.0 can directly generate product prototype screenshots and web dashboard mocks. It can also fold in live web info for more grounded outputs — that's the direct payoff of Qwen's ecosystem integration: search + retrieval + image closed inside one API.

Timeline context — Kimi K3 just open-sourced, Qwen3.8-Max Preview just launched, Qwen-Image-3.0 today. Alibaba and Moonshot have shipped three frontier releases in mid-late July, running at roughly twice the release cadence of OpenAI + Anthropic.

Behind the specs is a cadence contrast — Chinese labs are turning "one frontier release every 7 to 10 days" into a norm.

via QwenTeam