Skip to main content
Qwen Image 3.0 Pro is Alibaba’s third-generation image generation model, built around useful output: infographics, documents, and instruction-based editing rather than pure aesthetics. It is positioned as a tool for creating dense, information-rich visuals in a single pass, with the tagline “Rich Content, Authentic Details, Deep Knowledge.” The model accepts prompts of up to 4,500 tokens, roughly 4.5 times the input length of the previous generation. It uses that budget to compose newspaper pages, multi-panel infographics, academic papers with legible math notation, and other high-density layouts without compressing the brief into a few sentences. It also renders legible small text down to about 10 pixels across 12 languages and 20+ fonts, and integrates world knowledge for tasks like weather maps and live-stream interface replication.

What Qwen Image 3.0 Pro is good at

  • Ultra-long prompts: Accepts up to 4,500 tokens of instructions for detailed, information-dense compositions
  • Dense, practical layouts: Generates full newspaper pages, multi-panel infographics, and academic documents with formulas in one pass
  • Small text rendering: Produces legible text down to roughly 10px across 12 languages and 20+ fonts
  • World-knowledge integration: Generates content grounded in real-world information, such as weather maps for named cities
  • Instruction-based editing: Edits up to 3 reference images with targeted changes to objects, styles, or regions
  • Batch generation: Generates 1 to 6 images per run

Example outputs

Text-to-image generation, with the model handling layout, typography, and composition: Qwen Image 3.0 Pro text to image example Instruction-based image editing, preserving the original structure: Qwen Image 3.0 Pro image edit example

Use it in ComfyUI

Qwen Image 3.0 Pro workflows

Run the text-to-image and image edit workflows in ComfyUI, locally or on Comfy Cloud