Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
OmniGen v2
#57 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
OmniGen v2
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
OmniGen v2
- + Excellent adherence to lighting and placement instructions.
- + Realistic glass refraction and reflections on the table surface.
- + Clean and minimalist photographic aesthetic.
- − The glass cube appears to have an open top or lacks a visible top lid under the book.
Qwen Image Max
- + Strong visual quality with nice texture on the red book.
- + Good implementation of directional window light and shadows.
- + Shows the plant reflected clearly in the bottom of the cube.
- − The blue sphere is awkwardly floating in the center rather than resting on the bottom.
- − The plant is positioned more to the side than 'behind' the cube as requested.
Verdict: OmniGen v2 followed the prompt's spatial instructions more accurately, placing the sphere on the bottom of the cube and the plant directly behind it. While Qwen Image Max has slightly more interesting textures and shadows, the floating sphere makes the physics of the scene feel less grounded compared to OmniGen v2.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
OmniGen v2
- + Clean and aesthetic composition with a strong shallow depth of field.
- + Accurate braided hair with small beads as requested.
- − The character looks pristine and youthful rather than 'battle-worn'.
- − Dirt and scars look like cosmetic spots rather than realistic injuries.
Qwen Image Max
- + Excellent depiction of 'battle-worn' with realistic skin textures, deep scars, and grime.
- + Superior detail in the engraved plate armor and worn leather textures.
- + Very effective torchlight lighting and bokeh sparks that feel integrated into the scene.
- − The bokeh sparks are slightly overwhelmed by digital noise/grain in some areas.
Verdict: Qwen Image Max is the clear winner as it fully captures the 'battle-worn' aesthetic with grit and high-detail textures, whereas OmniGen v2 produced a character that looks too clean and polished for the prompt. Qwen Image Max also delivered much more intricate engraving on the armor and a more convincing torchlight atmosphere.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
OmniGen v2
- + Excellent text legibility and graphic design layout
- + Accurately includes the price starburst element
- + Clean, commercial aesthetic suitable for a fast-food ad
- − Fails the 'exploded' requirement as the burger is fully assembled
- − Texture is more illustrative than photorealistic
- − Text for 'Limited Time Only' is cut off slightly on the left
Qwen Image Max
- + Successfully depicts an exploded burger with suspended components
- + High level of photorealistic detail and dynamic motion effects
- + All text requirements are met with the requested fiery, glowing effect
- − The '€6.99' text has a slight visual artifact on the currency symbol
- − The overall composition feels a bit more cluttered compared to A
Verdict: Qwen Image Max followed the prompt more accurately by actually creating an 'exploded' burger with mid-air components and motion, whereas OmniGen v2 produced a standard static burger. Qwen Image Max also delivered more realistic food textures and superior fiery typography effects, despite OmniGen v2 having a cleaner graphic design.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
OmniGen v2
- + Strong image clarity and vibrant colors.
- + The woman is sitting in the back seat as requested.
- − Major anatomical failure: human hands are attached to the capybara's body.
- − The woman is in the front passenger seat, not the back seat.
- − The capybara's head looks like a poorly blended overlay on a human body.
Qwen Image Max
- + Excellent anatomical interpretation with capybara-like paws on the wheel.
- + Authentic vintage NYC taxi interior with realistic lighting and background details.
- + Better character placement and overall cinematic composition.
- − The passenger is technically in the passenger seat rather than the back seat.
- − The roof of the taxi has some digital artifacts and strange textures.
Verdict: Qwen Image Max is the superior choice because it correctly renders the capybara's paws on the steering wheel, whereas OmniGen v2 gives the animal human hands. While Qwen Image Max placed the woman in the front seat instead of the back, it captured the professional demeanor and realistic taxi aesthetic much more effectively than OmniGen v2.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent typography for the main title heading.
- + Clean parchment texture effect.
- + High contrast colors that make the text pop.
- − Failed to render the specific details 'The Arches' and 'of frights' correctly, with several typos.
- − The illustration style is more vector-like and less 'cinematic' than requested.
- − The scroll banner text is garbled and unreadable.
Qwen Image Max
- + Perfect text rendering for all requested details, including date, time, and location.
- + Rich, atmospheric illustrations with thorny borders and detailed bats.
- + Superior composition that balances the gothic font with the central jack-o-lantern.
- − The transition between the central illustration and the parchment background is slightly soft.
- − The scroll banner is slightly overlapping the pumpkin base.
Verdict: Qwen Image Max is the clear winner as it successfully rendered all the requested text details without a single typo, which is crucial for an invitation. While OmniGen v2 has a bold style, its failure to accurately spell the location and the scroll message makes it unusable as a party invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent 3D isometric perspective on a diorama base.
- + Clean, bold text rendering that matches the requested layout.
- + Vibrant colors and a very clean, minimal aesthetic.
- − The flag icon is generic and does not represent the Japanese flag.
- − The sushi pieces have a strange hybrid appearance with tails, looking somewhat AI-generated rather than realistic.
Qwen Image Max
- + Accurately depicts the Japanese flag icon.
- + Higher variety and realism in the sushi models (maki, nigiri, and tamago).
- + Excellent material rendering, particularly on the glass diorama base.
- − The 'JAPAN' text is slightly cut off at the top of the frame.
- − The 'SUSHI' text has minor shadow artifacts on the letter 'I'.
Verdict: Qwen Image Max is the clear winner as it followed all specific prompt instructions, including the Japanese flag icon and providing a realistic variety of sushi. While OmniGen v2 has a very pleasing 3D cartoon style, its failure to render the correct flag and the slightly confusing anatomy of its sushi pieces make it less successful overall.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
OmniGen v2
- + Good adherence to the requested butterfly count.
- + Vibrant lighting and colors that create a wholesome vibe.
- − Failed to include a rabbit, instead merging features of a fox and rabbit.
- − Stylization is very cartoonish and lacks the 'hyper-photorealistic' quality requested.
- − Composition is static and lacks the 'tumbling' and 'chasing' actions.
Qwen Image Max
- + Successfully captures a realistic, photographic style with high-quality fur textures.
- + Dynamic composition perfectly illustrates the animals tumbling and playing together.
- + Excellent execution of god rays and dew sparkles in a lush meadow.
- − Failed to include the baby bunny in the scene.
- − The fox's paws have slight anatomical blurring where they interact with the dog.
Verdict: Qwen Image Max is the clear winner as it delivered a sophisticated, professional-grade photograph that accurately captured the playful movement and realistic textures requested. OmniGen v2 produced a generic, clip-art style illustration that failed to meet the 'hyper-photorealistic' requirement and struggled with animal anatomy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
OmniGen v2
- + Clean vector-style execution
- + Good color palette adherence
- + Accurate date rendering
- − Significant spelling error in the primary text ('CAFFFLORIN')
- − Typography is somewhat overlapping and cramped within the ribbon
- − Minimalist style is a bit too simplified and lacks character
Qwen Image Max
- + Perfect text rendering including the accented 'É'
- + Excellent artistic texture on the background
- + Superior illustrative details on the cloche dome and steam
- − The 'Est. 1720' is on a banner but the main brand name is not
- − The steam is slightly more ornate than 'minimalist' usually suggests
Verdict: Qwen Image Max is the clear winner as it correctly spelled the brand name 'Caffè Florian' with the proper accent, whereas OmniGen v2 failed significantly on the text. Qwen Image Max also provided much better texture and artistic depth while still maintaining the requested vintage color scheme and logo structure.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.