Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
OmniGen v2
#55 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
OmniGen v2
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
OmniGen v2
- + Excellent photographic quality and realistic material textures
- + Accurate lighting and reflections on the sphere and glass surfaces
- + Perfect spatial arrangement following the prompt's logical constraints
- − The plant is quite blurred in the background, making it less 'partially visible through the glass' and more like a backdrop
Qwen Image 2.0
- + Successfully shows the plant clearly through the glass walls
- + Accurate color palette for the red book and blue sphere
- − The blue sphere is floating unnaturally in the center without support
- − Optical distortions in the glass are physically incorrect, creating phantom spheres
- − The glass cube has internal vertical partitions not requested
Verdict: OmniGen v2 produced a much more coherent and realistic image with physically accurate light interactions and placement, though the plant is very out of focus. Qwen Image 2.0 struggled with the physics of the scene, resulting in a floating sphere and confusing internal reflections that look like multiple objects.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
OmniGen v2
- + Excellent handling of reflections on the wet pavement
- + Colors are vibrant and the lighting is cinematic
- + Clear representation of light rain particles
- − The subject is standing and holding the bike rather than repairing it
- − Composition feels slightly generic and AI-generated rather than a 'candid street photo'
- − Anatomical issues where his foot blends into the bike pedal
Qwen Image 2.0
- + Strong adherence to the 'rapairing' action with a crouching pose
- + Captures the 'imperfect framing' and 'candid' feel much better
- + High realism in skin textures and clothing details
- − The motion blur on the passing car is quite subtle
- − The background character's hand on the right is an unnecessary distraction
- − Wet pavement reflections are less pronounced than in the other model
Verdict: Qwen Image 2.0 followed the specific intent of the prompt much better, capturing a realistic 'candid' moment of an elderly man actually performing a repair. OmniGen v2 produced a more aesthetically pleasing image with beautiful reflections, but the subject is simply posing with the bicycle rather than fixing it.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
OmniGen v2
- + Excellent soft lighting and clean facial features
- + Very high quality engraving details on the armor
- + Artistic and pleasing composition
- − Character looks too clean and 'pristine' for the battle-worn request
- − Dirt on face looks more like beauty marks than actual grime
Qwen Image 2.0
- + Perfectly captures the 'battle-worn' aesthetic with realistic scars and grit
- + Excellent interpretation of hair beads and textured cloth layers
- + Strong environment interaction with sparks and fire
- − Anatomy issues with the hand resting on the sword
- − The skin texture is slightly over-sharpened and coarse
Verdict: Qwen Image 2.0 followed the prompt much more accurately regarding the 'battle-worn' and 'highly detailed texture' requirements, providing a grit that OmniGen v2 lacked. While OmniGen v2 produced a more traditionally beautiful image with cleaner lines, it ignored the requested scars and dirt which were central to the character's description.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
OmniGen v2
- + Excellent structure that mimics a physical two-page menu layout
- + Strong use of vibrant colors and design accents to create a modern aesthetic
- + Well-defined sections for different food categories with clear headings
- − Significant spelling errors in prominent headings like 'RESTAURATED MENTS'
- − Body text is purely illegible scribbles
- − Font choice feels a bit dated compared to the modern minimalist prompt
Qwen Image 2.0
- + Perfect adherence to the requested grid layout for food photos
- + High-quality, appetizing food photography that looks professional
- + Accurate spelling of main category headings like 'APPETIZERS', 'PIZZA', and 'MAINS'
- − Individual item names and small text are mostly gibberish
- − The layout feels more like a digital app interface than a traditional restaurant menu
- − Pricing is repetitive and unrealistic across all items
Verdict: Qwen Image 2.0 followed the prompt's layout requirements much better, providing a clean grid and correctly spelling the primary category headers. OmniGen v2 struggled with basic spelling and created a more cluttered design, whereas Qwen Image 2.0 produced high-quality imagery that perfectly fits a modern casual dining theme.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
OmniGen v2
- + Excellent text legibility and clean graphic design layout
- + Perfect rendering of the starburst price tag as requested
- + Consistent lighting across all visual elements
- − Failed to provide the 'exploded' view where components are suspended
- − Missing the Euro symbol (€) in the price tag
- − The burger looks more like a 3D illustration than photorealistic food photography
Qwen Image 2.0
- + Successfully captured the 'exploded' view with suspended components
- + Beautifully rendered fiery/glow effect on the text as requested
- + High level of texture detail in the food and sauce droplets
- − The '6' in the price has some rendering artifacts
- − Text placement is slightly cluttered at the top
Verdict: Qwen Image 2.0 is the clear winner as it followed the complex structural prompt for an 'exploded burger' and included the specific '€' symbol and fiery text effects perfectly. OmniGen v2 failed to provide the exploded deconstruction and produced a more generic, static advertisement with missing currency symbols.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
OmniGen v2
- + The text has a clear chalk-like texture.
- + Follows the general layout of a menu board.
- − Numerous spelling errors including 'SPECALS' and garbled menu item names.
- − The handwriting looks more like a digital font than natural handwriting.
- − Fails the prompt's request for elegant cursive in the title.
Qwen Image 2.0
- + Excellent typography with realistic cursive and natural handwriting variations.
- + High adherence to the prompt including specific spelling, dates, and prices.
- + Fantastic environmental composition with a realistic café background and lighting.
- − None notable for this specific prompt.
Verdict: Qwen Image 2.0 perfectly executed the complex text requirements, including correct spelling of all items and the specific date requested. OmniGen v2 struggled significantly with text rendering, resulting in several typos and illegible words, and it failed to capture the 'cozy café' atmosphere by only showing the board. Qwen Image 2.0 is the clear winner for its superior realism, composition, and prompt adherence.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
OmniGen v2
- + Clean, vector-like aesthetic with high contrast.
- + Accurately represents an astronaut and horse in a space setting.
- − Failed the specific spatial instruction; the astronaut is on top, not the horse.
- − The background and lighting feel somewhat flat and simplified.
- − The horse's anatomy is slightly stiff.
Qwen Image 2.0
- + Highly detailed textures, especially on the horse's scales and the spacesuit.
- + Dynamic, cinematic composition with bubbles and a planetary background.
- + Superior lighting and sense of movement.
- − Failed the specific spatial instruction; the astronaut is riding the horse.
- − Minor artifacts in the reins and mane.
Verdict: Both OmniGen v2 and Qwen Image 2.0 failed the specific logic-defying constraint of 'horse on top, not vice versa,' opting instead for the more traditional 'astronaut riding a horse.' Qwen Image 2.0 is the superior image due to its stunning cinematic detail, textures, and dynamic composition, whereas OmniGen v2 looks much more like a basic digital illustration.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
OmniGen v2
- + Clean, high-resolution rendering of the taxi exterior and driver cap.
- + Good lighting on the capybara's face.
- + Accurate interpretation of the jacket and cap.
- − Anatomical failure: the capybara has human hands holding the steering wheel.
- − Composition error: the woman appears to be in the front passenger seat rather than the back seat.
- − The background lights are generic and lack the specific feel of Manhattan.
Qwen Image 2.0
- + Successfully rendered capybara paws on the steering wheel.
- + Correct spatial layout with the passenger clearly in the back seat.
- + The background environment and reflections look much more like a realistic NYC street.
- − The capybara's head is slightly oversized for the body/seat.
- − Minor artifacts in the passenger's facial details due to depth of field.
Verdict: OmniGen v2 fails significantly on anatomy by giving the capybara human hands and placing the passenger in the front seat. Qwen Image 2.0 captures the requested scene much more accurately, including the capybara's paws, the passenger's placement in the back, and a more convincing New York City atmosphere.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
OmniGen v2
- + Features a bold, high-contrast graphic style
- + Includes a clear stylized border with webs as requested
- − Several text errors including 'FRIGTS', 'ARCAS', and garbled text in the scroll
- − The layout is cluttered and the text overlays parts of the background awkwardly
- − Lacks the specific 'thorns' element in the border
Qwen Image 2.0
- + Excellent text rendering with near-perfect spelling and elegant gothic typography
- + Stronger adherence to the 'vintage' and 'parchment' aesthetic with thorns and webs in the border
- + Superior composition with a more realistic and cinematic Jack-o-lantern
- − The '7pm' text is slightly small compared to the other event details
- − The background trees are a bit repetitive in their placement
Verdict: Qwen Image 2.0 is the clear winner as it followed all complex text instructions perfectly, whereas OmniGen v2 struggled with spelling and scroll placement. Qwen also captured the vintage gothic atmosphere much more effectively with its use of textures, thorns, and lighting.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent 3D isometric diorama aesthetic following the 'miniature' prompt
- + Clean, bold 3D typography that fits the graphic design style
- + Very consistent cartoonish 3D rendering with soft lighting
- − The flag icon is generic and does not represent Japan
- − The sushi anatomy is slightly surreal, mixing nigiri and maki elements oddly
Qwen Image 2.0
- + Features a correct and clear Japanese flag icon
- + Higher variety of sushi types with realistic material textures
- + Clean typography and accurate placement
- − Leans more toward a realistic photo on a plate rather than a '3D cartoon' or 'miniature' style
- − The wood grain on the base looks like a flat photo rather than a 3D isometric model
Verdict: OmniGen v2 much better captures the requested '3D cartoon isometric' style, creating a cohesive diorama that feels like a stylized miniature. While Qwen Image 2.0 has more realistic sushi and a correct flag, it feels more like a standard food photograph composited with text, failing the 'cartoon' and 'isometric 3D' stylistic requirements of the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
OmniGen v2
- + Features vibrant, saturated colors that create a cheerful atmosphere.
- + Follows the request for big expressive eyes in a stylized way.
- − Fails to be hyper-photorealistic, appearing more like a 3D animation or digital illustration.
- − Missing one of the requested animals, showing only three instead of four.
- − The composition is static and lacks the 'playfully chasing' and 'tumbling' action requested.
Qwen Image 2.0
- + Successfully includes all four distinct animals: golden retriever, tabby kitten, bunny, and fox kit.
- + Achieves a high level of photorealism with professional-grade lighting and fur textures.
- + Captures the dynamic action of the animals tumbling and playing together as requested.
- − The fox kit is partially obscured and pinned at the bottom, making its details harder to see.
- − The butterfly on the right has slightly awkward placement against the kitten's head.
Verdict: Qwen Image 2.0 is the clear winner as it followed all prompt instructions, including the specific list of four animals and the requested action of 'tumbling together', whereas OmniGen v2 only included three animals and had a static composition. Furthermore, Qwen Image 2.0 achieved a true hyper-photorealistic aesthetic, while OmniGen v2 produced a stylized, cartoon-like image that ignored the realism requirement.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
OmniGen v2
- + Strong minimalist vector aesthetic
- + Clean balanced composition
- + Accurate establishment date rendering
- − Significant spelling error in the main brand name
- − The steam element is overly simplistic and disconnected
Qwen Image 2.0
- + Perfect text rendering for the brand name including accents
- + Beautiful vintage illustration style with subtle texture
- + Creative integration of the steam inside the cloche
- − The ribbon tail on the right is slightly awkward in its curl
- − Less minimalist than requested, leaning more towards illustrative
Verdict: Qwen Image 2.0 is the clear winner because it correctly spells 'Caffè Florian' and handles the typography beautifully, whereas OmniGen v2 fails on the primary text. Qwen Image 2.0 also provides a much more cohesive and visually appealing interpretation of the cloche and steam elements.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
OmniGen v2
- + Adheres to the color palette well.
- + Captures the modern vector layout with a clear grid structure.
- − Confuses Apollo 11 with 'Apollo 17'.
- − Contains significant gibberish text and spelling errors like 'NSA' instead of NASA.
- − Fails to include specific requested icons like the Saturn V rocket.
Qwen Image 2.0
- + Excellent adherence to the requested 6-step sequence with accurate icons for each.
- + High text legibility and mostly correct spelling including names of the astronauts.
- + Strong thematic arrangement that visualizes the trajectory from Earth to Moon surface.
- − Includes a slight misspelling of 'Translunar' as 'Translunjar'.
- − Shadowing on the lunar module and Moon surface is slightly more detailed than a strictly 'flat-vector' style.
Verdict: Qwen Image 2.0 followed the complex multi-step prompt almost perfectly, providing accurate icons and logical flow for the Apollo 11 mission. In contrast, OmniGen v2 failed on historical accuracy (labeling it Apollo 17), text rendering, and specific icon requirements, resulting in a generic and nonsensical infographic.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request