Head to head
Esc

Models · slot A

to navigate to pick

OmniGen v2 VectorSpaceLab Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

OmniGen v2

16.8 arena score

#57 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image

21.4 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

OmniGen v2

0.0%

win rate

Ties

0.0%

Qwen Image

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent prompt adherence with all objects correctly placed.
  • + Superior glass rendering with realistic thickness and refractive index.
  • + High visual clarity and vibrant colors.
  • The blue sphere appears slightly too large, taking up much of the internal space.
  • The plant's positioning relative to the glass is less impactful for demonstrating visibility through glass compared to Model B.

Qwen Image

  • + Follows all prompt instructions accurately.
  • + The composition feels more balanced with the 'small' sphere being appropriately sized.
  • + Includes a clear reflection of the blue sphere on the base of the glass cube.
  • The glass cube has strange internal vertical lines that don't match standard cube geometry.
  • The overall image is slightly softer and less crisp than OmniGen v2.

Verdict: Both models successfully followed a complex spatial prompt involving multiple objects and interactions. OmniGen v2 (Image A) produced a more polished and physically convincing glass material, while Qwen Image (Image B) did a better job with the relative scale of the 'small' sphere and included realistic reflections on the bottom surface. OmniGen v2 wins slightly due to higher detail and better rendering of the glass edges.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent handling of wet pavement reflections
  • + Vibrant colors on the red bicycle
  • The man is standing still rather than actively repairing the bike
  • Image looks overly clean and lacks the requested motion blur from cars

Qwen Image

  • + Successfully captures the man in a 'repairing' pose
  • + Includes the requested motion blur on the background vehicles
  • Some structural issues with the bicycle frame and pedal placement
  • The rain effect appears somewhat static and thick

Verdict: OmniGen v2 produces a cleaner, more aesthetically pleasing image but fails to capture the active 'repairing' action and motion blur requested. Qwen Image follows the prompt more accurately by showing the man working on the seat and incorporating background motion, despite more noticeable AI artifacts on the bike itself.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent soft lighting and bokeh effect
  • + High-quality skin texture and lifelike eyes
  • The character looks too pristine and clean for the 'battle-worn' description
  • The dirt on the face appears like stylized freckles rather than actual grime

Qwen Image

  • + Perfect adherence to the 'battle-worn' and scarred descriptor
  • + Highly detailed engraving on the plate armor and leather straps
  • + Strong creative interpretation of the beads in the braids
  • The torch flame looks somewhat artificial and illustrative compared to the rest of the image
  • Some messy hair-braid anatomy near the top of the head

Verdict: Qwen Image is the clear winner for its superior adherence to the narrative elements of the prompt, successfully capturing the grit, scars, and complex armor textures of a battle-worn soldier. While OmniGen v2 produced a beautiful portrait, it failed to convey the 'battle-worn' aspect, presenting a character that looks too clean and polished.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Features a comprehensive multi-column layout.
  • + Includes a wide variety of food photography examples.
  • Numerous spelling errors in headings like 'Apptetizes' and 'Pizzzzan'.
  • The layout feels slightly fragmented across the center fold.

Qwen Image

  • + Excellent typography for headers with a modern aesthetic.
  • + Clean, professional grid composition that aligns well with the prompt.
  • The placeholder text for menu items is illegible.
  • Grouping 'Pizza/mains' into one section simplifies the requested three-part structure.

Verdict: Qwen Image delivers a more cohesive and visually pleasing design with superior font choices and a clean grid layout. While OmniGen v2 attempts a more complex document structure, its frequent spelling errors and cluttered columns make it less functional than Qwen Image's streamlined design.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent text legibility for the main title
  • + Clean, professional graphic design aesthetic
  • + Clear 6.99 price display within a starburst
  • Completely failed the 'exploded' and 'suspended' burger prompt
  • Background lacks the requested glowing embers and fire details
  • Missing the Euro symbol (€) requested in the prompt

Qwen Image

  • + Successfully captured the 'exploded' and 'deconstructed' burger concept with suspended ingredients
  • + Excellent atmospheric lighting with embers and fire
  • + Accurately included the Euro symbol with the price
  • The 'MAGIC BURGER' text has a slight glow but is less impactful than Model A
  • Layout is a bit more cluttered compared to a standard advertisement

Verdict: Qwen Image followed the complex 'exploded burger' and 'suspended ingredients' prompt much better than OmniGen v2, which produced a static, fully assembled burger. While OmniGen v2 has cleaner title text, Qwen Image captured the specific Euro currency symbol and the requested fiery atmosphere far more effectively.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Features more realistic, textured chalk strokes
  • + Captures the wooden frame tightly for a centered focus
  • Several spelling errors in the menu items
  • Garbled text and nonsensical overlapping letters in the middle section
  • Missing the 'cozy café' background requested in the prompt

Qwen Image

  • + Excellent legibility and adherence to the specific menu text
  • + Successfully includes the 'cozy café' background for better context
  • + Logical layout that mimics a real-world restaurant chalkboard
  • The year is incorrectly rendered as '20026' instead of '2026'
  • The handwriting looks slightly more like a digital font than natural chalk

Verdict: Qwen Image followed the complex prompt instructions much better, correctly rendering the specific menu items and price points with a beautiful café background. While OmniGen v2 has a more realistic chalk texture, its text is riddled with hallucinations and spelling errors, making it unusable for a specific menu challenge.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Clean, vibrant colors with a strong cinematic glow effect.
  • + Clear and detailed rendered space elements like stars and moons.
  • Prompt adherence failure: the astronaut is riding the horse instead of the 'horse on top'.
  • The leg of the astronaut and the saddle area have fused, unnatural geometry.

Qwen Image

  • + Better cinematic composition with realistic lighting from the planet below.
  • + More natural integration of the horse and astronaut within the environment.
  • Prompt adherence failure: ignored the 'horse on top, not vice versa' instruction.
  • The horse's back right leg has a strange hoof orientation.

Verdict: Both models failed the specific instruction for a 'horse on top' of an astronaut, instead providing the standard astronaut-on-horse image. Qwen Image is the preferred winner because its composition, background depth, and lighting are significantly more cinematic and realistic than the flatter, more cartoonish aesthetic of OmniGen v2.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent capybara face rendering
  • + Strong cinematic lighting
  • Major anatomical failure with human hands on the capybara's body
  • The passenger is sitting in the front seat instead of the back seat

Qwen Image

  • + Follows spatial instructions with the passenger correctly placed in the back seat
  • + Correctly interprets the capybara's paws on the steering wheel
  • + Higher level of detail in the taxi exterior and driver uniform
  • The passenger's phone is rendered somewhat clumsily

Verdict: Qwen Image is the clear winner as it correctly followed the spatial prompt by placing the passenger in the back seat, whereas OmniGen v2 placed her in the front next to the driver. Additionally, OmniGen v2 suffered a significant logic failure by giving the capybara human hands, while Qwen Image successfully rendered capybara-like paws.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

OmniGen v2
Qwen Image
0% wins 0% ties 100% wins

AI Judge Analysis

OmniGen v2

  • + Excellent typography style that fits the gothic theme
  • + Vibrant cinematic lighting on the jack-o-lantern
  • + Clean parchment border effect
  • Significant spelling and layout issues in the body text
  • The secondary banner text is gibberish
  • The scroll banner is poorly integrated with the text below it

Qwen Image

  • + Successfully included the thorns in the border as requested
  • + Highly accurate text rendering for the event details
  • + Better composition with a more realistic 3D jack-o-lantern and twisted trees
  • Small typo in the title text ('Halle Party' instead of 'Halloween')
  • The top-most decorative text is illegible

Verdict: Qwen Image follows the prompt much more effectively by including the requested thorns in the border and accurately rendering the specific date, time, and location details. While OmniGen v2 has a strong aesthetic for the main title, its failure to generate legible body text makes it less functional as an invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Excellent typography with clean drop shadows
  • + High-quality textures on the salmon with realistic sheen
  • + Perfectly centered and clean composition
  • The flag icon does not represent the Japanese flag
  • Anatomically confused sushi that looks like a hybrid between nigiri and a roll

Qwen Image

  • + Includes a correct Japanese flag icon as requested
  • + Strong 3D miniature felt/plastic aesthetic that fits the 'cartoon scene' theme
  • + Better variety of sushi types presented on the diorama
  • Text is slightly less crisp and has a minor alignment glitch on the 'S' of SUSHI
  • Chopstick placement is slightly floating

Verdict: Both models followed the complex prompt requirements for an isometric diorama. Qwen Image is the preferred winner because it correctly rendered the Japanese flag icon and provided a more charming 'miniature' toy-like aesthetic, whereas OmniGen v2 failed the flag requirement and created confusing hybrid sushi structures despite having superior text rendering.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Captures the golden sunrise and god rays effectively.
  • + Vibrant colors and a very high-contrast, cheerful aesthetic.
  • Missed one of the four required animals (the baby bunny).
  • The style is very illustrative and cartoony, failing the 'hyper-photorealistic' requirement.
  • Animals have stiff, symmetrical poses rather than a playful 'tumbling' interaction.

Qwen Image

  • + Successfully included all four animals: puppy, kitten, bunny, and fox kit.
  • + Achieved a much higher level of photorealism with realistic fur textures and lighting.
  • + Dynamic composition with animals actually appearing to chase butterflies and interact.
  • Minor anatomical artifacts on the fox's paws.
  • The kitten's pose is a bit awkward in mid-air.

Verdict: Qwen Image is the clear winner as it followed the prompt's count and species requirements, whereas OmniGen v2 missed the bunny. Furthermore, Qwen Image adhered to the 'hyper-photorealistic' style request, while OmniGen v2 produced a generic, flat 3D animation style.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Clean vector emblem style with sharp lines
  • + Accurate rendering of the year '1720'
  • + Followed the color palette and texture requirements well
  • Misspelled the primary name as 'CAFFFLORIN'
  • Steam element is very simple and looks a bit disconnected

Qwen Image

  • + Includes the accent mark in 'CAFFÉ'
  • + Better visual representation of steam rising from the cloche
  • + Successfully incorporated subtle paper texture
  • Significantly garbled typography with overlapping and mismatched letters
  • The name appears as 'FL CAFÉ. Lorai AN' instead of 'Caffè Florian'
  • The cloche is filled-in rather than a minimalist line-based icon

Verdict: OmniGen v2 produced a much cleaner and professional-looking logo, though it committed a spelling error by merging the words and missing letters. Qwen Image followed the cloche design instructions well but failed significantly on typography and layout, resulting in a cluttered and illegible text block. OmniGen v2 is the winner because its cleaner aesthetic is more usable for a logo context despite the typo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

OmniGen v2
Qwen Image

AI Judge Analysis

OmniGen v2

  • + Clean layout with a consistent grid structure.
  • + Adheres to the requested color palette accurately.
  • + High visual clarity in the infographic icons.
  • Incorrect mission number specified as Apollo 17.
  • Severe text legibility issues and gibberish labels.

Qwen Image

  • + Successfully includes the Saturn V and crew names requested.
  • + Better logical flow showing the progression of the mission.
  • + Text is more recognizable and legible compared to the alternative.
  • Includes literal prompt instructions as text on the poster.
  • Icons and text labels are misaligned with their actual numbers.

Verdict: Qwen Image is the preferred choice because it captures the specific iconography of the Saturn V and crew names, whereas OmniGen v2 fails on the historical accuracy of the mission number. Although Qwen Image includes some Meta-text from the prompt, its overall composition and legibility are superior for an infographic task.

Next steps

Explore each model