Head to head
Esc

Models · slot A

to navigate to pick

Grok Imagine Image xAI Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Grok Imagine Image

23.4 arena score

#27 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image

21.4 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

Grok Imagine Image

0%

win rate

Ties

0%

Qwen Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent handling of glass refraction and reflections on the table.
  • + Realistic wood texture and lighting from the window.
  • + Correct inclusion of all requested elements with a high level of photorealism.
  • The blue sphere is levitating realistically but the cube's proportions are slightly tall.
  • The glass cube has very thick edges that feel a bit like acrylic.

Qwen Image

  • + Closer to a perfect geometrical cube shape.
  • + Clean, vibrant colors for the sphere and book.
  • + Good spatial arrangement of the objects on the table.
  • The plant is significantly less 'visible through the glass' than requested.
  • The lighting is flatter and lacks the directional shadow depth seen in the other image.
  • The bottom of the cube has an odd mirrored base effect that wasn't requested.

Verdict: Grok Imagine produced a superior result with much more convincing glass physics and lighting, accurately showing the plant distorted through the translucent medium. While Qwen followed the layout instructions well, its rendering is more sterile and it failed to capture the 'visible through the glass' aspect as effectively as Grok.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent depiction of motion blur on passing cars
  • + High realism with the inclusion of a face mask, common in Japan
  • + Very natural lighting and colors that feel like a real amateur photograph
  • The framing is very tight, obscuring the man's face entirely
  • The red bicycle is partially cut off at the bottom

Qwen Image

  • + Clearer view of the subject and his actions
  • + Good reflection on wet pavement
  • + Strong adherence to the 'elderly Japanese man' descriptor
  • Lacks motion blur on the background car despite the prompt request
  • The rain texture looks slightly artificial and uniform
  • The bicycle geometry is slightly warped near the pedals

Verdict: Grok Imagine Image captured the cinematic atmosphere better, particularly with the motion blur and the 'imperfect framing' that felt like a genuine candid street photo. Qwen Image followed the subject prompt well but failed to execute the motion blur and had more 'AI-typical' clean textures that lacked the gritty realism of Grok's output.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Breathtakingly detailed engravings on the plate armor with realistic lighting reflections.
  • + Extremely lifelike skin texture and convincing faint scars.
  • + Sophisticated composition with natural-looking bokeh and lighting integration.
  • The 'small beads' in the hair are somewhat understated and blend in a bit too much.

Qwen Image

  • + Excellent adherence to the 'hair braided with small beads' prompt with visible, colorful beads.
  • + Good representation of the leather straps and chainmail underlayer.
  • + Clear depiction of the torch being held in the scene.
  • The sparks have a very artificial, digital starburst look that lacks realism.
  • The scars look like painted-on makeup rather than actual skin tissue.
  • The armor engraving is less refined and looks more like a pattern overlay than sculpted metal.

Verdict: Grok Imagine produces a much more cinematic and high-fidelity image with superior skin textures, lighting, and armor detail. While Qwen Image followed the specific bead instruction more literally, the overall visual quality in Grok Imagine — particularly the realistic integration of sparks and the intricate etchings on the armor — makes it the far better interpretation of a 'battle-worn' character.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography with nearly perfect spelling of headings and items
  • + Structured sections for appetizers, pizza, and mains as requested
  • + Clean, high-quality food photography integrated into the layout
  • Repetitive menu items like 'Steak Frites' and 'Grilled Salmon' appearing multiple times
  • Layout is a bit cluttered compared to a truly minimalist aesthetic

Qwen Image

  • + Strong minimalist aesthetic with a clear grid system
  • + Vibrant, colorful background blocks for the food photos
  • + Excellent use of white space and bold sans-serif font
  • Poor text rendering with several illegible words and typos like 'Pizzaurant'
  • Lacks separate sections for pizza and mains as requested in the prompt
  • Food photos are somewhat repetitive in terms of subject matter

Verdict: Grok Imagine Image produced a much more functional and professional menu with mostly accurate text and distinct sections. While Qwen Image captured the 'minimalist' and 'grid' aspects of the prompt more artistically, its failure to generate legible text and its omission of distinct menu categories makes it less effective as a design piece.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent text integration with fire and glow effects that match the prompt perfectly.
  • + Highly photorealistic textures on the lettuce, patty, and tomatoes.
  • + Dynamic composition with splashes of sauce that enhance the 'exploded' feel.
  • The starburst for the price is a bit generic and flat compared to the rest of the 3D scene.

Qwen Image

  • + Strong neon-like 'MAGIC BURGER' text rendering.
  • + Good depth of field with the embers in the foreground and background.
  • The burger components are mostly stacked rather than 'exploded' as requested.
  • The starburst and price text lack the 'fiery, glowing effect' requested.
  • The textures look slightly more plastic compared to the other model.

Verdict: Grok Imagine followed the prompt much more effectively, particularly regarding the 'exploded' nature of the burger and the specific fiery effects on the text. Qwen produced a more traditional burger stack with some floating pieces, whereas Grok Imagine successfully integrated all design elements into a cohesive, high-energy advertisement.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Perfect adherence to text content, including the spelling and year.
  • + Highly realistic chalk texture with genuine-looking smudges and dust patterns.
  • + Consistent and elegant handwritten style that feels authentic to a café atmosphere.
  • The 'elegant cursive' requirement for the title is more of a decorative print style than true cursive.

Qwen Image

  • + Good use of layout space with varied text sizes.
  • + Warm, inviting café background with pleasant bokeh.
  • Includes significant typos such as '20026' for the year and 'Risoto'.
  • Text looks like a digital brush stroke rather than genuine chalk on a board.
  • The handwriting style is less consistent across lines.

Verdict: Grok Imagine Image followed the prompt with near-perfect accuracy, correctly rendering all text strings, prices, and the specific date requested. In contrast, Qwen introduced several spelling errors and a significant date error (20026), while its text appearance lacked the realistic chalk texture present in Grok's output.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 'horse on top' spatial instruction
  • + Beautiful cinematic lighting and vibrant nebula colors
  • + Dynamic, surreal composition that captures the prompt's intent
  • The horse's front hoof/leg merges oddly with the astronaut's hand

Qwen Image

  • + Clean execution of a traditional 'astronaut on a horse' concept
  • + Good lighting on the astronaut's suit
  • Failed the core spatial instruction of 'horse on top'
  • Anatomical issues with the horse's legs, particularly the extra jointed or floating hooves
  • More generic and less surreal than requested

Verdict: Grok Imagine successfully followed the difficult spatial constraint of placing the horse on top of the astronaut, resulting in a truly surreal and cinematic image. Qwen Image defaulted to a standard interpretation and failed the primary instruction, while also exhibiting several anatomical glitches in the horse's legs.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent photorealism in the capybara's fur and the taxi interior
  • + Accurately places the passenger in the back seat as requested
  • + Includes specific details like taxi fare stickers on the dashboard
  • The passenger is sitting in the passenger side front seat position instead of truly in the back
  • The capybara's claws are somewhat unnaturally elongated and sharp

Qwen Image

  • + Successfully positions the human passenger clearly in the back seat
  • + Strong color contrast and great lighting on the capybara's profile
  • + The driver's cap has a more formal and professional 'driver' look
  • The hands of the capybara appear more like monkey or human hands with fur rather than capybara paws
  • The steering wheel placement and car interior layout are slightly distorted

Verdict: Both models followed the prompt well, but Grok Imagine produced a significantly more realistic and immersive image with better textures. Qwen Image better followed the spatial instruction for the passenger being in the back seat, but the odd anatomy of the capybara's hands makes it less effective than Grok Imagine's high-quality render.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent text rendering with perfect spelling and high-quality gothic typography.
  • + Rich cinematic lighting on the jack-o-lantern and centered composition.
  • + Superior integration of the thorn and web border within the parchment texture.
  • The parchment edges are slightly less distressed than Image B.

Qwen Image

  • + Strong atmospheric depth with the positioning of the twisted trees.
  • + Good inclusion of the requested scroll banner element.
  • + Accurate adherence to the parchment and dark sky prompt.
  • Significant spelling errors in the main title ('Halle Party').
  • Unnecessary gibberish text at the very top of the hierarchy.
  • The jack-o-lantern light has a slightly blockier, less realistic glow compared to Image A.

Verdict: Grok Imagine is the clear winner as it successfully rendered all text elements with 100% accuracy and professional typography. While Qwen captured the atmospheric gothic vibe well, it failed significantly on the text rendering, misspelling the primary title as 'Halle Party Invitation'.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent 3D material rendering with realistic lighting and subsurface scattering on the fish.
  • + Superior text rendering with clean, centered typography.
  • + High visual clarity and adherence to the isometric perspective.
  • The 'miniature diorama' feel is slightly less pronounced compared to Model B.

Qwen Image

  • + Captures the 'miniature diorama' and 'cartoon scene' style perfectly with its playful proportions.
  • + Good use of the diorama base and inclusion of chopsticks to enhance the scene.
  • + Effective soft, clay-like textures.
  • Text rendering is slightly messy with a distorted secondary flag icon.
  • The flag placement is a bit cluttered compared to the clean layout of Model A.
  • Lighting is a bit flatter than the requested PBR style.

Verdict: Grok Imagine Image produced a much cleaner and more professional-looking image with superior typography and lighting that accurately reflects the requested PBR materials. While Qwen Image created a charming miniature scene, it struggled with the text quality and overall clarity that Grok Imagine Image achieved.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Dynamic sense of motion and 'tumbling' as requested in the prompt.
  • + Vibrant color palette with a strong 'joyful' vibe.
  • + Good integration of the sunrise and dramatic backlighting.
  • Has a very stylized, 'AI-art' look rather than the requested hyper-photorealism.
  • The fur texture looks more like digital brushstrokes than real hair.
  • Missing the butterflies explicitly mentioned in the prompt.

Qwen Image

  • + Successfully achieves a much more photorealistic look with natural fur textures.
  • + Includes all requested elements including the butterflies.
  • + Superior lighting effects with soft god rays and realistic dew sparkles.
  • The composition is a bit more static compared to the 'tumbling' requested.
  • The puppy's expression is somewhat vacant compared to the more expressive small animals.

Verdict: While Grok Imagine Image captures the playful energy of the prompt well, it fails on the 'hyper-photorealistic' requirement, appearing more like a 3D digital illustration. Qwen Image provides a much more realistic interpretation with better detail in the fur and dew, while also including all prompt requirements like the butterflies.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography and spelling of the brand name.
  • + High-quality vector aesthetic with professional-grade shading.
  • + Incorporates all requested text elements clearly.
  • Includes some strange additional shapes behind the cloche that look like a handle or spoon.
  • Repeats 'Est. 1720' twice, which deviates slightly from a standard logo layout.

Qwen Image

  • + Successfully captures the requested vintage banner style.
  • + Accurate interpretation of the 'minimalist' instruction.
  • Severely garbled and overlapping text in the main brand name.
  • The vector lines are somewhat shaky and less refined than a professional logo.
  • Typographical errors in the 'Caffè' accent mark and primary name.

Verdict: Grok Imagine Image is the clear winner because it produces legible, well-composed typography and a professional vector finish. Qwen Image fails significantly on the core requirement of text rendering, resulting in an unreadable brand name with overlapping characters.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Grok Imagine Image
Qwen Image

AI Judge Analysis

Grok Imagine Image

  • + Successfully included all six requested steps in a logical order.
  • + Followed the color palette and flat-vector style very accurately.
  • + Text rendering is surprisingly clean for most headings.
  • Step 3 text contains several spelling errors and overcrowding.
  • Includes 'NASA inspired' literally within the image body.

Qwen Image

  • + Clean vector aesthetic with good use of the requested color palette.
  • + Includes the prompt instruction '(Stop at landing)' as part of the header text.
  • Failed to include 6 distinct steps, skipping from 4 to Tranquility.
  • Significant text errors and labeling confusion, such as calling the rocket 'Sarth Orbit Wicon'.
  • Composition is more cluttered and less informative than the competitor.

Verdict: Grok Imagine is the clear winner as it successfully visualized all six requested chronological steps of the Apollo mission with appropriate iconography for each. Qwen failed the core task of sequential storytelling, skipping steps and producing nonsensical labels. While both models struggled slightly with fine text details, Grok Imagine's layout is far more professional and aligned with the infographic prompt.

Next steps

Explore each model