Head to head
Esc

Models · slot A

to navigate to pick

Grok Imagine Image xAI Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

Grok Imagine Image

23.4 arena score

#27 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

Grok Imagine Image

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 'small' sphere requirement
  • + Realistic glass refraction and depth of field
  • + Accurate rendering of light direction and shadows
  • The glass object is slightly rectangular rather than a perfect cube
  • The blue sphere appears to float without a support structure

Qwen Image Max

  • + Perfect cube geometry
  • + Clean aesthetic and high resolution
  • + Highly visible plant through the glass
  • The blue sphere is quite large, ignoring the 'small' descriptor
  • The black design on the red book was not requested
  • Lighting is a bit flat compared to Model A

Verdict: Grok Imagine Image followed the descriptive 'small blue sphere' prompt much better and captured a more natural photorealistic lighting. Qwen Image Max produced a cleaner cube and sharper details, but it failed to follow the scale instructions for the sphere and added unnecessary details to the book.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Exquisite floral engraving detail on the armor.
  • + Perfectly captures the 'warm torchlight' and bokeh sparks requested.
  • + Highly detailed facial texture including subtle scars and dirt.
  • The character looks somewhat young and clean for a 'battle-worn' description.

Qwen Image Max

  • + Excellent 'battle-worn' characterization with aged skin and deep scars.
  • + Superior texture on the leather straps and tattered cloth underlayer.
  • + Great representation of hair braided with colorful beads.
  • Lighting is a bit more diffused and less 'warm' than requested.
  • Small anatomical glitch in the right eye (viewer's left).

Verdict: Grok Imagine produces a more polished and aesthetically pleasing image with stunning metal engravings and beautiful lighting. However, Qwen Image Max better captures the 'battle-worn' spirit of the prompt with more convincing aging, grit, and detailed textures on the non-metal materials like the worn leather and cloth.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent 'exploded' layout with clearly suspended components as requested
  • + Clean and highly legible text rendering for all elements
  • + Dynamic splash effects with sauces add to the sense of motion
  • The starburst for the price looks like a flat clip-art element rather than integrated into the 3D scene

Qwen Image Max

  • + Price starburst is beautifully rendered with light rays and glowing effects
  • + High texture detail on the patty and bun
  • + Very atmospheric lighting and ember effects
  • Failed the 'exploded' instruction as the burger is mostly assembled
  • The secondary text is slightly less polished compared to the main title

Verdict: Grok Imagine followed the complex 'exploded burger' layout much better than Qwen Image Max, showing a clear separation of all ingredients. While Qwen Image Max produced a more integrated and visually impressive price starburst, the failure to separate the burger components makes it a less accurate response to the specific prompt instructions.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent photorealism in the textures of the capybara's fur and the woman's clothing.
  • + Captures the 'bored' expression of the passenger perfectly as requested.
  • + Realistic cinematic lighting consistent with a night scene in a taxi.
  • The passenger is sitting in the front passenger seat rather than the back seat.
  • The capybara's paws look slightly sharp and claw-like compared to real capybara anatomy.

Qwen Image Max

  • + Correctly places the human passenger in the back seat as requested.
  • + The profile angle provides a great sense of depth and interior detail.
  • + Realistic rendering of the capybara's fur and the leather taxi seats.
  • The capybara's hands look more like monkey or human hands covered in fur than capybara paws.
  • The taxi ceiling has strange cracked/damaged artifacts that look unintentional.

Verdict: Grok Imagine Image provides a more photorealistic and high-quality image, but it fails a key spatial prompt by placing the passenger in the front seat. Qwen Image Max correctly follows the layout instructions by placing the passenger in the back, though it suffers from anatomical issues with the capybara's hands and minor artifacts on the car ceiling. Overall, Grok is preferred for its superior clarity and lighting, even with the positional error.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent text rendering with perfect spelling and clear gothic fonts
  • + Features a distinct thorn and spiderweb border that blends well with the parchment style
  • + Balanced composition with clear hierarchy of Information
  • The transition between the central illustration and the parchment background is a bit sharp in the corners

Qwen Image Max

  • + Atmospheric integration of the central scene with the background
  • + Accurate text rendering for both the banner and the body text
  • + Highly detailed twisted trees and border designs
  • The 'You' in the banner is slightly warped
  • The composition feels a bit more cramped with the larger border elements

Verdict: Both models performed exceptionally well, following all complex text and stylistic requirements. Grok Imagine Image is selected as the winner for its cleaner text layout and superior font choices, particularly with the elegant gothic title and the legible body text. Qwen Image Max is also of very high quality but has slightly more distortion in the cursive banner text.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography rendering with clean, flat design.
  • + Strong adherence to the 'cartoon' and 'soft texture' style requested.
  • + Precise 45° isometric perspective with a professional diorama feel.
  • The rice grains look more like rounded bumps than realistic grains.
  • The Japanese flag icon is at the very top rather than below the text as is common in layouts.

Qwen Image Max

  • + High material quality with realistic PBR subsurface scattering on the fish.
  • + Creative use of a glass/acrylic diorama base.
  • + Detailed food rendering, especially the eel and egg textures.
  • Typography is slightly cramped and has awkward drop shadows.
  • The sushi pieces are arranged in a less balanced circular fashion.
  • The perspective feels slightly lower than a true 45-degree isometric view.

Verdict: Grok Imagine Image followed the graphic design aspects of the prompt more effectively, delivering clean typography and a cohesive cartoon-isometric aesthetic. Qwen Image Max achieved higher realism in the food textures (PBR materials), but was less successful with the text layout and the specific 'miniature 3D cartoon' style.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Grok Imagine Image
Qwen Image Max

AI judge analysis unavailable for this challenge.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Grok Imagine Image
Qwen Image Max

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography rendering with perfect spelling of 'Caffè Florian'
  • + Clean vector aesthetic that feels modern yet vintage
  • + Bold contrast and good use of the requested color palette
  • Confusing illustrative elements attached to the dome (coffee cup/spoon hybrid)
  • Double-up of the 'Est. 1720' text feels redundant and cluttered

Qwen Image Max

  • + Successfully incorporates the banner element requested in the prompt
  • + Beautifully textured vintage paper background
  • + Consistent wood-cut/engraved illustration style for the cloche and steam
  • Small spelling error in the name (extra E on 'CAFFÉ')
  • Less 'minimalist' than Model A due to heavy shading and background texture

Verdict: Both models followed the prompt well, but they excelled in different areas. Grok Imagine produced a cleaner, more usable vector logo with perfect text, while Qwen Image Max captured the requested 'banner' element and vintage atmosphere more effectively despite a minor spelling error in the accent mark.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.