An image generation model by xAI designed to generate highly aesthetic images from text descriptions.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Grok Imagine Image
#25 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
Grok Imagine Image
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Grok Imagine Image
- + Excellent photographic realism with natural lens blur and lighting.
- + Correctly interprets the 'small' scale of the blue sphere within the cube.
- + The reflection/refraction of the plant through the glass is very convincing.
- − The sphere appears to be floating without visible support, which feels slightly unnatural.
Qwen Image 2.0
- + Strong adherence to all spatial requirements of the prompt.
- + The 'soft window light from the left' is clearly represented on the wood and sphere.
- − The 'small blue sphere' is actually quite large relative to the cube.
- − The glass cube has strange internal reflections that look like multiple spheres exist inside.
- − The plant behind the glass does not refract naturally, appearing more like a print on the back glass.
Verdict: Grok Imagine produces a much more aesthetically pleasing and realistic image with superior handling of light and transparency. While both models struggled with the physics of a floating sphere, Grok Imagine's interpretation of the plant through the glass and the scale of the objects is more accurate to the prompt and more visually coherent than Qwen Image 2.0.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Grok Imagine Image
- + Excellent depiction of motion blur with the passing car
- + Strong adherence to 'imperfect framing' for a candid street feel
- + Captures a very realistic urban atmosphere with wet pavement reflections
- − The subject's face is obscured and less detailed
- − The bicycle has some structural inconsistencies
- − Light rain is barely visible
Qwen Image 2.0
- + Exceptional natural skin texture and facial detail
- + Very clear and structurally sound red bicycle
- + Good interaction between the subject and the object being repaired
- − Lacks significant motion blur for the passing car
- − Composition is a bit too clean for 'imperfect framing'
- − The car in the background looks slightly like a cutout
Verdict: Both models followed the prompt well, but Grok Imagine (Image A) achieved a more convincing 'candid' and cinematic film look through its more authentic motion blur and lighting. Qwen Image 2.0 (Image B) produced much higher detail in the subject's face and the bicycle, though it felt slightly more like a staged portrait than a raw street photo.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Grok Imagine Image
- + Exceptional intricate detail on the ornate engraved plate armor
- + Beautiful soft lighting and bokeh sparks that create a cinematic mood
- + High-fidelity texture on the skin and fly-away hair strands
- − The character looks relatively clean and youthful for the 'battle-worn' description
- − Beads in the hair are present but less colorful and distinct than in Model B
Qwen Image 2.0
- + Conveys the 'battle-worn' and 'dirt' aspects of the prompt more effectively through skin texture and scarring
- + Excellent inclusion of the beads and braided hair as requested
- + Stronger narrative composition with the addition of the sword pommel and hand
- − The hand has anatomical issues with several merging and oddly shaped fingers
- − The bokeh sparks appear as solid orange blobs rather than realistic lighting artifacts
- − Armor engraving is less detailed and consistent compared to Model A
Verdict: Grok Imagine produces a much cleaner, more aesthetically pleasing image with superior technical rendering of armor and light. Qwen Image 2.0 captures the gritty 'battle-worn' essence and beads better, but it is let down by significant anatomical errors in the hand and less refined digital effects.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography with clean, legible heading fonts.
- + Highly professional layout that follows real-world menu design principles.
- + Sophisticated use of negative space and scattered food photography.
- − Smaller text is gibberish upon close inspection.
- − The food items are scattered rather than in a strict grid as requested.
Qwen Image 2.0
- + Perfect adherence to the grid-based layout request.
- + Very high-quality photography with vibrant colors.
- + Uniform and coherent sectioning.
- − The text has significant rendering errors and AI artifacts.
- − Composition feels a bit repetitive, with too many pizza shots in non-pizza sections.
Verdict: Grok Imagine produces a more realistic and professional-looking menu design that could almost be used in a real restaurant, despite the nonsense body text. Qwen Image 2.0 followed the specific grid instruction more accurately and provided more appetizing food close-ups, but the text rendering is distorted and the layout is less sophisticated. Grok Imagine is the likely winner for its superior graphic design sense and typeface selection.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography with clean, readable text that perfectly matches the requested fiery effect.
- + Superior 'exploded' effect with individual ingredients clearly separated and dynamic sauce splashes.
- + Vibrant lighting and high-quality rendering of the burger texture.
- − The starburst for the price is a bit clean and graphic compared to the photorealistic burger.
Qwen Image 2.0
- + Strong atmospheric smoke and ember effects around the burger.
- + Good integration of the price starburst into the fiery theme.
- − The burger is mostly stacked rather than 'exploded' with suspended components.
- − Typography for the secondary text is smaller and less impactful than Model A.
- − The '€6.99' text has slight rendering artifacts and lacks the requested fiery glow.
Verdict: Grok Imagine followed the prompt more accurately, specifically capturing the 'exploded' nature of the burger where components are suspended in mid-air. While both models handled the main title well, Grok Imagine provided much cleaner typography and better dynamic motion with the sauce splashes, making for a more professional-looking advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Grok Imagine Image
- + Excellent text rendering with no spelling errors.
- + The chalk texture and smears on the board look highly realistic.
- + Strong adherence to the layout prompt by keeping all text on single lines.
- − The title is not in 'elegant cursive' as requested, but rather a print style similar to the body text.
Qwen Image 2.0
- + Natural lighting and depth of field create a very believable café atmosphere.
- + Excellent chalk texture and realistic handwriting variations.
- + Accurately rendered all requested menu text.
- − Text layout is somewhat cramped, forcing line breaks not suggested in the prompt.
- − The 'cursive' request for the title was mostly ignored in favor of a standard print style.
Verdict: Both models performed exceptionally well on a difficult text-integration task. Grok Imagine Image is slightly superior for its clean composition and Ability to keep long menu items on a single line while maintaining perfect legibility. While Qwen Image 2.0 has perhaps a more realistic photographic quality, the text layout is a bit more cluttered.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Grok Imagine Image
- + Successfully followed the difficult logic of the horse on top of the astronaut
- + Highly cinematic lighting and nebula background
- + Superior anatomical detail on the horse
- − The connection between the horse's hoof and the astronaut is slightly awkward
- − The composition feels a bit like they are floating separately rather than 'riding'
Qwen Image 2.0
- + High visual clarity and sharp rendering
- + Beautiful atmospheric effect with the Earth and floating droplets
- − Failed the primary prompt instruction of having the horse on top
- − Common AI artifacting where the horse's rear leg merges with the tail and body
- − Very generic interpretation of 'astronaut riding a horse' despite the specific constraint
Verdict: Grok Imagine was the only model to successfully interpret the complex logic of the prompt, placing the horse on top of the astronaut. Qwen Image 2.0 defaulted to a standard 'astronaut riding a horse' image, completely ignoring the negative constraint 'horse on top, not vice versa'.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the 'back seat' instruction for the passenger
- + High level of photographic realism and sharp textures on the capybara fur
- + Accurate depiction of a New York night atmosphere through the glass
- − The passenger is seated in the front passenger seat rather than the back seat as requested
- − The taxi light is placed inside the car/on the dashboard inappropriately
Qwen Image 2.0
- + Natural-looking capybara driver's cap
- + Dynamic angle that captures the essence of a New York street
- + Accurate depiction of both paws on the steering wheel
- − The passenger is sitting in the front passenger seat, ignoring the 'back seat' instruction
- − Noticeable claw artifacts and slightly muddy textures on the capybara's hands
- − The capybara's head scale is slightly too large for the car cabin
Verdict: Both models failed the specific instruction to place the passenger in the back seat, placing her in the front instead. Grok Imagine Image is the preferred output due to its superior photographic clarity, better rendering of the woman's face/hands, and a more realistic taxi interior despite the strange placement of the roof light.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Grok Imagine Image
- + Excellent text rendering with no spelling errors.
- + Greater visual depth and more atmospheric lighting.
- + The border with thorns and webs is more detailed and three-dimensional.
- − The parchment edges are partially cut off at the top and bottom.
Qwen Image 2.0
- + Strong adherence to all prompt elements including the scroll banner.
- + Good layout balance for a square format invitation.
- + The twisted trees are well-integrated into the composition.
- − The word 'Invitation' has a slight character overlap or artifact.
- − Lighting is flatter compared to the other model.
Verdict: Both models followed the complex prompt very well, including all requested text and objects. Grok Imagine Image is slightly superior due to its more cinematic lighting and cleaner typography, whereas Qwen Image 2.0 has minor text rendering issues and a flatter visual style.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the '3D cartoon' and 'isometric' style instructions.
- + Clean, bold text rendering with perfect vertical center alignment.
- + High-quality soft lighting and refined, rounded 3D textures.
- − The sushi and soy sauce are floating slightly above the wooden board rather than resting on it.
Qwen Image 2.0
- + More realistic PBR materials for the food items.
- + Captures the requested small raised diorama base and plate correctly.
- + Accurate text rendering and flag inclusion.
- − Missed the 'cartoon' styling request, opting for a photographic look.
- − The text and flag are not at the top-center as requested, but shifted slightly.
Verdict: Grok Imagine Image followed the stylistic prompt much more effectively, delivering a cohesive isometric 3D cartoon aesthetic with excellent typography. Qwen Image 2.0 produced a high-quality photographic image but largely ignored the 'cartoon' requirement, resulting in a less stylized final product.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Grok Imagine Image
- + Dynamic sense of movement and energy
- + Effective use of warm, glowing backlighting
- + Cohesive illustrative style
- − Fails the 'hyper-photorealistic' requirement by appearing highly stylized and cartoonish
- − Anatomy is slightly deformed, especially on the fox and kitten's limbs
- − Missing the butterflies explicitly requested in the prompt
Qwen Image 2.0
- + Successfully achieves a hyper-photorealistic look with realistic fur textures
- + Includes all requested elements including butterflies and god rays
- + Excellent interaction between the animals that actually captures 'tumbling together'
- − The fox kit's face is slightly obscured and looks a bit awkward in its pose
- − The bunny's scale relative to the other animals is a bit large
Verdict: Qwen Image 2.0 is the clear winner as it adhered to the 'hyper-photorealistic' instruction, whereas Grok Imagine produced a stylized 3D animation look. Qwen Image 2.0 also correctly included the butterflies and created a more believable interaction between the four different animals in a natural meadow setting.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography including the correct grave accent on 'Caffè'
- + Superior vector aesthetic that feels like a professional logo
- + Well-composed arrangement of text and icon
- − Redundant text (Est. 1720 appears twice)
- − Anatomically confusing handle shapes at the side of the cloche
Qwen Image 2.0
- + Strong 'vintage banner' execution for the date
- + Cohesive illustrative style with nice metallic gradients
- + Simple and centered composition
- − Incorrect accent on 'Caffè' (uses an acute accent instead of grave)
- − Typography placement inside the cloche feels cramped and less like a logo emblem
- − Steam appears to be behind or inside the cloche rather than coming from it
Verdict: Grok Imagine Image produced a much more professional and believable logo design with superior typography and vector execution, despite the redundant text. Qwen Image 2.0 followed the banner instruction well, but the font choices and incorrect accent mark on the brand name made it feel less authentic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the requested color palette and vector style.
- + Highly legible and mostly correct text for a generated image.
- + Well-organized layout with distinct iconography for each step.
- − Step 3 contains some garbled text ('3rajcoory').
- − The Saturn V icon lacks some vertical scale characteristic of the rocket.
Qwen Image 2.0
- + Features a very clean, vertical flow that feels like a professional poster.
- + Includes a clever lunar module icon for the landing phase.
- + Accurately represents the lunar orbit and translunar arc visually.
- − Contains a spelling error in a primary heading ('Translunjar').
- − The crew silhouettes are less integrated into the main poster design.
- − The scale of icons is somewhat inconsistent compared to the text size.
Verdict: Grok Imagine Image is the winner for its superior layout, color usage, and adherence to the 'infographic' aesthetic. While both models struggled slightly with text, Grok Imagine Image felt more like a complete, balanced poster, whereas Qwen Image 2.0 felt a bit sparser and had a more distracting typo in a large font.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request