xAI's premium image generation model offering higher fidelity output and stronger performance on single-image editing benchmarks compared to the standard Grok Imagine model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Grok Imagine Image Pro
#17 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
Grok Imagine Image Pro
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent photorealism with a highly believable wood texture and soft lighting.
- + Perfectly rendered glass physics, including accurate refraction of the plant pot behind it.
- + Includes legible and clever text on the book spine that relates to the prompt.
- − The plant pot is visible through the glass but the main plant body extends far above the cube.
Qwen Image 2.0
- + Follows all prompt elements including the glass cube, blue sphere, and red book.
- + Good use of soft window light from the left side.
- − The blue sphere appears to be floating unnaturally inside the cube rather than resting on the surface.
- − The refraction and reflections in the glass are messy, creating duplicate floating spheres that don't align with reality.
- − The composition feels a bit cramped with the cube very close to the window edge.
Verdict: Grok Imagine Image Pro is the clear winner due to its superior photorealistic rendering and physically accurate handling of glass and light. While both models followed the prompt successfully, Qwen Image 2.0 struggled with the internal physics of the scene, resulting in a floating sphere and confusing reflections, whereas Grok produced a professional-quality photograph with charming details like the title on the book.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent full-body composition and spatial grounding
- + Captures the motion blur of passing cars clearly as requested
- + Highly realistic textures on the road and pavement environment
- − The hands and wrench interaction is slightly anatomically awkward
- − The lighting on the man's face is a bit flat compared to the backgrounds
Qwen Image 2.0
- + Superb skin texture and facial realism
- + Accurately represents the 'imperfect framing' and '50mm' feel with a tighter crop
- + Stronger puddles and rain-on-pavement details
- − The man's right hand is mangled with extra fingers and fused forms
- − The bicycle geometry is slightly warped near the pedal and chain area
Verdict: Both models followed the prompt well, but Grok Imagine Image Pro provides a more coherent overall scene with better adherence to the 'motion blur' and 'street photo' aesthetic. While Qwen Image 2.0 has superior skin texture and lighting on the subject's face, the significant anatomical errors in the hands make it less effective as a realistic photograph compared to the more structurally sound output from Grok.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent execution of the engraved plate armor and readable Latin text.
- + Superior detail in facial features, skin texture, and lifelike eyes.
- + Masterful use of lighting and depth of field with realistic bokeh sparks.
- − The hair is somewhat styled in a way that looks more like modern dreadlocks than a classic braid.
Qwen Image 2.0
- + Successfully captures a more rugged, older character that feels 'battle-worn'.
- + Effective use of beads in the hair braids and vibrant secondary colors.
- − The hand modeling is poor, with distorted fingers and anatomy.
- − The armor engraving is significantly less detailed and more generic than Model A.
- − The image quality is grainier with more visible artifacts.
Verdict: Grok Imagine Image Pro produced a significantly more refined and technically sound image, with incredible detail in the metalwork and facial textures. Qwen Image 2.0 captured a unique character design but failed on technical execution, particularly regarding the anatomy of the hand and the clarity of the armor engravings.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent adherence to the requested grid and section layout.
- + Features readable and semi-coherent English text for item names and descriptions.
- + High-quality food photography with consistent lighting and vibrant colors.
- − The placeholder text for several pizza items is repeated exactly.
- − Some minor text artifacts in the smaller descriptions.
Qwen Image 2.0
- + Features a trendy, clean aesthetic with rounded corners on images.
- + Includes price indicators which enhance the menu feel.
- − Complete failure to organize by sections (labeled 'Mains' contains a pizza).
- − The text is entirely illegible garble.
- − Poor grid consistency with some images spanning columns and others not.
Verdict: Grok Imagine Image Pro is the clear winner as it successfully organized the menu into the requested categories (Appetizers, Pizza, Mains) with relevant imagery for each. Qwen Image 2.0 struggled with layout logic, placing several pizzas under the 'Mains' header and producing completely unreadable text.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent dynamic motion with high-quality splashes and floating components.
- + Perfect text rendering for all three requested elements.
- + High level of photographic texture on the meat patty and melting cheese.
- − The starburst element looks a bit like a flat clip-art sticker compared to the rest of the image.
- − Price uses a comma instead of a period, though common in some regions.
Qwen Image 2.0
- + Text has a very impressive 'fiery' glow effect that integrates into the scene.
- + Effective use of vertical space and flame effects.
- + Clean and clear price starburst.
- − The burger is not clearly 'exploded' as much as Model A; the bottom half remains mostly stacked.
- − The text 'LIMITED TIME ONLY' is somewhat small and less prominent.
Verdict: Grok Imagine Image Pro followed the prompt more accurately by providing a truly 'exploded' burger with well-separated components and dynamic motion. While Qwen Image 2.0 had superior flaming text effects, it failed to fully separate the burger layers as requested in the prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent text rendering with perfect spelling and placement.
- + Realistic chalk texture and natural variations in handwriting style.
- + High contrast and clear visibility of all requested menu items.
- − The composition is a bit flat and strictly centered, lacking a bit of environmental depth.
Qwen Image 2.0
- + Beautiful environmental lighting and composition that conveys the 'cozy café' atmosphere well.
- + Good use of chalk smudges and texture to enhance the realism of the board.
- − The text layout is slightly messy, particularly around the prices.
- − Multiple instances of small text artifacts and slightly less consistent handwriting compared to Model A.
Verdict: Grok Imagine Image Pro produced a near-perfect rendition of the specific text requested with high clarity and consistent handwriting style. While Qwen Image 2.0 offered a more visually interesting composition and better lighting, it struggled slightly with the organization of the text and price alignment, making Grok the more successful model for this text-heavy prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Grok Imagine Image Pro
- + Perfect adherence to the unusual 'horse on top' request
- + Vibrant, cinematic colors and lighting
- + High detail in the nebula and planetary background
- − The horse appears to be floating just above rather than 'riding' in a traditional physical sense
Qwen Image 2.0
- + High textural detail on the horse and space suit
- + Good composition with Earth in the background
- − Completely failed the negative constraint/specific instruction of 'horse on top'
- − The scales on the horse's neck look a bit muddy and inconsistent
Verdict: The main differentiator was the specific prompt instruction to have the 'horse on top'. Grok Imagine Image Pro followed this surreal request perfectly, creating an interesting and literal interpretation, whereas Qwen Image 2.0 ignored the specific instruction and generated a standard astronaut riding a horse.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent composition from the dashboard perspective that clearly places the human in the back seat.
- + Very high level of photorealism and texture, particularly on the capybara's fur and the woman's clothing.
- + Accurate text rendering on the hat with 'NYC TLC' identifying the taxi context.
- − The capybara's hands look slightly more like paws with human-like finger articulation which is a bit uncanny.
- − The lighting is slightly flat compared to the cinematic night feel of model B.
Qwen Image 2.0
- + Dynamic angle that gives a great sense of being inside the vehicle with the driver.
- + The blurred city lights reflection on the window adds a high degree of realism to the night setting.
- + Great adherence to the 'bored expression' prompt for the passenger.
- − The passenger is physically sitting in the front passenger seat next to the driver, failing the prompt's request for her to be in the back seat.
- − The capybara's hat is generic and lacks the specific 'taxi driver' branding requested.
- − Lower overall resolution and clarity compared to Model A.
Verdict: Grok Imagine Image Pro is the winner because it correctly followed the spatial instruction to place the passenger in the back seat, whereas Qwen Image 2.0 placed her in the front next to the driver. Grok also produced a much sharper image with superior detail in the capybara's fur and legitimate-looking taxi branding.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Grok Imagine Image Pro
- + Perfect text rendering for all requested details including date, time, and location.
- + Dynamic and moody cinematic lighting with smoke effects and varying color tones.
- + Excellent composition with a detailed thorn and web border integrated into the parchment design.
- − The text at the bottom is slightly crowded compared to the top title.
Qwen Image 2.0
- + Clear and legible gothic typography for the header.
- + A clean, symmetrical layout that works well for a square invitation.
- + Good adherence to the request for twisted trees and a thorn border.
- − The jack-o-lantern lighting is a bit flat compared to the other model.
- − The text in the small banner is slightly distorted and contains a comma where none was requested.
- − The overall image lacks the rich, atmospheric depth of the first image.
Verdict: Grok Imagine Image Pro is the winner as it flawlessly rendered every piece of text requested, including the specific date and location. Its lighting and atmospheric effects are significantly more polished and professional than Qwen Image 2.0, which struggled slightly with the banner text and had a flatter visual style.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent 3D miniature 'cartoon' aesthetic as requested.
- + Precise 45-degree isometric perspective.
- + Very clean typography and creative flag icon integration.
- − Texture of the nigiri rice looks slightly uniform and synthetic.
Qwen Image 2.0
- + High realism in food textures, specifically the eel and salmon.
- + Clear, bold text rendering.
- + Good adherence to the square format and diorama base.
- − Missed the 'cartoon' style request, leaning towards a photographic look.
- − The flag icon is a standard rectangle rather than a stylized icon.
- − Perspective is a bit flatter than a true 45-degree isometric view.
Verdict: Grok Imagine Image Pro followed the stylistic cues perfectly, delivering a high-quality isometric 3D cartoon render that matches the prompt's aesthetic. Small differences in text placement and icon style make it feel more design-oriented. Qwen Image 2.0 produced a realistic and appetizing image, but failed to capture the specific 'miniature 3D cartoon' style requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent depiction of god rays and sunrise lighting consistent with the prompt.
- + Sharp focus on all subjects with very clean 'masterpiece' aesthetic.
- + Included all requested animal types clearly.
- − Generated two kittens instead of the requested one tabby kitten.
- − The butterflies look somewhat static and pasted-on compared to the environment.
Qwen Image 2.0
- + Correctly included exactly one of each animal requested.
- + The 'tumbling' interaction is much more natural and dynamic than in Image A.
- + The fur textures appear softer and more integrated with the lighting.
- − The fox's face/eye area shows slight anatomical warping during the tumble.
- − The bunny looks a bit detached from the main action in the center.
Verdict: Grok Imagine Image Pro produces a cleaner, more vibrant image with beautiful lighting, but it fails the count requirement by adding an extra kitten. Qwen Image 2.0 followed the prompt details more accurately and captured the 'tumbling' interaction much better, even if some fine details on the fox are slightly messy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent typography and precise text rendering
- + Clean vector emblem style that looks like a professional logo
- + Well-balanced composition with an appropriate 'Est. 1720' banner
- − Steam effect is a bit simple/swirly rather than natural
Qwen Image 2.0
- + Stronger vector texture and gradients on the cloche dome
- + Creative integration of steam inside the dome's reflection area
- + Maintains warm brown and cream tones as requested
- − The ribbon/banner tails are asymmetrical and messy
- − The text alignment inside the ribbon is slightly off-center
- − Typography is less elegant and 'classic' compared to Model A
Verdict: Grok Imagine Image Pro produced a much more realistic and professional-looking logo with perfect typography and balanced composition. Qwen Image 2.0 followed the prompt well but struggled with the symmetry of the banner and the refinement of the fonts, making it less suitable for an actual brand identity.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Grok Imagine Image Pro
- + Excellent typography with no spelling errors.
- + Consistent and clean flat-vector iconography for every step.
- + Perfect adherence to the requested NASA-inspired color palette and layout.
- − Step 3 (Translunar) icon is a bit abstract compared to the others.
Qwen Image 2.0
- + Strong composition with a sense of scale for the landing phase.
- + Good adherence to the requested steps.
- − Includes a spelling error ('Translunjar').
- − Inconsistent icon styles, mixing silhouette humans with detailed line-art modules.
- − Poor text placement where orbit rings overlap the 'Lunar Orbit' text.
Verdict: Grok Imagine Image Pro produced a professional-grade infographic with perfect spelling, consistent vector styling, and a clean layout that feels like a finished product. Qwen Image 2.0 struggled with text legibility and included a typo, and the icon styles were not as cohesive. Grok Imagine Image Pro is the clear winner for its adherence to the 'clean, modern vector' aesthetic requested.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request