OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
Grok Imagine Image Pro
#17 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Grok Imagine Image Pro
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent photographic texture on the book and table surface
- + Superior glass refraction and edge rendering
- + Higher image resolution and clarity
- − The glass cube is missing its front face, appearing more like a frame or open display case
Grok Imagine Image Pro
- + Successfully renders all faces of the glass cube
- + Clever creative touch with the book title 'Reflections in Glass'
- + Includes realistic reflections of the sphere on the inner glass surfaces
- − Noticeable artifact where a duplicate plant pot appears reflected or glitched on the right
- − Lower resolution with softer details on the book and plant leaves
Verdict: GPT Image 2 provides a much cleaner, more high-end photographic result but fails the basic geometry of a 'cube' by omitting the front glass panel. Grok Imagine Image Pro follows the prompt accurately by creating a fully enclosed glass box and adds a thematic easter egg on the book, though it suffers from lower visual fidelity and some spatial artifacts.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent depiction of raining texture on the jacket and pavement.
- + Highly realistic skin textures and natural interaction with tools.
- + Includes specific details like Japanese text on the sign and a toolbox that add to the narrative.
- − The composition is a bit cluttered with the post and sign obstructing the subject.
- − The motion blur on the background car is slightly subtle.
Grok Imagine Image Pro
- + Better execution of motion blur from passing cars.
- + Clean composition that focuses clearly on the subject.
- + Effective usage of reflections on the wet pavement.
- − The man is using a wrench to tighten a tire, but the tool is floating and disconnected from his hand.
- − The bicycle's kickstand and rear wheel physics appear slightly warped.
- − Skin looks slightly more smoothed/AI-generated compared to Model A.
Verdict: Model A (GPT) creates a much more convincing and realistic scene with incredible attention to natural textures and logical hand-tool interactions. While Model B (Grok) captures the motion blur and lighting reflections slightly better, it suffers from significant anatomical and physical errors in how the man holds the wrench.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Exceptional photographic realism in the facial skin texture and hair strands.
- + Naturalistic lighting integration with warm highlights and subtle bokeh.
- + High-quality engraving detail on the armor with aged, battle-worn patina.
- − The beads in the hair are very subtle and blend into the texture of the braids.
- − The overall mood is slightly soft for a 'battle-worn' paladin.
Grok Imagine Image Pro
- + Strong adherence to the 'beads' and 'scars' requirements with clear visual representation.
- + Excellent text rendering of Latin on the gorget, enhancing the paladin theme.
- + Dramatic bokeh sparks and high contrast lighting effectively set the scene.
- − The skin texture appears slightly more 'CG' or airbrushed compared to Model A.
- − The scar over the eye is stylized and looks slightly painted on rather than integrated into the skin.
Verdict: GPT Image 2 produces a more lifelike and photographically convincing portrait with superior lighting and texture. Grok Imagine Image Pro follows the specific prompt details like beads, scars, and specific armor engravings more literally, but has a more digital, less naturalistic finish. GPT Image 2 is preferred for its higher artistic and technical visual quality.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with perfect spelling and legibility.
- + Well-organized professional layout with consistent iconography and sections.
- + Accurate adherence to 'bold sans-serif fonts' and 'vibrant accents'.
- − The logo 'NOVA' has a slightly distressed texture that clashes slightly with the clean minimalist prompt.
Grok Imagine Image Pro
- + Strong image quality and lighting within the food photography.
- + Captures the basic grid layout and minimalist white background.
- − Failed to render legible item descriptions, containing many spelling errors and nonsensical text.
- − Layout feels repetitive as food descriptions for different pizzas are identical copies of each other.
- − Typography is thin and stylized rather than the 'bold sans-serif' requested.
Verdict: GPT Image 2 (Model A) is vastly superior as it produces a functional, professional-grade menu with perfect text and thoughtful graphic design elements. In contrast, Grok Imagine Image Pro (Model B) suffers from significant legibility issues and repetitive placeholder-style text that renders the design unusable for its intended purpose.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with a consistent fiery neon effect as requested.
- + Highly detailed textures on the bun, seeds, and grilled patty.
- + Dynamic composition with realistic sauce splashes and well-organized ingredients.
- − The 'LIMITED TIME ONLY' text is slightly cramped inside its border.
Grok Imagine Image Pro
- + Good sense of depth with smaller embers scattered throughout the frame.
- + Clear, legible price starburst that stands out well.
- + Creative use of melting cheese connecting the bun to the burger.
- − The text doesn't truly follow the 'fiery, glowing effect' prompt, appearing more like a 2D graphic overlay.
- − The tomato slices and lettuce are less physically convincing in their arrangement compared to the other model.
Verdict: GPT Image 2 is the clear winner as it perfectly followed the complex text rendering instructions, creating a cohesive fiery aesthetic across all titles and the starburst. Grok Imagine Image Pro produced a high-quality food image, but the text elements felt like disjointed stickers rather than being integrated into the fiery atmosphere of the scene.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'elegant cursive' title request.
- + High-quality chalk texture with realistic grain and smudging.
- + Perfect spelling and formatting for all menu items and prices.
- − The slant in the handwriting is very subtle, bordering on perfectly upright.
Grok Imagine Image Pro
- + Strong chalk texture and realistic varied line weights.
- + Good layout that fills the board space effectively.
- + Followed all text instructions with accurate spelling and pricing.
- − Failed to provide 'elegant cursive' for the title, using a block print style instead.
- − Handwriting style fluctuates significantly between items, losing the 'same handwriting' requirement.
Verdict: GPT Image 2 is the superior overall image because it successfully followed the stylistic requirement for an elegant cursive title while maintaining a consistent handwriting style throughout the board. Grok Imagine Image Pro failed to use cursive for the title and the handwriting styles between the listed items felt disjointed rather than originating from the same hand.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 2
- + Successfully replicates the specific character from Image 2 including face, hair, and accessories.
- + Matches the complex pose and background from Image 1 nearly perfectly.
- + Integrates clothing elements from both images effectively into the final composition.
- − The fingers on the upper right hand are distorted and unnatural.
- − There is some slight blurring where the hair meets the neck.
Grok Imagine Image Pro
- + Maintains high visual clarity and a clean aesthetic.
- + Preserves the original environment and lighting of Image 1.
- − Completely failed the character reference instruction, ignoring the person in Image 2.
- − The face generated does not match either the source images or the requested character.
- − Failed to incorporate any clothing details (scarf, sunglasses, black sweater) from Image 2.
Verdict: GPT Image 2 followed the complex instructions excellently, successfully transplanting the specific character from Image 2 into the difficult pose from Image 1 while maintaining his likeness and accessories. Grok Imagine Image Pro completely ignored the character reference, simply generating a generic woman in the original pose. GPT Image 2 is the clear winner for its high level of adherence to the multi-step editing prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'horse riding astronaut' concept with functional-looking saddlery.
- + Highly detailed textures on the spacesuit and horse fur.
- + Strong cinematic lighting and realistic lunar surface.
- − The astronaut's glove anatomy on the left hand is slightly distorted.
- − The composition is a bit tight, cropping the horse's head.
Grok Imagine Image Pro
- + Vibrant, colorful nebula background and planet for a more 'space' feel.
- + Clean, artistic composition with a sense of zero-gravity weightlessness.
- − Failed the specific prompt instruction of 'riding'; the horse is floating above the astronaut rather than being on top in a riding posture.
- − Anatomy of the horse's legs is awkward, and the astronaut's fingers are poorly defined.
Verdict: GPT Image 2 followed the complex spatial instructions perfectly, depicting a horse literally riding a saddled astronaut on the moon. Grok Imagine Image Pro produced a more vibrantly colored image, but failed the unique 'horse on top' riding instruction, instead showing two separate entities floating near each other.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 2
- + Successfully transferred the exact clothing items from Image 2.
- + Preserved the person's face, hair, and background with high fidelity.
- − The clothing texture is slightly smoother than the original pea coat.
Grok Imagine Image Pro
- + Preserved the face and background elements correctly.
- − Completely failed to use the clothing from Image 2.
- − Visual artifacts present on the hands and rings.
Verdict: GPT Image 2 followed the edit instructions perfectly, accurately transferring the pea coat, scarf, and jeans from the reference image onto the target person. Grok Imagine Image Pro failed the core task by generating a generic royal outfit instead of the specific clothing requested in Image 2.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture on the capybara fur.
- + Effective use of depth of field and bokeh lighting.
- + Natural interaction with the steering wheel.
- − The passenger is heavily blurred and less distinct.
- − The composition is a bit tight, losing some city context.
Grok Imagine Image Pro
- + Strong composition showing both subjects and the city clearly.
- + Accurate rendering of taxi meters and interior details.
- + Captures the bored expression of the businesswoman perfectly.
- − The capybara's hands/paws have slightly unnatural long claws.
- − The lighting is somewhat flat and less cinematic than the competitor.
Verdict: Both models followed the complex prompt well, with GPT Image 2 offering superior fur textures and cinematic atmosphere. However, Grok Imagine Image Pro is preferred because it provides a better side-by-side view of the driver and passenger, clearer city background, and includes more specific taxi-related details.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with a professional vintage gothic feel.
- + Highly detailed and cohesive background featuring a moon, castle, and city skyline.
- + Perfect integration of the border, webs, and thorns within the composition.
- − The jack-o-lantern is quite large relative to the other elements, slightly overpowering the text.
Grok Imagine Image Pro
- + Good clarity on the jack-o-lantern's glow and smoke effect.
- + Accurate inclusion of all requested text elements, including the banner and details.
- − The 'thorns' in the border look more like repetitive wire than organic vines.
- − The trees are less detailed and feel a bit more caricatured than cinematic.
- − The layout feels a bit cramped at the bottom with the text overlapping the border.
Verdict: GPT Image 2 provides a much more sophisticated and atmospheric design that truly feels like a 'cinematic' vintage poster, with superior textures and better-integrated background elements. While Grok Imagine Image Pro follows the text instructions accurately, its illustration style and border design feel less polished and more like a standard digital graphic rather than a gothic invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent PBR material rendering with high-quality textures on the fish and wood.
- + Strong adherence to the 'diorama base' request with miniature details like the lantern.
- + Perfect execution of the text layout and flag icon.
- − The composition is slightly busier than 'minimal' due to the additional miniature elements.
Grok Imagine Image Pro
- + Successfully captures a clean 'cartoon' miniature style.
- + Accurate text rendering and centered composition.
- + Follows the request for 'minimal garnish' better than the other model.
- − Materials look more like plastic or play-dough than realistic PBR.
- − The diorama base is very simple, lacking the requested details and texture depth.
Verdict: GPT Image 2 is the superior output because it successfully balances the 'cartoon' aesthetic with 'realistic PBR materials,' resulting in much higher visual fidelity. While Grok Imagine Image Pro is clean and minimal, its materials appear flat and toy-like compared to the refined textures and lighting found in GPT Image 2.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent anatomical accuracy for all animals
- + Naturalistic lighting with soft god rays and realistic depth of field
- + Superior fur texture and expressive, believable eyes
- − The kitten's raised paw has slightly blurry toe definition
Grok Imagine Image Pro
- + Dynamic 'tumbling' pose for the fox kit
- + Good interpretation of the sunrise and meadow atmosphere
- + Clearer background landscape
- − Included two kittens instead of the requested one
- − The puppy's paws have anatomical issues and strange digit counts
- − The fur texture looks more like digital brushwork than photorealistic fur
Verdict: GPT Image 2 is the superior image due to its impressive anatomical accuracy and realistic lighting, which captures the 'hyper-photorealistic' part of the prompt effectively. Grok Imagine Image Pro failed the specific count requirements by including two kittens and suffered from significant anatomical defects in the puppy's paws and the fox's legs.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent classic typography with ornate serif detailing.
- + Intricate hand-drawn texture that perfectly fits the vintage aesthetic.
- + High-quality vector emblem composition with sophisticated framing.
- − The 'FLORIAN' text is slightly less minimalist than requested.
Grok Imagine Image Pro
- + Successfully follows the minimalist requirement.
- + Clean, clear layout that is very legible.
- + Accurate text rendering for both the name and the date.
- − The 'Est. 1720' banner is very basic and lacks the requested vintage flair.
- − The steam effect is a bit too simple, feeling more like a modern clip-art.
Verdict: GPT Image 2 provides a much more convincing 'vintage' interpretation with rich texture and superior typography that feels authentic to the 1720 era requested. While Grok Imagine Image Pro is more 'minimalist', it lacks the artistic depth and high-end branding feel found in GPT's output, appearing more like a modern digital reconstruction.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography and nearly perfect text rendering throughout the poster.
- + Dynamic and engaging layout with clear information hierarchy.
- + High level of detail in the spacecraft illustrations while maintaining a cohesive style.
- − Includes more illustrative shading than the requested 'flat-vector' style.
- − Some technical inaccuracies in the icons, such as the command module facing away from the moon in the trajectory step.
Grok Imagine Image Pro
- + Strictly adheres to the 'flat-vector' style and iconographic requirement.
- + Follows the NASA-inspired color palette very accurately.
- + Clean, minimalist aesthetic that feels like a modern infographic.
- − Text is slightly blurry or poorly rendered in the lower sections.
- − Composition has large amounts of empty gray space, making it feel less impactful.
- − Icons for descent and landing are a bit abstract and less recognizable than the ones in the other model.
Verdict: GPT Image 2 is much more visually impressive and polished, featuring professional-grade typography and a rich layout that fills the space effectively. While Grok Imagine Image Pro adhered more strictly to the 'flat-vector' prompt instruction, its layout feels sparse and the text rendering is significantly lower quality compared to the crisp execution of GPT Image 2.
Explore each model
xAI's premium image generation model offering higher fidelity output and stronger performance on single-image editing benchmarks compared to the standard Grok Imagine model