OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
Grok Imagine Image
#26 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Grok Imagine Image
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent photographic texture on the book and table.
- + Perfectly follows all spatial instructions, including position of the plant and light source.
- + Realistic reflections of the sphere on the bottom of the cube.
- − The glass cube looks slightly like a frame rather than a solid volume in some areas.
Grok Imagine Image
- + Beautiful soft lighting and bokeh effect.
- + Good material contrast between the glossy sphere and matte book.
- − The sphere appears to be levitating in the air rather than resting on the bottom of the cube.
- − The perspective of the cube is slightly distorted at the edges.
- − The plant pot overlaps with the sphere's visual space in a slightly confusing way.
Verdict: GPT Image 2 is the superior image as it adheres perfectly to the spatial requirements and maintains high physical realism, particularly with the sphere resting naturally on the surface. Grok Imagine Image creates a more dreamlike aesthetic, but the levitating sphere and slightly warped glass geometry make it less accurate to the prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent detail in the man's face and hands with realistic skin textures
- + Effective use of shallow depth of field and natural-looking wet pavement reflections
- + Includes a relevant toolbox and Japanese signage that adds to the narrative
- − The bike's frame geometry is slightly warped near the rear wheel
- − The 'motion blur' on the background car looks more like static blur than movement
Grok Imagine Image
- + Stronger sense of cinematic atmosphere and 'candid' feel
- + The motion blur on the passing vehicle is very realistic and matches the prompt well
- + Composition feels more like a spontaneous street photograph
- − The subject's face is obscured and poorly defined
- − Anatomical issues with the man's hands which appear mangled and poorly formed
- − Lower level of fine detail compared to the other model
Verdict: GPT Image 2 provides a much higher quality render with impressive detail in the subject's face and skin texture, whereas Grok Imagine Image excels at capturing the requested 'motion blur' and cinematic street photography feel but fails significantly on anatomical correctness. GPT Image 2 is the better overall image because it maintains clarity and realism in the main subject while still adhering to the majority of the environmental cues.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Exceptional photorealistic skin texture and fine facial details
- + The armor engraving and wear look authentic and heavy
- + Subtle and natural integration of the braided hair and beads
- − The warm torchlight is mostly behind her, missing the strong reflective quality on the metal face-forward
- − Slightly less 'battle-worn' in terms of visible injuries compared to Model B
Grok Imagine Image
- + Excellent use of warm torchlight and glowing bokeh sparks
- + Strong adherence to the 'faint scars' prompt with clear marks on the face
- + Highly intricate and clean scrollwork on the plate armor
- − Face looks a bit too smooth/airbrushed for the requested gritty theme
- − The torches in the background are somewhat distracting and lack depth
- − The leather strap buckle is slightly distorted
Verdict: GPT Image 2 provides a more grounded and realistic interpretation with superior texture work on the skin and worn armor, making it feel more like a 'battle-worn' character. Grok Imagine Image excels in dynamic lighting and specifically capturing the scar details, but suffers from a more artificial, digital look in the facial rendering. GPT Image 2's overall composition and cinematic quality make it the more successful image.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with coherent descriptions and accurate spelling
- + Highly professional grid layout that mimics a real-world graphic design product
- + Clear categorization with functional pricing and contact information
- − None significant, though it adheres strictly to a standard corporate style
Grok Imagine Image
- + Good use of vibrant food photography and negative space
- + Captures the vibrant and casual aesthetic requested in the prompt
- − Garbled, unreadable placeholder text for descriptions
- − Repetitive menu items like multiple 'Grilled Salmon' and 'Steak Frites' entries
- − Poor image alignment and floating food elements
Verdict: GPT Image 2 (Model A) is vastly superior as it produces a fully functional, professional-grade menu with legible text and a logical layout. Grok Imagine (Model B) suffers from significant AI artifacts, including illegible text and repetitive content that makes the design unusable for its intended purpose.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture on the burger patty and lettuce.
- + Perfect adherence to the fiery, glowing text effect requested.
- + Superior composition that creates a professional advertising layout.
- − The 'LIMITED TIME ONLY' box is slightly cluttered with sparks, making it a bit busy.
Grok Imagine Image
- + Strong sense of movement with the splash effect of the sauces.
- + Clean, legible text for the secondary message.
- − Lighting on the burger is inconsistent with the fiery background.
- − The burger components look more like 3D assets than photorealistic food.
- − The '€6.99' starburst lacks the fiery, glowing effect requested in the prompt.
Verdict: GPT Image 2 is much more successful, delivering a high-end commercial aesthetic with incredible textures and perfect adherence to the complex lighting and text requirements. In contrast, Grok Imagine Image feels more like clip-art with flatter lighting and fails to apply the requested fiery effect to all text elements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the elegant cursive requirement for the title.
- + Superior chalk texture with realistic Pressure and tapering in the strokes.
- + Stronger composition with better spacing and a more authentic wooden frame.
- − Small text at the bottom is slightly less legible than the main body.
Grok Imagine Image
- + Perfect spelling of all menu items and additional text.
- + Very clear and legible handwriting style.
- + Good use of chalk smudges on the board for added realism.
- − Failed the requirement for an 'elegant cursive' title, using print instead.
- − The text layout feels a bit cramped toward the bottom of the board.
- − The chalk texture looks slightly more digital and uniform compared to Model A.
Verdict: GPT Image 2 is the clear winner as it followed the specific stylistic instruction for an 'elegant cursive' title, whereas Grok Imagine used a standard print style. Additionally, GPT Image 2 captured a much more authentic chalk texture on the 'Today's Specials' heading, making it look genuinely hand-drawn.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the character reference, including face, sunglasses, scarf, and black clothing.
- + Perfectly recreates the complex dynamic pose from the reference image.
- + Matches the lighting and yellow studio background of the source image correctly.
- − The fingers on the raised hand are distorted and lack detail.
- − Some artifacts are visible where the scarf meets the clothing.
Grok Imagine Image
- + Perfectly preserves the original Image 1 without any changes.
- − Completely failed to perform the edit requested.
- − Did not incorporate any elements from the character reference (Image 2).
Verdict: GPT Image 2 (Model A) successfully followed the complex instruction to merge the character from Image 2 with the pose and environment of Image 1. Grok Imagine Image (Model B) failed the task entirely, simply returning the first source image without any modifications.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'horse on top' role reversal prompt.
- + Highly detailed textures on the spacesuit and lunar surface.
- + Creative use of a saddle and reigns to reinforce the riding concept.
- − The horse's front legs are anatomically awkward as they grip the reigns.
- − The composition is a bit tight with the top of the horse's head cut off.
Grok Imagine Image
- + Beautiful, cinematic color palette and lighting within the nebula.
- + Great sense of movement and dynamic posing.
- − Failed the specific prompt instruction of the horse riding the astronaut.
- − The horse and astronaut are floating near each other rather than in a riding interaction.
- − Anatomical errors in the astronaut's hands.
Verdict: GPT Image 2 followed the complex instructional prompt perfectly, depicting a horse literally riding an astronaut with appropriate gear like a saddle. Grok Imagine Image produced a visually stunning and artistic space scene, but it ignored the key 'horse on top' role-reversal requirement, showing them as separate entities.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 2
- + Perfectly replicates the outfit from Image 2, including the specific scarf pattern and coat texture.
- + Retains the person's exact face and hair from Image 1 with high fidelity.
- + Seamlessly blends the new clothing into the lighting and pose of the original beach scene.
- − None notable; it successfully completed all parts of the multi-image instruction.
Grok Imagine Image
- + Matches the person's face and hair from Image 1 well.
- + Maintains the background environment accurately.
- − Completely ignored the clothing in Image 2, generating a generic 'elaborate' royal outfit instead.
- − The added right hand contains anatomical glitches and does not match the person's skin patterns correctly.
Verdict: GPT Image 2 followed the complex multi-step instructions perfectly, accurately transferring the specific clothing from the reference image while keeping the subject's identity and background intact. Grok Imagine Image failed the task by ignoring the visual reference for the outfit and inventing a different style of clothing, while also introducing anatomical errors in the added hand.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photographic texture and lighting within the car
- + Highly accurate capybara anatomy and 'professional' expression
- + Realistic use of depth of field focusing on the protagonist
- − The passenger is slightly out of focus compared to the requested importance
- − Only one paw is clearly visible on the steering wheel
Grok Imagine Image
- + Perfect adherence to showing both paws on the steering wheel
- + Very clear depiction of the bored human businesswoman in the same focal plane
- + Strong background composition that effectively screams New York City
- − Physical layout error where the passenger appears to be in the front seat instead of the back seat
- − The taxi light is strangely placed on the interior ceiling or low-profile roof
Verdict: GPT Image 2 (Model A) provides superior photorealism and character detail, capturing the 'professional' vibe of the capybara perfectly. However, Grok Imagine Image (Model B) followed specific instructions better regarding the paws and the passenger's expression, though it failed the spatial logic by placing the passenger in the front passenger seat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with a vintage, high-quality aesthetic
- + Superior textural detail on the parchment, border, and background
- + Cinematic composition with integrated background elements like the Gothic castle and bridge
- − The text at the bottom is slightly small compared to the decorative elements
Grok Imagine Image
- + Clear and legible central text rendering
- + Clean, graphic-style elements that make for a distinct poster layer
- + Solid adherence to the specific banner and border requests
- − The background is very sparse and lacks the 'moody night sky' depth requested
- − Visual elements look more like digital clip-art than a cohesive vintage poster
Verdict: GPT Image 2 (Model A) significantly outperforms Grok Imagine in terms of artistic quality, atmosphere, and 'vintage' feel. While Grok Imagine produced a clean and readable layout, GPT Image 2 created a rich, cinematic world with intricate details that perfectly match the gothic aesthetic requested.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to text hierarchy and requested font styles.
- + Very high detail in textures, particularly the grain of the fish and the wood grain of the board.
- + Superior lighting and shadows that create a convincing 3D miniature effect.
- − Slightly more than 'minimal' garnish with the inclusion of the stone lantern and foliage.
- − The flag icon is styled as an emoji rather than a standard flag graphic.
Grok Imagine Image
- + Clean, minimalist aesthetic that very closely follows the 'soft refined textures' prompt.
- + Perfectly centered composition with a very clear isometric perspective.
- + Accurate text rendering and simple, clean flag icon.
- − The sushi rice appears as uniform spheres rather than realistic grains.
- − Less detail in the 'PBR materials' compared to the other model.
- − Text is flat 2D rather than the 'bold' 3D-adjacent look of the first image.
Verdict: GPT Image 2 is the superior image due to its impressive attention to material detail and sophisticated lighting, which gives it a premium 3D render feel. While Grok Imagine Image correctly interpreted the minimalism, its textures (especially the rice) look overly simplified and less realistic. GPT Image 2 also followed the text styling prompts more effectively by providing large, bold, bordered lettering.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent anatomical realism for all four animal species.
- + Beautiful interaction with butterflies and realistic golden-hour lighting/god rays.
- + High level of fur detail and clear textures in the wildflowers.
- − The fox kit has slightly unusual dark paws that look a bit stylized compared to the rest of the body.
Grok Imagine Image
- + Capture the 'playfully chasing' and 'tumbling' dynamic very well with energetic poses.
- + Vibrant colors and a very high 'cutesy' factor with large expressive eyes.
- − Has a very stylized, 'AI-typical' digital look rather than the requested hyper-photorealism.
- − The fur texture appears overly smoothed and lacks realistic individual strands.
- − Missed the butterfly requirement entirely.
Verdict: GPT Image 2 is the clear winner as it successfully delivered a hyper-photorealistic scene that matches every detail of the prompt, including the specific animals and butterflies. Grok Imagine Image produced a much more stylized, cartoonish output and failed to include the requested butterflies.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling and accent marks.
- + Highly detailed engraving style with excellent use of texture and shading.
- + Sophisticated composition that perfectly captures the vintage, high-end cafe aesthetic.
- − The 'minimalist' instruction was interpreted more as 'vintage ornate' rather than simple vector.
Grok Imagine Image
- + Clean vector style lines which lean closer to a modern minimalist logo.
- + Accurate colors and text rendering.
- − Includes redundant 'Est. 1720' text appearing twice in the logo.
- − The cloche clutters the design by trying to integrate a coffee cup and spoon into its shape, which looks awkward.
- − The steam is overly thick compared to the elegance of the typeface.
Verdict: GPT Image 2 (Model A) is the clear winner as it produces a professional, cohesive, and historically appropriate logo that perfectly matches the 'Caffè Florian' brand identity. While Grok Imagine (Model B) attempts a more minimalist vector approach, it suffers from redundant text and a cluttered, poorly integrated central icon.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with no spelling errors.
- + Sophisticated grid-based layout that reads like a professional poster.
- + Highly accurate illustrations of the Saturn V and Lunar Module.
- − The style leans slightly more toward 3D rendering than the requested 'flat-vector' style.
- − The astronaut icons are very detailed, making them feel less like consistent icons and more like portraits.
Grok Imagine Image
- + Successfully captured the 'flat-vector' style and icon-centric aesthetic requested.
- + Adheres strictly to the color palette and simplified iconography.
- + Good balance of negative space for a modern clean look.
- − Significant text rendering failures including '3rajcory' and 'Transluiory'.
- − The layout is cluttered and less chronologically intuitive compared to Image A.
- − The 'Saturn V' rocket illustration is overly generic and lacks the iconic look of the real vehicle.
Verdict: GPT Image 2 is the superior choice because it provides a professional, ready-to-use infographic with perfect typography and high-quality illustrations. While Grok Imagine captured the 'flat-vector' style more accurately, its severe spelling errors and disjointed layout make it unusable as an informative poster.
Explore each model
An image generation model by xAI designed to generate highly aesthetic images from text descriptions.