Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 OpenAI Grok Imagine Image xAI

Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.

GPT Image 1

23.2 arena score

#28 of 62 in Text-to-Image

Skill signature · Text-to-Image

Grok Imagine Image

23.4 arena score

#26 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1

0%

win rate

Ties

0%

Grok Imagine Image

0%

win rate

Shared challenges 20

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to lighting instructions with natural window lighting from the left
  • + Clean and realistic textures on the wood, book, and glass
  • + Accurate spatial arrangement with the sphere sitting on the base inside the cube
  • The glass cube is missing a front pane, appearing more like a frame than a solid cube

Grok Imagine Image

  • + Successfully renders the cube as a solid object with clear refractive qualities
  • + Good depth of field and atmospheric quality in the background
  • The blue sphere is levitating unnaturally in the center of the cube
  • The scale of the sphere is slightly smaller than implied by the prompt

Verdict: GPT Image 1 followed the lighting and placement instructions more faithfully, creating a very grounded and realistic scene, though the cube looks more like an open glass box. Grok Imagine produced a more convincing solid glass material but failed on physics by having the sphere levitate in the center of the cube. GPT Image 1 is the winner for its superior texture quality and natural lighting.

Man and Car in California

Editing
Edit instruction

“Make a photo of the man driving the car down the California coastline”

Source
GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent preservation of the specific man's face and unique hairstyle from the source image
  • + High fidelity to the Rolls-Royce model and its specific features
  • + Effective motion blur on the wheels and road suggests actual driving
  • The driver appears slightly too large for the interior scale of the car

Grok Imagine Image

  • + Beautiful background composition of the California coastline
  • + Dynamic lighting and reflections on the car body
  • + Clean integration of the car onto the coastal highway
  • Failed to use the man from the source image, replacing him with a generic older white male
  • Altered the car's headlight design, deviating from the source image's square housing

Verdict: GPT Image 1 successfully followed the complex instruction of combining both specific source images, accurately placing the correct man behind the wheel of the correct car. Grok Imagine failed the image editing task by ignoring the source photo of the man entirely and changing the vehicle's model details.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent skin texture and hyper-realistic facial details
  • + Strong atmospheric lighting with raindrops visible on the bike
  • + High fidelity and sharpness on the primary subject
  • Lack of motion blur on the background cars despite the prompt and shutter speed context
  • The background cars look static and parked rather than passing by
  • The framing is quite central rather than imperfectly candid

Grok Imagine Image

  • + Successfully captured motion blur from a passing car
  • + Better adherence to the 'imperfect framing' and 'candid' street photography style
  • + Highly realistic bicycle and background details capturing a true city vibe
  • Face is partially hidden and obscured by a mask, reducing the impact of 'natural skin texture'
  • Lacks the fine detail and sharpness found in the first image

Verdict: GPT Image 1 excels in technical portraiture and fine detail, particularly in the face and skin texture, but fails to capture the requested motion blur of passing cars. Grok Imagine creates a more authentic 'candid' street photography feel with genuine motion blur and framing, though it sacrifices some of the fine detail and subject focus found in the other image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Exceptional realism in facial textures, skin imperfections, and lifelike eyes.
  • + Excellent engraving detail on the armor with a realistic metallic finish.
  • + Very effective lighting that feels integrated into the scene.
  • The 'battle-worn' appearance feels slightly more like simple dirt than a paladin post-combat.
  • Composition is very tight, cutting off some of the requested leather and cloth details.

Grok Imagine Image

  • + Successfully includes all requested elements, including leather straps and cloth underlayers.
  • + Strong atmospheric bokeh andชัดเจน sparks that add to the fantasy aesthetic.
  • + Detailed hair braiding with visible wooden beads.
  • The facial features look overly smoothed and 'perfect,' lacking the gritty realism of a battle-worn character.
  • The torchlight is slightly overexposed in the background, drawing focus away from the subject.

Verdict: GPT Image 1 produces a far more convincing and realistic portrait with incredible skin texture and nuanced lighting that fits the 'battle-worn' description perfectly. While Grok Imagine Image does a better job of including every technical prompt detail like the leather straps and specific beads, its overall 'AI-smooth' aesthetic lacks the gravitas and photographic quality found in the GPT output.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent high-resolution food photography that looks very appetizing.
  • + Perfectly legible text with clear sections.
  • + Strong layout with a professional, clean aesthetic.
  • The placeholder text is repetitive and nonsensical.
  • The layout is heavily cropped, showing only a portion of the menu.

Grok Imagine Image

  • + Provides a complete menu layout including a variety of items and actual food names.
  • + Creative use of circular food photos to break up the text blocks.
  • + Good adherence to the 'grid' and 'multi-section' prompt requirements.
  • The text is largely unreadable and full of gibberish characters.
  • Visual quality of the food images is lower than Model A.
  • A bit cluttered, losing some of the 'minimalist' feel.

Verdict: GPT Image 1 produces significantly higher quality food photography and clearer typography, though it only shows a small segment of the page. Grok Imagine creates a more comprehensive menu design with a wider variety of elements, but it suffers from significant AI text artifacts and lower-fidelity imagery. GPT Image 1 is the better choice for a professional aesthetic, despite the repetition in placeholder text.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent typography with a consistent fiery glow effect.
  • + Beautifully rendered food textures, especially the lettuce and the sear on the patty.
  • + Clean, professional composition suitable for a high-end food advertisement.
  • The price in the starburst is missing the '6', reading only '€.99'.
  • Less 'exploded' feel, as the ingredients are mostly stacked vertically rather than scattered.

Grok Imagine Image

  • + Successfully captured a chaotic, dynamic 'exploded' motion with ingredients flying out.
  • + Fully accurate text rendering, including the correct price with no missing digits.
  • + Stronger sense of action and energy with the addition of flames and sauce splashes.
  • The '€6.99' starburst looks like a flat clip-art sticker rather than being integrated into the scene.
  • Slightly less photorealistic lighting on the individual vegetable slices compared to the other model.

Verdict: Model B (Grok Imagine) followed the prompt more accurately by including all the requested text and providing a more dynamic, scattered 'exploded' layout. While Model A (GPT Image 1) achieved a more sophisticated aesthetic and better food textures, the missing digit in the price and the more static composition make it less effective as a direct answer to the prompt.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent text legibility and accuracy
  • + Convincing granular chalk texture
  • The '9' for the cookies is missing the dollar sign
  • The handwriting style is a bit too uniform, bordering on a digital font look
  • The composition is a tight crop that lacks the requested cafe background

Grok Imagine Image

  • + Successfully captured the requested 'cozy cafe' environment
  • + Handwriting has more natural variation and personality
  • + Perfect adherence to all requested text including the dollar signs
  • The title is not in 'elegant cursive' as requested
  • Some letters in the bottom sentences are slightly smudged

Verdict: GPT Image 1 has superior chalk texture and clarity, but it fails to include the cafe background and misses a symbol in the text. Grok Imagine captures the requested environment perfectly and follows the text requirements more accurately with a much more natural, handwritten feel.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Successfully integrated character features from Image 2 including the sunglasses, scarf, and black sweater.
  • + Successfully replicated the complex body pose from Image 1.
  • + Matched the yellow studio lighting and backdrop of Image 1.
  • The anatomical transition between the character's face and original neck area is poorly blended.
  • The right hand is rendered as a distorted nub.
  • Low overall image resolution and some visible artifacting.

Grok Imagine Image

  • + High visual quality and resolution.
  • + Perfectly captures the lighting and environment of Image 1.
  • Completely failed to incorporate the character reference from Image 2.
  • Modified the pose by raising a leg rather than keeping the crossed legs from Image 1.
  • Ignored the request to change the person/clothing to match Image 2.

Verdict: GPT Image 1 followed the complex multi-image instruction by combining the character's identity from Image 2 with the pose from Image 1, despite having lower technical image quality. Grok Imagine Image effectively ignored the character reference entirely, essentially providing a slight variation of the original Image 1 without changing the identity or clothing as requested.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent cinematic lighting and texture on the horse's coat.
  • + High-quality rendering of the space suit and Earth in the background.
  • + Strong overall composition and realistic shadows.
  • Completely failed the semantic instruction of the prompt (horse on top).
  • Generates a standard 'astronaut on horse' cliché.

Grok Imagine Image

  • + Followed the specific instruction of 'horse on top' through a surreal interaction.
  • + Vibrant, high-contrast colors with a deep nebula background.
  • + Dynamic pose that feels truly surreal and ethereal.
  • The astronaut's anatomy and physical connection to the horse are slightly awkward.
  • Some AI artifacts in the horse's mane and the background nebulae.

Verdict: GPT Image 1 followed a traditional interpretation, resulting in a high-quality but incorrect scene of an astronaut riding a horse. Grok Imagine followed the difficult 'horse on top' instruction perfectly by creating a surreal composition where the horse floats above and interacts with the astronaut, making it the clear winner for prompt adherence.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Matches the outfit from Image 2 accurately
  • + Preserves the face and general background elements
  • The lighting on the person is too moody compared to the bright beach background
  • Subtle changes were made to the hair texture and facial proportions

Grok Imagine Image

  • + Successfully preserves the original subject's exact face, hair, and lighting
  • + Maintains the bright, original beach background perfectly
  • Completely ignored the reference outfit in Image 2
  • Generated a random royal costume instead of the requested attire

Verdict: GPT Image 1 correctly followed the core instruction to use the clothing from Image 2, though it slightly altered the subject's face and lighting. Grok Imagine Image failed the task by creating a generic elaborate outfit unrelated to the reference image, despite preserving the person better.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent photorealistic texture on the capybara's fur
  • + Accurate low-light cinematic color grading
  • + Captured the 'bored' expression of the passenger perfectly despite the blur
  • The passenger is technically in the back seat but the perspective makes the car feel smaller and cramped
  • Slightly less background detail in the city lights

Grok Imagine Image

  • + Successfully shows more of the Manhattan street environment
  • + Great detail on the businesswoman and her professional attire
  • + Composition clearly separates the front and back of the vehicle
  • The capybara's hands look more like paws with prominent claws, which is slightly less 'humanized' as requested by professional expression
  • The passenger appears to be sitting in the front passenger seat rather than the back seat as requested
  • The lighting on the capybara is a bit flat compared to the surrounding environment

Verdict: GPT Image 1 followed the prompt more accurately by placing the businesswoman in the back seat, whereas Grok Imagine Image placed her in the front passenger seat. GPT Image 1 also achieved a more convincing photorealistic look with cinematic lighting, although Grok Imagine Image featured a more vibrant and detailed background of Manhattan.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent atmospheric lighting and vintage aesthetic.
  • + Highly accurate text rendering for the main titles.
  • + Strong composition with a cohesive, dark color palette.
  • Failed to include the specific time '7pm' and fused the location into the 'Time' line.
  • The border elements like thorns are less defined compared to Model B.

Grok Imagine Image

  • + Included every specific text detail, including the '7pm' time.
  • + Excellent adherence to all visual elements like thorns, webs, and the parchment texture.
  • + Clear, legible typography for the secondary details.
  • The parchment edges appear somewhat clipped against the black background.
  • The lighting is slightly flat compared to the cinematic quality of Model A.

Verdict: Both models followed the prompt well, but Grok Imagine (Image B) is the superior choice for an actual invitation as it included all the specific text details requested, including the time and separate location lines. While GPT Image 1 (Image A) has a more atmospheric and cinematic art style, its failure to correctly render the 'Time' line makes it less functional as a piece of graphic design.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
GPT Image 1
Before After
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent preservation of the original image's background, clothing, and lighting.
  • The added hair has an unrealistic, wig-like volume that doesn't naturally match the subject's head shape.
  • Loss of detail on the forehead with a muddy-looking hairline.

Grok Imagine Image

  • + Provides a highly realistic hair texture and a believable natural hairline.
  • + Maintains the anatomical proportions of the subject's head while adding density.
  • + Seamlessly blends the new hair with the existing sideburns and beard.
  • Very minor change to the shape of the glasses frame near the right temple.

Verdict: Grok Imagine is the clear winner as it provides a subtle, realistic, and anatomically correct head of hair that integrates perfectly with the subject's original features. In contrast, GPT Image 1 adds an overly voluminous, artificial-looking mass of hair that obscures the forehead and looks disconnected from the rest of the head.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent 3D miniature toy-like aesthetic with soft, high-quality textures.
  • + Text and flag icon are integrated neatly with consistent typography.
  • + Better lighting and shadows that create a sense of depth and realism within the cartoon style.
  • The diorama base is a bit large compared to the plate size.

Grok Imagine Image

  • + Successfully includes a variety of sushi types and a soy sauce bowl.
  • + Strict adherence to the 45-degree isometric projection.
  • + Clear and legible text at the top.
  • The flag icon is placed above the text rather than below it as implied by the layout request.
  • The rendering of the rice grains looks slightly less refined and more repetitive than Model A.

Verdict: GPT Image 1 (Model A) delivers a more polished, professionally rendered 3D scene with superior texture work and a more artistic interpretation of the 'miniature' prompt. Grok Imagine (Model B) follows the isometric layout strictly and provides more variety in the sushi, but the overall image quality and composition feel slightly more generic than Model A.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent hand-drawn caricature art style that feels traditional and authentic.
  • + Maintains the subject's casual denim outfit from the source image within the new context.
  • + Cleverly integrates the dog and hockey themes into the background news graphic and the desk.
  • The facial resemblance is slightly distorted by the extreme caricature style.
  • The news desk perspective is a bit flat compared to the character.

Grok Imagine Image

  • + Strong facial resemblance to the source image while maintaining the 'big head' caricature aesthetic.
  • + Highly creative and humorous incorporation of the hockey theme with the dog wearing skates and a helmet.
  • + Clean, professional digital illustration style with dynamic background elements.
  • The transition to a formal news suit loses the source preservation of the original denim shirt.

Verdict: Both models successfully interpreted the prompt, but Grok Imagine (Model B) delivered a more humorous and polished execution by putting a dog in hockey skates and keeping a closer likeness to the original person's face. GPT Image 1 (Model A) followed the instructions well with a charming hand-drawn style and preserved the original clothing, but Model B's composition was more dynamic and creative.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the 'hyper-photorealistic' part of the prompt with natural lighting and fur textures.
  • + Realistic anatomy and scale for all four animals.
  • + Beautiful implementation of golden hour god rays and shimmering atmosphere.
  • The butterfly in the top left corner has a slightly simplified, graphical look compared to the animals.

Grok Imagine Image

  • + High saturation and vibrant colors create a very joyful, whimsical vibe.
  • + Clear presence of dew sparkles as requested in the prompt.
  • The style is highly stylized/digital illustrative rather than 'hyper-photorealistic'.
  • Anatomical issues, particularly the fox's face looking more like a red panda or stylized raccoon.
  • The 'tumbling' motion looks more like floating animals than physical interactions.

Verdict: GPT Image 1 followed the prompt much more effectively, delivering a hyper-photorealistic result with convincing fur textures and natural lighting. Grok Imagine failed the realism requirement, producing a highly stylized, almost AI-cartoonish image with anatomical inconsistencies. GPT Image 1's composition feels grounded and expertly lit, whereas Grok's output feels overly saturated and flat.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent hand-painted texture resembling colored pencil or pastel work.
  • + Captures the dreamy, warm, and nostalgic mood perfectly.
  • + Faces are transformed into a genuine Ghibli-esque anime style while maintaining the distinct character expressions.
  • Loses some of the background detail from the source image due to heavy soft focus.

Grok Imagine Image

  • + High fidelity to the original source image's composition and clothing details.
  • + Clean line art and clear background elements.
  • + Good lighting that reflects a classic anime aesthetic.
  • The faces, particularly the man's, look like a filter over the original photo rather than a hand-drawn illustration.
  • Lacks the requested 'soft pastel' and 'hand-painted texture' found in Model A.

Verdict: GPT Image 1 (Model A) is the winner as it successfully reimagined the scene as a piece of art with the specific textures and mood requested in the prompt. While Grok Imagine (Model B) preserved the source image structure more accurately, it failed to provide the hand-painted, nostalgic Ghibli atmosphere, resulting in a look that feels more like a standard photo-to-cartoon filter.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
GPT Image 1
Before After
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent wind-blown hair effect that looks natural and chaotic.
  • + Maintains very high source preservation across the entire image.
  • + Subtle and realistic floating leaves that match the park's lighting.
  • The dog's tail has a slightly messy, feathered artifacting at the tip.

Grok Imagine Image

  • + Successfully adds a large volume of falling leaves and wind-blown hair.
  • + Stronger sense of 'energy' as requested by the prompt.
  • Leaves look like a flat overlay and lack depth or interaction with the scene's lighting.
  • The dog's ears have been altered into a 'floppy' position that looks slightly unnatural compared to the source.

Verdict: GPT Image 1 is the winner because its edits feel integrated into the physics of the scene; the hair has realistic movement and the leaves are subtle. While Grok Imagine Image added more motion, the leaves look like 2D stickers placed on top of the image and it modified the dog's ears unnecessarily.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the 'banner' requirement for the date.
  • + Clean, professional vector typography and layout.
  • + Accurate spelling including the grave accent on Caffè.
  • Failed the 'light background' requirement, providing a black background instead.
  • The steam element is a bit thick/heavy for a minimalist logo.

Grok Imagine Image

  • + Successfully used a light background with subtle vintage texture as requested.
  • + Elegant typography and dynamic steam illustration.
  • + Good use of the brown and cream color palette.
  • Redundant text, repeating 'Est. 1720' twice.
  • Clutter in the icon with a cup and spoon handle merging awkwardly into the cloche.
  • Banner was placed tucked under the cloche rather than being a distinct design element.

Verdict: GPT Image 1 followed the conceptual layout instructions for the banner and typography more accurately but completely failed the background color requirement. Grok Imagine Image captured the 'light background' and 'texture' requests much better, but suffered from repetitive text and a more cluttered icon design. GPT Image 1 is the likely winner for its superior professional logo composition and precision.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1
Grok Imagine Image

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the color palette with a clean, professional aesthetic.
  • + High-quality vector styling that feels like a cohesive design piece.
  • + Excellent text legibility for the main categories.
  • Included a typo 'EARLLUNAR' and misaligned labels with icons.
  • Failed to include all six requested steps in a logical sequence.

Grok Imagine Image

  • + Successfully included all six requested steps with corresponding numbered icons.
  • + Followed the layout request for a full step-by-step infographic more comprehensively.
  • + Included a recognizable NASA logo and clear heading.
  • Severe text rendering issues with secondary labels like '3rajcory' and 'Moom'.
  • The 'Translunar' icon is duplicated or confusingly split into multiple labels.

Verdict: Grok Imagine provided a much better infographic structure by including all six requested steps in order, whereas GPT Image 1 missed steps and misaligned its labels. However, GPT Image 1 has superior visual polish and much cleaner typography. Grok Imagine is the winner for better prompt adherence regarding the specific content requested.

Next steps

Explore each model