OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#28 of 62 in Text-to-Image
Grok Imagine Image Pro
#17 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Grok Imagine Image Pro
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to lighting instructions with natural shadows and highlights.
- + Solid object geometry and clean textures.
- + Successful rendering of all spatial requirements including sphere placement and plant visibility.
- − The glass cube looks more like a frame as the side faces lack sufficient refractive or reflective properties.
- − The blue sphere has a matte texture that lacks the specular highlights expected in bright window light.
Grok Imagine Image Pro
- + Highly realistic textures on the distressed wooden table.
- + Superior glass rendering with convincing reflections and refractions of the interior and background.
- + Excellent text rendering on the book spine.
- − The front left edge of the cube appears to missing or completely invisible, compromising the 3D shape.
- − The sphere has a slight misalignment with its own reflection on the bottom glass surface.
Verdict: Both models followed the prompt instructions perfectly, including the specific spatial arrangements and lighting. GPT Image 1 is more consistent in its geometric construction of the cube, whereas Grok Imagine Image Pro provides a much more photorealistic render with impressive glass reflections and a beautiful table texture, despite a slight lack of definition on the cube's front edge.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the specific person's likeness, including his unique hairstyle and accessories.
- + High fidelity to the source vehicle's design and color.
- + Dynamic motion blur on the wheels and road creates a realistic sense of driving.
- − The steering wheel placement looks slightly off relative to the driver's hands.
Grok Imagine Image Pro
- + Beautiful composition of the California coastline and winding road.
- + Good preservation of the vehicle model and details.
- − Completely failed to use the specific man from the source image, replacing him with a generic white male.
- − The driver's scale and positioning within the cabin seem slightly small.
Verdict: GPT Image 1 is the clear winner as it successfully integrated both source images by placing the specific man from the second photo into the vehicle from the first. Grok Imagine Image Pro failed the core task of identity preservation, replacing the subject with a different person entirely. GPT Image 1 also conveys a better sense of speed and action through the use of motion blur.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and hyper-realistic facial details
- + Authentic shallow depth of field and soft background bokeh
- + Compelling wet-weather atmosphere and cinematic lighting
- − Physical bicycle geometry is nonsensical around the rear wheel and spokes
- − The car in the background lacks the requested motion blur effect
Grok Imagine Image Pro
- + Accurately depicts light motion blur on the passing cars in the background
- + Includes a wrench tool which adds to the narrative of repairing the bike
- + The bicycle structure is more coherent and recognizable
- − The man's scale relative to the curb and background feels slightly off
- − The lighting on the man feels more artificial than the natural look requested
- − Skin textures appear smoother and more 'generated' compared to Model A
Verdict: GPT Image 1 produces a far superior portrait with stunningly realistic skin and facial features, though it fails to render a functional-looking bicycle. Grok Imagine Image Pro better follows the secondary prompt requirements like 'motion blur' and includes storytelling details like a tool, but it lacks the photorealistic 'no stylization' quality and depth of GPT Image 1.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent atmospheric lighting with realistic reflections on the metal armor
- + Very lifelike skin texture and natural facial scarring
- + Masterful use of shallow depth of field for emotional focus
- − The beads in the hair are less prominent than requested
- − Leather and cloth underlayers are mostly obscured by shadows and the framing
Grok Imagine Image Pro
- + Highly legible and relevant Latin text ('Lux in tenebris') engraved on the armor
- + Exceptional detail on the leather straps and multiple braids with visible beads
- + Strong adherence to all specific requested elements, including the underlayer
- − The sparks look somewhat digitally overlaid rather than naturally integrated into the focal plane
- − The skin texture appears slightly more 'airbrushed' and less gritty than Model A
Verdict: Model A (GPT) excels at atmospheric realism and emotional depth, providing a more cinematic portrait with realistic lighting. Model B (Grok) is the clear winner for prompt adherence, effectively including every specific detail from the beads and leather straps to the engraved armor text, all while maintaining high visual clarity.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent high-resolution food photography that looks appetizing and professional.
- + Clean, bold typography that is very legible and matches the requested style.
- − The placeholder text contains many typos ('Apperoiation descrigion').
- − The grid layout is cut off at the bottom and feels incomplete.
Grok Imagine Image Pro
- + Successfully includes all three requested sections: Appetizers, Pizza, and Mains.
- + Utilizes a complete 3x3 grid layout that feels like a finished menu page.
- + Includes realistic footer details like an address and website.
- − The food descriptions are repetitive and nonsensical (e.g., describing Avocado Toast as 'thin crust with prosciutto').
- − The font used for the descriptions is slightly messy and harder to read than the headers.
Verdict: GPT Image 1 produces much higher quality food photography and cleaner typography, but it fails to include the full range of requested sections. Grok Imagine Image Pro follows the layout instructions better by including all sections (Appetizers, Pizza, and Mains) in a full grid, even though the descriptive text is repetitive and nonsensical. GPT Image 1 is the winner because the visual quality is significantly more professional and closer to a real-world design, despite the typos.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent glowing fiery text effect as requested.
- + Clean and symmetrical exploded composition.
- + Captures the embers and dark background atmosphere perfectly.
- − Failed the price text, showing € .99 instead of €6.99.
- − The burger layers feel a bit static compared to the requested motion.
Grok Imagine Image Pro
- + Successfully rendered the specific price of €6.99.
- + Dynamic composition with more convincing motion and drips.
- + High level of texture on the patty and bun.
- − The text doesn't fully capture the requested fiery, glowing effect, appearing more like 2D graphic overlays.
- − The starburst is a generic graphic rather than a glowing integrated element.
Verdict: Grok Imagine Image Pro was more successful in accurate text rendering, correctly displaying the price where GPT Image 1 failed. However, GPT Image 1 followed the aesthetic requirements for 'fiery, glowing' text effects much more closely, resulting in a more cohesive advertisement visual even with the typo.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent chalk texture within the letters
- + Perfect spelling and adherence to the character limit
- − Text looks more like a printed font than natural handwriting
- − The 'elegant cursive' requirement for the title was not met
Grok Imagine Image Pro
- + Natural handwriting style with varied letter slants and sizes
- + Authentic café environment with better lighting and board framing
- + More accurate interpretation of the cursive title request
- − Slightly less realistic chalk grain compared to Model A
Verdict: Grok Imagine Image Pro is the winner as it successfully captured the requested 'handwritten' aesthetic and 'elegant cursive' title, whereas GPT Image 1 produced text that looks like a static digital font. Grok Imagine Image Pro also provided a better sense of place with the visible cafe background and wooden frame.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1
- + Successfully integrated the man from Image 2 into the pose from Image 1.
- + Transferred character details like the sunglasses and scarf effectively.
- + Matched the yellow background and red ottoman from the source image.
- − Anatomy of the feet and legs is distorted and lacks natural realism.
- − The facial expression is significantly altered from the character reference.
- − The scarf physics and placement look superimposed rather than wrapped naturally in the pose.
Grok Imagine Image Pro
- + Perfectly preserved the original background, lighting, and high image quality of Image 1.
- − Failed to perform the edit requested, returning only a cropped version of Image 1.
- − Completely ignored the character reference from Image 2.
Verdict: GPT Image 1 followed the instructions by attempting to combine the two images, successfully placing the character from the second image into the pose of the first, despite some anatomical flaws. Grok Imagine Image Pro completely failed the task, simply returning a version of the first source image without any character changes. GPT Image 1 is the clear winner for actually performing the requested image editing operation.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and textured realism
- + Clear, high-quality rendering of the space suit and horse's coat
- + Follows the general theme of an astronaut and horse in space
- − Completely failed the specific inversion instruction (horse on top)
Grok Imagine Image Pro
- + Successfully interpreted the difficult 'horse on top' prompt instruction
- + Vibrant colors and a sense of surrealism
- + Good composition with a clear background element
- − The astronaut's anatomy and hand rendering are slightly distorted
- − Alignment of the horse feels more like it's floating above rather than 'riding'
Verdict: GPT Image 1 produced a much higher quality cinematic image but completely missed the tricky core of the prompt: that the horse should be on top. Grok Imagine Image Pro successfully interpreted the surreal 'horse riding astronaut' instruction, making it the superior choice for prompt adherence despite slightly lower realistic quality.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Excellent replication of the outfit from Image 2, including the specific coat, plaid scarf, watch, and jeans.
- + Near-perfect facial and background consistency compared to Image 1.
- + Realistic lighting and shadows that match the environment of the source image.
Grok Imagine Image Pro
- + Keeps the background and person's head very close to the original source image.
- − Completely ignored the clothing in Image 2, generating a generic royal outfit instead.
- − The hands are a different skin tone than the face, leading to a major anatomical inconsistency.
- − Failed the primary instruction to use the 'exact elaborate outfit' from the reference.
Verdict: GPT Image 1 followed the instructions almost perfectly, successfully transferring the specific pea coat, plaid scarf, and watch from Image 2 while maintaining the identity of the person in Image 1. Grok Imagine Image Pro completely failed the prompt by generating a random royal costume that was not present in any source image and suffered from inconsistent skin tones on the hands.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Features a high-quality close-up with excellent fur texture on the capybara.
- + The blurred city bokeh in the background is very cinematic and realistic.
- + Matches the 'bored expression' of the passenger perfectly.
- − The capybara's front paws look more like human hands in gloves than actual paws.
- − The taxi cap is a bit generic compared to the detail in Model B.
Grok Imagine Image Pro
- + Excellent environmental storytelling with the 'NYC TLC Medallion' text on the cap.
- + The composition shows more of the taxi interior and the city streets.
- + The capybara's paws look more like realistic animal claws/paws on the wheel.
- − The passenger is sitting in the front seat or a strangely configured middle seat rather than clearly in the back.
- − The perspective is slightly inconsistent, feeling more like a dashboard camera than a natural viewpoint.
Verdict: Both models followed the prompt well, but GPT Image 1 (Model A) provides a more cinematic and photorealistic close-up that successfully captures the contrast between the surreal driver and the bored passenger. Grok Imagine Image Pro (Model B) excels at specific details like the taxi medallion text and paw anatomy, but the seating arrangement of the passenger feels less like a traditional back-seat ride.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with a clean, legible gothic font
- + Strong atmosphere with dark, moody lighting and a cohesive color palette
- + Accurate rendering of the jack-o-lantern and bats in a classic style
- − Failed to include the specific time '7pm' and merged the time label with the location
- − Text hierarchy is a bit flat with large dates/locations competing with the title
Grok Imagine Image Pro
- + Perfect adherence to all text requirements including date, time, and location
- + Highly detailed border featuring the requested thorns and spiderwebs
- + Excellent use of the parchment texture and scroll banner for a vintage feel
- − The transition from the parchment invitation to the dark background border is a bit cluttered
- − The main title text has some slight kerning issues on the word 'Invitation'
Verdict: Grok Imagine Image Pro is the winner because it successfully included every piece of text requested, including the time and specific location, whereas GPT Image 1 missed the time and formatted the bottom text poorly. While GPT Image 1 has a very cohesive cinematic mood, Grok Imagine Image Pro's attention to detail regarding the border elements (thorns and webs) and the parchment aesthetic makes it a more functional invitation.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Excellent full head of hair that meets the 'thick' requirement well.
- + Preserves the original facial structure and jacket details perfectly.
- − The hairline looks slightly unnatural where it meets the forehead.
- − The hair volume is arguably a bit too high, making it look slightly like a wig.
Grok Imagine Image Pro
- + Highly realistic hairline integration with natural-looking hair follicles.
- + Exceptional preservation of the source image's lighting, skin texture, and background.
- − The hair is somewhat thin/receding, not fully meeting the 'full, thick head of hair' instruction.
Verdict: GPT Image 1 followed the instruction for thick hair more literally, but the result looks somewhat bolted-on and slightly alters the forehead shape. Grok Imagine Image Pro achieved a much more photorealistic and seamless integration of the hair, though it opted for a more conservative and slightly receding hairstyle rather than the 'full, thick' request. Grok is the winner for its superior blend and preservation of the original image's fidelity.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'cartoon' aesthetic with soft, clay-like textures.
- + Perfectly executed diorama base following the isometric perspective.
- + Superior layout and typography that feels integrated into the scene.
- − Small inaccuracies in Japanese flag proportions.
Grok Imagine Image Pro
- + Provides a wider variety of sushi items including Nigiri and Maki.
- + Clean, professional lighting with nice subsurface scattering on the salmon.
- + Accurate text rendering and flag icon placement.
- − The diorama base is circular, which feels less 'isometric' than Model A's square block.
- − The placement of the text feels slightly higher and less well-balanced within the composition.
Verdict: GPT Image 1 (Model A) better captures the requested 'isometric miniature' and 'cartoon scene' feel with its chunky, stylized geometry and a square diorama base that fits the 45-degree perspective perfectly. While Grok Imagine Image Pro (Model B) offers more food variety and realistic material translucency, its composition feels less curated than Model A's graphic design approach.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Excellent hand-drawn watercolor aesthetic typical of traditional caricatures.
- + Strong preservation of the original person's clothing and facial structure within the caricature style.
- + Clever integration of all themes in a single, balanced composition.
- − The caricature's expression is slightly more 'menacing' than 'humorous' compared to the source.
- − Some minor anatomical issues where the arm meets the desk.
Grok Imagine Image Pro
- + Very high level of detail and multiple humorous elements like the dog in a hockey helmet.
- + Perfectly legible and thematic text (e.g., 'Pups & Pucks').
- + Dynamic composition with a professional news studio feel.
- − The person feels less like a caricature of the specific source image and more like a generic stylized character.
- − The image is somewhat cluttered with repetitive dog assets.
Verdict: Both models followed the complex prompt extremely well. GPT Image 1 feels more like a genuine personalized caricature, successfully translating the subject's face and original denim shirt into a charming watercolor style. Grok Imagine Image Pro creates a more polished, high-energy digital scene with excellent text rendering and more obvious hockey references like the Stanley Cup, but it loses more of the specific likeness of the woman in the source image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent dynamic composition that captures the requested 'tumbling' movement
- + Atmospheric lighting with soft god rays and beautiful backlit fur
- + High level of detail in the fur texture and expressive facial features
- − The fox's front paws look a bit blurry/unstructured
- − Slightly tighter crop limits the scale of the meadow
Grok Imagine Image Pro
- + Clearer landscape view of the wildflower meadow and the sunrise
- + Distinctly colorful butterfly species and a wider variety of flowers
- + Playful posing of the fox kit on its back adds to the 'joyful' vibe
- − Included two kittens instead of the requested single tabby kitten
- − The golden retriever's head shape and ear placement look anatomically awkward
- − The lighting feels flatter and more artificial/HDR-like compared to the natural feel of Image A
Verdict: GPT Image 1 followed the specific count of animals requested and provided a much more realistic and artistically cohesive scene with superior backlit lighting. Grok Imagine Image Pro struggle with the prompt instructions by including an extra kitten and produced a more 'digital' looking image with anatomical inconsistencies in the dog.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Successfully captures a dreamy, textured illustration style with soft pastel colors
- + Accurately replicates the character poses and composition of the original meme
- + Creates a warm and nostalgic atmosphere as requested
- − The faces look somewhat generic and less expressive than the original characters
- − The heavy texture obscures some of the finer details of the clothing
Grok Imagine Image Pro
- + Excellent preservation of the original image's structure and background details
- + Maintains the specific facial expressions and likeness of the people in the meme
- + Successfully blends a Ghibli-esque watercolor style with the real photo
- − The lighting is a bit flatter compared to the 'dreamy' request
- − The man's beard and face look slightly like a filter rather than a fully redrawn illustration
Verdict: GPT Image 1 produces a more cohesive 'illustration' that feels like a standalone piece of art but loses some of the specific character identity. Grok Imagine Image Pro performs a superior job as an image editor by perfectly preserving the iconic expressions and background of the original meme while still applying the requested Ghibli watercolor aesthetic. Grok is the winner for its better balance of style transformation and source preservation.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Successfully added wind-blown hair and flying foliage.
- + Improved the dog's posture and facial expression, adding to the energetic feel.
- + Preserved the background bridge and landscape accurately.
- − The flying debris looks more like small twigs or dried bits rather than defined leaves.
- − Significant changes were made to the dog's face and the woman's hair texture, reducing source preservation.
Grok Imagine Image Pro
- + Excellent source preservation, maintaining the woman's face and the dog's appearance perfectly.
- + Added clearly defined yellow maple leaves that create a vibrant sense of motion.
- + Wind effect on hair is realistic while keeping the original hair color and style.
- − Some leaves appear to be floating in front of the subject without depth blurring.
Verdict: Grok Imagine Image Pro is the winner because it successfully applied all elements of the prompt while maintaining a much higher level of identity preservation for both the woman and the dog. GPT Image 1 generated a more 'energetic' dog but fundamentally changed the facial features of the subjects, whereas Grok Imagine Image Pro seamlessly integrated the wind and leaves into the existing scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with correct accents
- + Subtle sandpaper-like texture on the logo itself
- + Consistent monochromatic brown tone
- − Failed the light background requirement (output is black)
- − The cloche is a bit bulky and less elegant
Grok Imagine Image Pro
- + Adhered perfectly to the light background with subtle texture
- + Excellent vector emblem composition with circular framing
- + Elegant steam and cloche illustration
- − The accent mark on 'Caffè' is oriented incorrectly
- − The banner ribbon ends are slightly clunky in their vector execution
Verdict: Grok Imagine Image Pro followed the background and style instructions much better than GPT Image 1, creating a professional circular emblem on a textured cream background. While GPT Image 1 had better typography accuracy regarding the accent mark, its failure to provide a light background makes it less successful for the specific prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent flat-vector aesthetic that matches the 'modern vector infographic' style perfectly.
- + Bold, readable typography with high-contrast layout elements.
- + Captures the NASA-inspired specific color palette effectively.
- − Confused layout where text labels do not align with their corresponding icons.
- − Includes spelling errors like 'EARLLUNAR'.
- − The logical flow of the infographic is disjointed and hard to follow.
Grok Imagine Image Pro
- + Impeccable logical structure with a clear vertical timeline matching all requested steps.
- + Very accurate text rendering with no spelling errors.
- + Consistent iconography and professional layout that feels like a real educational poster.
- − The scale of some elements is quite small, leaving a lot of empty gray space.
- − The color palette is slightly washed out compared to the 'navy and muted red' requested.
Verdict: Model B (Grok Imagine Image Pro) is the clear winner because it successfully organized all six requested steps into a logical, numbered timeline with perfect spelling and relevant icons for each stage. Model A (GPT Image 1) has a visually striking art style but fails as an infographic due to poor alignment between the text and the icons, along with several typos.
Explore each model
xAI's premium image generation model offering higher fidelity output and stronger performance on single-image editing benchmarks compared to the standard Grok Imagine model