OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
LongCat-Image
#61 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
LongCat-Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to lighting instructions (soft light from the left)
- + Clean and realistic book texture
- + Good plant placement with clear visibility through glass
- − The sphere appears to be floating unnaturally above the base
- − The blue sphere has a matte texture which was unspecified but looks slightly less like a classic 'sphere' object than Model B
LongCat-Image
- + Realistic physics with the sphere resting on the bottom surface
- + Beautiful light refraction and reflection on the glass cube and sphere
- + Strong prompt adherence including the window background
- − The plant is slightly less 'behind' the cube than in Model A, appearing more to the side
Verdict: Both models followed the prompt perfectly. LongCat-Image (Model B) is the winner because it handled the physics and optical properties of glass much more realistically, specifically with the blue sphere resting on the surface and showing credible reflections. GPT Image 1 featured a floating sphere which felt like a compositional error given the realistic setting.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and realistic portrayal of an elderly person.
- + Accurate shallow depth of field that feels professional and cinematic.
- + Higher quality rendering of textures like wet fabric and metal.
- − Missed the request for motion blur on passing cars.
- − Framing is very tight, losing some of the street atmosphere.
LongCat-Image
- + Successfully captured motion blur from passing cars.
- + Good use of reflections on the wet pavement.
- + Includes visible rain streaks that add to the atmosphere.
- − Anatomical and structural issues with the bicycle (multiple wheels/frames merging).
- − Lower visual quality with a slightly plastic, AI-generated look to the faces.
- − The rain looks somewhat like a filter rather than a natural part of the scene.
Verdict: GPT Image 1 is the superior image due to its high level of realism, excellent skin textures, and coherent bicycle structure. While LongCat-Image followed more of the atmospheric prompts like motion blur and visible rain, the technical failures in the bicycle's geometry and the less realistic face rendering make it less successful overall.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and realistic, battle-worn appearance with convincing dirt and subtle scars.
- + The lighting and bokeh sparks feel integrated and atmospheric.
- + Ornate engraving on the armor is incredibly detailed with realistic weathering.
- − The beads in the hair are very subtle and blend in too much with the hair color.
- − Composition is a bit tight, losing some of the promised 'leather straps and cloth underlayer' details.
LongCat-Image
- + Very clear interpretation of the beads in the hair with vibrant colors.
- + Excellent visibility and detail on the leather straps, buckles, and chainmail/cloth transition.
- + Strong cinematic lighting with clear torchlight reflection.
- − The scars look a bit 'painted on' rather than being part of the skin texture.
- − The character looks a bit too pristine and 'pretty' for a battle-worn description.
- − The sparks in the top right are a bit distracting and look like digital artifacts.
Verdict: GPT Image 1 captures the 'battle-worn' aesthetic much better, featuring lifelike eyes and realistic skin textures that tell a story of hardship. While LongCat-Image excels at showcasing all the technical elements of the prompt (beads, leather straps, and cloth) with great clarity, the overall composition feels more like a costume photo than a grit-driven portrait. GPT Image 1 is the preferred winner for its superior realism and mood.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Strong text legibility with clean, professional sans-serif fonts.
- + Consistent high-quality food photography that fits a professional menu style.
- + Strict adherence to the 2-column minimalist layout requested.
- − Repetitive placeholder text with spelling errors ('Apperoiation descrigion').
- − Limited variety in the grid layout, focusing only on two main columns.
LongCat-Image
- + Dynamic and colorful layout with vibrant blue and yellow accents.
- + Includes a larger variety of food imagery across the design.
- − Text is illegible and contains severe AI artifacts/gibberish.
- − Layout feels cluttered and does not follow a clean professional grid as well as the competitor.
- − Source imagery has some warped elements and inconsistent lighting.
Verdict: GPT Image 1 is the clear winner as it provides a functional, high-quality design with readable text and professional food photography. While LongCat-Image is more adventurous with its color palette, the illegible text and disorganized layout make it unusable for a menu design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'exploded' layout with all components suspended separately.
- + Consistent fiery, glowing effect applied to all text elements as requested.
- + High level of texture detail on the bun, tomato, and patty.
- − The price text is missing the number '6', displaying only '.99'.
- − The starburst shape for the price is slightly irregular.
LongCat-Image
- + Accurate rendering of the numerical price and secondary message within the starburst.
- + Dynamic background featuring literal coals and flames that enhances the theme.
- + Glassy, high-quality lighting on the main title text.
- − Failed the 'exploded' layout requirement, as the burger is mostly assembled rather than suspended in mid-air.
- − The secondary message is integrated into the starburst rather than being a separate element as implied by the layout.
Verdict: GPT Image 1 followed the structural layout prompt much better, successfully creating the 'exploded' view of the burger, whereas LongCat-Image provided a mostly stacked burger. However, LongCat-Image was more accurate with the specific text '6.99', while GPT Image 1 omitted the first digit of the price.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and spelling accuracy.
- + The chalk texture is highly realistic with believable stroke variations.
- + The composition is clean and centered, making it easy to read.
- − The font lacks an 'elegant cursive' style as requested, appearing more like standard print.
- − The 'cozy café' background is missing, showing only the board.
LongCat-Image
- + Successfully captured the 'cozy café' setting in the background.
- + The chalk handwriting has a more artistic, hand-drawn flair to it.
- − Significant spelling errors throughout almost every word.
- − Failures in text rendering, with characters blending into illegible scribbles.
- − The layout is cluttered and the pricing does not align correctly with items.
Verdict: GPT Image 1 is the clear winner because it correctly spells all requested menu items and numbers, whereas LongCat-Image suffers from severe illegibility and nonsensical spelling. While LongCat-Image did a better job including the café background, its failure to generate readable text specifically requested in the prompt makes it unsuccessful for this task.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and dark atmosphere
- + Anatomically coherent horse and astronaut interaction
- + Clean, high-resolution rendering with realistic textures
- − Fails the specific negative constraint to have the horse on top of the astronaut
LongCat-Image
- + Includes interesting background elements like space debris and planets
- + Vibrantly colored assets with a clear lunar-style landscape
- − Fails the specific negative constraint to have the horse on top of the astronaut
- − Messy composition with floating artifacts and incoherent background objects
- − Anatomical issues with the horse's back legs and the astronaut's scale
Verdict: Both GPT Image 1 and LongCat-Image failed the specific logical constraint to place the horse on top of the astronaut (horse-riding-astronaut), instead providing the literal astronaut-riding-horse. However, GPT Image 1 is the superior image due to its cohesive cinematic lighting, detailed textures, and anatomy, whereas LongCat-Image contains numerous visual artifacts and a cluttered background.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the capybara's fur
- + Cinematic lighting and high-quality bokeh in the background
- + Captures the bored expression of the passenger perfectly
- − The capybara's paws look more like human-monkey hybrid hands than rodent paws
- − The viewpoint is from the front hood looking in, rather than a side profile of the interior
LongCat-Image
- + Shows more of the taxi exterior and environment for context
- + Large, detailed capybara with accurate-looking fur and ears
- + Includes the uniform jacket with a dress shirt as requested
- − Anatomical failure with the passenger, featuring two faces for one body
- − The capybara's hand/claw is poorly rendered and not properly gripping the wheel
- − The hat is floating weirdly on the head rather than being worn
Verdict: GPT Image 1 is the superior image because it maintains high visual quality and avoids the significant anatomical errors present in LongCat-Image, such as the double-faced passenger and floating hat. While LongCat-Image captures more of the taxi's scale, GPT Image 1 creates a more believable, cinematically professional scene with better lighting.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Perfect text accuracy for all requested fields including the date and location.
- + Sophisticated vintage aesthetic with atmospheric lighting.
- + Clean, professional layout that matches the gothic invitation theme.
- − The scrolls and banner elements are a bit subtle compared to the dark background.
LongCat-Image
- + Creative use of a torn parchment effect and physical thorn borders.
- + Stronger 'gothic' typeface for the main header.
- − Significant spelling errors in the event details such as 'The Armiees' and '7umm'.
- − Layout feels cluttered with inconsistent bat styles and jagged cutouts.
- − Failed to place the small banner text in a scroll as requested.
Verdict: GPT Image 1 is the clear winner as it successfully rendered all the requested text with perfect spelling and high legibility. While LongCat-Image had a creative approach to the thorn border, the composition was messy and the text generation failed significantly at the bottom of the poster.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with clean, centered alignment.
- + Consistent soft-clay 3D aesthetic that matches the 'cartoon scene' prompt perfectly.
- + High-quality lighting and shadows that give the diorama realistic depth.
- − The chopsticks are slightly merged into the diorama base.
- − The fish texture is somewhat uniform and lacks the detail seen in organic materials.
LongCat-Image
- + Beautiful subsurface scattering and textures on the salmon and tuna pieces.
- + Accurate interpretation of a traditional wooden sushi plinth as the diorama base.
- + Dynamic Flag icon with a waving effect.
- − The text alignment is slightly off-center and tilted.
- − The 'JAPAN' text has a slight visual artifact on the letter 'N'.
- − The scale of the garnish (wasabi) is a bit small compared to the sushi.
Verdict: GPT Image 1 followed the instructions for a clean, perfectly centered isometric diorama with superior typography. While LongCat-Image captured more realistic textures on the fish itself, GPT Image 1 provided a more polished and professional 'ultra-clean' composition that aligns better with the requested graphic design style.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the prompt with all four distinct animals clearly represented.
- + Very convincing interaction and dynamic 'tumbling' motion between the animals.
- + Subtle and natural-looking golden hour lighting with realistic god rays.
- − The fox kit has dark front paws that almost look like detached silhouettes in the grass.
LongCat-Image
- + Bright and colorful composition with prominent butterflies.
- + Strong 'god ray' effect coming from the top left corner.
- − Failed to include a separate baby bunny, instead merging features to create a cat with rabbit ears.
- − Anatomical issues including the dog having five legs/paws visible in the gait.
- − The dew sparkles look like floating glass orbs rather than natural moisture.
Verdict: GPT Image 1 successfully rendered all four requested animals with natural anatomy and a convincing sense of movement. LongCat-Image failed significantly on the prompt adherence by merging the cat and bunny into a single hybrid creature and exhibited several anatomical logic errors. GPT Image 1 is the clear winner for its realistic fur textures and superior composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Clean vector illustration style
- + Excellent typography and accurate spelling
- + Strong adherence to the minimalist aspect of the prompt
- − Solid black background ignores the 'light background' request
- − Slightly too simplified for a vintage style
LongCat-Image
- + Successfully uses a light background with subtle texture
- + Rich illustrative style that feels more authentically 'vintage'
- + Dynamic steam and banner design
- − Significant spelling errors with 'Caffè' appearing twice and incorrectly merged
- − Cluttered composition compared to a true logo
- − Line art on the cloche is slightly messy
Verdict: GPT Image 1 is the superior choice for a logo, offering clean vector lines, perfect typography, and a cohesive design, even though it missed the light background request. LongCat-Image captured the requested texture and background color well, but failed significantly on the text rendering and minimalism, resulting in a cluttered and unusable emblem.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with mostly legible text such as 'ARMSTRONG' and 'TRANQUILITY'.
- + Followed the color palette and flat-vector style request very well.
- + Includes most mission steps logically across the layout.
- − Has one significant spelling error for the step 'EARLLUNAR' and skips 'Lunar Orbit' as its own labeled step.
- − The layout is a bit cluttered with mismatched labels and icons.
LongCat-Image
- + Features a very clean and professional layout with good use of negative space.
- + Iconography is sharp and consistent across the design.
- + Excellent interpretation of the 'NASA-inspired' palette.
- − Failed significantly on text rendering with gibberish words like 'Aalneiitn'.
- − Missed the specific sequential mission steps requested in the prompt.
- − The rocket icon looks more like a shuttle than a Saturn V.
Verdict: GPT Image 1 is the preferred output because it successfully followed the complex instruction to detail specific mission steps and provided mostly readable, accurate text. While LongCat-Image has a more aesthetic modern layout, its failure to include the correct steps and the presence of nonsensical text makes it less useful as an infographic.
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images