OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
LongCat-Image
#61 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
100.0%
win rate
Ties
0.0%
LongCat-Image
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photographic texture on the red book cover and pages.
- + Superior representation of plants visible through the glass with realistic distortion.
- + Perfect adherence to all spatial requirements in the prompt.
- − The sphere is quite large relative to the 'small' descriptor.
- − The bottom edge of the glass cube looks slightly thick, almost like a tray.
LongCat-Image
- + Beautiful translucent material for the blue sphere.
- + Clear lighting direction from the window as requested.
- + Clean, minimalist composition.
- − The glass refractive logic is a bit weak, as the plant behind it looks almost unchanged through the glass.
- − The book's placement and texture feel slightly more synthetic than Model A.
Verdict: Both models followed the prompt perfectly, but GPT Image 1.5 is the winner due to its superior rendering of textures and complex refractive physics. Specifically, GPT Image 1.5 captures the way the green leaves distort through the glass cube more realistically than LongCat-Image, which feels slightly flat behind the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent handling of wet textures and reflections on the pavement.
- + Captures highly realistic skin texture and clothing details with a candid feel.
- + Accurate bike mechanics and authentic rain particles.
- − The car in the background lacks the specific motion blur requested in the prompt.
LongCat-Image
- + Successfully incorporates motion blur into the passing cars.
- + Good use of cinematic lighting from headlights and street storefronts.
- + Matches the 'imperfect framing' request by including a foreground pole.
- − The red bicycle involves nonsensical geometry with three wheels and overlapping frames.
- − The man's hands and the bike parts have significant AI artifacts.
- − The rain looks like vertical streaks rather than a natural mist.
Verdict: GPT Image 1.5 delivers a far more grounded and realistic image with superior anatomical and mechanical accuracy, despite missing the motion blur on the cars. LongCat-Image attempts more of the complex prompt instructions like motion blur and imperfect framing, but fails significantly on basic physical coherence, resulting in a distorted three-wheeled bicycle and messy details. GPT Image 1.5 is the clear winner for its photographic quality and adherence to 'no stylization'.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Exquisite micro-textures on the engraved metal and worn leather straps.
- + Masterful use of lighting and depth of field that feels professional and cinematic.
- + Authentic portrayal of 'battle-worn' with realistic dirt and grimy skin texture.
- − The hair beads are somewhat subtle and blend into the hair color.
LongCat-Image
- + Clearly follows the request for braided hair with visible beads.
- + Good contrast between the polished armor and the fire effects in the background.
- + Correctly incorporates the cloth underlayer and leather straps.
- − The scarring looks like a digital stamp rather than a natural part of the skin.
- − The lighting on the character feels flat and disconnected from the bright fire behind them.
- − Overall image has a slightly smoothed, CGI aesthetic compared to the first.
Verdict: GPT Image 1.5 delivers a high-fidelity, cinematic portrait with incredible attention to the texture of the weathered armor and skin, which perfectly captures the 'battle-worn' prompt. LongCat-Image follows all prompt instructions but suffers from a more artificial, less integrated lighting style and lacks the lifelike detail found in the competitor's work.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling and pricing consistency.
- + Clean, professional layout that matches the 'modern minimalist' aesthetic perfectly.
- + High-quality, realistic food photography that aligns with the specific menu items.
- − The grid is slightly asymmetrical, focusing more on a two-column split than a modular grid.
LongCat-Image
- + Dynamic use of color blocks and a more 'vibrant' accent style as requested.
- + Good variety in the food photo grid layout.
- − Text is completely illegible and garbled, failing basic professional menu standards.
- − Failed to include a specific 'Mains' section that is readable.
- − Visual artifacts and warped plate shapes in several photos.
Verdict: GPT Image 1.5 is the clear winner as it produces a fully functional, professional-grade menu with perfect text rendering and high-quality food photography. LongCat-Image fails the primary requirement of a menu by providing garbled, nonsensical text and lower-quality image assets that lack the clean minimalism requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent 'exploded' effect with clearly separated, dynamic components.
- + Perfect execution of the fiery, glowing background and text effects.
- + Highly detailed food textures that look appetizing and photorealistic.
- − The crown of the 'M' in the title text is slightly cut off at the top edge.
LongCat-Image
- + Clean and legible typography for all requested text elements.
- + Good lighting on the food items making them pop against the dark background.
- − Failed the 'exploded burger' prompt; the ingredients are stacked normally rather than suspended.
- − The overall image feels more like a static graphic than a dynamic, motion-filled ad.
- − The bun texture looks slightly artificial compared to Model A.
Verdict: GPT Image 1.5 is the clear winner as it perfectly captured the 'exploded' motion requested in the prompt, whereas LongCat-Image provided a standard stacked burger. GPT Image 1.5 also excelled in creating a cohesive, high-energy atmosphere with superior photorealistic textures and better integration of the fiery theme.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors.
- + Highly realistic chalk texture with dusty smudges and natural handwriting variations.
- + Consistent and elegant cursive style throughout the entire board.
- − The background context of the 'cozy café' is barely visible compared to the close-up of the board.
LongCat-Image
- + Good environmental storytelling, showing the menu board within a nicely blurred café setting.
- + The board hardware and easel look physically realistic.
- − Significant spelling errors on almost every word (e.g., 'Toays Srays', 'Arlil').
- − Text rendering is messy and lacks the 'elegant cursive' requested in the prompt.
- − Incorrect formatting and layout compared to the prompt instructions.
Verdict: GPT Image 1.5 is the clear winner as it perfectly rendered the complex text requirements of the prompt with 100% accuracy and a very convincing chalk-on-blackboard texture. LongCat-Image struggled significantly with text generation, producing numerous spelling hallucinations and failing to follow the specified layout.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic lighting and texture
- + Highly detailed environment with realistic lunar dust effects
- + Superior integration of the astronaut and horse into the scene
- − The astronaut's leg positioning is slightly awkward against the horse's flank
LongCat-Image
- + Clean, bright color palette
- + Good legibility of the subject against the dark space background
- − Floating artifacts and strange aircraft shapes in the sky
- − Horse's anatomy is distorted, particularly the elongated front legs
- − Overall composition feels less cinematic and more like a collage
Verdict: GPT Image 1.5 is the clear winner as it produces a professional, cinematic image with cohesive lighting and deep textures that match the 'highly detailed' prompt. LongCat-Image suffers from significant anatomical distortions in the horse and contains several nonsensical artifacts in the background that detract from the surrealism.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic texture on the capybara fur and clothing.
- + Cinematic lighting and high-quality background bokeh.
- + Perfectly captures the 'bored' expression of the passenger as requested.
- − The capybara's paws look slightly more like human fingers with hair than actual paws.
- − The taxi interior is a bit dark, making it hard to see dashboard details.
LongCat-Image
- + Clear side-view composition showing the full exterior and interior dynamic.
- + Accurately places the capybara in a professional jacket with dress shirt cuffs.
- + Good rendering of the city environment outside the window.
- − The capybara's hand/paw anatomy is distorted and looks like sharp claws.
- − The passenger is duplicated/hallucinated into two very similar people.
- − The hat is floating awkwardly above the capybara's head.
Verdict: GPT Image 1.5 is the clear winner due to its superior photorealistic quality and adherence to the character descriptions. LongCat-Image suffers from significant anatomical distortions in the capybara's hands and unnecessarily duplicates the passenger in the back seat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors in any of the requested fields.
- + Highly cohesive 'vintage gothic' aesthetic with a beautiful dark parchment texture.
- + Great composition that integrates all elements within a unified, atmospheric scene.
- − The jack-o-lantern is central but slightly smaller than the text in weight.
- − The 'parchment' effect is applied to the whole scene rather than just a poster element.
LongCat-Image
- + Good use of negative space to make the central jack-o-lantern and title stand out.
- + Vibrant colors and high contrast between the parchment and the night sky window.
- − Multiple spelling and text errors including '3uulie', '7mm', and 'The Armiees'.
- − The thorns are floating awkwardly around the paper rather than forming a natural border.
- − The composition feels fragmented with the scene appearing inside a cut-out hole in the paper.
Verdict: GPT Image 1.5 is the clear winner as it successfully rendered every piece of text without a single spelling error, which is crucial for an invitation prompt. While LongCat-Image had a creative layout, it suffered from hallucinated text and poor integration of the thorn border. GPT Image 1.5 perfectly captured the 'vintage gothic' request through its lighting, texture, and font choice.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent PBR material rendering with realistic wood grain, ceramic, and liquid textures.
- + Strict adherence to the diorama base requirement and isometric perspective.
- + Clean, high-quality text rendering and professional layout.
- − Included many extra elements like a teapot and soy sauce bottle not explicitly requested in the 'minimal garnish' instruction.
LongCat-Image
- + Successfully captured the 'cartoon' and 'miniature' aesthetic with soft, clay-like textures.
- + Very clean and minimalist composition that focuses on the sushi.
- + Handled the text and small flag icon correctly.
- − The 'diorama base' is very simple and lacks the detail of a true 3D miniature scene.
- − The fish textures look more like plastic or rubber than 'realistic PBR materials' requested.
Verdict: GPT Image 1.5 produced a superior technical result with high-fidelity materials and a well-defined diorama world that feels truly 3D and isometric. While LongCat-Image captured the 'cartoon' aspect well, it failed to provide the realistic material quality and environmental depth that GPT Image 1.5 achieved.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'tumbling together' part of the prompt with dynamic posing.
- + Superior texture rendering of fur and dew sparkles.
- + High anatomical accuracy for all four requested animals.
- − The fox kit and puppy share very similar facial structures, making them look related.
LongCat-Image
- + Strong 'god rays' and lighting effects.
- + Clear, vibrant colors in the wildflower meadow.
- − Failed to include the baby bunny, instead merging ears onto the kitten to create a 'cabbit' hybrid.
- − The animals are static and posing rather than 'playfully chasing and tumbling'.
- − The fox has an unnaturally long, stiff tail that lacks realistic physics.
Verdict: GPT Image 1.5 is the clear winner because it correctly depicts all four requested animals with high anatomical accuracy and a realistic 'tumbling' interaction. LongCat-Image failed a core part of the prompt by omitting the bunny and instead generating a kitten with rabbit ears, while also featuring less realistic fur textures.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent vector emblem aesthetic with clean, professional typography.
- + Accurately represents the cloche dome with subtle steam and lighting.
- + High contrast and clean layout suitable for an actual logo.
- − Failed the light background requirement by using a solid black background.
- − The 'Est. 1720' text is slightly off-center within the banner.
LongCat-Image
- + Successfully incorporated the light background with subtle texture.
- + Creative integration of the text into the cloche shape.
- + Captures the warm brown and cream tones perfectly.
- − Redundant text rendering with 'Caffè' appearing twice.
- − Typographic execution is messy with overlapping letters and inconsistent line weights.
- − Composition feels cluttered compared to a professional logo.
Verdict: GPT Image 1.5 produced a much more professional and viable logo design with superior typography and vector clarity, though it failed the background color prompt. LongCat-Image followed the background and texture prompt more closely but suffered from significant text errors and a cluttered composition. GPT Image 1.5 is the preferred choice for a logo design due to its clean execution and readability.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the sequential steps requested.
- + Perfect text rendering for all mission phases.
- + Clean, professional vector aesthetic that matches the NASA-inspired palette.
- − Simple, almost clip-art style for the crew silhouettes.
- − A small typo in 'Tranquillity' (usually spelled with one 'l' in US English, though double is acceptable in other regions).
LongCat-Image
- + Creative use of a map-pin style icon for the lunar landing.
- + Follows the requested color palette reasonably well.
- − Failed significantly on text rendering with gibberish words.
- − Did not follow the specific 6-step sequence requested.
- − Icons like the rocket resemble a shuttle rather than a Saturn V, and the lunar module design is messy.
Verdict: GPT Image 1.5 produced a highly usable and accurate infographic that perfectly followed all six requested steps with crisp, legible text. In contrast, LongCat-Image failed to include the correct steps and suffered from significant text hallucinations and incoherent layout. GPT Image 1.5 is the clear winner for its functional design and adherence to the structured prompt.
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images