OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
HiDream I1 Fast
#51 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
100.0%
win rate
Ties
0.0%
HiDream I1 Fast
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of materials, especially the glass thickness and the leather-bound book texture.
- + Includes realistic refractions and reflections of the blue sphere and the plant through the glass.
- + Accurate lighting direction and intensity from the left window.
- − The plant is quite dominant in the background, making it slightly more than 'partially' visible.
HiDream I1 Fast
- + Successfully follows the spatial arrangement of the prompt.
- + Clean, minimalist composition with a light, airy feel.
- − The sphere and book look somewhat flat and lacks realistic depth/texture compared to the table.
- − The glass cube lacks thickness and realistic internal reflections for its volume.
- − The sphere appears to be floating rather than resting on the bottom of the cube.
Verdict: GPT Image 1.5 is the clear winner due to its superior rendering of light and materials, particularly the heavy glass base and the detailed texture of the red book. While both models followed the prompt accurately, HiDream I1 Fast produced a flatter image where the sphere appeared to hover weightlessly inside the cube.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'repairing' aspect of the prompt
- + Captures the 'candid' and 'imperfect framing' requested through a tighter, lower-angle crop
- + Highly realistic skin textures and wet surfaces that look like a genuine 50mm photograph
- − The motion blur of the car is present but subtle compared to the sharpness of the background lights
HiDream I1 Fast
- + Strong bokeh and reflections on the pavement
- + Includes the motion blur and cars clearly in the background
- + Cinematic color palette and atmosphere
- − The subject is sitting on the bike rather than repairing it
- − Anatomy issues where the man's left leg seems to vanish or merge into the bike frame
- − Floating effect under the man's feet where they don't quite connect to the ground and its reflection
Verdict: GPT Image 1.5 is the clear winner as it successfully depicts the man actually repairing the bicycle, whereas HiDream I1 Fast simply shows a man sitting on one. GPT Image 1.5 also achieves a significantly higher level of photorealism and physical coherence, avoiding the floating artifacts and anatomical inconsistencies found in the HiDream output.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Exceptional photographic realism with lifelike skin texture and eyes.
- + Highly detailed engraving on the plate armor and visible texture on the leather straps.
- + Perfectly executed warm torchlight lighting and atmospheric bokeh sparks.
- − The braided hair with beads is slightly less prominent than the rest of the armor details.
HiDream I1 Fast
- + Strong emphasis on the braided hair with colorful beads as requested.
- + Clear, ornate engraving and good contrast in the lighting.
- + Good adherence to the 'battle-worn' prompt with visible facial scarring.
- − Skin texture looks overly smooth and synthesized compared to a real photograph.
- − The armor looks somewhat cleaner and more 'costume-like' rather than functional battle-worn metal.
- − Lighting on the face feels a bit flat compared to the dramatic lighting on the armor.
Verdict: GPT Image 1.5 is the clear winner due to its superior photorealism and textural detail; the rendering of the skin, the micro-scratches on the armor, and the atmospheric lighting are professional grade. While HiDream I1 Fast followed the specific detail of the beads in the hair more vibrantly, the overall image quality feels more like a high-quality 3D render rather than the lifelike portrait achieved by GPT Image 1.5.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with highly legible and appropriate fonts.
- + Smart and professional layout that organizes sections clearly.
- + High-quality food photography that accurately represents the described menu items.
- − The food photos/grid take up a significantly large portion of the right side, leaving little breathing room for longer descriptions.
HiDream I1 Fast
- + Follows the request for a grid of food photos efficiently.
- + Includes a clear header section for the restaurant name.
- − Text is completely illegible and contains gibberish characters.
- − Failed to create separate sections for pizza and mains as distinct categories.
- − Poor overall image quality with significant artifacts and low resolution.
Verdict: GPT Image 1.5 produced a professional, production-ready menu layout with perfect text legibility and high-quality imagery. In contrast, HiDream I1 Fast failed on almost every technical level, producing garbled text and low-resolution graphics that do not meet professional standards. GPT Image 1.5 is the clear winner for its adherence to functional design principles and the prompt's structural requirements.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic texture on food items like the patty and bun.
- + Perfect text rendering for all requested messages.
- + Highly dynamic composition with a strong sense of an 'exploded' view.
- − The background is very busy, slightly distracting from the central product.
HiDream I1 Fast
- + Clean layout with a clear focus on the product.
- + Good use of the fiery glowing effect on the title text.
- − Failed to properly create an 'exploded' burger, with components mostly stacked.
- − Poor text rendering on secondary messages with spelling errors.
- − Visual artifacts present on the floating onion rings and starburst.
Verdict: GPT Image 1.5 follows the prompt with much higher fidelity, successfully rendering all requested text perfectly and creating a truly dynamic 'exploded' burger effect. HiDream I1 Fast struggles with the composition, failing to separate the burger layers and producing garbled text for the secondary messages.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling and no artifacts
- + Authentic chalk texture and natural handwriting variations
- + Strict adherence to the formatting and content of the prompt
- − Simple composition focusing only on the chalkboard surface
- − Lacks the 'cozy café' environmental context mentioned in the prompt
HiDream I1 Fast
- + Great environmental context showing a coffee shop interior
- + Handwritten aesthetic for the title text is pleasing
- − Severe text garbling and overlapping characters on menu items
- − Spelling errors such as 'Buter' and '$2 $9'
- − Frequent digital artifacts and messy letter rendering
Verdict: GPT Image 1.5 is the clear winner as it flawlessly executes the complex text requirements of the prompt with realistic chalk textures and perfect spelling. While HiDream I1 Fast provides a better environmental composition of a café, the actual text is largely illegible and filled with digital hallucinations.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'in space' setting with detailed planets and stars
- + High level of texture detail in the astronaut suit and horse fur
- + Cinematic composition with dynamic lighting and a sense of motion
- − The horse's front legs and hooves have anatomical irregularities
- − The dust/ground interaction is a bit nonsensical for deep space
HiDream I1 Fast
- + Clean and clear render of the horse and astronaut
- + Good posture and lighting on the main subjects
- − Fails to follow the 'in space' prompt, placing the subjects in a desert instead
- − The sky and background are mundane rather than surreal or cinematic
- − Lacks the high-detail complexity requested
Verdict: GPT Image 1.5 is the clear winner as it successfully interprets the 'in space' setting and cinematic style requested in the prompt. While HiDream I1 Fast produces a clean image, it completely fails to place the subjects in space, showing them in a desert with a blue sky.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with shallow depth of field and convincing taxi meter details
- + Capybara hands/paws are correctly placed on the steering wheel as requested
- + Strong adherence to the requested 'bored' expression for the passenger
- − Slightly tight framing cuts off edges of the scene
HiDream I1 Fast
- + Captures the exterior street lights and bokeh well
- + Clearer view of the overall taxi structure
- − Anatomical failure with human hands protruding from the capybara's body to hold the wheel
- − The passenger is sitting in the front seat instead of the back seat as requested
- − The capybara is positioned as a passenger rather than the driver in relation to the steering wheel
Verdict: GPT Image 1.5 is the clear winner as it accurately places the capybara in the driver's seat with its paws on the wheel and the woman in the back seat. HiDream I1 Fast fails the spatial prompts, placing the woman in the front seat and giving the capybara human hands that are disconnected from its body.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Perfect text rendering for all requested strings.
- + Excellent artistic atmosphere with a detailed thorn and web border.
- + Cinematic lighting that perfectly matches the 'moody night sky' and 'dark parchment' request.
- − None notable for this prompt.
HiDream I1 Fast
- + Follows the general composition and requested pumpkin placement.
- + Vibrant colors that stand out against the background.
- − Severe text artifacts and spelling errors in the scroll and event details.
- − The 'dark parchment' is flat and lacks the vintage gothic texture of the other version.
- − Composition feels like a clip-art collage rather than a 'polished' cinematic poster.
Verdict: GPT Image 1.5 is the clear winner as it flawlessly rendered all requested text, including the complex scroll and location details. While HiDream I1 Fast struggled with text legibility and artistic detail, GPT Image 1.5 delivered a high-quality, professional-looking invitation with superior parchment textures and atmospheric lighting.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering and placement that feels integrated into the design.
- + Highly detailed PBR materials with realistic textures on the wood, fish, and ceramic.
- + Rich composition with diverse sushi types, a teapot, and soy sauce bottle.
- − The scene is slightly more complex than 'minimal garnish' requested.
HiDream I1 Fast
- + Captures the 'minimal' aspect of the prompt very well.
- + Follows the isometric layout and diorama base requirement accurately.
- + Clean, soft cartoon aesthetic with smooth surfaces.
- − The sushi design is illogical, featuring a piece of fish on top of a maki roll.
- − Text is slightly cluttered with extra symbols and lines that weren't requested.
- − Material textures are very flat compared to the PBR request.
Verdict: GPT Image 1.5 produced a much higher quality image with sophisticated material rendering and perfect text spelling. While HiDream I1 Fast followed the 'minimal' constraint more closely, it failed on the anatomical logic of the sushi (placing nigiri toppings on a roll) and had much lower texture detail. GPT Image 1.5 is the clear winner for its professional finish and adherence to the PBR and high-clarity requirements.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the prompt by including all four specified animals (puppy, kitten, bunny, fox).
- + High level of detail in fur textures and dew sparkles, creating a very tactile feel.
- + Dynamic composition that successfully captures the 'tumbling together' and 'playful' aspect of the prompt.
- − The fox kit has a slightly distorted paw area on the bottom right.
- − The sun rays are a bit digitally over-enhanced, bordering on artificial.
HiDream I1 Fast
- + Beautiful, soft bokeh and lighting that creates a warm, wholesome atmosphere.
- + Clean rendering of the animals' faces with very expressive eyes.
- + Good placement of butterflies to lead the viewer's eye through the frame.
- − Failed to include the baby bunny as requested in the prompt.
- − Included two kittens instead of one kitten and one bunny.
- − The pose is relatively static and lacks the 'tumbling' energy described.
Verdict: GPT Image 1.5 is the clear winner as it followed all instructions, including the specific list of four different animals. While HiDream I1 Fast produced a visually pleasing and clean image, it failed on a key prompt requirement by omitting the bunny and duplicating the kitten, whereas GPT Image 1.5 captured the chaotic, playful energy requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling and accent marks
- + Sophisticated use of stippling and gradients for a vintage texture
- + Strong emblem composition that feels professional and balanced
- − Ignored the request for a light background, opting for solid black
- − The steam effect is a bit chunky compared to the fine lines of the logo
HiDream I1 Fast
- + Successfully followed the instructions for a light background with subtle texture
- + Good minimalist vector style that feels clean and modern-retro
- − Significant typo in the year, rendering it as '17210' instead of '1720'
- − The steam graphics are messy and inconsistent with the circular frame
- − Text alignment inside the cloche is slightly off-center
Verdict: GPT Image 1.5 produced a much higher quality logo with perfect typography and sophisticated shading, though it failed to use the requested light background. HiDream I1 Fast followed the background color instructions but failed significantly on the text elements, including a major typo in the established date and poor integration of the steam icons. GPT Image 1.5 is the clear winner for its professional finish and accuracy in text rendering.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the six-step sequence with perfectly legible text.
- + Highly consistent flat-vector style that perfectly matches the 'modern infographic' request.
- + Accurate iconography representing the Saturn V and Lunar Module.
- − The inclusion of a red horizon in the Earth/Moon orbit panels is slightly confusing geographically.
- − The three silhouettes at the top are identical rather than distinct.
HiDream I1 Fast
- + Captures the requested NASA color palette effectively.
- + Includes a central Saturn V rocket that acts as a focal point.
- − Severe spelling errors and garbled text throughout the infographic.
- − The sequence of steps is illogical and does not match the prompt's instructions.
- − Visual artifacts and messy lines detract from the flat-vector aesthetic.
Verdict: GPT Image 1.5 produced a professional, publication-ready infographic that followed every step of the prompt with perfect text rendering and style consistency. In contrast, HiDream I1 Fast failed to follow the logical sequence of the mission and suffered from significant text corruption and incoherent iconography. GPT Image 1.5 is the clear winner for its clarity, composition, and prompt adherence.
Explore each model
Distilled version of HiDream AI's 17B parameter text-to-image model