OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
OmniGen v2
#57 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
0%
win rate
Ties
0%
OmniGen v2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent refraction and reflection on the glass surfaces
- + Highly realistic wood texture and lighting
- − The sphere appears slightly large relative to the 'small' prompt descriptor
OmniGen v2
- + Perfect adherence to all spatial relationships in the prompt
- + Clean, vibrant colors and sharp focus
- − Physics error where the book appears to be floating slightly above the glass
- − The sphere looks a bit like it is floating or not quite seated on the bottom surface
Verdict: Both models followed the prompt perfectly in terms of object placement and color. GPT Image 1.5 is the winner due to its superior rendering of realistic materials and physics, particularly the way the book sits realistically on the glass and the complex reflections within the cube, whereas OmniGen v2 has a slight 'floating' effect for the book.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'repairing' action with tools visible and realistic posture.
- + Highly detailed and realistic texture on the clothes, skin, and bicycle surface.
- + Successfully captures the bokeh and cinematic street photography aesthetic requested.
- − The motion blur on the car in the background is subtle rather than pronounced.
OmniGen v2
- + Beautiful, clear reflections on the wet pavement.
- + Good color contrast between the red bike and the muted street tones.
- − The subject is just standing with the bike rather than repairing it as requested.
- − The background cars are static and lack the requested motion blur.
- − The rain effect looks like a simple overlay rather than interacting with the scene.
Verdict: GPT Image 1.5 followed the prompt much more accurately, depicting a man actually engaged in repairing the bicycle with a realistic, candid framing. OmniGen v2 failed the primary action of the prompt, showing a man simply standing next to a bike, and lacked the cinematic depth and realistic textures found in the other image.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of rough, battle-worn textures on both the skin and the weathered armor.
- + Intricate engraving and realistic materials including leather and cloth.
- + Superior composition with lifelike eyes and convincing facial scarring.
- − The braids are a bit messy, though this fits the 'battle-worn' theme.
OmniGen v2
- + Clean, symmetrical braids with visible beads.
- + Good lighting contrast between the torch and the subject.
- − The skin and armor look too clean and smooth for a 'battle-worn' character.
- − The 'scars and dirt' appear like neat, painted-on freckles rather than actual grime.
- − The armor engraving lacks the depth and realism seen in the competitor.
Verdict: GPT Image 1.5 is the clear winner as it perfectly captures the 'battle-worn' aesthetic with grit, realistic scars, and highly detailed textured materials. OmniGen v2 produced a much softer, cleaner image that failed to convey the rugged experience of a paladin, with dirt that looks artificial and armor that lacks the requested engraving detail.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors.
- + Realistic food photography that matches the text descriptions.
- + Professional and functional layout suitable for a real restaurant.
- − The grid for photos is slightly irregular in alignment at the bottom.
OmniGen v2
- + Strong use of vibrant accent colors as requested.
- + Clean, minimalist aesthetic with plenty of white space.
- − Significant spelling errors in all headers (e.g., 'Restaurated Ments', 'Apptetizes').
- − The text content is illegible jibberish below the headers.
- − Poor layout coherence with repeating 'Pizzzan' and 'Pizzzas' sections.
Verdict: GPT Image 1.5 is the clear winner as it produces a fully functional, professional menu with perfectly rendered text and high-quality food photography. OmniGen v2 fails significantly on text generation, producing numerous spelling errors and gibberish that render the design unusable.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'exploded burger' requirement with all components suspended.
- + Highly detailed photorealistic textures for the meat, vegetables, and bun.
- + Dynamic composition with embers and a fiery atmosphere that matches the prompt perfectly.
- − The text 'MAGIC BURGER' is slightly distorted by the fiery effect but still legible.
OmniGen v2
- + Clean, readable graphic design for the logo and text elements.
- + Good lighting on the subject against a dark background.
- − Failed the primary prompt instruction for an 'exploded burger' with suspended components.
- − Lacks the photorealistic detail requested, appearing more like a 3D digital render.
- − Missing the currency symbol '€' and 'LIMITED TIME ONLY' text is cut off and poorly placed.
Verdict: GPT Image 1.5 is the clear winner as it perfectly captured the 'exploded' motion and photorealistic detail requested in the prompt. OmniGen v2 failed the main structural requirement of the image, provided a standard static burger, and had significant clipping issues with the secondary text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling and realistic chalk texture.
- + Highly accurate adherence to the cursive title and specific menu items requested.
- + Exceptional realism in the chalkboard surface, including smudges and dust.
- − The 'Brown Butter Chocolate Chip Cookies' item text is cut off on the right edge, though it follows the prompt's truncated text.
OmniGen v2
- + Natural variation in handwriting styles across the board.
- + Good chalkboard frame representation.
- − Contains numerous spelling errors and garbled text (e.g., 'Specals', 'Lemont', 'glite free').
- − Failed to follow the cursive title instruction, using a blocky sans-serif style instead.
- − The layout is cluttered and messy with overlapping text and incorrect price placement.
Verdict: GPT Image 1.5 is the clear winner, demonstrating near-perfect text rendering and highly realistic chalk textures that match the prompt's aesthetic requirements. OmniGen v2 struggled significantly with spelling, legibility, and following the specific formatting instructions for the title and menu items.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic detail with realistic texture on the spacesuit and horse fur.
- + Complex, visually engaging composition including a lunar surface, planets, and nebulae.
- + Strong atmospheric lighting and dynamic action pose.
- − Failed the negative constraint; the astronaut is riding the horse instead of the horse being on top.
OmniGen v2
- + Clean, illustrative style with bold colors.
- + Correctly interprets the space setting with a minimalist backdrop.
- − Failed the negative constraint; the astronaut is riding the horse instead of the horse being on top.
- − Flat lighting and lack of 'cinematic' or 'highly detailed' quality requested.
- − Anatomical issues where the horse's legs meet the bottom of the frame.
Verdict: Both models failed to follow the specific surreal instruction for the horse to be 'on top' of the astronaut. However, GPT Image 1.5 is significantly better in terms of quality, offering a cinematic and highly detailed aesthetic, whereas OmniGen v2 produced a generic, flat illustration that ignored most of the stylistic keywords.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with realistic texture on the fur, jacket, and taxi interior.
- + Perfect adherence to the 'both paws on the steering wheel' instruction.
- + Conveys the requested bored/normal atmosphere through the woman's posture and facial expression.
- − The taximeter in the bottom left is slightly blurry and lacks clear digits.
OmniGen v2
- + Accurately depicts a businesswoman in the background looking at a phone.
- + The yellow taxi exterior is vibrant and clear.
- − Anatomical failure where human hands are growing out of the capybara's sleeves.
- − The composition places the passenger in the front seat despite the prompt asking for her to be in the back seat.
- − The overall image quality has a plastic, artificial look compared to a photorealistic scene.
Verdict: GPT Image 1.5 is the clear winner as it masterfully blends a surreal concept with high-fidelity photorealism, correctly placing the paws on the wheel and the passenger in the back. OmniGen v2 fails significantly by giving the animal human hands and incorrectly placing the passenger in the front seat, which ruins the logic of a taxi ride.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling of all requested text.
- + High-quality 'vintage gothic' aesthetic with realistic textures and cinematic lighting.
- + Seamless integration of the thorn and web border within the composition.
- − The parchment texture is very dark, which may slightly reduce the readability of the subtext.
OmniGen v2
- + Clean, vector-style illustration with high contrast.
- + Good usage of the parchment scroll effect for the overall poster shape.
- − Significant text errors and garbled characters in the banner and event details.
- − Fails to follow the gothic/vintage aesthetic, looking more like a modern cartoon.
- − Incorrectly applies the thorn and web border as a thin yellow outline rather than a spooky textured element.
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the complex text requirements, flawlessly. It captured the 'vintage gothic' atmosphere with impressive textural detail, whereas OmniGen v2 struggled with spelling and produced a generic cartoonish style that did not match the requested mood.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering and perfect adherence to the flag icon request.
- + Highly detailed PBR materials with realistic textures on the wood, ceramic, and fish.
- + Great composition that feels like a professional 3D diorama.
- − Technically slightly more crowded than 'minimal garnish', though it adds to the aesthetic value.
OmniGen v2
- + Strong isometric 3d cartoon aesthetic with bold colors.
- + Clean and simple diorama base.
- + Clear and legible text.
- − The flag icon is incorrect, showing a red and yellow design instead of the Japanese flag.
- − The sushi design is biologically confusing, mixing nigiri tails with maki-style centers.
- − Lacks the 'refined textures' and 'realistic PBR' quality requested.
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the specific Japanese flag icon and the request for realistic PBR materials. OmniGen v2 failed on the flag icon and produced a more generic, plastic-looking style that didn't capture the 'refined textures' mentioned in the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'hyper-photorealistic' part of the prompt with intricate fur textures.
- + Includes all four requested animals (puppy, kitten, bunny, fox).
- + Naturalistic lighting with effective use of god rays and dew sparkles.
- − The fox's paw on the right looks slightly deformed and merged with the background.
OmniGen v2
- + Bright, vibrant colors and high contrast.
- + Clean, cute character designs for a child-friendly aesthetic.
- − Failed to include the bunny, missing 25% of the requested subjects.
- − Ignored the 'hyper-photorealistic' instruction, opting for a 3D cartoon/CGI style.
- − The butterfly on the left is floating unnaturally with no depth integration.
Verdict: GPT Image 1.5 strictly followed all aspects of the prompt, delivering a highly detailed, realistic scene with all four requested animals. In contrast, OmniGen v2 failed to include the baby bunny and completely ignored the request for photorealism, producing a stylized cartoon image instead. GPT Image 1.5 is the clear winner for its superior prompt adherence and technical execution of textures and lighting.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'hand-painted' texture and 'dreamy' lighting requested.
- + Perfectly captures the facial expressions from the source image in the new art style.
- + Preserves the depth of field and color palette of the original photo effectively.
- − The image has some digital noise/sparkle artifacts that look a bit cluttered.
OmniGen v2
- + Clean, bold lines that clearly mimic modern anime styles.
- + Good structural preservation of the characters' positions and outfits.
- − Completely fails to capture the 'jealous' and 'checking out' expressions, making everyone look happy.
- − Lacks the requested 'hand-painted textures' and 'soft pastel colors', opting for flat digital cel-shading instead.
- − The background is significantly simplified compared to the source.
Verdict: GPT Image 1.5 is the clear winner as it successfully translates the specific emotional nuances (the man's gaze and the woman's anger) of the 'Distracted Boyfriend' meme into the requested Ghibli-esque painterly style. OmniGen v2 fails the core of the editing task by removing the characters' expressions and ignoring the stylistic instructions for soft textures and lighting.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1.5
- + Excellent preservation of the original subjects and background.
- + Highly effective use of wind in the hair that looks natural yet dynamic.
- + Abundant flying leaves create a strong sense of energy.
- − The leash handle has changed slightly from the original brown leather loop.
OmniGen v2
- + Successfully added motion to the hair in a single direction.
- + Preserved the identity of the woman and the dog well.
- − Very few flying leaves were added compared to the prompt's request for a lively feel.
- − The background and overall image quality look slightly more processed and 'AI-smooth' than the original.
- − The dog's face became slightly more generic and less like the specific source dog.
Verdict: GPT Image 1.5 is the clear winner as it perfectly captured the 'energetic and lively' instruction by adding many leaves and complex hair motion while keeping the source image's texture and details intact. OmniGen v2 was much more conservative with the requested changes, adding only a few static-looking leaves and simplifying the overall image quality.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with correct spelling and sophisticated ligatures.
- + Rich, detailed textures on the cloche and banner that enhance the vintage feel.
- + Effective use of warm brown and cream tones with professional shading.
- − Failed the requirement for a light background, providing a black one instead.
- − The style is more illustrative than a 'minimalist' vector emblem.
OmniGen v2
- + Successfully followed the requirement for a light background with subtle texture.
- + Accurately captured the minimalist vector emblem style requested.
- − Significant spelling error in the main text ('CAFFFLORIN').
- − The composition of the banner and cloche feels less balanced and somewhat generic.
- − The steam element is overly simplistic and looks like a stray mark.
Verdict: GPT Image 1.5 produced a much higher quality logo with perfect typography and attractive textures, though it failed the background color requirement. OmniGen v2 correctly identified the minimalist style and background color but failed significantly on text accuracy and overall aesthetic polish. GPT Image 1.5 is the winner as the typography and professional design outweigh the background discrepancy.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Perfectly follows all 6 requested steps in order with accurate labels
- + High-quality vector illustrations that maintain a consistent art style
- + Excellent text rendering for names and phase descriptions
- − Includes a small red element at the bottom of panels that could be misinterpreted as Mars rather than Earth/Moon horizons
OmniGen v2
- + Clean minimalist color palette that matches the NASA-inspired request
- + Uses simple geometric icons consistent with flat-vector design
- − Major text errors including 'APOLO 17' instead of Apollo 11 and 'NSA' instead of NASA
- − Fails to include all 6 requested steps, providing only a vague selection
- − Gibberish text throughout the infographic labels
Verdict: GPT Image 1.5 is the clear winner as it perfectly adheres to the complex sequential instructions, providing all six specific panels with accurate text and cohesive visual storytelling. OmniGen v2 fails significantly on prompt adherence, getting the mission number wrong, misspelling basic words like 'Apollo' and 'NASA', and failing to visualize the requested steps.
Explore each model
Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on