OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
OmniGen v2
#57 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
0%
win rate
Ties
0%
OmniGen v2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'partially visible through the glass' instruction for the plant.
- + High-quality texture on the red book and wooden table.
- + Realistic refraction and reflections within the glass cube.
- − The glass cube has a slightly thick, frame-like edge that makes it look like a display case rather than a solid glass object.
OmniGen v2
- + Successfully includes all requested elements in the correct spatial arrangement.
- + Good soft lighting effect from the left side.
- + Clean, modern aesthetic with vibrant colors.
- − The plant is not visible through the glass cube, which fails a specific part of the prompt.
- − The sphere appears to be floating slightly rather than resting on the bottom surface.
- − The cube has some geometric inconsistencies where the book meets the top surface.
Verdict: GPT Image 2 followed the prompt much more accurately, particularly the difficult instruction to show the plant visible through the glass cube. OmniGen v2 failed to render the plant behind the glass, and its overall composition feels less grounded in terms of physics and lighting compared to GPT Image 2.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'repairing' aspect of the prompt with realistic tools and posture.
- + Highly realistic skin textures and lighting that feel like a real photograph.
- + Successfully captures the busy urban atmosphere with motion blur on cars and Japanese signage.
- − The composition is a bit cluttered with the foreground post and sign.
OmniGen v2
- + Strong, clean reflections on the wet pavement.
- + Good color contrast between the red bike and the green/grey background.
- − The man is just standing with the bike rather than repairing it.
- − The rain effect looks like a simple Photoshop filter overlay rather than natural atmosphere.
- − Anatomical issues with the hands and the bike's mechanical structure (chain/pedal area).
Verdict: GPT Image 2 is the clear winner as it actually depicts the man repairing the bike, whereas OmniGen v2 shows him simply standing next to it. GPT Image 2 also delivers a much more convincing 'candid street photo' aesthetic with natural skin textures and realistic environmental details, while OmniGen v2 feels more like an artificial 3D render with a rain filter.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Exceptional photographic realism in skin texture, dirt, and pores.
- + Masterful engraving detail on the armor with realistic weathering and rust.
- + Superior interpretation of 'battle-worn' with messy, complex braiding.
- − The beads in the hair are small and less distinct than in Model B.
OmniGen v2
- + Clear and distinct beads in the braided hair as requested.
- + Strong, dramatic use of 'bokeh sparks' and warm torchlight in the background.
- + Correctly identifies the leather straps and cloth underlayer.
- − The dirt on the face looks like procedural splatters rather than realistic grime.
- − The armor has a smooth, plastic-like finish that lacks the 'battle-worn' texture requested.
- − The skin is overly airbrushed and lacks the lifelike micro-texture seen in Model A.
Verdict: GPT Image 2 is significantly more lifelike, capturing the 'battle-worn' aesthetic through intricate skin textures, weathered armor engraving, and realistic dirt. While OmniGen v2 interprets the background elements like sparks and torchlight with more saturation, its subject appears overly clean and lacks the high-fidelity detail found in the armor and facial features of GPT Image 2.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with perfect spelling and coherent menu items.
- + Highly professional and realistic layout that looks ready for print use.
- + Clear and consistent categorization of Appetizers, Pizza, and Mains.
- − The grid is linear rather than a dense square grid, though it fits the categories well.
OmniGen v2
- + Successfully uses a grid layout as requested in the prompt.
- + Uses vibrant color blocks behind photos to create pop.
- − Severely garbled text and typos in headers like 'APPTETIZES' and 'RESTURATED MENTS'.
- − The text bodies are illegible filler scribbles.
- − The layout is disorganized with repeated categories and poor spacing.
Verdict: GPT Image 2 is significantly superior, producing a functional, aesthetically pleasing menu with perfectly legible text and professional food photography. OmniGen v2 fails on almost every technical level regarding text and layout, providing nonsensical words and illegible descriptions that render the menu useless.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Perfect adherence to the 'exploded burger' layout with suspended components.
- + High-quality photorealistic textures on the food and excellent fiery typography.
- + Full inclusion of all requested text elements with the correct currency symbol and styling.
- − The composition is very busy, though it fits the 'dynamic' request.
OmniGen v2
- + Clean layout with clear title rendering.
- + Good use of the starburst element for the price point.
- − Failed to create an 'exploded' burger, showing a static stacked burger instead.
- − Missing the Euro symbol and the fiery effect on the text.
- − Lower level of photorealism, appearing more like a 3D digital render than a photo.
Verdict: GPT Image 2 followed the prompt across every metric, successfully delivering an exploded food assembly with realistic textures and complex fire-styled typography. OmniGen v2 failed the primary layout instruction of an exploded burger and ignored several stylistic text requirements, resulting in a much simpler and less dynamic advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Perfectly rendered text that follows the prompt exactly without spelling errors.
- + Very realistic chalk texture and authentic handwriting style.
- + Excellent environmental lighting and café context in the composition.
- − The text is slightly thin, which might make it harder to read from a distance in a real café.
OmniGen v2
- + High contrast between the white text and the blackboard makes it very legible.
- + The wooden frame is clean and well-defined.
- − Severe spelling and character rendering issues throughout the menu items.
- − The handwriting looks more like a digital brush than natural chalk.
- − The text layout is messy and fails to follow the structured menu format requested.
Verdict: GPT Image 2 followed the prompt perfectly, delivering accurate text, realistic chalk textures, and a cozy atmosphere. In contrast, OmniGen v2 failed significantly with image-to-text rendering, resulting in garbled words and a lack of realistic handwriting style.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the complex prompt requirement of the horse being on top
- + Remarkable level of texture and detail on the space suit and lunar surface
- + Strong cinematic lighting and realistic photographic style
- − The leather straps and reins have some anatomical and physical clipping issues
OmniGen v2
- + Clean, vibrant colors and balanced layout
- + Accurate rendering of a standard horse and astronaut
- − Completely failed the specific prompt instruction for the horse to be on top
- − Style is more illustrative and less 'cinematic' or 'highly detailed' compared to Model A
- − Lack of background depth or interesting environment
Verdict: GPT Image 2 followed the very specific and difficult prompt instruction to invert the traditional rider-mount relationship, resulting in a surreal and detailed image. OmniGen v2 ignored the core 'horse on top' instruction entirely, providing a standard, cliché interpretation of the prompt. GPT Image 2's superior technical detail and prompt adherence make it the clear winner.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism and texture detail in the capybara's fur and jacket.
- + Captures the professional, calm expression requested in the prompt perfectly.
- + Accurate spatial arrangement with the passenger clearly in the back seat.
- − The passenger's face is slightly blurry and lacks fine detail.
- − One of the capybara's paws is merged with the driver's jacket sleeve.
OmniGen v2
- + Bright, saturated colors that emphasize the yellow taxi theme.
- + Clear rendering of the human passenger's features.
- − Serious anatomical failure where the human passenger appears to be reaching across to steer the car.
- − The capybara's head is awkwardly photoshopped onto a human body with human hands.
- − The perspective is cramped, making it look like the passenger is in the front seat or overlapping the driver.
Verdict: GPT Image 2 is the clear winner as it successfully handles the complex prompt with realistic lighting, textures, and proper spatial depth. OmniGen v2 fails significantly on the composition, resulting in human hands coming out of the driver's seat and a confusing arrangement of the human passenger.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with no spelling errors.
- + Sophisticated, high-detail gothic border and cinematic lighting.
- + Perfect adherence to all prompt details including the specific location and date.
- − The parchment texture is very dark, which may reduce readability of fine footer text.
OmniGen v2
- + Includes the requested twisted trees and glowing jack-o-lantern elements.
- + Good use of color contrast between the silhouettes and the background.
- − Several spelling errors in the scroll banner and event details.
- − The 'border with webs and thorns' is simplified and lacks the requested gothic intricacy.
- − Layout is cluttered with overlapping text and nonsensical lettering.
Verdict: GPT Image 2 followed the prompt perfectly, producing a professional-grade invitation with flawless text and a rich, atmospheric gothic aesthetic. OmniGen v2 struggled significantly with the text rendering and provided a much more basic, clip-art style illustration that lacked the 'vintage gothic' polish requested.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with stylization
- + Highly detailed 3D assets with realistic textures
- + Accurate Japanese flag icon
- − Lighting is a bit flat compared to the requested 'gentle' 3D aesthetic
OmniGen v2
- + Stronger isometric 3D miniature style
- + Clean and simple composition
- + Effective use of the 'gentle lighting' prompt
- − The flag icon is incorrect
- − Sushi anatomy is inconsistent (nigiri-maki hybrids)
- − Text alignment and drop shadows are slightly less refined
Verdict: GPT Image 2 followed almost every detail of the prompt, including the specific text and the correct flag, while providing a rich variety of sushi assets. OmniGen v2 captured the 'miniature diorama' lighting and simplified aesthetic better but failed on the flag icon and the logical structure of the sushi pieces. GPT Image 2 is the clear winner for its superior detail and accuracy.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'hyper-photorealistic' part of the prompt with realistic fur textures and anatomy.
- + Accurately includes all four requested animals (puppy, kitten, bunny, and fox).
- + Dynamic and natural composition showing movement and playfulness in the meadow.
- − The fox kit has black paws and features that lean slightly more towards an adult fox's markings than a typical kit.
- − The god rays are a bit overwhelming in the upper center, obscuring some background detail.
OmniGen v2
- + Bright, cheerful, and wholesome color palette.
- + Clean and simple composition suitable for a children's book style.
- − Failed the 'hyper-photorealistic' requirement, delivering a 3D digital illustration/cartoon style instead.
- − Missing one of the four animals requested (the baby bunny is absent).
- − The butterflies have unrealistic anatomy and appear pasted onto the scene.
Verdict: GPT Image 2 is the clear winner as it successfully followed all instructions, including the specific list of four animals and the 'hyper-photorealistic' style. OmniGen v2 failed the primary stylistic prompt by producing a cartoonish illustration and also missed the bunny animal requirement.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography including the correct name and date.
- + High level of detail with classic woodcut-style shading and texture.
- + Superior composition within a formal emblem frame.
- − May be slightly more ornate than the requested 'minimalist' descriptor.
OmniGen v2
- + Captures a more minimalist silhouette style.
- + Strict adherence to the requested color palette.
- − Spelling error in the main text ('CAFFFLORIN').
- − The 'Est. 1720' text is not on the banner as requested.
- − The cloche illustration is very basic and lacks the 'retro' charm of the competitor.
Verdict: GPT Image 2 is the clear winner as it successfully renders all text elements accurately and provides a professional-grade vintage aesthetic. OmniGen v2 fails on basic text rendering, with a glaring typo and a very simplistic illustration that lacks the requested texture and style.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with no spelling errors in major labels and crew names.
- + Follows all requested steps (1-6) with highly accurate and consistent iconography.
- + Superior visual quality and composition that looks like a professional infographic.
- − Some of the detailed illustrations of the Lunar Module and Saturn V deviate slightly from a 'flat vector' style, appearing more like full illustrations.
- − Included a small eagle crest that was not specifically requested, though it fits the theme.
OmniGen v2
- + Adheres strictly to the 'flat-vector' style with simple geometry and icons.
- + Uses the requested color palette effectively.
- − Severe spelling errors throughout the image (e.g., 'APOLO 17', 'NSA', 'EARTHH ORLDT').
- − Failed to follow the requested 6-step sequence, providing a cluttered and confusing layout.
- − Low visual clarity with nonsensical symbols and overlapping text.
Verdict: GPT Image 2 (Model A) is the clear winner, producing a professional-grade infographic with perfect text rendering and accurate adherence to the 6-step mission sequence. OmniGen v2 (Model B) failed significantly on both typography and instruction following, producing garbled text and ignoring the specific steps requested in the prompt.
Explore each model
Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on