OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Z-Image Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'soft window light from the left' instruction.
- + High visual clarity and realistic textures on the sphere and book.
- + Strong composition with a clearly defined plant visible through the glass.
- − The sphere appears slightly large for the description of 'small'.
- − The bottom of the glass cube has a heavy metallic base not explicitly requested.
Z-Image Turbo
- + Accurately depicts a 'small' blue sphere as requested.
- + The glass cube looks more like a standard hollow glass vessel.
- + Good surface reflections on the wooden table.
- − The plant is extremely blurry and barely recognizable in the background.
- − The lighting is somewhat flat and lacks the clear directional quality requested.
- − There are some minor artifacts on the edges of the book.
Verdict: GPT Image 1 succeeded in capturing the specific lighting and spatial relationship between the plant and the glass cube much more effectively than Z-Image Turbo. While Z-Image Turbo followed the scale of the 'small' sphere better, GPT Image 1's superior focal depth and textural details make it the more visually striking and accurate representation of the prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent shallow depth of field and bokeh effects
- + Highly realistic skin textures and fine details on the bicycle and clothing
- + Strong cinematic lighting and atmospheric mood with wet pavement reflections
- − The 'imperfect framing' requested is somewhat subtle as the composition is quite balanced
- − Lacks significant motion blur for the passing cars
Z-Image Turbo
- + Successfully captures an older Japanese man with a red bicycle
- + Good rendering of the wet pavement and rain streaks
- − The man is pushing or standing by the bike rather than repairing it
- − Composition lacks the shallow depth of field and 'cinematic' feel requested
- − Anatomical issues with the right hand merging into the handlebars
Verdict: GPT Image 1 far exceeds Z-Image Turbo in terms of visual quality, realism, and adherence to the cinematic requirements of the prompt. While Z-Image Turbo provides a more literal 'street photo' composition, it fails on key technical aspects like depth of field and the specific action of 'repairing' the bike, whereas GPT Image 1 creates a compelling, professional-looking image with natural textures.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI judge analysis unavailable for this challenge.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with almost perfect spelling.
- + High-quality, appetizing food photography that looks professional.
- + Clean and balanced layout that adheres strictly to the minimalist prompt.
- − The grid structure is a bit basic, feeling more like a cropped page than a full menu sheet.
- − Missing the 'Mains' category header despite showing a main course image.
Z-Image Turbo
- + Strong grid layout with a high volume of food images.
- + Bold use of vibrant orange accents as requested.
- + Effective use of negative space around the menu borders.
- − Significant spelling errors and garbled text (e.g., 'PIZZA MANS', 'SE IIION').
- − Food photography is slightly less realistic with some AI artifacts in the pizza toppings.
- − Composition is a bit cluttered compared to the minimalist request.
Verdict: GPT Image 1 is the superior choice because it delivers high-quality food photography and clear, legible text that actually functions as a menu. While Z-Image Turbo has a more interesting 3x3 grid composition, its inability to render sensible english text and 'Pizza Mans' typo makes it unusable for a professional design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'exploded' concept with space between all ingredients.
- + High-quality photorealistic food textures, especially the lettuce and bun.
- + Text is perfectly integrated with the requested ember/fiery glow effect.
- − The price in the starburst is missing the '6' digit, reading only '€ .99'.
- − The composition is a bit tight at the top and bottom edges.
Z-Image Turbo
- + Clean, readable text that includes all requested characters correctly.
- + Dynamic environment with physical fire and depth-of-field effects.
- + Vibrant colors and a polished commercial aesthetic.
- − Failed the 'exploded' instruction as the ingredients are mostly stacked normally.
- − The burger looks more 'tilted' than 'suspended' or 'exploding'.
- − The starburst effect is a yellow glow rather than the requested 'fiery, glowing effect' seen in the title.
Verdict: GPT Image 1 followed the conceptual layout instructions much more accurately, successfully creating an 'exploded' view where all components are separated in air. However, it failed on the specific numeric text for the price. Z-Image Turbo produced a cleaner, more typical advertisement with correct text, but largely ignored the primary 'exploded' burger prompt requirement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent chalk texture throughout the lettering
- + Followed the prompt to complete the 'Brown Butter' text intelligently
- + Perfectly legible and aesthetically pleasing layout
- − Missed the request for 'elegant cursive' for the title
- − The handwriting is a bit too uniform, bordering on looking like a font
Z-Image Turbo
- + Realistic smudging and dust on the chalkboard surface
- + Varied letter sizes and slanting feel very human
- + Bold, clear presentation
- − Spelling error: 'Mustroom' instead of 'Mushroom'
- − Missed the request for elegant cursive in the title
- − The chalk strokes look a bit more like a digital brush than actual chalk on board
Verdict: GPT Image 1 followed the complex text requirements more accurately, including the completion of the truncated 'Brown Butter...' item, and maintained perfect spelling. Z-Image Turbo captures a more realistic chalkboard atmosphere with smudges, but failed on basic spelling with the word 'Mustroom'.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and atmospheric depth in the space background.
- + Highly detailed texture on the astronaut suit and horse's coat.
- + Clearer visual of Earth and distant celestial bodies enhancing the setting.
- − The reins seem to disappear or merge oddly with the astronaut's hand.
- − The horse's back hooves have a slightly strange shape/orientation.
Z-Image Turbo
- + Bright, clear focal point with good color contrast between the white suit and brown horse.
- + The horse's mane and tail have a dynamic, flowing appearance.
- − Failed the specific prompt instruction 'horse on top, not vice versa' by placing the astronaut on top of the horse.
- − The background is relatively flat and less cinematic than the competitor.
- − The lighting on the astronaut doesn't naturally match the deep space environment.
Verdict: Both models failed the negative constraint and literal interpretation of 'horse on top, not vice versa', instead providing a standard 'astronaut riding a horse' image. GPT Image 1 is the superior image due to its stunning cinematic lighting, higher texture detail, and better integration of the subjects into the space environment, whereas Z-Image Turbo feels more like a composite with flatter lighting.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic lighting and shallow depth of field.
- + Accurate text rendering on the hat.
- + High-quality fur texture and realistic passenger expression.
- − The capybara's paws look slightly human-like in their grip.
- − Visible exterior roof light is partially cut off at the top.
Z-Image Turbo
- + Good inclusion of safety elements like the seatbelt.
- + Clear side profile of the capybara showing the professional expression.
- + Accurate placement of both front paws on the steering wheel.
- − The passenger appears to be in the front passenger seat rather than the requested back seat.
- − The hands/paws on the steering wheel have a distorted, primate-like anatomy.
Verdict: GPT Image 1 followed the prompt much more accurately, correctly placing the passenger in the back seat and providing a more cinematic, photorealistic atmosphere. While Z-Image Turbo had good lighting, its failure to place the passenger in the rear and the unnatural anatomy of the driver's hands makes GPT Image 1 the superior choice.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Exceptional typographic layout and legibility
- + Atmospheric cinematic lighting on the jack-o-lantern
- + Clean and professional graphic design style
- − Confused the Time and Location fields in the bottom text
- − The parchment texture is very subtle compared to the request
Z-Image Turbo
- + Stronger adherence to the 'dark parchment' and 'thorns' prompt elements
- + Included all required text fields correctly
- + Dynamic and layered composition
- − Typo in the location text ('Archves' instead of 'Arches')
- − The 'You are invited...' banner is floating awkwardly without a scroll behind it
- − Overall image is a bit cluttered
Verdict: GPT Image 1 offers a much more polished and professional aesthetic with superior lighting and font choices, though it fails on the logic of the bottom text fields. Z-Image Turbo captures more of the specific prompt details like the thorns and the literal parchment look, but suffers from a typo and less cohesive design. GPT Image 1 is the likely winner for its visual sophistication despite the text field error.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Successfully added a thick, full head of hair as requested
- + Preserved the glasses and majority of the original facial features
- + Maintained the original background and lighting well
Z-Image Turbo
- + Improved the texture and color of the beard
- + High overall image resolution and clarity
- − Failed the primary edit instruction by leaving the person largely bald
- − Altered the face shape and removed the person's glasses
- − Failed to preserve the original person's identity and facial features
Verdict: GPT Image 1 followed the instructions accurately by adding a full head of hair while preserving the identity and glasses of the subject in the source image. Z-Image Turbo failed the prompt entirely, failing to add substantial hair and significantly changing the person's face while removing his spectacles.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Perfectly followed the text instructions, including the correct Japanese flag.
- + Higher complexity in the 3D scene with multiple sushi types and condiments.
- + Superior rendering of textures, particularly the subsurface scattering effect on the salmon.
- − The text 'JAPAN' is slightly off-center compared to the 'SUSHI' text.
Z-Image Turbo
- + Clean, minimalist aesthetic that fits the 'miniature' prompt well.
- + Good use of material depth on the diorama base.
- + Excellent centering of the text elements.
- − Used the Chinese flag instead of the requested Japanese flag for a Japan-themed image.
- − The sushi composition is very basic compared to the other model.
- − Visible artifacting/blurriness on the top of the salmon piece.
Verdict: GPT Image 1 is the clear winner as it accurately rendered the Japanese flag and provided a much more detailed and professional-looking 3D scene. Z-Image Turbo failed the specific prompt requirements by displaying a Chinese flag for a dish and label explicitly identified as Japanese, and the overall image quality was lower with less refined textures.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Perfectly follows the caricature instruction with exaggerated features and a watercolor art style.
- + Successfully incorporates all requested elements: news anchor desk, hockey equipment, and a companion dog.
- + Maintains recognizable features from the source image, such as the hair color and denim shirt.
- − The transition to a hand-drawn style means total loss of the original photographic background.
- − The hockey stick on the desk is slightly warped in perspective.
Z-Image Turbo
- + Excellent preservation of the original image's lighting, skin texture, and background composition.
- + Subtly adds a small dog in the background, maintaining a realistic photographic style.
- − Completely failed the main instruction to create a 'caricature' and make it 'exaggerated'.
- − Missing the 'tv show anchor' and 'hockey' elements entirely.
- − The added dog is blurry and lacks detail.
Verdict: GPT Image 1 followed the creative brief perfectly, transforming the photo into a classic caricature that balanced all requested thematic elements (hockey, news, dogs) while keeping the subject recognizable. Z-Image Turbo largely ignored the edit instructions, failing to change the style to a caricature or include the anchor and hockey themes, resulting in a nearly identical photo to the source.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent dynamic motion with animals appearing to run and tumble.
- + Superior lighting effects with clear god rays and a warm golden hour glow.
- + Better integration of the animals into the environment with accurate shadows and fur interactions.
- − The fox's front paws look a bit structurally odd or blurry.
- − The butterflies are somewhat simplified in detail compared to the foreground subjects.
Z-Image Turbo
- + High clarity on the textures of the animals' fur.
- + Successfully includes all requested subjects in a clear, centered group.
- − The composition feels more static and staged than 'playfully chasing'.
- − The lighting is flatter with less pronounced god rays compared to Image A.
- − The kitten's facial structure appears a bit distorted and unnatural.
Verdict: GPT Image 1 captures the 'playful chasing' aspect of the prompt much more effectively through dynamic posing and a sense of movement. While Z-Image Turbo has good texture, it feels like a static photoshoot, whereas GPT Image 1 feels like a cohesive, lived-in scene with superior atmospheric lighting and better composition.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'Studio Ghibli' style with hand-painted watercolor textures.
- + Perfectly captures the character expressions and layout of the original meme in an illustrative format.
- + Subtle use of soft pastel colors and dreamy lighting as requested.
- − The plaid pattern on the shirt is simplified to horizontal lines, losing part of the original detail.
Z-Image Turbo
- + High preservation of the original subjects' realistic faces and the specific plaid pattern of the shirt.
- − Completely failed the stylistic edit; the image remains a photograph rather than a Ghibli-inspired illustration.
- − The colors are slightly muted but do not meet the 'soft pastel' or 'hand-painted' criteria.
- − The woman on the right has her expression changed from angry to neutral, losing the context of the meme.
Verdict: GPT Image 1 successfully transformed the photo into a beautiful Ghibli-inspired illustration while maintaining the instantly recognizable composition and character dynamics of the memorialized 'distracted boyfriend' meme. Z-Image Turbo almost entirely ignored the core stylistic instruction, producing a photorealistic image that looks like a slightly desaturated version of the original.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the hair motion request with strands visibly wind-swept.
- + Abundant flying leaves create a strong sense of dynamic motion throughout the frame.
- + Strong preservation of the original subjects' faces and general appearance.
- − The dog's leash has been shortened and the handle has become a messy loop disconnected from the hand.
- − The overall lighting is slightly harsher and more saturated than the source image.
Z-Image Turbo
- + Preserves the original structure of the image, including the dog's leash and handle, much better than Model A.
- + Maintains a natural, soft lighting that matches the source image closely.
- − The hair motion is very subtle, failing to fully capture the 'blowing in the wind' instruction.
- − Fewer flying leaves result in a less 'energetic and lively' feel compared to the other model.
- − Significant alterations to the face of the woman compared to the original source.
Verdict: GPT Image 1 succeeded best at the core task of adding dynamic motion, particularly with the hair and the quantity of leaves, though it struggled with the logical consistency of the dog's leash. Z-Image Turbo preserved the general layout better but failed to deliver the level of energy requested and noticeably changed the woman's facial features. GPT Image 1 is the winner for its superior interpretation of the 'energetic and lively' prompt instruction.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with correct inclusion of the accent mark
- + Strong vector logo composition with a centered banner
- + Good texture within the brown elements
- − Prompt requested a light background; it provided a black one
- − The banner shape is slightly irregular at the ends
Z-Image Turbo
- + Adhered perfectly to the request for a light background with subtle texture
- + Clean, professional typography and iconography
- + High icon clarity and symmetry
- − Missing the 'Est. 1720' banner, opting for horizontal lines instead
- − The 'è' accent mark is stylized as a leaf/drop, which may hinder legibility
Verdict: GPT Image 1 followed the banner and typographical details more closely but failed to use a light background as requested. Z-Image Turbo captured the overall aesthetic and color palette better, and while it missed the requested banner, it succeeded in creating a professional minimalist logo that fits the 'Vintage' description perfectly.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the requested NASA-inspired color palette and flat-vector style.
- + Contains all the requested textual elements for mission steps, even with minor spelling errors.
- + Includes creative and relevant supporting details like the crew names and sillhouettes.
- − Has several typos including 'EARLLUNAR' and inconsistent placement of labels.
- − The layout is a bit cluttered and the flow of the steps is confusing to follow.
Z-Image Turbo
- + Clean, modern layout with plenty of white space that fits the minimalist infographic aesthetic.
- + High-quality vector icons, particularly for the Earth and Moon.
- + Very crisp lines and clear typography.
- − Significant spelling errors in every headline including 'APOLIO E 11', 'Translurian', and 'Descenty'.
- − Fails to include all requested steps and icons, missing the specific 'Lunar Orbit' and 'Landing' surface icons.
Verdict: GPT Image 1 followed the complex prompt instructions much more closely, including the specific crew names and all phases of the mission, whereas Z-Image Turbo simplified the content significantly and missed several required steps. Although Z-Image Turbo has a cleaner aesthetic, its frequent and glaring spelling errors ('Apolio', 'Descenty') and lack of detail make GPT Image 1 the superior choice for an infographic.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering