OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#4 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent photographic realism and soft, natural lighting.
- + Correct placement and scale of all objects relative to each other.
- + Highly detailed textures on the book cover and the wooden table.
- − The plant is slightly more behind the cube than visible through it compared to Model B.
Vidu Q2
- + Strong adherence to the spatial requirements of the prompt.
- + Interesting play of light and shadow on the table surface.
- + Captures the refraction of the plant through the glass well.
- − The glass cube has illogical reflections, such as a second blue sphere appearing on its base.
- − The book has some minor distortion on the spine text and edges.
Verdict: Both models followed the prompt instructions perfectly in terms of object placement. GPT Image 2 is the superior image due to its higher level of photorealism and clean rendering of the glass, whereas Vidu Q2 contains a phantom reflection of the blue sphere that breaks the physical logic of the scene.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent handling of motion blur on passing cars as requested.
- + Authentic street photography composition with 'imperfect framing' that feels natural.
- + Superior skin textures and realistic environmental details like the toolbox and Japanese signage.
- − The light rain is very subtle and barely visible.
- − Minor logical issue with the man's seat appearing to hover over a paint bucket.
Vidu Q2
- + Stronger emphasis on wet pavement reflections.
- + Dynamic close-up framing that highlights the textures of the bike.
- − Total failure on motion blur for the car in the background.
- − Significant anatomical and structural issues with the bike (the chain and frame are nonsensical).
- − Over-sharpened, high-contrast look that contradicts the 'no stylization' instruction.
Verdict: GPT Image 2 is the clear winner as it successfully incorporates almost every technical prompt requirement, including the difficult motion blur and naturalistic candid framing. Vidu Q2 fails on the motion blur requirement and produces a highly stylized, AI-looking image with significant structural errors in the bicycle's anatomy.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Exceptional photorealistic skin texture and lifelike eyes.
- + Masterful handling of lighting, with soft torchlight glints and realistic metal weathering.
- + Superior hair detailing, including the small beads and intricate braiding requested.
- − The 'battle-worn' aspect is slightly more subtle than expected, appearing more like surface dirt than heavy scarring.
Vidu Q2
- + Strong adherence to the armor engraving and leather strap details.
- + Effective use of bokeh sparks in the background to create atmosphere.
- + Clearly visible facial scarring that aligns with the 'battle-worn' prompt.
- − The skin and hair have a slightly plastic, CGI-like finish compared to Model A.
- − Lighting on the face feels a bit flat and less integrated with the background environment.
Verdict: GPT Image 2 is the superior image due to its incredible photorealism and sophisticated lighting, which perfectly captures the requested 'warm torchlight' and 'lifelike eyes'. While Vidu Q2 does an excellent job with the mechanical details of the armor and leather, its overall texture quality feels more like a video game render than a realistic photograph, making GPT Image 2 the more convincing and visually appealing portrait.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Perfect text rendering with coherent menu items and descriptions
- + Professional, realistic food photography that matches the labels
- + Excellent layout with clear sections for Appetizers, Pizza, and Mains as requested
- − Slightly busy header with the logo and side text, though typical for the industry
Vidu Q2
- + Clean pastel color palette with vibrant accents
- + Solid attempt at a grid-based layout
- − Text is completely illegible gibberish
- − Food images are messy and don't match the pizza category
- − Failed to properly delineate the requested sections logically
Verdict: GPT Image 2 is a production-ready menu design with perfect typography, logical layout, and high-quality food photography that strictly adheres to the prompt. Vidu Q2 fails significantly in visual quality, producing illegible text and confusing food depictions that do not follow the categories provided in the prompt.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture on the meat patty, bun, and vegetables.
- + Perfect execution of the 'fiery, glowing effect' for all requested text.
- + Dynamic composition with realistic sauce splashes and ember particles.
- − The '6' in the price has an unusual, slightly stylized shape compared to standard fonts.
Vidu Q2
- + Clean layout with the title centered at the top.
- + Good rendering of fresh lettuce and tomato slices.
- + Includes all requested text elements in the specified hierarchy.
- − The currency symbol is incorrect, showing a blurred hybrid mark instead of the Euro sign.
- − The fiery background feels more like a flat texture than a dynamic environment.
- − Lower level of detail in the food textures compared to the competitor.
Verdict: GPT Image 2 is the clear winner as it perfectly captures the high-end commercial aesthetic with superior textures and lighting. It followed the stylistic instructions for 'fiery, glowing' text much more effectively than Vidu Q2, which struggled specifically with the Euro currency symbol and overall image depth.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent text accuracy, spelling every word correctly including the complex third item.
- + Perfectly captures the authentic texture of dry chalk on a blackboard.
- + Maintains consistent handwriting style across all lines while following layout instructions.
- − The cursive for the title is relatively simple rather than highly 'elegant' cursive.
Vidu Q2
- + Dynamic chalk texture with varied thickness and pressure.
- + Bold, legible title design with good contrast.
- − Significant spelling errors throughout the menu items such as 'Truffe Musshoom' and 'Octopd'.
- − The text for the third item becomes gibberish ('Browd Botter Chotoouttey').
- − Mismatched price for the first item compared to the prompt ($34 vs $24).
Verdict: GPT Image 2 followed the prompt perfectly, rendering all text with 100% accuracy and maintaining a very realistic chalk texture and handwriting style. Vidu Q2 struggled significantly with text rendering, resulting in numerous spelling errors and garbled words in the bottom half of the menu.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the complex crossed-leg pose from Image 1.
- + High character consistency, successfully transferring the face, hair, sunglasses, and scarf from Image 2.
- + Matches the atmospheric lighting and soft shadows of the original scene perfectly.
- − The right hand (upper) has an extra finger and some anatomical blurring at the fingertips.
- − The left foot on the box has slightly distorted toe anatomy.
Vidu Q2
- + Great facial similarity to the character in Image 2.
- + Captures the scarf and clothing style effectively while maintaining the sharp lighting of the yellow background.
- + The high-resolution rendering makes the textures of the clothes and skin very clear.
- − The left hand (lower) has significant anatomical issues, appearing with six fingers and an unnatural palm shape.
- − The pose is slightly less accurate to Image 1's leg positioning, losing the 'tightness' of the wrap.
- − The fingers on the right hand (upper) are strangely elongated and distorted.
Verdict: GPT Image 2 is the clear winner as it more accurately replicates the specific, difficult pose from Image 1 while maintaining very high character fidelity from Image 2. While both models struggled with hand anatomy, Vidu Q2 produced more glaring structural errors in the lower hand and failed to match the exact leg wrap seen in the source pose.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the specific prompt instruction to have the horse on top
- + High realism and detail in the texture of the space suit and horse fur
- + Convincing cinematic lighting and lunar environment background
- − The horse's front legs are morphing into the harness/reins in a confusing way
Vidu Q2
- + Beautiful vibrant colors and cosmic aesthetic
- + Dynamic and graceful horse pose
- + High visual appeal in the nebula and galaxy background
- − Completely failed the negative constraint to have the horse on top
- − Followed the traditional 'astronaut on horse' trope instead of the surreal instruction
Verdict: GPT Image 2 followed the complex spatial instructions of the prompt perfectly, creating a surreal and high-quality image of a horse riding an astronaut. Vidu Q2 ignored the specific instruction about the positioning, delivering a standard astronaut-on-horse image which, while visually pretty, failed the primary challenge of the prompt.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 2
- + Excellently preserves the base person's face, hair, and vitiligo patterns without distortion.
- + Accurately replicates the scarf pattern and coat style from Image 2.
- + Maintains high source preservation for the background and facial expression.
- − The lighting on the person's face feels slightly flattened compared to the original image.
- − Missed the sunglasses and the specific gold ring accessory from Image 2.
Vidu Q2
- + Successfully includes the sunglasses and gold watch/ring accessories from Image 2.
- + Creates a high-contrast, sharp image that matches the fashion aesthetic.
- − Failed to keep the person's face unchanged, significantly altering the facial structure and adding a mustache.
- − Introduced severe artifacts where the vitiligo pattern is mistakenly applied on top of the coat's fabric.
- − Altered the person's hair and modified the background significantly.
Verdict: GPT Image 2 is the clear winner as it strictly followed the negative constraints to keep the person's face, hair, and background completely unchanged while successfully transferring the complex outfit. Vidu Q2 failed the primary editing constraint by changing the subject's face and adding a mustache, and it also suffered from logical errors where the subject's skin texture was rendered on the outside of the clothing.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with shallow depth of field and realistic lighting.
- + Very high detail on the capybara's fur and the apparel textures.
- + Captures the 'bored' expression of the passenger perfectly.
- − The composition is quite tight, making it a bit harder to see the full taxi interior.
Vidu Q2
- + Clearer view of the overall taxi interior and the street scene through the windows.
- + Accurate adherence to all major prompt elements, including the steering wheel placement.
- + Great use of vibrant city lights to establish the New York night atmosphere.
- − Lighting on the capybara's head is a bit flat and looks slightly composited.
- − Passenger's hands and phone interaction have minor anatomical warping.
Verdict: GPT Image 2 provides a more cinematic and photorealistic result with superior texture rendering on the capybara. While Vidu Q2 offers a wider composition that shows more of the taxi environment, GPT Image 2 is preferred for its higher image quality and more convincing lighting integration between the subjects and the background.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Perfect text rendering for all requested details, including date and location.
- + High-quality gothic aesthetic with intricate details like thorns, spiderwebs, and lanterns.
- + Creative integration of 'The Arches' and a NYC-style skyline into the background.
- − None notable; the image follows all prompt instructions perfectly.
Vidu Q2
- + Strong parchment effect and central jack-o-lantern glowing effect.
- + Includes requested bats and twisted tree branches.
- − Severe spelling errors in the title, banner text, and event details.
- − Incorrect date (2025 instead of 2026) and messy character rendering (30.70).
- − Lower overall visual quality with some muddy textures in the background.
Verdict: GPT Image 2 is the clear winner as it flawlessly executes the complex text requirements, including the specific date and location. Vidu Q2 struggles significantly with typography, producing numerous spelling errors and failing to accurately render the requested event details. GPT Image 2 also offers a much more sophisticated and atmospheric composition with its cinematic lighting and well-integrated background elements.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent PBR materials with realistic textures on the fish and rice.
- + Sophisticated 3D diorama base with detailed modeling and lighting.
- + Flawless text rendering and icon placement.
- − The scene is slightly more complex than 'minimal' with the added stone lantern and foliage.
Vidu Q2
- + Successfully captures the requested 'cartoon' aesthetic with soft, rounded shapes.
- + Adheres to the diorama and plate requirement.
- + Correct text content and placement.
- − The sushi details are slightly messy and less defined compared to A.
- − The flag icon is floating awkwardly on a pole from the letter 'N'.
- − Lower overall resolution and clarity in textures.
Verdict: GPT Image 2 is the superior output, providing professional-grade PBR materials and a much higher level of detail in the miniature diorama. While Vidu Q2 captures the 'cartoon' style well, it fails to match the high-clarity and realistic texture quality specifically requested in the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent fur texture and realistic lighting that catches individual hairs.
- + Dynamic and coherent composition with all animals interacting naturally.
- + Perfectly captures individual animal species features like the fennec-like ears of the fox and the golden retriever's expressive face.
- − The kitten's paw reaching up has a slightly clumped claw structure.
Vidu Q2
- + Includes a wide variety of colorful butterflies and a very lush flower field.
- + Bright, energetic color palette that fits the 'joyful' prompt requirement.
- − The fox looks more like a small dog or a fox-puppy hybrid rather than a distinct fox kit.
- − Duplicated a golden retriever puppy, which results in five animals instead of the four requested.
- − Anatomical issues with the rabbit's body shape and the tails of the puppies.
Verdict: GPT Image 2 followed all prompt instructions perfectly, including the specific count and type of animals, while maintaining superior photorealism and fur textures. Vidu Q2 struggled with the count by adding an extra dog and had softer, less detailed textures on the animals' faces and fur.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling of 'Caffè Florian' and 'Est. 1720'.
- + Superb vector emblem style with sophisticated line work and shading.
- + High-quality subtle paper texture and well-balanced composition.
- − The detail level is quite high for a 'minimalist' request, bordering on ornate.
Vidu Q2
- + Features the requested cloche and steam elements.
- + Uses the correct warm brown and cream color palette.
- − Failed significantly on text rendering with multiple misspellings like 'Farmiin' and 'Fopli20'.
- − The steam creates a strange, inconsistent visual inside the dome.
- − The composition is bottom-heavy and lacks the professional finish of a vector logo.
Verdict: GPT Image 2 is the clear winner as it perfectly executes the textual requirements and provides a professional, high-fidelity vector emblem with no errors. Vidu Q2 fails on basic text generation and produces a much lower quality illustration with amateurish design choices.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography and nearly perfect spelling throughout the graphic.
- + Strict adherence to the requested NASA-inspired color palette and layout.
- + High-quality, detailed icons that accurately reflect the specific mission stages requested.
- − The style leans more toward a detailed illustration than a strictly flat vector style.
- − Small graphical glitch on the landing module legs in step 6.
Vidu Q2
- + Successfully captures a clean 'flat vector' aesthetic.
- + Included the requested palette of light gray, blue, and red.
- − Extreme spelling errors making the text unreadable (e.g., 'ALFONCH', 'Laup 1').
- − Failed to follow the requested 6-step chronological sequence accurately.
- − Generic icons lack the specific Saturn V and mission-specific details requested.
Verdict: GPT Image 2 is significantly superior as it successfully follows every instruction, including the specific 6-stage sequence and accurate mission text. Vidu Q2 fails on basic prompt adherence regarding the steps and contains severe text gibberish, making it unusable for an infographic. GPT Image 2's composition and professional finish are much more aligned with the requested NASA-inspired theme.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing