Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 Mini OpenAI Vidu Q2 ShengShu Technology

Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.

GPT Image 1 Mini

25.0 arena score

#13 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1 Mini

0%

win rate

Ties

0%

Vidu Q2

0%

win rate

Shared challenges 20

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent soft lighting consistent with the 'window light' prompt.
  • + Clean, minimalist composition with realistic textures on the book and table.
  • + Accurate glass refraction and transparency.
  • The plant is more to the side than strictly 'behind' the cube as requested.
  • The glass cube edges are very thin, looking more like a frame than a solid glass object.

Vidu Q2

  • + Successfully positions the plant directly behind the cube so it is visible through the glass.
  • + Brilliant handling of light and shadows, including the reflection of the blue sphere.
  • + The glass cube has realistic thickness and weight.
  • The lighting is quite harsh and 'hard' rather than the 'soft window light' requested.
  • Slightly more cluttered composition compared to Model A.

Verdict: Both models followed the spatial prompts accurately, but Vidu Q2 captured the complexity of the scene better by placing the plant specifically behind the glass as requested and rendering realistic reflections for the sphere. While GPT Image 1 Mini handled the 'soft' lighting request better, its plant placement was less direct, making Vidu Q2 the slightly stronger choice for overall prompt adherence and visual depth.

Man and Car in California

Editing
Edit instruction

“Make a photo of the man driving the car down the California coastline”

Source
GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Successfully integrated both source images with high fidelity to the original subjects.
  • + Excellent preservation of the man's clothing, hairstyle, and facial features.
  • + Realistic lighting and shadow integration within the car's cockpit.
  • The scale of the man relative to the car is slightly too large, making the car look small.

Vidu Q2

  • + Dynamic composition and angle that showcases the full car and the coastline effectively.
  • + Good motion blur on the road adds a sense of speed.
  • Poor preservation of the man's specific identity and clothing from the source image.
  • Visible AI artifacts on the man's hands and the steering wheel area.
  • The man appears miniaturized within the car's cabin.

Verdict: GPT Image 1 Mini is the clear winner for its superior ability to preserve the identity and details of the man from the source image, including his unique coat and hairstyle. While Vidu Q2 creates a nice landscape, it fails to maintain the specific character requested and exhibits more technical flaws in the rendering of the human subject.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent adherence to the 'imperfect framing' and '50mm' look with a tight, cinematic crop.
  • + Highly realistic skin textures and natural-looking hands.
  • + Effective use of bokeh and muted colors to create a moody atmosphere.

Vidu Q2

  • + Features more prominent characterful reflections on the wet pavement.
  • + Captures the aged texture of the bicycle frame well.
  • + Good use of motion blur in the background car.
  • Anatomical errors in the hands and arms, with fingers appearing distorted.
  • The composition is cluttered and the subject's face is cut off awkwardly.
  • Over-saturated colors and harsh lighting reduce the requested cinematic realism.

Verdict: GPT Image 1 Mini produces a much more coherent and aesthetically pleasing image, successfully capturing the 'candid' and 'cinematic' feel requested. Vidu Q2 fails on anatomical details, particularly the hands, and its composition lacks the focus and professional quality of the first image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent depiction of battle-worn texture with realistic dirt and weathering
  • + Ornate engravings on the armor are intricate and historically plausible
  • + Superior lighting and color palette that feels cinematic and warm
  • Hair lacks the specific small beads requested in the prompt
  • Shallow depth of field is a bit too blurry in some foreground areas of the armor

Vidu Q2

  • + Successfully includes the requested beads in the hair braids
  • + Very sharp and detailed texture on the leather straps and underlayer fabric
  • + Good armor reflections with clear light sources
  • The character looks slightly too young and clean for a 'battle-worn' description
  • The armor design feels somewhat more like a digital painting than a lifelike photo
  • The composition is a bit more generic than the atmospheric Model A

Verdict: GPT Image 1 Mini captures the 'battle-worn' atmosphere much more effectively with realistic skin textures, grit, and cinematic lighting. While Vidu Q2 followed the specific instruction for beads in the hair better, it lacks the depth of character and the lifelike grit found in GPT Image 1 Mini's interpretation.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text legibility and font choice
  • + Clean, perfectly organized grid of high-quality food photography
  • + Followed the specific header requests exactly
  • Design is a bit too sparse, lacking descriptive text or prices
  • Empty space on the left feels unfinished

Vidu Q2

  • + Contains representative pricing and item descriptions
  • + Good use of vibrant accent colors and energetic layout
  • + Fits more information into a single view
  • Text is largely gibberish or misspelled
  • Layout is cluttered and lacks white space
  • The food images are repetitive and don't match the specific headers like pizza

Verdict: GPT Image 1 Mini produced a clean, professional, and readable layout that followed all category instructions, though it lacks simulated body text. Vidu Q2 included body text and pricing, but the text is illegible and the overall composition is too chaotic for a minimalist design.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with a consistent fiery glow effect.
  • + Clean and professional layout that feels like a polished advertisement.
  • + Accurate currency symbol and starburst design.
  • The burger feels somewhat static rather than dynamic or 'exploded'.
  • The background is less 'fiery' and more of a subtle ember texture.

Vidu Q2

  • + Strong sense of motion with sauce droplets and high-intensity fire.
  • + The 'exploded' burger effect is more dynamic with several layers and scattering ingredients.
  • + Vibrant colors and high visual energy.
  • The currency symbol is incorrect, showing a pound-like hybrid instead of a Euro symbol.
  • The text lacks the requested 'fiery, glowing' effect compared to Model A.
  • The 'LIMITED TIME ONLY' text is slightly cramped near the top.

Verdict: GPT Image 1 Mini produced a more professional-looking advertisement with superior typography and perfect adherence to the text requirements, including the correct currency symbol. While Vidu Q2 captured the 'exploded' and 'fiery' motion much better, it failed on the specific details of the text effects and the currency symbol. GPT Image 1 Mini is preferred for its overall polish and accuracy.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text legibility and spelling accuracy for all menu items.
  • + Consistent chalk texture across all characters.
  • + Clean and professional composition that fits the 'handwritten-style' prompt perfectly.
  • The 'elegant cursive' requested for the title is rendered as all-caps print instead.
  • The distribution of letters is perhaps too uniform, making it look slightly like a font rather than organic handwriting.

Vidu Q2

  • + Successfully attempted cursive elements for the title as requested.
  • + Realistic chalk smudge effects and varied stroke weights add to the 'cozy café' aesthetic.
  • Extreme spelling errors and garbled text across every menu item.
  • Incorrect price for the first item ($34 instead of $24).
  • Visual clutter and overlapping characters make the bottom section unreadable.

Verdict: GPT Image 1 Mini is the clear winner due to its superior text rendering and adherence to the specific menu content. while Vidu Q2 captured a more authentic cursive 'chalkboard' feel, the text is riddled with severe spelling errors and illegible characters that make it unusable for its intended purpose.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Successfully captures the facial features and accessories of the character in Image 2.
  • + Matches the yellow studio background and red ottoman from Image 1.
  • + Correctly adapts the clothing style to the character while keeping the color scheme.
  • Fails to replicate the specific crossed-leg pose from Image 1, settling for a generic crouching pose.
  • Anatomical issues with the foot on the ottoman which appears strangely merged or distorted.
  • The hair texture and style are slightly altered from the source character.

Vidu Q2

  • + Excellent adherence to the complex, specific body pose and crossed-leg position from Image 1.
  • + High degree of character preservation for the face, sunglasses, and the patterned scarf.
  • + Great natural integration of the character's clothing into the dynamic pose.
  • The hand on the left has an extra finger/anatomical distortion.
  • Minor rendering artifacts around the lower legs where they intersect.

Verdict: Vidu Q2 is the clear winner as it successfully recreated the precise, difficult pose from Image 1 while maintaining the character identity from Image 2. GPT Image 1 Mini failed the primary task of matching the pose, providing a much simpler crouching stance that ignored the leg positioning of the reference image.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent cinematic lighting and atmosphere
  • + Realistic textures on the spacesuit and horse fur
  • + Sophisticated and balanced composition
  • Failed the negative constraint; the astronaut is riding the horse instead of the horse on top

Vidu Q2

  • + Vibrant, colorful surreal aesthetic
  • + Intricate details in the galaxy patterns on the horse's body
  • Failed the negative constraint; the astronaut is riding the horse
  • The astronaut's hands and the reins are messy and poorly rendered

Verdict: Both models failed the negative constraint to place the 'horse on top' of the astronaut, defaulting to the standard 'astronaut on horse' trope. GPT Image 1 Mini is the better image overall due to its superior cinematic lighting, realistic textures, and anatomical consistency, whereas Vidu Q2 has significant artifacts in the hands and reins.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent integration of clothing textures and lighting.
  • + Preserves the overall composition and color palette of the beach background.
  • + Captures the essence of the vitiligo pattern across multiple limbs.
  • Failed to keep the person's exact face and hair, changing the facial structure significantly.
  • The scarf pattern is simplified compared to the source image.
  • Omitted the sunglasses from the source outfit.

Vidu Q2

  • + Successfully included the sunglasses and watch from the reference outfit.
  • + Preserved the specific vitiligo pattern on the forehead more accurately than Model A.
  • + Attempted to maintain the sand textures on the skin.
  • Fused the coat sleeve into the arm, creating a major anatomical artifact.
  • Significantly altered the facial features and added a mustache appearing to mix both source people.
  • The background beach detail is heavily modified and less realistic than the original.

Verdict: Both models failed to perfectly preserve the base person's face as requested, but GPT Image 1 Mini produced a much more coherent and high-quality image. Vidu Q2 struggled significantly with the 'clothing adaptation' part of the prompt, resulting in a disturbing artifact where the coat sleeve is merged with the subject's arm, whereas GPT Image 1 Mini handled the layers and fabric behavior realistically despite the face change.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent photorealistic texture on the capybara's fur
  • + Cinematic lighting that accurately reflects a night scene
  • + Captures the professional, calm expression perfectly
  • Only one paw is clearly visible on the steering wheel instead of both front paws
  • Overall composition is quite dark, obscuring some interior details

Vidu Q2

  • + Successfully shows both front paws on the steering wheel
  • + Brighter lighting allows for more visible Manhattan street details in the background
  • + Great composition showing more of the car's interior architecture
  • The capybara's head looks somewhat pasted onto the body
  • A slight 'uncanny' quality to the businesswoman's facial features

Verdict: GPT Image 1 Mini produces a much more photorealistic and moody image with superior texture work on the capybara, creating a more believable scene despite the absurd subject. While Vidu Q2 follows the anatomical instructions of the prompt more closely (showing both paws), its lighting and blending feel more like a digital composite than a real photograph. GPT Image 1 Mini is the winner for its superior artistic cohesion and realism.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with no spelling errors.
  • + Perfect atmospheric and cinematic lighting that matches the 'moody night sky' prompt.
  • + Strong layout that feels like a polished, professional invitation.
  • The 'parchment' texture is very subtle and looks more like a modern poster grain.
  • The twisted trees are somewhat obscured by the dark shadows.

Vidu Q2

  • + Captures the 'torn parchment' aesthetic much more literally and creatively.
  • + The border elements like webs and thorns are very high contrast and detailed.
  • Contains multiple significant spelling errors in all text fields.
  • Failed to include the correct year (2025 instead of 2026).
  • The composition feels a bit cluttered with chaotic elements.

Verdict: GPT Image 1 Mini is the clear winner because it successfully followed all text instructions with 100% accuracy, providing a professional and usable invitation. While Vidu Q2 had a creative interpretation of the parchment and border, the widespread spelling errors and failure to capture the correct date make it unusable for its intended purpose.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
GPT Image 1 Mini
Before After
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent full head of hair that perfectly matches the instruction.
  • + Natural-looking texture and volume.
  • + Maintains the overall aesthetic and color palette of the scene.
  • Significantly alters the person's facial features, making him look noticeably younger or like a different person.
  • The hairline and forehead integration looks slightly smoothed over compared to the rugged original skin texture.

Vidu Q2

  • + Superb preservation of the original facial features, wrinkles, and skin texture.
  • + Successfully adds hair while keeping the identity of the person perfectly intact.
  • + The hair texture and lighting match the source image flawlessly.
  • The hairline on the right side (near the forehead) has a minor digital blending artifact.
  • The volume of the hair is slightly less 'full' than Model A, though more realistic.

Verdict: While GPT Image 1 Mini creates a very lush head of hair, it fails the task of preserving the person's facial features by essentially generating a new face. Vidu Q2 is the clear winner as it masterfully adds realistic hair while keeping the subject's identity, skin texture, and background completely intact as requested.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with clean, bold characters.
  • + High-quality 3D materials that perfectly match the 'soft refined textures' and PBR request.
  • + Superior isometric composition and lighting.
  • The sushi models are slightly simplistic compared to the complex request.

Vidu Q2

  • + Successfully included a 3D flag icon as requested.
  • + Good variety of sushi types on the plate.
  • The text has minor kerning and alignment issues.
  • The lighting is a bit harsh, creating a more plastic look than 'soft refined textures'.
  • The diorama base has odd small artifacts/blobs at the corners.

Verdict: GPT Image 1 Mini followed the aesthetic instructions much more closely, delivering a professional-grade 3D render with clean typography and excellent texture work. Vidu Q2 succeeded in including the specific flag icon, but the overall image clarity and material quality were lower, with noticeable artifacts on the base.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Features a classic colored-pencil caricature art style with highly exaggerated facial features.
  • + Strong source preservation of the subject's denim outfit and general hair structure.
  • + Clearly incorporates all requested elements: news anchor desk, microphone, dog, hockey stick, and puck.
  • The transition between the arm and the hand holding the dog is a bit muddy.
  • The hockey stick placement felt slightly like an afterthought at the edge of the frame.

Vidu Q2

  • + Excellent integration of the hockey theme by placing the news desk inside a rink.
  • + High-quality vector art style with vibrant colors and clear, readable text.
  • + Includes multiple dogs and clever details like paw prints on the script paper.
  • Anatomical issues with the hands, including an extra finger on the right hand.
  • The microphone is being held in a slightly awkward, floating manner.

Verdict: Both models performed exceptionally well in translating the source image into a caricature while adhering to all thematic requirements. GPT Image 1 Mini captures the traditional 'exaggerated feature' essence of a caricature more effectively, whereas Vidu Q2 creates a more complex and visually interesting scene by merging the hockey rink and news desk environments. However, Vidu Q2 suffers from minor anatomical artifacts in the hands, making GPT Image 1 Mini the slightly more polished result.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent anatomical accuracy for all four animals.
  • + Beautiful lighting with soft god rays and realistic fur texture.
  • + Clean composition with a focused, cinematic feel.
  • Simple background with fewer flowers compared to the other model.

Vidu Q2

  • + Lush, busy environment with many flowers and butterflies.
  • + Captures a high energy, playful scene.
  • Anatomical errors including a dog with five legs and a kitten with a rabbit-like tail.
  • Lower level of realism with a more 'digital painting' look than photographic.
  • The fox kit has inconsistently colored limbs and a slightly distorted face.

Verdict: GPT Image 1 Mini is the clear winner as it produces high-fidelity, anatomically correct animals that perfectly match the prompt's request for hyper-photorealism. Vidu Q2 struggles with basic anatomy, resulting in extra limbs on the dog and hybrid features on the kitten, which breaks the immersion of an 8K masterpiece.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent adherence to the 'hand-painted textures' and 'warm mood' instructions
  • + High degree of source preservation with character poses and backgrounds
  • + Beautiful soft pastel color palette
  • The colored pencil texture is slightly more reminiscent of children's book illustration than cinematic Ghibli anime style

Vidu Q2

  • + Perfectly captures the Studio Ghibli cel-shaded character style
  • + Maintains clear lines while adding subtle watercolor washes to the background
  • + Excellent preservation of the original image's composition and details
  • The lighting is a bit flat compared to the 'gentle lighting' requested

Verdict: Both models successfully interpreted the prompt, but Vidu Q2 captured the specific 'Studio Ghibli' aesthetic more accurately with its clean line art and anime-style shading. GPT Image 1 Mini created a lovely, warm illustration, but its heavy crayon/pencil texture feels more like a storybook than a Ghibli film frame.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
GPT Image 1 Mini
Before After
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Successfully added wind-blown hair and flying leaves.
  • + Maintained the general layout and composition of the source image.
  • Significantly altered the facial features of the woman, losing her likeness.
  • The dog's face and fur texture became less sharp and slightly distorted compared to the original.

Vidu Q2

  • + Excellent source preservation, keeping the woman's face and the dog's appearance nearly identical to the original.
  • + Convincing hair motion that integrates well with the original hairstyle.
  • + Applied vibrant, colorful leaves that add to the energetic feel.
  • One or two leaves have slight transparency artifacts where they overlap the dog.

Verdict: Vidu Q2 is the clear winner as it successfully applied all parts of the edit prompt—hair motion and flying leaves—while perfectly preserving the identities of the woman and the dog. GPT Image 1 Mini captured the motion well but failed as an image editor by significantly changing the woman's facial features and degrading the quality of the dog.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with accurate spelling and accents.
  • + High-quality vector emblem style with clean lines.
  • + Perfect adherence to the banner and established date request.
  • Ignored the request for a light background, providing solid black instead.
  • Minimalist aesthetic is slightly busy due to the hatch-style shading.

Vidu Q2

  • + Followed instructions for a light background with subtle texture.
  • + Good color palette using the requested warm brown and cream tones.
  • + Elegant cloche dome illustration.
  • Severe spelling errors in the restaurant name.
  • Repetitive and redundant text with multiple versions of the name and date.
  • Poor text rendering on the bottom line with distorted characters.

Verdict: GPT Image 1 Mini produced a professional, usable logo with flawless typography and composition, though it failed to use a light background. Vidu Q2 followed the background and color instructions better but failed significantly on text accuracy and name consistency, making the logo unusable. GPT Image 1 Mini is the clear winner for its technical precision in logo design.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1 Mini
Vidu Q2

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with no spelling errors.
  • + Strict adherence to the requested 6-step logical flow.
  • + Clean, professional flat-vector aesthetic with a consistent NASA-inspired palette.
  • The 'Translunar' icon is a bit abstract and messy compared to the others.
  • The layout is slightly unbalanced with a large gap on the middle right.

Vidu Q2

  • + Visual style includes nice subtle gradients and depth.
  • + Includes consistent iconography across all depicted steps.
  • None of the text is legible or spelled correctly.
  • The numbering system is incoherent and does not follow the requested 1-6 sequence.
  • Displays five astronauts instead of the three associated with the Apollo 11 mission.

Verdict: GPT Image 1 Mini is the clear winner as it successfully rendered all text perfectly and followed the complex 6-step prompt instructions. Vidu Q2 failed significantly on text legibility and factual accuracy, producing garbled labels and an incorrect number of crew members.

Next steps

Explore each model