OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
100.0%
win rate
Ties
0.0%
Vidu Q2
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic clarity and soft lighting
- + Clean minimal composition
- + Accurate representation of the blue sphere inside the cube
- − The cube geometry is slightly warped at the corner joints
- − Reflection of the sphere on the base is missing or inconsistent
Vidu Q2
- + Detailed textures on the book and table
- + Accurate implementation of shadows and reflections within the glass cube
- + Rich, dynamic lighting effect from the window
- − The plant leaves are slightly blurred and look less realistic than the foreground
- − The plant's pot is hidden in a way that makes it look like it is floating
Verdict: Both models followed the prompt perfectly, including the placement of the sphere, book, and plant. Vidu Q2 is the winner because it handles the complex physics of glass much better, featuring realistic reflections and caustic shadows on the table, whereas GPT Image 1 feels slightly more digital and lacks a reflection for the internal sphere.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1
- + Excellent character preservation of the man's face and hairstyle.
- + Natural-looking lighting and motion blur that blends the car into the environment.
- + Maintains the specific proportions and model details of the source car.
- − The driver side is on the right, whereas both the source car and American roads typically use a left-hand drive configuration.
Vidu Q2
- + Successfully places the driver on the left side of the vehicle.
- + Produces a vibrant and clear landscape for the California coastline.
- − Significant loss of character likeness; the man's hair and facial features are generic.
- − Noticeable artifacts around the front grill and headlights of the car.
- − The perspective of the road lines and the car's motion feels slightly distorted.
Verdict: GPT Image 1 is the superior edit because it maintains a high degree of fidelity to both the source car and the specific man from the reference image, including his unique hairstyle. While it inadvertently creates a right-hand drive layout, Vidu Q2 fails to preserve the man's likeness and introduces several visual artifacts on the car's body.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and realistic face weathering
- + Beautifully rendered rain droplets on the man's hair and clothing
- + Strong atmospheric perspective with soft bokeh in the background
- − The rear wheel of the bicycle lacks a chain and has nonsensical gear structures
- − The car in the background is static with no motion blur as requested
Vidu Q2
- + Successfully captures 'imperfect framing' with a close-up, candid feel
- + Shows complex bicycle mechanics like the chain and derailleur
- + Reflections on the wet pavement are vibrant and detailed
- − Anatomical issues with three hands appearing to work on the bike
- − The background car lacks the requested motion blur
- − The subjects are partially cropped out in an awkward way
Verdict: GPT Image 1 produces a far more aesthetically pleasing and high-quality image with exceptional skin and atmospheric details, though it fails on bicycle mechanical accuracy. Vidu Q2 attempts a more complex composition and captures the 'imperfect framing' well, but is severely compromised by anatomical errors including an extra hand.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exceptional skin texture with realistic dirt and bloodstains
- + Highly ornate and weathered engraving on the plate armor
- + Atmospheric lighting that perfectly captures the warm torchlight and bokeh sparks
- − The beads in the hair are very subtle and blend in with the braids
- − The leather straps mentioned in the prompt are largely obscured
Vidu Q2
- + Excellent depiction of the leather straps and garment textures
- + Clearly visible beads and ornaments in the braided hair
- + Sharp focus on the face and scars
- − The skin and hair look slightly too clean for a battle-worn character
- − The lighting feels more like studio lighting with a warm filter rather than natural torchlight
- − Armor textures lack the fine-etched detail seen in the competitor
Verdict: GPT Image 1 captures the 'battle-worn' aesthetic much better through gritty skin textures and realistic grime, creating a more convincing atmosphere. While Vidu Q2 excels at following the more specific object instructions like beads and leather straps, its execution feels more like a clean digital render compared to the cinematic realism of GPT Image 1.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with legible bold sans-serif fonts
- + High-quality, realistic food photography
- + Clean minimalist layout that looks professional
- − Repetitive placeholder text with typos like 'descrigion'
- − Limited layout complexity compared to a full menu
Vidu Q2
- + More comprehensive menu layout with multiple columns
- + Good use of color accents throughout the design
- + Captures the 'vibrant' and 'casual dining' atmosphere well
- − Garbled, unreadable text throughout the entire image
- − Messy, incoherent food images with many digital artifacts
- − Chaotic layout that lacks the requested 'minimalist' feel
Verdict: GPT Image 1 is the clear winner as it provides a professional, clean, and usable design with high-quality food photography and readable (though repetitive) text. Vidu Q2 fails on basic visual quality, producing garbled text and messy, incoherent food images that do not meet the professional standard requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with a consistent glowing, fiery texture.
- + Photorealistic food textures, especially the sear on the patty and the freshness of the lettuce.
- + Clean composition with balanced space for all required text elements.
- − Incorrect price rendering, showing '€.99' instead of '€6.99'.
- − The starburst shape is a bit jagged and lacks the 'fiery' rendering of the other text.
Vidu Q2
- + Dynamic and exciting background with intense fire effects and motion.
- + Accurate numerical price rendering within a comic-style starburst.
- + Good integration of the 'LIMITED TIME ONLY' text with the main title.
- − Poor character rendering; the Euro symbol and several letters in 'LIMITED' are mangled.
- − The bottom bun texture appears blurry and low-resolution compared to the top half.
- − The lettuce and tomato layers look somewhat artificial and flat.
Verdict: GPT Image 1 offers superior photorealistic detail in the burger itself and displays high-quality text rendering, though it failed to correctly include the '6' in the price. Vidu Q2 captures more 'motion' and energy with its fiery background, but suffers from significant text artifacts and lower overall clarity. GPT Image 1 is the preferred choice for a professional advertisement due to its cleaner layout and superior image quality.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and spelling accuracy
- + Consistent chalk texture across all letters
- + Followed the specific menu items and prices almost perfectly
- − Handwriting feels slightly 'font-like' and uniform compared to real chalk art
- − Missed the 'elegant cursive' requirement for the title
Vidu Q2
- + Features a very authentic, artistic 'hand-drawn' chalk style with natural smudges
- + Captures the cozy café background environment
- + Good use of varying chalk weights for emphasis
- − Significant spelling errors throughout the menu items
- − The handwriting is messy to the point of being difficult to read
- − Failed to maintain the requested price for the first item
Verdict: GPT Image 1 followed the prompt's content requirements much more effectively, producing legible text and accurate spelling for the complex menu items. While Vidu Q2 captured a more realistic 'chalk art' aesthetic with better environmental context, its inability to spell words correctly like 'Mushroom' or 'Octopus' makes it less useful as a menu.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1
- + Successfully replicates the complex cross-legged pose on the red box
- + Captures the lighting and yellow background of source image 1 well
- − Face and features bear little resemblance to the character reference in Image 2
- − The hands and lower arm anatomy are distorted and lack detail
- − The scarf texture is flat and lacks the realistic draping of the original
Vidu Q2
- + Excellent character preservation with high facial similarity to Image 2
- + Accurately replicates the scarf pattern and clothing details from the character reference
- + Precise adherence to the dynamic pose/anatomy while maintaining higher image resolution
- − None significant; captures all aspects of the multi-image instruction effectively
Verdict: Vidu Q2 is the clear winner as it successfully blends the character identity from Image 2 with the difficult pose of Image 1. GPT Image 1 fails on character likeness and has significant anatomical issues with the hands, whereas Vidu Q2 maintains high-fidelity details on the face, accessories, and overall composition.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and realistic textures on the horse's coat and spacesuit.
- + Clean composition with a consistent mood.
- + High level of anatomical detail in the horse's musculature.
- − Failed the specific spatial instruction 'horse on top, not vice versa'.
- − Interpretation is conventional rather than truly surreal.
Vidu Q2
- + Vibrant, imaginative use of nebulae themes within the horse's skin.
- + Captured the 'surreal' aspect of the prompt more effectively through color and lighting.
- + Dynamic composition with a sense of motion.
- − Failed the specific spatial instruction 'horse on top, not vice versa'.
- − Some minor artifacts where the astronaut's hands interact with the reins.
- − Colors are somewhat oversaturated, leaning more toward digital art than cinematic realism.
Verdict: Both models failed the specific instruction to place the horse on top of the astronaut, instead opting for the standard 'astronaut on horse' trope. GPT Image 1 is technically superior in terms of realistic lighting and detail, while Vidu Q2 offers a more creative and colorful interpretation of the surreal space theme. GPT Image 1 is preferred because its execution of the subject matter is more polished and fits the 'cinematic' requirement better.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the person's facial features and vitiligo patterns
- + High-quality rendering of the clothing and scarf textures
- + Maintains the lighting and atmospheric feel of the original beach scene
- − The scarf pattern, while similar, is not an exact match to the source
- − The watch added is a different style than the gold one in Image 2
Vidu Q2
- + Successfully includes the sunglasses and gold watch from Image 2
- + Adds the scarf and coat with reasonable accuracy
- − Significantly alters the person's face and hair, failing the 'completely unchanged' constraint
- − Major anatomical errors including a missing right arm and sand texture floating over clothing
- − Poor lighting integration making the figure look like a cutout
Verdict: GPT Image 1 is the clear winner as it successfully preserves the person's identity and the background while realistically applying the new clothing. Vidu Q2 fails on multiple levels, including severe anatomical distortions (missing arm), failing to preserve the subject's face, and poor overall image coherence.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the capybara's fur
- + Cinematic lighting and composition that creates a moody atmosphere
- + Captures the bored, non-plussed expression of the passenger perfectly
- − The passenger appears to be in the front passenger seat rather than the back seat
- − The capybara's paws look more like primate hands than capybara paws
Vidu Q2
- + Accurately places the passenger in the back seat as requested
- + Bright, clear details of the taxi interior and city lights
- + The capybara's expression is very professional and alert
- − The capybara appears to be floating or disconnected from the headrest
- − The passenger's phone and fingers have some minor distortion
- − The lighting feels a bit more artificial compared to Model A
Verdict: GPT Image 1 offers superior photorealism and a more cinematic quality, though it fails on the spatial instruction by placing the passenger in the front. Vidu Q2 follows the prompt's layout instructions more accurately by placing the woman in the back seat, but the overall image has a slightly more saturated, 'AI-processed' look with some compositing artifacts around the capybara's neck.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Perfect text rendering for all requested fields with no spelling errors.
- + Excellent composition that feels professional and balanced.
- + Atmospheric lighting and textures that evoke a true vintage gothic aesthetic.
- − The fonts, while elegant, are more serif-classic than distinctly 'gothic' compared to Model B.
- − The border is a bit subtle, blending into the dark background.
Vidu Q2
- + Captures the 'gothic' typography style beautifully for the main title.
- + Includes a clear parchment-style background as requested in the prompt.
- + Vibrant lighting on the Jack-o-lantern creates a strong focal point.
- − Extremely poor text accuracy with numerous spelling errors like 'Intovztion' and 'might of fiigts'.
- − Incorrect date (2025 instead of 2026) and 'Tmm' instead of 'Time'.
- − The composition feels a bit cluttered with the thorns overlapping the bright parchment.
Verdict: GPT Image 1 is the clear winner due to its flawless execution of text and professional layout, adhering to every detail of the prompt perfectly. Vidu Q2, while possessing a charming gothic font style, fails significantly in spelling and factual accuracy, rendering the invitation unusable.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Successfully adds a thick head of hair
- + Matches the dark color of the beard
- + Maintains the background and main jacket details correctly
- − The hair looks like a separate layer or a wig placed on top
- − Noticeable softening of facial details and skin texture compared to the source
- − The hairline is too sharp and lacks realistic transition
Vidu Q2
- + Excellent hair texture and realistic blending with the existing skin
- + Perfectly preserves facial identity and fine skin details
- + The hairline looks incredibly natural and matches the head shape well
- − None notable for this specific request
Verdict: Vidu Q2 is the clear winner as it seamlessly integrates the new hair with the existing subject, preserving the original skin texture and facial sharpness. GPT Image 1 adds the hair but significantly degrades the quality of the face, making it look blurry and poorly blended.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Perfectly clean and legible typography follow the specific prompt layout.
- + Very high-quality PBR materials with soft, clay-like textures.
- + Excellent composition with a professional, balanced 3D isometric look.
- − The flag icon is slightly simplified compared to the rest of the high-quality render.
Vidu Q2
- + Good variety of sushi types on the plate.
- + Adheres well to the requested 45-degree isometric angle.
- − The text rendering is inconsistent, especially around the edges of 'SUSHI'.
- − Visible artifacts and 'hallucinated' shapes around the base of the diorama.
- − The flag icon is positioned awkwardly on a pole growing out of the letter 'N'.
Verdict: GPT Image 1 is the clear winner as it provides a professional-grade 3D render with clean typography and excellent material work. Vidu Q2 fails to match the level of polish, showing significant artifacts on the diorama base and messy edges on the text.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Excellent caricature style with exaggerated facial features that remain recognizable.
- + Thoughtful integration of all prompt elements including a news desk, dogs, and hockey graphics.
- + Maintains the subject's clothing style and hair color from the source image.
- − The 'TV 13' screen in the background has a slightly distorted box shape.
Vidu Q2
- + Successfully merges a hockey rink with a news set background.
- + Captures the subject's likeliness well in a clean digital illustration style.
- + Includes creative details like paw prints on the news scripts.
- − Anatomical issues with the left hand (extra finger/distorted palm) and right hand (poor grip on microphone).
- − The caricature style is less 'exaggerated' and more like a standard comic book illustration.
Verdict: GPT Image 1 followed the instructions for an 'exaggerated caricature' much better than Vidu Q2, resulting in a funnier and more stylistically appropriate image. While Vidu Q2 had a clever background composition, it suffered from significant anatomical errors in the hands, whereas GPT Image 1 felt more polished and cohesive.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of god rays and warm golden sunrise lighting
- + Very high fur detail and realistic textures
- + Superior composition with a more focused and intimate 'tumbling' interaction
- − The fox kit has a slightly awkward paw/leg anatomy in the foreground
- − Fewer butterflies compared to the other model
Vidu Q2
- + Includes a large number of colorful butterflies and a dense wildflower meadow
- + Captures a very bright and joyful color palette
- + Good variety of animal poses across the frame
- − Included two golden retriever puppies instead of the requested one
- − Visual quality is less realistic with flatter lighting and less defined fur texture
- − Anatomical errors such as the kitten having three front legs
Verdict: GPT Image 1 is the superior choice as it provides a much more photorealistic result with beautiful volumetric lighting and high-quality textures. While Vidu Q2 includes more butterflies, it suffers from several anatomical glitches (like an extra cat leg) and fails to follow the count instruction by including two puppies.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Captures the specific 'Studio Ghibli' painterly aesthetic with soft textures and watercolor-like blending.
- + Maintains the warm, nostalgic mood and color palette requested in the prompt.
- + Provides a cohesive artistic transformation while keeping the recognizable pose and characters.
- − The characters' faces are significantly simplified and lose the likeness of the original subjects.
- − The plaid pattern on the man's shirt is largely lost in the painterly texture.
Vidu Q2
- + Preserves the original composition and character details with high accuracy.
- + Successfully translates the scene into a clean anime style with clear line work.
- + Excellent retention of the background elements and the specific patterns on the man's shirt.
- − The style feels more like a generic modern anime or 'manhwa' rather than the specific Ghibli aesthetic.
- − The lighting is flat and lacks the 'dreamy' quality requested in the prompt.
Verdict: GPT Image 1 offers a much better interpretation of the Studio Ghibli brand, accurately providing the soft textures, dreamy lighting, and warm mood requested. While Vidu Q2 is far better at preserving the original image's specific details and character likenesses, its final style is too 'crisp' and modern to be considered Ghibli-inspired.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Excellent source preservation, keeping the light and colors consistent with the original
- + Dynamic hair motion looks natural and integrated
- + Successfully adds subtle motion to the dog's fur and tail
- − The flying leaves are quite small and many look like blurry specs or artifacts
- − The hand holding the leash has become slightly distorted compared to the original
Vidu Q2
- + Strong visual impact with vibrant, colorful flying leaves
- + Hair motion is very clear and well-defined
- + Adds a nice depth of field effect to the leaves in the foreground
- − Leaves look a bit like a flat overlay rather than being part of the environment
- − Changed the color temperature of the entire image, making it warmer/more orange than the source
Verdict: Both models handled the edit well, but they took different approaches to the 'energetic feel'. GPT Image 1 focused on realism and motion within the existing subject (like the dog's tail), whereas Vidu Q2 focused on a more cinematic, high-contrast look with prominent foreground leaves. Vidu Q2 is slightly more successful because the 'flying leaves' instruction is more clearly executed and aesthetically pleasing.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Perfect text rendering of 'Caffè Florian' and 'Est. 1720'.
- + Excellent minimalist vector aesthetic with a clear silhouette.
- + Correct inclusion of all requested elements including the cloche and banner.
- − Failed the prompt instruction for a 'light background' by using solid black.
- − Lack of 'cream' tones mentioned in the color palette.
Vidu Q2
- + Followed the background instruction with a subtle texture on a light cream background.
- + Good use of the requested warm brown and cream color palette.
- − Major text hallucinations and typos (e.g., 'Farmiin', 'Esttt', 'Fopli20').
- − Steam is rendered inside the glass cloche, which looks visually confusing.
- − Cluttered composition that lacks the requested minimalism.
Verdict: GPT Image 1 is the clear winner due to its perfect text accuracy and clean minimalist design, despite failing the background color instruction. Vidu Q2 followed the background and color instructions better, but the significant spelling errors and cluttered design make it unusable as a logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with mostly correct spelling of mission stages and crew names.
- + Strict adherence to the NASA-inspired color palette and flat vector style.
- + Clear, high-contrast layout that functions well as an infographic.
- − One spelling error ('EARLLUNAR' vs Lunar Orbit).
- − The sequence flow is slightly cluttered with icons and text not aligning perfectly in a linear path.
Vidu Q2
- + Clean aesthetic with modern iconography.
- + Good use of spacing and a central vertical symmetry in the layout.
- − Severely garbled text and nonsensical words like 'ALFONCH' and 'Laup 1'.
- − Incorrect number of crew members (5 silhouettes instead of 3).
- − Fails to follow the specific 6-step prompt, mixing up labels and icons.
Verdict: GPT Image 1 is the clear winner as it successfully follows the functional requirements of an infographic, including legible and correct crew names and mission stages. Vidu Q2 produces a visually pleasing aesthetic but fails significantly on typography and factual accuracy regarding the Apollo 11 mission.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing