OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#5 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photographic realism and natural soft lighting
- + Accurately represents the transparency of the glass cube
- + Clean and simple composition that highlights all prompt elements
- − The sphere is slightly larger than 'small' would imply
- − The top of the cube appears to be open or missing under the book
Vidu Q2
- + High level of detail in the wooden table texture and plant leaves
- + Good handling of complex shadows and caustics
- + The sphere is a smaller, more 'small' size as requested
- − The book is hovering awkwardly above the cube with a noticeable gap
- − Perspective on the cube's base feels slightly distorted
Verdict: Both models followed the prompt instructions perfectly, including the spatial relationships between the objects. GPT Image 1.5 is preferred for its superior realism and natural lighting, whereas Vidu Q2 has a floating artifact where the book meets the cube.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1.5
- + Excellent preservation of the man's facial features and specific hairstyle
- + The lighting on the man and inside the car realistically matches the outdoor coastal environment
- + The cinematic perspective feels more like a passenger taking a photo.
- − The car is a right-hand drive model in California
- − The steering wheel and man's hand look slightly unnatural and low-detail
Vidu Q2
- + Successfully shows almost the entire car from the source image
- + High-quality rendering of the California coastline and winding road
- + Correct placement for a driver in a US-based setting (left-hand drive)
- − The man's facial details and distinctive hair are noticeably changed from the source
- − The car's proportions feel slightly stretched compared to the reference
- − Poor hand/steering wheel integration
Verdict: Both models handled the complex task of merging two distinct subjects into a new environment well. GPT Image 1.5 is preferred because it better preserved the specific identity and hairstyle of the man from the source image, whereas Vidu Q2 simplified his appearance significantly. GPT Image 1.5 also achieves better lighting integration, even though it incorrectly placed the steering wheel on the right side.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of raining conditions with visible droplets and wet textures.
- + Coherent and logical bicycle anatomy including the chain and cassette.
- + Authentic skin textures and believable 'candid' crouching posture.
- − Missing the requested motion blur on the passing car.
- − The composition feels a bit too centered and clean for 'imperfect framing'.
Vidu Q2
- + Effective shallow depth of field and beautiful ground reflections.
- + Strong adherence to the 'imperfect framing' prompt with a tight, cropped composition.
- − Significant anatomical errors including three hands and a disjointed bicycle frame.
- − Poorly rendered bicycle mechanics with the chain appearing as a disconnected wire.
- − Face lacks the same level of natural skin texture found in Model A.
Verdict: GPT Image 1.5 is the clear winner due to its superior anatomical coherence and realistic rendering of the scene. While Vidu Q2 captures the requested 'imperfect framing' and reflections well, it fails significantly on technical details, producing an image with three hands and a broken bicycle geometry.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of cinematic torchlight and bokeh sparks
- + Highly realistic skin texture with lifelike dirt and scarring
- + Exceptional detail on the engraved armor and weathered leather straps
- − The braiding with beads is a bit messy and lacks clear definition in some strands
Vidu Q2
- + Clear and distinct hair braids with visible beads
- + Strong, clean composition with balanced lighting
- + Good interpretation of the cloth underlayer texture
- − The skin and armor look slightly too clean for the 'battle-worn' description
- − The lighting feels more studio-like rather than atmospheric torchlight
- − Less detailed texture on the leather straps compared to Image A
Verdict: GPT Image 1.5 is the clear winner as it captures the 'battle-worn' and 'torchlight' elements of the prompt with significantly more realism and atmospheric depth. While Vidu Q2 produces a sharp and clean image, it lacks the gritty detail and high-fidelity textures found in GPT Image 1.5.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Perfect text rendering with zero spelling errors and logically categorized sections.
- + High-quality food photography that looks professional and appetizing.
- + Clean, modern layout that adheres strictly to the minimalist casual dining prompt.
- − Slightly standard or safe aesthetic that lacks a unique artistic flair.
Vidu Q2
- + Vibrant color palette and energetic layout that feels youthful.
- + Creative use of icons and graphic accents.
- − Unreadable gibberish text throughout the entire design.
- − Food images are cluttered and lack the clarity of professional food photography.
- − Disorganized layout that fails to clearly separate the requested categories.
Verdict: GPT Image 1.5 produced a highly usable, professional menu with flawless text and excellent photography that perfectly matches all prompt requirements. In contrast, Vidu Q2 failed significantly on text generation and organizational clarity, resulting in an unreadable design. GPT Image 1.5 is the clear winner for its functional and aesthetic superiority.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to all text requirements including the specific currency symbol.
- + Superior texture rendering on the patty and bun for a photorealistic food-photography look.
- + Highly dynamic composition with embers and flying debris that enhances the 'exploded' theme.
- − The composition is a bit crowded near the top text.
Vidu Q2
- + Clean layout with well-defined fire elements in the background.
- + Good use of the 'fiery, glowing effect' on the main title text.
- − Failed to render the Euro symbol correctly, using a stylized hash instead.
- − The food textures appear somewhat plastic or artificial compared to the other model.
- − The explosion of ingredients feels less 'dynamic' and more like a simple vertical hover.
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the difficult Euro symbol and the specific starburst price tag. While Vidu Q2 produced a clean image, it failed on the currency symbol and had much flatter lighting and textures on the food items.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling across all items.
- + Highly realistic chalk texture with dusty smudges and authentic line weights.
- + Followed the truncated prompt for the third item by logically completing it as 'Brown Butter Chocolate Chip Cookies'.
- − The composition is a bit tight at the top and bottom edges.
Vidu Q2
- + Dynamic chalk lettering with high contrast and artistic flair.
- + Good background bokeh that suggests a café setting.
- − Extremely poor spelling and legibility, with several words becoming gibberish.
- − Inconsistent prices and symbol rendering.
- − Messy composition with overlapping text and chaotic lines at the bottom.
Verdict: GPT Image 1.5 is the clear winner as it produced flawless, legible text that perfectly adhered to the prompt's request for a realistic handwritten style. Vidu Q2 struggled significantly with text generation, resulting in numerous spelling errors like 'Octopd' and 'Lemeun', and failed to maintain the professional aesthetic of a menu.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1.5
- + Excellent preservation of facial identity and accessories like the sunglasses and scarf.
- + Correctly matched the red stool and yellow background environment.
- + Smooth integration of the clothing details onto the new pose.
- − The character is wearing pants which contradicts the source pose containing bare legs.
- − The anatomy of the right arm and hand is slightly stiff compared to the reference.
Vidu Q2
- + Successfully captured the bare legs from the reference pose while maintaining the character's upper body.
- + Highly accurate recreation of the dynamic hand and finger positions from Image 1.
- + Excellent lighting and shadow work on the yellow background.
- − The facial features are slightly distorted and lose some of the specific identity of the man in Image 2.
- − Structural issues with the right hand showing an extra finger/atypical anatomy.
Verdict: Both models performed remarkably well on this complex task, but Vidu Q2 followed the prompt more precisely regarding the specific body position by including the bare legs shown in Image 1. While GPT Image 1.5 maintained the character's facial likeness better, Vidu Q2 achieved a more dynamic and anatomically accurate representation of the difficult pose despite some minor digital artifacts in the hands.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic lighting and texture on the horse's coat.
- + Detailed space environment with diverse elements like planets and a lunar lander.
- + High degree of realism in the mechanical parts of the spacesuit.
- − The prompt specifically asked for 'horse on top', which this model failed to interpret, placing the astronaut on top instead.
- − The composition feels a bit crowded with many competing background elements.
Vidu Q2
- + Successfully captured a more surreal aesthetic as requested in the prompt.
- + Vibrant color palette that blends the horse and the cosmic background artistically.
- + Cleaner focus on the central subject without distracting background props.
- − Failed the negative constraint/instruction 'horse on top, not vice versa'.
- − Anatomically awkward front legs on the horse.
Verdict: Both models completely failed the negative constraint instruction to place the 'horse on top' of the astronaut, instead providing the standard astronaut-riding-horse trope. GPT Image 1.5 is the preferred image because it delivers a much higher level of detail, better lighting, and more 'cinematic' textures compared to the slightly generic digital art style of Vidu Q2.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1.5
- + Excellent preservation of the original wooden structure on the left
- + Realistic fabric texture and drape of the coat and scarf
- + Successfully adapted the pose from Image 2 onto the character's lower half
- − Failed to preserve the head and face of the person from Image 1
- − Did not include the sunglasses from Image 2
- − The vitigilo patterns on the chin and hands do not match the source Image 1
Vidu Q2
- + Keeps the person's exact face, hair, and vitiligo patterns visible
- + Includes the sunglasses and jewelry from Image 2
- + Maintains the background beach environment accurately
- − The coat is rendered with strange sand-like stains on the sleeve that weren't requested
- − The pose is slightly stiff compared to the reference images
- − Minor distortion on the hand and watch detail
Verdict: GPT Image 1.5 failed the most basic instruction of keeping the person's face and hair unchanged, essentially replacing the subject with a generic character. Vidu Q2 succeeded in preserving the identity of the person from Image 1 while accurately transferring the elaborate outfit, accessories, and sunglasses from Image 2, despite adding some odd sand textures to the clothing.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with shallow depth of field typical of high-end photography
- + The capybara's paws are placed correctly on the steering wheel with realistic texture
- + The lighting inside the cabin feels organic and well-integrated
- − The composition is a bit tight, cutting off much of the vehicle interior
- − The passenger's face is slightly blurry
Vidu Q2
- + Dynamic composition showing more of the car's interior and the city environment
- + Perfect adherence to the bored expression requested for the passenger
- + Clear and sharp details throughout the entire frame
- − The capybara's paws look more like long human fingers, which is anatomically incorrect
- − The lighting feels slightly more artificial or 'rendered' compared to Image A
Verdict: Both models followed the prompt very well, but GPT Image 1.5 achieves a much higher level of photorealism that looks like a real film still. While Vidu Q2 offers a better wide-angle composition and better facial expression on the human passenger, the hand-like paws on the capybara are unsettling and break the realism established by GPT Image 1.5.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Perfect text rendering for all requested details, including date and location.
- + Atmospheric and cohesive vintage gothic aesthetic with a dark parchment feel.
- + Strong composition where every element like the thorns, webs, and moon feels integrated.
- − The dark, moody lighting makes the 'Location' text at the bottom slightly harder to read against the dark background.
Vidu Q2
- + Bright, clear parchment texture that provides high contrast for the central items.
- + Good variety in the thickness of the thorny border.
- − Extreme spelling errors in almost every text field, including 'Intovztion' and 'invieed'.
- − Hallucinated incorrect date information and failed to separate time from location.
- − Art style feels more like a digital illustration than a 'vintage' gothic poster.
Verdict: GPT Image 1.5 is the clear winner as it followed every text instruction perfectly, rendering the specific date, time, and location accurately. Vidu Q2 struggled significantly with typography, resulting in numerous spelling errors and incorrect dates. GPT Image 1.5 also captured the 'vintage gothic' mood much more effectively with its cinematic lighting and textured parchment.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1.5
- + Excellent texture matching the beard
- + Highly realistic hairline and integration with sideburns
- + Near-perfect preservation of facial structures and background details
- − Slightly changes the bridge of the glasses where it meets the nose
Vidu Q2
- + Successfully adds a full head of hair
- + Maintains original glasses and facial features well
- − The hair texture appears overly smooth and lacks the coarse realism of the existing beard
- − The hairline looks artificial and slightly floating
- − Slight distortions in the background lighting around the hair edges
Verdict: GPT Image 1.5 is the clear winner as it creates a believable, coarse hair texture that perfectly matches the subject's beard and skin tone. Vidu Q2 provides a decent volume of hair, but the texture is too silky and the transition at the hairline lacks the realism and skin integration seen in the and GPT Image 1.5 output.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering and layout.
- + High-detail PBR materials with realistic wood and ceramic textures.
- + Perfect adherence to the 45-degree isometric diorama request.
- − The scene is slightly less 'minimal' than requested due to the teapot and soy sauce bottle.
Vidu Q2
- + Matches the 'cartoon' aesthetic well with soft colors.
- + Clean, minimal composition.
- + Includes the flag icon as requested.
- − Text rendering is slightly less crisp than the competitor.
- − The 3D modeling of the sushi is somewhat mushy and lacks the clarity of high-end PBR materials.
- − The diorama base is very simple compared to the detailed scene in the other image.
Verdict: GPT Image 1.5 is the clear winner as it perfectly executes the PBR material request and high-detail isometric diorama style while maintaining crisp typography. Vidu Q2 follows the 'cartoon' instruction well, but the overall image quality and texture definition are significantly lower.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1.5
- + Expertly captures the news anchor aesthetic with studio lighting, cameras, and a ticker.
- + Strong caricature style with exaggerated features that still maintain the subject's likeness.
- + Very high visual quality and creative integration of the dog wearing a hockey helmet.
- − Changes the subject's outfit from the original denim to a red dress.
- − The hockey player in the background feels a bit disconnected from the main scene compared to the other elements.
Vidu Q2
- + Excellent source preservation, maintaining the denim jacket and black shirt from the original image.
- + Creative integration of hockey elements, such as the puck on the desk and the rink backdrop.
- + Includes clear 'News' text and a professional microphone.
- − The hands are poorly rendered with extra or distorted fingers.
- − The caricature of the face is less distinctive than model A, leaning more towards a generic cartoon style.
Verdict: GPT Image 1.5 is the stronger overall choice because it successfully creates a high-quality, professional-looking caricature with great humor, such as the dog in the hockey helmet. While Vidu Q2 does a better job of preserving the subject's original clothing, its technical execution is marred by significant anatomical errors in the hands and a less polished art style.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Excellent fur texture rendering and realistic dew sparkles on the grass.
- + Emotional resonance through expressive, high-detail eyes and adorable poses.
- + Superior lighting with beautiful god rays and a consistent golden hour glow.
- − The fox's paw in the bottom right corner is slightly malformed or merged into the grass.
Vidu Q2
- + Stronger sense of 'chasing' and 'tumbling' movement throughout the scene.
- + Effective use of vertical space with butterflies spread across the entire frame.
- − Includes two golden retriever puppies instead of the requested one.
- − Noticeable artifacts such as the kitten missing hind legs and the rabbit having an extra tail/leg.
- − Lower photorealistic quality with a more digital, over-saturated illustrative feel.
Verdict: GPT Image 1.5 is the clear winner as it provides high-fidelity, masterpiece-level rendering with incredibly soft fur textures and a cohesive lighting scheme. While Vidu Q2 captures the 'energy' of the prompt by showing the animals in motion, it suffers from several anatomical errors and fails to follow the count of the requested subjects accurately.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'soft pastel' and 'warm, nostalgic mood' requirements.
- + Successfully captures the dreamy, painterly background texture associated with Ghibli films.
- + Translates facial expressions into a charming anime style that fits the theme.
- − The excessive glow and high saturation pull it slightly toward a generic shoujo manga style rather than pure Ghibli realism.
- − The foreground character is quite blurry, which loses some structural detail.
Vidu Q2
- + Stronger preservation of the original source image's geometry and composition.
- + Clean line work and color palette that feels very close to Studio Ghibli's character cel animation style.
- + Preserves subtle details like the plaid pattern on the shirt accurately while stylizing them.
- − The background is a bit more generic and less 'dreamy' compared to Model A.
- − Lighting is relatively flat compared to the 'gentle lighting' requested in the prompt.
Verdict: Both models did an excellent job of stylizing the famous meme into an anime aesthetic. GPT Image 1.5 leaned more into the atmospheric and painterly qualities of the prompt, creating a warmly lit, nostalgic scene, while Vidu Q2 focused on clean lines and accurate preservation of the original image's forms. GPT Image 1.5 is the preferred winner as its color palette and lighting better encapsulate the specific 'dreamy' and 'warm' mood requested in the instruction.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of wind-blown hair with high hair-strand detail.
- + Adds a large quantity of leaves varied in size and motion blur.
- + Maintains original identity and facial features perfectly.
- − The leaf placement looks a bit cluttered across the subject's body.
- − Motion blur on the leaves is inconsistent.
Vidu Q2
- + Natural integration of leaves with a better sense of depth.
- + Very high fidelity preservation of the original source image colors and lighting.
- + Clean hair movement that feels realistic without being overly chaotic.
- − The hair on the subject's right side (viewer's left) appears slightly more static than the other side.
- − Fewer leaves compared to Model A, making the effect slightly more subtle.
Verdict: Both models did an excellent job of following instructions while preserving the underlying source image. GPT Image 1.5 provides a more 'energetic' feel with significant hair movement and many flying leaves, whereas Vidu Q2 offers a more refined, realistic edit with better spatial depth and cleaner image quality. Vidu Q2 is slightly preferred for its more professional-looking integration of the added elements.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling of 'Caffè Florian' and 'Est. 1720'.
- + Strong vector-style illustration with sophisticated shading and texture.
- + Correct application of the requested banner element and classic typography.
- − Failed the request for a 'light background' by using a solid black background.
- − The steam element is a bit thick and stylized compared to the rest of the realism in the cloche.
Vidu Q2
- + Followed the background instruction by using a light, textured cream background.
- + The color palette adheres well to the 'warm brown and cream' request.
- − Significant spelling errors in every piece of text including 'Farmiin' and 'Esttt'.
- − The composition is cluttered with redundant text lines.
- − The cloche handle and steam lines are poorly defined and lack professional finish.
Verdict: GPT Image 1.5 is the clear winner due to its professional execution and perfect text spelling, despite missing the instruction for a light background. Vidu Q2 followed the background and color instructions but failed significantly on text rendering and overall logo design coherence, producing multiple misspellings.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors.
- + Perfect adherence to all six requested steps and iconography.
- + Clean, professional flat-vector aesthetic with a cohesive color palette.
- − The rocket in the first panel is partially cut off at the top.
- − Composition is slightly crowded toward the bottom panels.
Vidu Q2
- + Successfully captured the requested color palette.
- + Good use of white space for a modern clean look.
- + Accurate icons for the Earth and Moon orbit steps.
- − Severe spelling and gibberish text throughout the entire image.
- − Inconsistent number of astronauts (five instead of three).
- − The 'Saturn V' icon is a generic cartoon rocket shuttle, not a heavy lift launch vehicle.
Verdict: GPT Image 1.5 produced a highly professional infographic with perfect spelling and exact adherence to the multi-step prompt. In contrast, Vidu Q2 failed significantly on text rendering and provided inaccurate icons and astronaut counts, resulting in an unusable image.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing