Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [max] Black Forest Labs Vidu Q2 ShengShu Technology

Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [max]

24.0 arena score

#21 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [max]

0%

win rate

Ties

0%

Vidu Q2

0%

win rate

Shared challenges 17

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent handling of caustics and light refraction through the glass.
  • + Sophisticated photographic realism with soft, natural depth of field.
  • + Accurate placement of all requested elements.
  • The blue sphere has a textured, glittery appearance rather than a smooth finish.
  • The plant in the background is quite dark and slightly obscured.

Vidu Q2

  • + Strong prompt adherence with vibrant, clear colors.
  • + The glass cube has realistic thickness and greenish edge tinting typical of glass.
  • + Excellent visibility of the green plant through the glass panels.
  • The book appears slightly small relative to the cube.
  • The lighting is a bit harsh, creating high-contrast shadows that feel less 'soft'.

Verdict: Both models followed the prompt perfectly, including the spatial relationships between the cube, sphere, book, and plant. FLUX.1 Kontext [max] produced a more professional, cinematic photograph with beautiful light play, whereas Vidu Q2 provided a clearer, more literal interpretation with better visibility of the plant through the glass.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent depiction of rain with realistic streaks and a hazy atmosphere.
  • + Strong cinematic composition with a centered subject and convincing bokeh.
  • + Naturally rendered skin textures and hair on the subject.
  • The bike chain and mechanics are physically impossible, merging into the wheel.
  • The rain streaks appear layered over the lens rather than interacting with the 3D space.

Vidu Q2

  • + Effectively captures a candid feel with 'imperfect' close-up framing.
  • + Better representation of motion blur from the passing car in the background.
  • + Detailed skin textures and realistic hand anatomy.
  • The main bicycle structure is nonsensical, having two front forks and two front wheels.
  • Does not show 'light rain' falling in the air, only wet surfaces.
  • The bicycle chain is disconnected and floating near the pedal.

Verdict: FLUX.1 Kontext [max] creates a much more atmospheric and cinematic image that better captures the mood of rain and street photography reflections. However, Vidu Q2 followed the 'imperfect framing' prompt more closely even though it suffered from severe anatomical errors in the bicycle's construction, resulting in a nonsensical two-front-wheeled bike.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent close-up skin texture and realistic depth of field
  • + Strong adherence to the 'battle-worn' descriptor through rugged features
  • + Very high detail in the metallic engravings
  • The eyes appear slightly glowing/supernatural rather than just 'lifelike'
  • Missing the beads in the braids requested by the prompt

Vidu Q2

  • + Perfect inclusion of beads in the braided hair
  • + Excellent portrayal of scars, dirt, and blood on the skin
  • + Complex and highly detailed leather straps and cloth underlayers
  • The composition is a medium shot rather than the 'close portrait' requested
  • The skin texture on the face looks slightly smoother and more digital compared to Model A

Verdict: FLUX.1 Kontext [max] provides a superior close-up portrait with incredible skin and metal textures, though it misses the specific detail about beads in the hair. Vidu Q2 follows more of the smaller details like the beads and the leather straps but fails to provide a true 'close portrait' composition, resulting in a less intimate image. FLUX.1 Kontext [max] is the preferred choice for its photographic quality and better interpretation of the character's 'battle-worn' essence.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent structure that favors a realistic menu layout
  • + Clean and professional header with a coherent logo
  • + High-quality, appetizing food photography
  • Repetitive food images (mostly pizza) despite prompt for variety
  • Smaller body text is largely illegible gibberish
  • The script font used for section headers is difficult to read

Vidu Q2

  • + Includes all requested categories like Mains and Appetizers clearly labeled
  • + Uses bold sans-serif fonts as requested
  • + Vibrant food photography with good color balance
  • Layout feels somewhat cluttered and chaotic
  • Text contains many typos and nonsensical word fragments
  • Borders and graphic accents are a bit inconsistent

Verdict: Both models struggle with creating legible, realistic text, but FLUX.1 Kontext [max] produces a much more professional and realistic menu layout that looks like a finished product. While Vidu Q2 followed the prompt's structural sections (Appetizers/Mains) more literal-mindedly, the overall aesthetic is cluttered compared to the clean, minimalist design of FLUX.1 Kontext [max].

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent typography rendering and spacing
  • + Highly realistic textures on the patty and bun
  • The burger is not truly 'exploded' as it remains largely assembled in the center
  • Missing the requested starburst element for the price

Vidu Q2

  • + Successfully creates the 'exploded' look with vertically separated layers and flying sauce
  • + Strong adherence to all prompt elements including the starburst and fiery background
  • + Dynamic energy and motion feel more pronounced
  • Minor text artifact in the price where the currency symbol is distorted
  • The lighting on the lettuce looks a bit flat compared to the fiery surroundings

Verdict: Vidu Q2 is the winner because it followed the specific layout instructions much better, featuring a vertically exploded burger and a starburst for the price. While FLUX.1 Kontext [max] had cleaner text and realistic textures, it failed to provide the 'exploded' composition requested, keeping the burger mostly stacked.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Perfect text accuracy for all requested menu items and dates
  • + Authentic board texture with realistic chalk smudges and dust
  • + Clean, legible handwriting that maintains a natural chalk appearance
  • The title is in print-style block letters rather than the requested elegant cursive

Vidu Q2

  • + The title uses a more cursive/fluid style as requested in the prompt
  • + Excellent chalk texture with realistic color variations and pressure depth
  • Severe spelling errors throughout the menu items such as 'Truffe Musshoom' and 'Lemepun'
  • The text becomes unintelligible toward the bottom of the board

Verdict: FLUX.1 Kontext is the clear winner due to its superior text rendering capabilities, producing an almost perfect menu with the exact items and prices requested. While Vidu Q2 captured a more stylistic 'cursive' title and nice chalk grain, it failed significantly on spelling and legibility.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent anatomical rendering of the horse and astronaut gear.
  • + Cinematic lighting with a realistic light source from a distant planet.
  • + High-fidelity textures on the horse's fur and the suit material.
  • The composition is a bit static and conventional.

Vidu Q2

  • + Strong surrealist interpretation with the nebula-skin horse.
  • + Dynamic and colorful composition that feels more imaginative.
  • + Excellent use of the background galaxy to frame the subject.
  • The astronaut's face is visible through the visor but lacks detail.
  • The horse's anatomy is slightly distorted in the rear legs.

Verdict: FLUX.1 Kontext [max] produces a highly realistic and grounded 'cinematic' image with superior technical detail in the suit and horse hair. Vidu Q2 leans more into the 'surreal' aspect of the prompt, creating a visually striking galaxy-horse that feels more artistic, though it sacrifices some anatomical accuracy.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent photorealistic fur texture and lighting on the capybara.
  • + Successfully includes 'TAXI' text on the driver's cap.
  • + High-quality blurred background bokeh that feels like a real city night.
  • The passenger is holding a phone to her ear like a call, rather than looking at it as requested.
  • The perspective chosen makes it difficult to see both paws clearly on the steering wheel.

Vidu Q2

  • + Perfectly captures the passenger looking at her phone with a bored expression.
  • + Clearly shows both paws on the steering wheel as requested.
  • + Strong composition that shows more of the taxi interior and the street outside.
  • The capybara's hands look slightly more humanoid/primate-like than actual capybara paws.
  • The lighting on the capybara's face is a bit flat compared to Model A.

Verdict: While FLUX.1 Kontext [max] offers superior textures and photorealistic lighting, Vidu Q2 followed the specific behavioral prompts much more accurately, showing the passenger looking down at her phone and the capybara with both paws on the wheel. Vidu Q2 also provided a more comprehensive view of the interior and the scene layout requested.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Perfect text accuracy for all requested fields
  • + Excellent atmospheric lighting and cohesive color palette
  • + Professional and clean composition formatted perfectly for an invitation
  • The date format uses commas instead of dots
  • Repeated the location text in the bottom section

Vidu Q2

  • + Stronger parchment texture and torn edge effect
  • + Higher contrast in the background elements like thorns and webbing
  • + Captures the vintage aesthetic effectively
  • Significant spelling errors in the title, banner, and date
  • The location and time details are merged and include gibberish text
  • Lighting on the pumpkin feels less integrated with the environment

Verdict: FLUX.1 Kontext [max] is the clear winner due to its superior text rendering and adherence to complex instructions. While Vidu Q2 captures a nice vintage parchment feel, it fails significantly on spelling and date accuracy, whereas FLUX.1 Kontext [max] produces a professional, usable invitation with atmospheric cinematic lighting.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
FLUX.1 Kontext [max]
Before After
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Successfully added a thick, full head of hair that matches the prompt.
  • + Maintains high resolution and sharp details throughout the image.
  • Significantly altered the man's facial features, making him look like a different person.
  • The hairline looks slightly artificial and lacking in fine transitional hairs.

Vidu Q2

  • + Excellent preservation of the original man's identity and facial features.
  • + Realistic hair texture and a very convincing, natural-looking hairline.
  • + Perfectly preserved the original lighting, clothing, and background.
  • Small stray hairs on the right side of the head appear slightly digital/stretched when viewed closely.

Verdict: FLUX.1 Kontext [max] succeeded in adding hair but failed the primary editing goal of preserving the source image's identity, resulting in a face that looks like a different person. Vidu Q2 followed the instructions perfectly, providing a realistic head of hair while keeping the original man's face and the environment identical to the source.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent typography with a clean, playful 3D font style.
  • + Superior material rendering for the nigiri and plate.
  • + Very clean composition that perfectly follows the diorama base request.
  • Failed to include the requested flag icon.
  • The chopsticks are merged into the board surface slightly awkwardly.

Vidu Q2

  • + Includes all requested elements including the flag icon.
  • + Clean and legible text rendering.
  • + Good variety of sushi types on the plate.
  • Lighting is a bit harsh and flat compared to Model A.
  • The textures look more like generic plastic and lack the 'soft refined' quality.
  • Visual artifacts present on the edges of the diorama base.

Verdict: FLUX.1 Kontext [max] produced a much more professional and aesthetically pleasing image with high-quality textures and soft lighting that feels truly 'miniature.' While Vidu Q2 followed more of the specific prompt details (like the flag icon), its overall visual quality is lower and the textures appear less refined.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Successfully incorporates all elements: TV anchor desk, a dog, and a hockey stick.
  • + Maintains the denim shirt and black top from the source image effectively.
  • + Strong caricature style with exaggerated facial features and clean line art.
  • The hockey stick is barely visible and cut off at the top edge.
  • The person's face looks significantly different from the source image, losing some identity recognition.
  • Randomly added glasses that weren't in the source image.

Vidu Q2

  • + Excellent preservation of the subject's facial features while still achieving a caricature effect.
  • + Superb integration of the hockey theme with a rink background, goal, and puck.
  • + High visual variety including two different dog breeds and paper with paw prints.
  • The hand holding the microphone has anatomical issues with finger placement.
  • The left hand (resting on a stack of papers) has an extra-long ring finger and simplified anatomy.

Verdict: Vidu Q2 is the clear winner as it better preserves the subject's likeness while creating a much more detailed and thematic scene. While FLUX.1 Kontext [max] creates a decent cartoon, it misses the 'hockey' aspect almost entirely and changes the person's face by adding glasses, whereas Vidu Q2 fully immerses the character in a hockey-themed TV studio with multiple dogs.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent fur texture rendering and soft lighting.
  • + Captures the 'big expressive eyes' and 'wholesome vibe' perfectly.
  • + Includes all four distinct animals requested.
  • The animals are sitting still rather than 'playfully chasing and tumbling'.
  • The butterflies appear somewhat flat and lack detail.

Vidu Q2

  • + Excellent adherence to the 'playfully chasing' and 'tumbling' part of the prompt.
  • + Superior butterfly detail and variety.
  • + Dynamic composition with a strong sense of movement.
  • Duplicated the golden retriever puppy, including two instead of one.
  • The rabbit has a somewhat unusual spotted pattern and tail that looks slightly hybrid.
  • Artifacts on the puppy's paws in the background.

Verdict: FLUX.1 Kontext [max] creates a more polished, high-quality portrait that excels in fur detail and lighting, though it fails to depict the requested action. Vidu Q2 much better captures the energy of animals playing and chasing butterflies, but it suffers from counting errors by doubling a puppy and exhibits more anatomical artifacts. FLUX.1 Kontext [max] is the preferred choice for its superior photorealism and clean execution of the subjects.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent adherence to the Ghibli art style with soft, watercolor-like textures
  • + Perfectly captures the 'warm, nostalgic mood' requested in the prompt
  • + Preserves the composition and poses of the original meme perfectly
  • The faces look a bit generic and lose some of the specific expressions from the source

Vidu Q2

  • + Successfully translates the image into an illustrative style
  • + Retains more of the original character features, especially in the man's face
  • + Good use of pastel colors and soft lighting
  • The line work is a bit too clean and lacks the hand-painted, textured feel of Ghibli backgrounds
  • The background characters are slightly more distorted than in Image A

Verdict: Both models did an excellent job of preserving the iconic composition of the 'Distracted Boyfriend' meme while applying the requested style. FLUX.1 Kontext [max] is the winner because its aesthetic much more closely aligns with the specific 'hand-painted' and 'dreamy' qualities of Studio Ghibli films, whereas Vidu Q2 feels like a more modern, digital anime illustration.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
FLUX.1 Kontext [max]
Before After
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent wind physics applied to the hair and the dog's tail
  • + Subtle but effective placement of small green leaves matching the local trees
  • + Changes to the pose, like the arm outreach, enhance the energetic feel
  • Anatomical error with the left hand being slightly mangled
  • Significant changes to the appearance of the face compared to the source image

Vidu Q2

  • + Strong adherence to the 'flying leaves' request with vibrant autumn colors
  • + Excellent preservation of the woman's facial features and the overall scene structure
  • + Good motion effect in the hair
  • The orange leaves feel a bit disconnected from the summer-green trees in the background
  • Less dynamic change to the character's posture compared to Model A

Verdict: FLUX.1 Kontext [max] creates the most convincing 'dynamic motion' effect by naturally blowing the hair and tail while slightly adjusting the pose for energy, though it fails to preserve the original face well. Vidu Q2 does a much better job at source preservation and follows the leaves instruction more visibly, though the orange leaves clash slightly with the environment. Vidu Q2 is the preferred model for maintaining facial consistency while delivering on the core prompt requirements.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent typography with perfect spelling of 'Caffè Florian'.
  • + Authentic vintage woodblock/vector style with appropriate texture.
  • + Clean, balanced composition that follows all prompt instructions.
  • The steam is very minimalist, bordering on abstract.

Vidu Q2

  • + Elegant warm cream and gold color palette.
  • + Good illustration of the cloche dome.
  • Severe spelling errors including 'Caffe Farmiin' and 'Fopli20'.
  • Redundant and cluttered text layout.
  • Illogical steam placement inside and outside the glass.

Verdict: FLUX.1 Kontext [max] perfectly executed the text and the 'vintage minimalist' aesthetic, delivering a professional-grade logo. In contrast, Vidu Q2 failed significantly on text rendering, resulting in nonsensical words and a cluttered composition that does not meet the requirements of a minimalist logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [max]
Vidu Q2

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent vector illustration style that feels professional and cohesive.
  • + Matches the requested navy, muted red, and light gray palette perfectly.
  • + Stronger execution of complex graphical elements like trajectory arcs and the lander.
  • Internal labels are confused (e.g., Earth is labeled 'Moon').
  • Includes Saturn as a background element which was not requested.
  • The 'Saturn V' looks more like a 1950s atomic rocket than the actual NASA vehicle.

Vidu Q2

  • + Follows the step-by-step layout much more logically than Model A.
  • + Includes better iconography for Earth and Lunar orbits.
  • + Maintains a cleaner, more minimalist infographic aesthetic.
  • Text rendering is mostly gibberish (e.g., 'ALFONCH' instead of APOLLO).
  • The 'Saturn V' icon is very generic and lacks detail.
  • Visual quality is slightly softer/lower resolution compared to Model A.

Verdict: FLUX.1 Kontext [max] produces a much more beautiful and artistic illustration with high-quality vector details, but it fails significantly on data accuracy by mislabeling the planets. Vidu Q2 follows the requested infographic structure and step-by-step logic much better, though its text generation is poor and the overall visual polish is lower. FLUX.1 Kontext [max] is the winner for its superior aesthetic and adherence to the NASA-inspired color palette and 'flat-vector' style.

Next steps

Explore each model