Black Forest Labs' 12-billion parameter flow transformer for high-quality text-to-image generation, suitable for personal and commercial use with streaming support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [dev]
#16 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [dev]
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent handling of glass transparency and internal reflections.
- + High aesthetic quality with soft, realistic lighting.
- + Accurate interpretation of 'partially visible through glass' for the plant.
- − The sphere appears to be floating rather than sitting on a surface.
- − The cube structure has some internal glass stems that weren't requested.
Vidu Q2
- + Perfect adherence to object placement with the sphere resting on the bottom.
- + Highly detailed wood grain and book texture.
- + Good use of shadows to ground the objects in the scene.
- − The glass cube appears slightly distorted with uneven edge thicknesses.
- − The lighting is harsher than the requested 'soft window light'.
Verdict: Both models followed the prompt perfectly in terms of object inclusion. FLUX.1 [dev] produced a more artistic, soft-lit image with superior glass rendering, while Vidu Q2 focused on physical realism and sharper textures, accurately placing the sphere on the bottom of the cube. Vidu Q2 is slightly preferred for its more grounded composition and realistic proportions of the objects.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent composition and atmospheric depth.
- + High visual quality with realistic lens effects like bokeh and light reflections.
- + Accurately represents the Japanese street setting and the requested lens characteristics.
- − The man is holding the handlebars rather than actively 'repairing' the bike.
- − Lacks the specific 'motion blur from passing cars' requested in the prompt.
Vidu Q2
- + Stronger adherence to the 'repairing' action with hands engaged on the bike chain.
- + Includes the 'imperfect framing' and 'candid' feel with a tight, off-center crop.
- + Great natural skin texture and weathered look on the bicycle.
- − Poor anatomical correctness with three hands visible.
- − Low background quality with severe artifacts on the car and blurred background elements.
- − The bicycle geometry is distorted and lacks mechanical coherence.
Verdict: FLUX.1 [dev] produces a much higher quality image with superior lighting and atmospheric effects, though it captures a 'standing with' rather than 'repairing' pose. Vidu Q2 follows the 'repairing' and 'imperfect framing' instructions better, but fails significantly on a technical level with major anatomical glitches (extra hands) and messy background artifacts. FLUX.1 [dev] is the preferred choice for its photographic realism and coherent subject matter.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [dev]
- + Exceptional photographic realism in skin texture and eyes
- + Strong bokeh effect with pleasing spark shapes
- + Soft, naturalistic transitions in lighting
- − Missed the request for beads in the braids
- − Skin looks too pristine and youthful for a 'battle-worn' character
- − Armor is relatively plain with minimal engraving visible
Vidu Q2
- + Strong adherence to the 'battle-worn' aesthetic with clear scars and dirt
- + Highly detailed engraving on the plate armor and visible leather strap textures
- + Successfully included small beads within the braided hair
- − The character has a slightly 'digital' or CGI appearance compared to the photographic look of A
- − The torchlight reflection is a bit harsh on the face
Verdict: While FLUX.1 [dev] produces a more lifelike photographic portrait, Vidu Q2 followed the specific details of the prompt much better, including the beads, scars, and ornate engravings. Vidu Q2's interpretation captures the 'battle-worn' theme and the material variety (leather, cloth, metal) with greater accuracy.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent typographic hierarchy with clean, professional font rendering
- + Realistic and appetizing food photography that integrates well with the layout
- + Balanced use of white space that reflects a modern minimalist aesthetic
- − Failed to provide the requested 'grid' of food photos
- − Missing a dedicated 'Pizza' section heading in the layout
Vidu Q2
- + Successfully incorporated a grid-like layout for the food photos
- + Included headings for all requested sections including Pizza
- + Vibrant and colorful color palette as requested
- − Poor typography with significant illegibility and garbled text
- − Cluttered composition that feels busy rather than minimalist
- − Low visual quality with artifacts in the food imagery
Verdict: FLUX.1 [dev] produced a much higher quality, professional menu design that looks ready for use, though it missed the specific grid layout and pizza header. Vidu Q2 followed the grid and section requirements more closely but failed significantly on visual quality, text rendering, and minimalist aesthetic. FLUX.1 [dev] is the preferred choice for a clean, professional result.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [dev]
- + High image clarity and photorealistic textures on the burger buns and meat.
- + Interesting vertical composition that feels clean and magical.
- + Good integration of fire at the base.
- − Completely failed to include the primary text 'MAGIC BURGER'.
- − Lacks the requested 'fiery, glowing effect' on the text and the starburst element.
- − The background is relatively empty compared to the 'fiery background' request.
Vidu Q2
- + Successfully included all requested text: 'MAGIC BURGER', 'LIMITED TIME ONLY', and the price.
- + Captured the 'fiery' and 'glowing' aesthetic perfectly for both the background and the text.
- + Excellent dynamic motion with sauce splashes and embers.
- − The currency symbol in the price is slightly warped/incorrect.
- − The composition is a bit crowded compared to the clean layout of Model A.
Verdict: Vidu Q2 is the clear winner as it followed every part of the complex prompt, including all text strings and the specific 'fiery' stylistic requirements. FLUX.1 [dev] failed to render the main title and missed several key design elements like the starburst, despite having higher resolution textures.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent text legibility and spelling throughout the entire menu.
- + Captures the request for a hand-written style that still feels clean and organized.
- + Completes the third menu item logically while maintaining the chalkboard aesthetic.
- − The 'chalk' texture looks a bit too clean and digital, lacking the dusty smudging requested.
- − The handwriting style is very uniform, leaning slightly toward a digital font appearance.
Vidu Q2
- + Features a very realistic chalk texture with dusty smears and layered strokes.
- + Strong adherence to the 'natural variations' and 'slant' aspects of the prompt.
- − Numerous spelling errors including 'Truffe Musshoom', 'Octopd', and 'Browd Botter'.
- − The prices and text layout become increasingly messy and difficult to read toward the bottom.
Verdict: FLUX.1 [dev] is the clear winner due to its superior text rendering and perfect spelling, making it a functional menu graphic. While Vidu Q2 captured the authentic, messy texture of a physical chalkboard much better, the significant typos and garbled text at the bottom make it unusable for professional purposes.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent anatomical rendering of the horse and astronaut suit.
- + Features a cinematic, minimalist composition with realistic lighting.
- − Failed the negative constraint to have the horse on top of the astronaut.
Vidu Q2
- + Vibrant, surreal color palette with interesting celestial textures on the horse's coat.
- + High level of detail in the nebulae and background elements.
- − Failed the negative constraint to have the horse on top of the astronaut.
- − The horse's legs and hooves have anatomical issues and strange blurring.
Verdict: Both FLUX.1 [dev] and Vidu Q2 completely failed the specific logic constraint of placing the 'horse on top' of the astronaut. While both models produced a standard astronaut riding a horse, FLUX.1 [dev] is the superior image due to its clean composition and much higher anatomical accuracy compared to the distorted limbs in Vidu Q2.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent texture on the capybara's fur and whiskers
- + High photorealism with realistic bokeh and cinematic lighting
- + Accurate rendering of the woman's hands and the smartphone
- − The woman appears to be sitting in the passenger seat rather than the back seat
- − The capybara's paws are somewhat small and lack definition on the wheel
Vidu Q2
- + Perfect composition with the woman clearly in the back seat as requested
- + Full-body view of the capybara showing the jacket and hands on the wheel
- + Dynamic city background that clearly conveys the New York setting
- − The capybara's hands look slightly more like primate or human hands than capybara paws
- − The lighting on the woman's face is a bit flat compared to the driver's side
Verdict: Both models followed the prompt well, but Vidu Q2 is the winner because it successfully placed the woman in the back seat, whereas FLUX.1 [dev] placed her in the front passenger seat. While FLUX.1 [dev] has slightly superior textural realism, Vidu Q2 captured the narrative and spatial requirements of the scene more accurately.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent graphic design and composition.
- + Very high clarity and professional cinematic lighting.
- + Clean and legible text rendering for the event details.
- − Major spelling error in the title text ('Falloween Rantcy').
- − Includes minor repetitive text artifacts ('7pm, 7pm').
Vidu Q2
- + Captures the 'vintage parchment' texture effectively.
- + Includes spider webs in the border as requested in the prompt.
- + Atmospheric and chaotic gothic aesthetic.
- − Significant spelling errors throughout all text fields.
- − The year in the date is incorrect (2025 instead of 2026).
- − Lower overall image resolution and clarity compared to the competitor.
Verdict: FLUX.1 [dev] produced a much more professional and visually striking invitation with superior lighting and layout, though it failed significantly on the main title text. Vidu Q2 followed the prompt's specific details like spider webs and parchment better, but suffered from poor image quality and severe typos in every line of text. FLUX.1 [dev] is the preferred choice for its polish and compositional strength.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent soft refined textures and miniature diorama feel.
- + Accurate isometric 45-degree perspective.
- + Pleasant, gentle lighting that creates a realistic PBR candy-like effect.
- − Failed the text prompt, displaying gibberish like 'SUSH CATON' and extra symbols.
- − The sushi variety is repetitive, using only one type of nigiri.
Vidu Q2
- + Perfect text rendering for 'JAPAN' and 'SUSHI'.
- + Great variety in sushi design within the miniature style.
- + Includes the requested flag icon clearly.
- − The lighting is a bit flat compared to the requested soft refined texture.
- − The isometric scale feels slightly less 'miniature' than Model A.
Verdict: Each model followed the prompt's aesthetic style closely, but Model B is the clear winner for its perfect execution of the textual requirements. While FLUX.1 [dev] produced a more sophisticated lighting environment and texture, its failure to correctly spell the words and the inclusion of nonsense characters makes it unusable for the specific request compared to Vidu Q2.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [dev]
- + Successfully includes all four animal types requested.
- + Consistent lighting and soft color palette.
- + Good rendering of soft fur textures.
- − Stylized, cartoonish appearance rather than the requested hyper-photorealistic style.
- − Anatomically awkward limbs on the standing bunny-like creatures.
- − Lacks the dynamic action of 'chasing and tumbling'.
Vidu Q2
- + Strong adherence to the hyper-photorealistic style with natural lighting.
- + Captures the action of chasing and movement much better than the competitor.
- + Excellent use of god rays, dew sparkles, and a lush meadow environment.
- − Includes two golden retriever puppies instead of one.
- − The tabby kitten and bunny are slightly merged in the foreground composition.
- − Some butterfly anatomy is simplified.
Verdict: Vidu Q2 is the clear winner for its superior adherence to the 'hyper-photorealistic' instruction and its ability to capture the dynamic energy of animals chasing butterflies. FLUX.1 [dev] produced a very cute, high-quality image, but it opted for a 3D-animation/Disney style that ignored the core photorealism requirement.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent vector logo aesthetic with clean lines.
- + Very sophisticated use of typography and banner element.
- + Correct colors and beautifully executed steam effect.
- − Spelling error in primary name ('Flarilaan' instead of 'Florian').
- − Spelling error in subtext ('Reseaurant').
- − Random numbers '11011' and '1941' added without prompt instruction.
Vidu Q2
- + Successfully captured the cloche dome as a central graphic element.
- + Includes the requested 'Est. 1720' text more prominently.
- + Good subtle paper-like texture on the background.
- − Severe spelling hallucinations ('Caffe Farmiin', 'Esttt', 'Caffce Fopli20').
- − Composition is cluttered with redundant text lines.
- − Less 'minimal' than requested with thicker, heavier lines.
Verdict: FLUX.1 [dev] produces a much more professional-looking vector emblem that fits the 'vintage minimalist' aesthetic perfectly, despite some spelling errors. Vidu Q2 struggles significantly with text rendering and provides a cluttered composition that fails to meet the minimalist requirement of the prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [dev]
- + Successfully captures a sophisticated NASA-inspired color palette and modern aesthetic.
- + Follows the request for circular arcs and consistent infographic styling at the bottom.
- − The main text content is mostly gibberish despite having the correct title.
- − Included nonsensical extra planets like Saturn which were not requested and break the theme.
Vidu Q2
- + Includes higher quality, more detailed icons for the Lunar Module.
- + Displays a clearer step-by-step progression that is easier to follow at a glance.
- − Completely failed the main title, spelling it 'ALFONCH' instead of Apollo 11.
- − Suffers from significant text artifacts and incoherent labeling across all steps.
Verdict: FLUX.1 [dev] followed the aesthetic and layout instructions more closely, producing a visually balanced poster that feels like a professional infographic, despite the gibberish text. Vidu Q2 provided better individual icons for the descent stages but failed significantly on the primary title and overall composition. FLUX.1 [dev] is preferred for its adherence to the vector style and color palette.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing