Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
Wan 2.5 (Preview)
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent surface rendering and clarity of the glass material
- + Vibrant colors on the plant and spheres
- + Clean, modern aesthetic
- − Included an extra blue sphere on top of the red book which was not requested
- − The blue sphere inside is floating mid-air, which looks slightly unnatural
Wan 2.5 (Preview)
- + Followed all prompt instructions accurately without adding extra elements
- + Realistic lighting with dust motes and shadows from the soft window light
- + Excellent texture on the worn red book cover
- − The glass cube has slightly more distortion in its reflections
- − The composition feels slightly tighter and more cropped than Model A
Verdict: Wan 2.5 (Preview) is the winner because it adhered strictly to the prompt without adding unnecessary extra objects, whereas FLUX.1 [schnell] added a second blue sphere on top of the book. Wan 2.5 also handled the lighting and environmental details, like the dust in the sunlight and the shadows on the table, with superior realism.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent handling of wet pavement reflections and bokeh.
- + Captures the requested 'motion blur from passing cars' in the background.
- + Strong color contrast with the red vest and red bicycle.
- − The man's hands are mangled and poorly rendered.
- − The bicycle anatomy is incorrect, with the frame tubes not connecting realistically at the pedals.
- − The man's pose looks more like he is leaning on the bike rather than actively repairing it.
Wan 2.5 (Preview)
- + Natural skin texture and convincing elderly facial features.
- + The 'repairing' action is far more believable with tools out and a focused posture.
- + The bicycle and kickstand physics are more coherent than Model A.
- − Failed to include the requested motion blur for passing cars.
- − The hands, while better than Model A, still show some anatomical merging with the tools.
- − The rain effect is slightly less atmospheric compared to Model A.
Verdict: Wan 2.5 (Preview) produced a far more convincing 'candid' scene where the man is actually repairing the bicycle, showing superior character detail and structural coherence. While FLUX.1 [schnell] excelled at the background atmosphere and motion blur, it failed significantly on the human anatomy (hands) and the mechanical structure of the bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely high skin texture detail and realistic iris rendering.
- + Intense, cinematic lighting and color grading.
- + Captures a gritty 'battle-worn' expression effectively.
- − Fails to include visible beads in the hair braids as requested.
- − The armor is very dark and lacks the requested ornate engravings.
- − Very tight crop obscures the 'battle-worn' attributes of the full outfit.
Wan 2.5 (Preview)
- + Excellent adherence to all prompt elements, including beads, engraved plate, and tattered underlayers.
- + Highly realistic contrast between the warm torchlight and cool shadows.
- + Great texture work on the leather straps and frayed cloth.
- − The facial features are slightly soft compared to the sharpeness of the armor.
- − The 'scars' look more like fresh dirt/smudges rather than faint skin tissue damage.
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully incorporated every specific detail of the prompt, including the beads in the hair and the ornate engraving on the plate armor, which FLUX.1 [schnell] ignored. While FLUX.1 [schnell] has slightly more impressive skin pore detail, it failed to deliver on the 'paladin' aesthetic and the specific material request for the armor.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent font legibility for header text like 'MENU' and 'PIZZA'.
- + Clean, minimalist layout that feels like a real restaurant menu.
- − Includes some gibberish text in the smaller descriptions.
- − The food photos lack variety, appearing to show multiple versions of the same dish.
- − Failed to correctly include a 'Mains' category header with appropriate dishes.
Wan 2.5 (Preview)
- + Higher quality and more vibrant food photography.
- + Perfectly adheres to the grid layout with distinct sections for Appetizers, Pizza, and Mains as requested.
- + Creative use of color coding for different sections.
- − Significant spelling errors throughout, including the title 'Restormalit Menue'.
- − Visual artifacts such as a pizza slice appearing in the 'Mains' section.
Verdict: Wan 2.5 (Preview) provided a much more vibrant and complete realization of the prompt, including all three requested categories and high-quality food photography in a clear grid. While FLUX.1 [schnell] had cleaner typography and a more realistic minimalist aesthetic, it failed the composition requirements of the grid and specific menu sections.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic texture on the burger bun and patty.
- + Vibrant and dynamic fire background with good depth.
- − Failed to render the 'M' in 'MAGIC BURGER', showing 'AGIC BURGER' instead.
- − The burger is mostly assembled rather than 'exploded' with suspended components as requested.
- − The price text is messy, repeating the value and adding an incorrect extra digit.
Wan 2.5 (Preview)
- + Perfect adherence to the 'exploded' burger layout with clear separation of ingredients.
- + Text rendering is flawless, creative, and follows the 'fiery/glowing' style precisely.
- + Composition is balanced with a strong sense of motion and better integration of the starburst.
- − The 'exploded' lettuce pieces look slightly more like digital assets than organic food.
Verdict: Wan 2.5 (Preview) is the clear winner as it followed every part of the prompt, particularly the 'exploded' layout and complex text requirements. FLUX.1 [schnell] failed on basic text accuracy and did not properly deconstruct the burger as requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Legible chalk texture on the individual letters.
- + Follows the general layout of a chalkboard menu.
- − Severely failed text rendering with numerous spelling errors like 'Taffle Mushmnctiomn' and 'Pril'.
- − The handwriting style feels more like a digital marker than elegant cursive.
Wan 2.5 (Preview)
- + Excellent text adherence with almost perfect spelling of complex menu items.
- + Realistic chalkboard aesthetics including chalk dust, smudges, and authentic varying handwriting.
- + Successfully rendered the specific date and price points requested.
- − The word 'Herbs' is cut off at the end of the second menu item line.
Verdict: Wan 2.5 (Preview) is the clear winner as it followed nearly every text instruction perfectly, while FLUX.1 [schnell] struggled with basic spelling and failed to produce the requested menu items. Wan 2.5 also captured the requested 'elegant cursive' and 'chalk texture' much more effectively than the simplified printing in the FLUX image.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Perfectly adhered to the difficult logic of placing the horse on top of the astronaut
- + Highly cinematic lighting and composition
- + Successful surrealistic interpretation of the prompt
- − The horse appears to have two heads or a strange anatomical distortion in the neck area
- − The astronaut's suit architecture is slightly messy around the joints
Wan 2.5 (Preview)
- + High resolution with vibrant colors and cosmic details
- + Great anatomical rendering of the horse and astronaut gear
- − Completely failed the negative constraint/positional logic by placing the astronaut on top
- − Lacks the surreal quality requested by ignoring the specific 'horse on top' instruction
Verdict: FLUX.1 [schnell] is the clear winner because it correctly interpreted the difficult positional instruction to have the horse riding the astronaut. While Wan 2.5 (Preview) produced a high-quality image, it defaulted to a standard trope and ignored the core creative requirement of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent texture on the capybara's fur and the passenger's face.
- + Cinematic lighting that effectively conveys a nighttime New York atmosphere.
- + The passenger's expression and posture perfectly match the 'bored' requirement.
- − The capybara only has one paw near the wheel, with the other resting low.
- − The perspective makes the capybara look disproportionately large relative to the car interior.
Wan 2.5 (Preview)
- + Strict adherence to the 'both front paws on the steering wheel' instruction.
- + More authentic 'taxi driver' style cap compared to a simple beanie.
- + Excellent composition showing both the interior and exterior environment simultaneously.
- − The capybara's paws look slightly more like human hands/fingers than rodent paws.
- − The passenger's face is slightly blurry and less detailed than in Model A.
Verdict: Both models followed the prompt well, but Wan 2.5 (Preview) captured the specific technical requirements better by including both paws on the wheel and a more traditional taxi driver cap. While FLUX.1 [schnell] has slightly higher skin and fur texture quality, Wan 2.5 provided a more convincing 'professional' scene and a better overall composition of the taxi cab.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully captures the central glowing jack-o-lantern and bats
- + Consistent gothic color palette with orange and dark tones
- − Significant text hallucinations and repetition in the event details
- − The banner text and layout are cluttered and messy
- − Fails to clearly depict the requested 'webs and thorns' border
Wan 2.5 (Preview)
- + Excellent typography with perfect adherence to the requested text and event details
- + High-quality, detailed illustration including clear webs and thorns border
- + Effective cinematic lighting and use of the scroll banner
- − The parchment background is a bit bright compared to the 'dark parchment' request
- − The gnarled tree is isolated to one side rather than framing the scene
Verdict: Wan 2.5 (Preview) significantly outperforms FLUX.1 [schnell] by rendering all the requested text perfectly, whereas FLUX.1 [schnell] struggled with spelling, repeated lines, and added nonsense fields. Wan 2.5 (Preview) also followed the stylistic prompts for webs, thorns, and a scroll banner much more effectively, resulting in a professional-looking invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the isometric perspective and raised diorama base concept.
- + Very clean, high-clarity rendering with a professional graphic design aesthetic.
- + Accurate flag icon representation.
- − Completely missing the word 'SUSHI' requested in the prompt.
- − The 'JAPAN' text has low contrast against the light blue background.
Wan 2.5 (Preview)
- + Includes all requested text elements ('JAPAN' and 'SUSHI') with great legibility.
- + Features very soft, refined cartoon textures and appealing 3D lighting.
- + Better color contrast and composition for a social media or sticker-style graphic.
- − Failed to follow the 'isometric' perspective, opting for a standard perspective view.
- − The sushi design is slightly more generic compared to the detailed materials in the other model.
Verdict: Wan 2.5 (Preview) followed the text instructions more accurately by including both 'JAPAN' and 'SUSHI', whereas FLUX.1 [schnell] missed the second word. However, FLUX.1 [schnell] adhered much better to the specific 'isometric' and 'diorama base' style requested, delivering a cleaner technical result despite the text omission.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Vibrant and warm lighting with excellent bokeh.
- + Soft, fluffy fur textures and expressive eye details.
- + Clean composition without floating artifact issues.
- − Failed to include a rabbit, instead generating a kitten/fox hybrid creature in the middle.
- − The 'fox' resembles a red cat more than a fox kit.
- − Static poses do not convey the 'chasing' or 'tumbling' action requested.
Wan 2.5 (Preview)
- + Correctly includes all four requested animals: puppy, kitten, rabbit, and fox.
- + Excellent sense of movement and 'playfully chasing' action.
- + Beautiful rendering of god rays and dew sparkles.
- − The fox's eyes appear slightly unnatural and overly blue/glowing.
- − Some odd floating water droplets that look like glass beads rather than dew.
- − The transition between the puppy's paws and the grass is a bit blurry.
Verdict: Wan 2.5 (Preview) is the winner because it successfully followed the complex prompt by including all four distinct animals (golden retriever, tabby kitten, bunny, and fox), whereas FLUX.1 [schnell] failed to include the rabbit entirely. Additionally, Wan 2.5 (Preview) captured the dynamic action of chasing and tumbling, while FLUX.1 [schnell] produced a more static, portrait-like group shot.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector aesthetic
- + Symmetric layout for a round emblem style
- − Major spelling errors in the brand name ('CAFEÉ FRAMILAN')
- − Incorrect date ('EST. 7720')
- − Missing requested steam element
Wan 2.5 (Preview)
- + Perfect text rendering of 'Caffè Florian' and 'Est. 1720'
- + Includes all requested elements including the steam and banner
- + Beautiful vintage texture and paper effect
- − Slightly less 'minimalist' than Model A due to the detailed paper background
Verdict: Wan 2.5 (Preview) significantly outperformed FLUX.1 [schnell] by following every instruction in the prompt, especially regarding the specific text and addition of steam. While FLUX.1 [schnell] produced a clean vector, it failed on all legible text and omitted key visual details like the steam and chronological date.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the 'flat-vector' style with subtle gradients.
- + Consistent minimalist iconography.
- + Accurate NASA-inspired color palette implementation.
- − Text is largely illegible gibberish.
- − The diagram logic is confusing and doesn't clearly follow the 6 requested steps in sequence.
Wan 2.5 (Preview)
- + High degree of prompt adherence for all 6 specific steps of the mission.
- + Legible and accurate text labels including the crew names and landing site.
- + Clean composition that functions effectively as an educational infographic.
- − The rocket designs are generic and do not accurately reflect the Saturn V (looks more like the Space Shuttle in some parts).
- − The Earth and Moon icons have different levels of detail, breaking the consistency slightly.
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully incorporated all six requested infographic steps and delivered legible, accurate text. While FLUX.1 [schnell] captures the 'flat vector' aesthetic more authentically, it fails to execute the informational content of the prompt, resulting in a confusing layout with unreadable text.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities