Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent surface reflections and lighting on the glass.
- + Vibrant, clean aesthetic with sharp focus.
- + Accurately represents the spatial arrangement of the objects.
- − Added an extra blue sphere on top of the book that was not requested.
- − The 'floating' nature of the sphere inside the cube looks slightly unnatural.
Wan 2.6
- + Followed all prompt instructions precisely without adding extra objects.
- + Highly realistic textures on the wooden table and weathered book.
- + Natural integration of the plant's visibility through the back of the glass cube.
- − The lighting is a bit harsh compared to the 'soft window light' requested.
- − The sphere inside the cube has a slight alignment issue with its base reflection.
Verdict: Both models handled the complex spatial relationships and transparency well. However, Wan 2.6 followed the prompt instructions the most accurately, while FLUX.1 [schnell] hallucinated an additional sphere on top of the stack.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent handling of wet pavement reflections and soft lighting
- + Good technical quality with a clean cinematic look
- − The man appears to be posing or holding the handlebars rather than actively repairing the bike
- − Lacks requested motion blur in passing cars
- − Missing visible rain drops despite wet surfaces
Wan 2.6
- + Highly accurate interpretation of 'repairing' with tools and a crouched posture
- + Captures the interaction with rain and wet textures on clothing very realistically
- + Good inclusion of motion blur in the background traffic
- − Rain drops on the jacket look slightly repetitive and oversized upon close inspection
- − Distorted hand/finger anatomy while holding the tool
Verdict: Wan 2.6 is the clear winner for its superior prompt adherence, particularly the depiction of the man actually performing a repair and the inclusion of motion blur. While FLUX.1 [schnell] creates a visually clean image, it feels more like a staged portrait than a candid street scene of a bike being fixed.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely high-resolution skin texture and eye detail
- + Excellent shallow depth of field with cinematic lighting
- + Intense, lifelike gaze that captures character emotion well
- − Missed the request for beads in the hair braids
- − The plate armor is barely visible in the crop
- − Lacks the requested 'battle-worn' dirt and scars on the skin
Wan 2.6
- + Perfect adherence to all prompt elements including beaded braids and dirt
- + Superb texture on the leather straps, cloth underlayers, and engraved armor
- + Great inclusion of bokeh sparks and torchlight ambience
- − Skin texture is slightly less sharp than Model A
- − Composition is a bit more standard for a portrait compared to the extreme close-up of A
Verdict: While FLUX.1 [schnell] creates a more visually striking and higher-resolution facial portrait, Wan 2.6 is the clear winner for prompt adherence. Wan 2.6 successfully incorporated every specific detail requested, from the beaded braids and 'battle-worn' dirt to the complex textures of the leather and armor underlayers.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist aesthetic with plenty of white space
- + Clean, professional sans-serif typography for main headers
- + Accurate grid layout for food photos
- − Nonsense section title 'ORFEFUS' instead of requested 'MAINS'
- − Menu text is largely illegible gibberish
- − Food photos lack variety, appearing repetitive
Wan 2.6
- + Better adherence to 'vibrant accents' with the colorful geometric borders
- + Includes all three requested sections: Appetizers, Pizza, and Mains
- + Higher quality food photography with more variety
- − Typography is slightly messy with overlapping characters in some prices
- − Layout feels a bit crowded compared to Model A
- − Text is still mostly nonsensical despite having the correct headers
Verdict: Wan 2.6 is the preferred model as it correctly followed all three section requirements (Appetizers, Pizza, Mains) and incorporated the 'vibrant accents' better than FLUX.1 [schnell]. While FLUX.1 [schnell] had a cleaner minimalist aesthetic, it failed the prompt by hallucinating a section called 'ORFEFUS' instead of including 'Mains'.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + High resolution burger textures with realistic melting cheese
- + Good use of embers and lighting for a fiery atmosphere
- − Typos in the main title ('AGIC BURGER') and price ('€699')
- − The burger is largely assembled rather than 'exploded' as requested
- − Extra redundant text and price elements create a cluttered bottom half
Wan 2.6
- + Excellent adherence to the 'exploded' instruction with clear separation of ingredients
- + Accurate and high-quality rendering of all requested text
- + Highly creative fiery/glowing text effects that match the prompt's theme
- − The meat patty texture is slightly less detailed than Model A
Verdict: Wan 2.6 is the clear winner as it followed every instruction, including the difficult 'exploded' layout and specific text rendering. While FLUX.1 [schnell] produced a high-quality burger, it failed significantly on the text by omitting the first letter of 'MAGIC' and providing a wildly incorrect price in the starburst.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Text is clear and highly legible
- + Includes the requested date of April 30, 2026
- − Numerous spelling errors including 'Pril', 'Taffle', and 'Octtoopus'
- − The text looks more like a digital font or marker than realistic chalk on a blackboard
- − Failed to render the full names of the items correctly
Wan 2.6
- + Excellent adherence to the 'chalk handwriting' style with realistic texture and smudging
- + Perfect text rendering with zero spelling errors for all requested complex menu items
- + Beautifully captures the cozy café atmosphere and composition requested
- − The date is slightly overlapping the line above, though still legible
- − The very bottom line of text is a bit small, though impressively accurate
Verdict: Wan 2.6 is the clear winner as it followed every instruction perfectly, including difficult spelling and specific chalk textures. FLUX.1 [schnell] struggled significantly with the text content, producing many spelling errors and failing to capture the realistic 'chalk' look, instead producing text that looks like a white gel pen or digital font.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Perfectly adhered to the complex spatial instruction of 'horse on top'
- + Strong cinematic lighting with a clear focal point against a planet horizon
- + Successfully captured the surreal nature of the prompt
- − Anatomical glitch where the horse appears to have two heads/necks merged together
- − Astronaut equipment detail is slightly ambiguous in terms of orientation
Wan 2.6
- + High resolution with vibrant colors and cosmic details
- + Great texture on the horse's coat and mane
- + Well-rendered astronaut suit and reflection in the visor
- − Failed the core prompt instruction of 'horse on top, not vice versa'
- − Typical interpretation rather than a surreal one
Verdict: FLUX.1 [schnell] followed the difficult spatial constraints of the prompt, placing a horse sitting on an astronaut, whereas Wan 2.6 defaulted to a standard astronaut-riding-a-horse image. Although Wan 2.6 has superior detail and color, FLUX.1 [schnell] is the clear winner for actually following the specific 'not vice versa' instruction.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and lighting on the capybara's face.
- + Accurately places the passenger in the back seat as requested.
- − The capybara only has one paw near the wheel, and it is not gripping it.
- − The capybara's 'expression' is a bit too wide-eyed and startled rather than professional.
- − Composition feels cramped with the car's exterior frame blocking much of the view.
Wan 2.6
- + Perfect adherence to the instruction for 'both front paws on the steering wheel'.
- + Effective use of rain on the windshield and vibrant city lights for a cinematic feel.
- + Great styling of the capybara's coat and driver's hat.
- − The passenger appears to be in the front passenger seat rather than the back seat.
- − The passenger holds the phone with three hands/too many fingers.
Verdict: Model B (Wan 2.6) captures the essence of the prompt much better, specifically the action of the capybara driving with both paws on the wheel. While Model B failed to place the passenger in the back seat and has some anatomical issues with her hands, the overall composition and atmosphere of the Manhattan night are superior to FLUX.1 [schnell], which feels more like a staged selfie than a cinematic scene.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features a stylized graphic design
- + Followed the color scheme and general spooky vibe
- − Significant text errors including typos and repeated lines
- − Poor text layout for an invitation
- − Jack-o-lantern and trees look like flat clip art
Wan 2.6
- + Excellent text rendering with zero typos
- + Strong adherence to the 'vintage gothic parchment' request with thorns and webs
- + High-quality cinematic lighting around the central jack-o-lantern
- − The text is slightly cramped near the top border
Verdict: Wan 2.6 followed the prompt instructions perfectly, particularly with the complex text requirements and the vintage parchment aesthetic. FLUX.1 [schnell] struggled significantly with the text, producing several typos (e.g., 'firiichts') and redundant lines of information, while its overall composition lacked the cinematic depth of the other model.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent clean 3D render quality with smooth textures
- + High clarity and minimalist aesthetic
- + Accurate flag icon representation
- − Failed to include the word 'SUSHI' in the text
- − The scale of the sushi piece is a bit large for the diorama base
- − Missed the 'SUSHI below it' instruction
Wan 2.6
- + Followed all text instructions including the word 'SUSHI'
- + Better diorama composition with a wooden plate and garnishes
- + Excellent 3D miniature style with pleasing shadows
- − The text is not perfectly centered as requested
- − Minor artifacts in the rice grain textures
Verdict: Wan 2.6 followed the prompt instructions much more accurately, successfully including both lines of text and the flag icon while maintaining a beautiful miniature 3D aesthetic. FLUX.1 [schnell] produced a very clean image but missed the 'SUSHI' text requirement entirely.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Warm and vibrant color palette
- + High facial detail on the animals
- + Solid composition with clear depth of field
- − Failed to include a distinct bunny character (merged rabbit/kitten hybrid features)
- − The scene feels static rather than 'playfully chasing'
- − Anatomy of the ears on the middle-right animal is ambiguous
Wan 2.6
- + Successfully included all four distinct animal types requested
- + Excellent interpretation of 'playfully chasing' and 'tumbling'
- + Beautifully rendered god rays and dew sparkles that match the prompt
- − The fox's front paw has a slightly distorted, blocky artifact
- − The cat's anatomy in mid-air is a bit stiff
- − The background plants are slightly noisy in the upper-left corner
Verdict: Wan 2.6 is the clear winner as it successfully rendered all four requested animals (dog, cat, fox, and bunny), whereas FLUX.1 [schnell] missed the bunny entirely. Additionally, Wan 2.6 captured the 'chasing' and 'tumbling' action described in the prompt, creating a much more dynamic and accurate scene compared to the static pose in the other image.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong vector emblem aesthetic with a balanced circular composition.
- + Clean, professional typography style.
- − Failed to spell the name correctly, rendering 'Cafeé Framilan' instead of 'Caffè Florian'.
- − Incorrectly rendered the year as '7720' instead of '1720'.
- − Missing the steam element requested in the prompt.
Wan 2.6
- + Perfect adherence to text requirements, including 'Caffè Florian' and 'Est. 1720'.
- + Includes all requested elements: cloche dome, steam, and banner.
- + Excellent use of subtle texture on the background to enhance the vintage feel.
- − The banner is slightly small and tucked to the side rather than being a central design element.
Verdict: Wan 2.6 is the clear winner as it followed every detail of the prompt, including perfect spelling and numbers. FLUX.1 [schnell] failed significantly on text accuracy, misspelling the name and providing an impossible date, while also omitting the steam element.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent flat-vector aesthetic that matches the infographic style requested.
- + Follows the NASA-inspired color palette accurately.
- + Attempts to visualize the orbital and descent steps mentioned in the prompt.
- − The text is largely illegible gibberish.
- − The rocket icon looks more like a shuttle than a Saturn V.
- − The layout is somewhat cluttered and lacks clear logical flow for the numbered steps.
Wan 2.6
- + Features clean, legible typography for the header and mission crew names.
- + Captures the NASA-inspired navy and white color scheme well.
- + Image clarity is high with no obvious vector artifacts.
- − Failed to include any of the six specific infographic steps requested.
- − The 'poster' looks more like a textured towel or fabric print than a vector infographic.
- − Contains very little relevant mission content beyond the header and three names.
Verdict: FLUX.1 [schnell] is the winner as it attempted the complex infographic structure and flat-vector style requested, even though the text is illegible. Wan 2.6 produced a high-quality final image with great text rendering, but it completely ignored the specific technical requirements for the six distinct mission steps and the 'infographic' layout.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English