Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the request for the red book sitting on top of the cube.
- + High visual quality with realistic glass refraction and material textures.
- + Clear implementation of the blue sphere inside the cube.
- − Added an extra blue sphere on top of the book that was not requested.
- − The plant is mostly above rather than behind the cube, though still visible.
Stable Diffusion 3.5 Large
- + Successfully placed a blue sphere inside a glass enclosure.
- + Good representation of the soft window light from the left.
- + Realistic wood texture on the table.
- − Failed the spatial instruction by placing the cube on top of the book instead of the book on the cube.
- − The 'cube' looks more like a display case or cover than a solid glass object.
- − The red book is much larger than depicted in typical prompt interpretations.
Verdict: FLUX.1 [schnell] followed the complex spatial instructions much better than Stable Diffusion 3.5 Large, correctly placing the book on top of the cube and the sphere inside. Although FLUX.1 added an extra sphere onto the book, Stable Diffusion 3.5 Large completely inverted the vertical arrangement of the objects, making it less accurate to the prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent anatomical accuracy in the hands and face
- + Superior rendering of wet pavement and realistic reflections
- + High level of detail in the bicycle mechanics
- − Missed the motion blur requirement for passing cars
- − Rain effect is very subtle to the point of being almost invisible
Stable Diffusion 3.5 Large
- + Successfully captures visible rain and atmospheric wetness
- + Better adherence to the motion blur request for the background vehicles
- + Good sense of 'imperfect framing' and candid composition
- − Severe anatomical issues with the man's hands merging into the bicycle
- − The bicycle's architecture is physically impossible with mismatched frame segments
- − The skin texture on the arms appears overly distorted and burnt
Verdict: FLUX.1 [schnell] produces a much higher quality image with realistic textures and correct anatomy, though it fails to include the requested motion blur. Stable Diffusion 3.5 Large adheres better to the atmospheric prompts like motion blur and rain intensity, but ultimately fails due to significant structural hallucinations in the person and the bicycle. FLUX.1 [schnell] is the winner for its professional-grade visual coherence.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely sharp facial textures and lifelike iris details.
- + Effective use of warm lighting and narrow depth of field for a cinematic feel.
- + Successfully captures the 'battle-worn' look through skin texture and expression.
- − The armor is largely cropped out, missing the 'engraved plate' focal point.
- − The beads in the hair are represented by metal rings rather than distinct beads.
Stable Diffusion 3.5 Large
- + Excellent depiction of ornate engraved plate armor as requested.
- + Better framing for a paladin, showing the full harness and underlayers.
- + Follows the hair braiding instruction more accurately with complex braids.
- − The facial skin texture is slightly smoother and less realistic than Model A.
- − Lighting feels a bit flatter despite the background fires.
Verdict: Stable Diffusion 3.5 Large is the winner as it interprets the full scope of the prompt, providing a clear view of the ornate engraved armor and complex braids that FLUX.1 [schnell] largely cropped out. While FLUX.1 [schnell] has superior facial realism and micro-textures, it failed to showcase the primary 'paladin in plate' requirement of the composition.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist aesthetic with a clean white professional layout.
- + Clear grid structure for food photos as requested.
- + Fonts are modern, legible, and consistent with the casual dining theme.
- − The word 'Mains' is replaced by the hallucinated word 'ORFEFUS'.
- − The food images lack high-resolution detail upon closer inspection.
Stable Diffusion 3.5 Large
- + High-quality, vibrant food photography that captures the 'casual dining' feel.
- + Excellent bold sans-serif typography for the main header.
- + Good use of color accents around the borders.
- − The grid layout is messy, with photos spilling off the edges of the frame.
- − Significant spelling errors in every single section header (e.g., 'MAIMAES', 'APPETIZRS').
- − The composition feels cluttered rather than minimalist.
Verdict: FLUX.1 [schnell] captures the 'minimalist' and 'professional' design brief much better than its competitor, providing a layout that looks like a real menu despite one minor text hallucination. Stable Diffusion 3.5 Large produced superior food imagery, but failed significantly on the layout structure and spelling, resulting in a cluttered and unusable design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic rendering of food textures.
- + Maintains a clean composition for an advertisement.
- + Includes most of the requested text elements in a legible way.
- − Missed the first letter of 'MAGIC' (shows 'AGIC').
- − The burger is largely assembled rather than a truly 'exploded' view of components.
- − The price is repeated incorrectly as €699 in the starburst.
Stable Diffusion 3.5 Large
- + Strong atmosphere with high-quality fire and ember effects.
- + Very detailed textures on the buns and grilled patties.
- + Good sense of suspension in mid-air above the fire.
- − Completely failed to generate any of the requested text.
- − Failed the 'exploded burger' requirement, showing a stacked double burger instead.
- − The lighting is a bit overly chaotic for a clear commercial ad.
Verdict: FLUX.1 [schnell] is the winner because it attempted and largely succeeded at the complex text requirements, even with a minor spelling error and pricing typo. Stable Diffusion 3.5 Large produced a high-quality, atmospheric image but failed to include any text or follow the specific architectural 'exploded' instruction for the burger components.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + The handwriting style is very realistic with convincing chalk textures
- + Text is centered and large, making it the clear focal point of the image
- + Follows the year '2026' as requested in the prompt
- − Significant spelling errors throughout the menu items
- − The date says 'Pril' instead of April
- − Lacks the 'elegant cursive' requested for the title
Stable Diffusion 3.5 Large
- + The internal environment of the café is well-composed and visually appealing
- + Layout includes decorative frames and more complex chalkboard art
- + Text is legible from a distance though contains errors
- − Failed the date request, showing '2024' instead of '2026'
- − Spelling errors such as 'TODAAY' and 'Muglrrom'
- − Handwriting looks more like a digital font than natural chalk
Verdict: FLUX.1 [schnell] captures the specific texture of chalk and the requested year 2026, though it struggles significantly with spelling the menu items correctly. Stable Diffusion 3.5 Large provides a better overall scene composition of a café, but it fails on the specific date and has a more 'stamped' look to its text. FLUX.1 [schnell] is the winner for better adhering to the tactile qualities and specific details of the prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Perfectly follows the specific role-reversal instruction of the horse riding the astronaut.
- + High quality cinematic lighting and clean subject isolation.
- + Surreal composition that avoids visual clutter.
- − The horse has two heads/necks, which appears to be a generation artifact rather than a stylistic choice.
- − The astronaut's anatomy is a bit abstract and jumbled under the horse.
Stable Diffusion 3.5 Large
- + Excellent atmospheric effects with the stardust and planetary background.
- + Very high level of detail on the space suit and horse textures.
- + Artistically pleasing composition and movement.
- − Fails the negative constraint of the prompt; the astronaut is riding the horse.
- − Not as surreal as requested, following a more traditional sci-fi trope.
Verdict: FLUX.1 [schnell] is the clear winner because it correctly interpreted the difficult prompt instruction for the horse to be riding the astronaut, despite a major anatomical glitch with the horse's heads. Stable Diffusion 3.5 Large produced a visually stunning and more detailed image, but completely failed to follow the primary conceptual constraint of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully includes both the capybara driver and the businesswoman in the back seat.
- + Accurately captures the 'bored' expression of the passenger and the 'professional' look of the driver.
- + High-quality fur texture and realistic lighting for a night-time taxi interior.
- − The capybara only has one paw near the wheel, not both as requested.
- − The perspective makes the taxi interior feel slightly cramped and distorted.
Stable Diffusion 3.5 Large
- + Strong lighting and vibrant 'New York at night' bokeh in the background.
- + Good clothing detail on the capybara, including the jacket and cap.
- − Completely failed to include the passenger/businesswoman in the back seat.
- − The capybara's anatomy is slightly off, with one paw appearing to grow out of the middle of his chest.
- − The capybara's face looks more like a mix between a rodent and a dog than a pure capybara.
Verdict: FLUX.1 [schnell] followed the complex multi-subject prompt much better, successfully placing both the capybara and the human passenger in the scene with the correct expressions. Stable Diffusion 3.5 Large failed to render the passenger entirely, resulting in a composition that ignored a large part of the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully included all the required event details (Date, Time, Location).
- + Excellent border detail with thorn and web motifs.
- + Clean and legible typography for the header.
- − Significant text errors including 'You invited to a a night of frights' and gibberish 'Fate' and 'Time' lines.
- − The composition feels like a modern graphic design rather than a vintage gothic poster.
- − The background is very sparse with minimal detail in the sky and trees.
Stable Diffusion 3.5 Large
- + Superior aesthetic interpretation of 'dark parchment' and vintage textures.
- + High-quality illustration style for the scroll banner and background scenery.
- + Better adherence to the 'twisted trees' and 'glowing jack-o-lantern' visual requirements.
- − Completely failed to include the event details (Date, Time, Location) at the bottom.
- − Text rendering on the scroll is slightly shaky and contains artifacts/hallucinated characters.
Verdict: FLUX.1 [schnell] followed the prompt more comprehensively by including all the specific event information, but failed significantly on grammar and text layout. Stable Diffusion 3.5 Large produced a much more visually appealing and 'vintage' illustration that fits the gothic theme perfectly, but it omitted nearly half of the text requirements. Stable Diffusion is the better choice for artistic quality, while FLUX is better for functional data inclusion.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean isometric perspective with a perfect diorama base
- + High-quality soft lighting and refined materials
- + Excellent text legibility and clean background
- − Missing the word 'SUSHI' from the text overlay
- − The sushi piece itself looks slightly flat and less realistic compared to the base
Stable Diffusion 3.5 Large
- + Includes all requested text elements
- + Higher level of detail and variety in the sushi types
- + Realistic textures on the rice and fish
- − Ignored the 'solid light blue background' request by adding texture
- − Text is on an object rather than centered at the top as a graphic overlay
- − Diorama base is less 'miniature' in style compared to Model A
Verdict: FLUX.1 [schnell] followed the artistic style and layout instructions much better, producing a clean, professional-looking graphic, though it missed the word 'SUSHI'. Stable Diffusion 3.5 Large captured the specific text and food details well but failed to execute the minimalist, isometric graphic design requested, resulting in a cluttered composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and crispness in the animals.
- + Rich, vibrant colors that enhance the 'wholesome' vibe.
- + Artistic lighting with a beautiful bokeh effect.
- − Failed to include a rabbit, instead generating two cat-like creatures.
- − The anatomy of the middle-right animal is a hybrid of a kitten and a fox, lacking clear distinction.
- − The scene is static rather than showing them 'playfully chasing' or 'tumbling'.
Stable Diffusion 3.5 Large
- + Successfully included all four requested animal types (puppy, kitten, rabbit, fox).
- + Captured the dynamic 'playfully chasing' action requested in the prompt.
- + Included prominent 'god rays' and dew sparkles as specified.
- − The fox in the background has slightly messy facial features.
- − The butterflies have somewhat simplified wings compared to the animals.
- − The kitten's ears are slightly asymmetrical.
Verdict: Stable Diffusion 3.5 Large is the winner because it successfully followed the complex prompt requirements, including all four specific animal species and the dynamic movement of chasing butterflies. FLUX.1 [schnell] produced a high-quality image but failed to generate the bunny and opted for a more static, posed composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector-style execution
- + Strong balanced composition suitable for a logo
- + Nice use of subtle texture on the background
- − Failed the primary text prompt, spelling it 'Cafeé Framilan'
- − Incorrect established date '7720' instead of '1720'
- − Missing the steam element requested in the prompt
Stable Diffusion 3.5 Large
- + Accurate spelling of the name and the date 'Est. 1720'
- + Successfully included the steam and cloche icons
- + Excellent vintage paper texture and ornamental details
- − The 'Caffèé' includes an extra accented 'é' not in the prompt
- − The cloche illustration is slightly floating/disconnected in its center section
- − The logo elements are a bit spaced out for a cohesive emblem
Verdict: Stable Diffusion 3.5 Large is the clear winner because it successfully incorporated almost every specific prompt element, including the correct establishment date and the name 'Florian,' which FLUX.1 [schnell] completely failed. While Stable Diffusion added one extra letter to 'Caffè', its overall adherence to the vintage aesthetic and logic makes it much more useful as a logo draft.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully follows the requested flat-vector style with crisp lines.
- + Layout clearly represents a path from launch to landing.
- + Adheres well to the requested navy, white, and muted red color palette.
- − Text is largely illegible gibberish.
- − The rocket icon looks more like a cartoon plane than a Saturn V.
- − Fails to clearly label the specific 6 steps requested.
Stable Diffusion 3.5 Large
- + Higher level of detail in the lunar surface and planet textures.
- + Captured a more technical, 'blueprint' aesthetic.
- − Includes a space shuttle-like vehicle instead of a Saturn V rocket.
- − Extremely cluttered composition that fails to show a clear step-by-step infographic flow.
- − Includes nonsensical visual elements like a planet with rings that do not belong in a moon mission chart.
Verdict: FLUX.1 [schnell] followed the stylistic instructions much better, delivering a clean flat-vector infographic that visually maps out a journey, despite the garbled text. Stable Diffusion 3.5 Large produced a cluttered and confusing layout that ignored the step-by-step structure and included incorrect spacecraft and planetary bodies.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency