Black Forest Labs' distilled 9 billion parameter image generation model with sub-second inference and multi-reference support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [klein] 9B
#13 of 32 in Image Editing
Stable Diffusion 3.5 Large
#30 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [klein] 9B
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Follows all spatial instructions perfectly, including the book on top and the sphere inside.
- + Excellent rendering of refraction and reflections in the glass cube.
- + Beautiful bokeh effect on the background plant and natural-looking window light.
- − The blue sphere appears slightly floating or fused with the glass base rather than resting naturally.
- − The shape of the cube is slightly rounded, resembling a vase more than a perfect geometric cube.
Stable Diffusion 3.5 Large
- + Sharp, precise geometric edges on the glass cube.
- + Clear, vibrant colors and high-resolution textures on the wooden table.
- − Failed the primary spatial prompt: the sphere is on top of the book, and the book is inside the cube.
- − The lighting direction is inconsistent with the 'light from the left' instruction, appearing more overhead/frontal.
- − The plant is mostly obscured by the sofa and not clearly visible through the glass as requested.
Verdict: FLUX.2 [klein] 9B followed the complex spatial instructions perfectly, placing the sphere inside the cube and the book on top, while maintaining a high level of photographic realism. In contrast, Stable Diffusion 3.5 Large failed to follow the positional requirements, placing the sphere on the book and the book inside the cube, which contradicted the prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent skin texture and facial details
- + Accurate 50mm lens perspective with shallow depth of field
- + High level of realism in the reflections and wet pavement
- − Misses the motion blur requirement for passing cars
- − The bike geometry has some minor structural inconsistencies at the handlebars
Stable Diffusion 3.5 Large
- + Successfully incorporates motion blur in the background vehicle
- + Effective use of rain particles to create atmosphere
- + Includes a very red, distinct bicycle
- − Anatomical issues with the man's hands and arms
- − The facial features are less realistic and slightly muddy
- − Lighting feels a bit more artificial compared to the other model
Verdict: FLUX.2 [klein] 9B produces a far more realistic image with superior detail in the man's face and hands, adhering well to the 50mm lens and photo-realistic requirements. While Stable Diffusion 3.5 Large successfully captures the motion blur of the passing car, it fails on anatomy and fine detail, making FLUX the clear winner for its photographic quality.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent adherence to the 'beads in hair' prompt
- + Intricate textures on leather straps and gambeson
- + Strong use of warm torchlight and bokeh sparks
- − The symmetrical placement of torches in the background feels a bit artificial
- − Skin texture is slightly over-sharpened
Stable Diffusion 3.5 Large
- + Very realistic and natural facial expression
- + Exceptional intricate engraving on the plate armor
- + Good sense of depth with blurred background figures
- − Missed the 'small beads' in the hair braids entirely
- − The lighting is less 'warm torchlight' and more neutral daylight with some fire
- − Ornate armor patterns look a bit like stamped textures rather than custom engraving
Verdict: FLUX.2 [klein] 9B followed the prompt details much more closely, specifically including the hair beads and the prominent torchlight. While Stable Diffusion 3.5 Large produced a high-quality cinematically composed image, it failed several specific descriptors in the prompt like the beads and the lighting mood.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Successfully included all requested sections (Appetizers, Pizza, Mains).
- + Excellent use of bold sans-serif fonts that are highly legible.
- + Authentic grid layout that mimics a functional menu.
- − Text contains several spelling errors and nonsensical characters.
- − Food images are repetitive, mostly showing four variations of the same pizza.
Stable Diffusion 3.5 Large
- + Strong aesthetic composition with a unique side-grid approach.
- + High visual quality and vibrant colors in the food photography.
- + Professional minimalist vibe appropriate for a casual dining brand.
- − Failed to include the 'Pizza' section heading specifically requested.
- − The 'grid' format surrounds the text rather than being integrated within a standard menu layout.
- − Text is heavily garbled and difficult to read.
Verdict: FLUX.2 [klein] 9B followed the structural requirements of the prompt more closely by including all requested sections and a standard menu grid layout. While Stable Diffusion 3.5 Large produced more vibrant and varied food photography, its failure to include the 'Pizza' category and its less practical layout make it less successful for this specific design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent typography with perfect adherence to the requested text and formatting.
- + Great implementation of the 'exploded' burger concept with suspended ingredients.
- + Strong commercial aesthetic with clear hierarchy and professional lighting.
- − The 'starburst' for the price is a bit jagged compared to typical graphic design standards.
Stable Diffusion 3.5 Large
- + Atmospheric lighting with realistic fire and charcoal elements.
- + High level of detail on the burger patties and melting cheese.
- − Failed to include any of the requested text elements.
- − The burger is not 'exploded' or separated as requested; it is a static stack.
- − Lack of dynamic motion in the composition.
Verdict: FLUX.2 [klein] 9B is the clear winner as it followed all prompt instructions, including complex text rendering and the 'exploded' burger layout. Stable Diffusion 3.5 Large produced a high-quality image of a burger on fire, but completely ignored the text requirements and the specific structural request for suspended components.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI judge analysis unavailable for this challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent anatomical rendering of both horse and astronaut.
- + High contrast and vibrant color palette with clear celestial bodies.
- + Very sharp focus and clean textures on the astronaut's suit.
- − Failed the negative constraint; the astronaut is on top of the horse.
- − Includes nonsensical AI-generated text at the bottom.
- − Composition is a bit cluttered with too many small planets.
Stable Diffusion 3.5 Large
- + Atmospheric and cinematic lighting that blends the subjects into the environment.
- + Strong sense of motion with the gaseous trail following the horse.
- + Better handle on the 'surreal' aspect of the prompt through lighting and haze.
- − Failed the negative constraint; the astronaut is on top of the horse.
- − Lower level of detail on the astronaut's face and the horse's anatomy.
- − Image quality appears slightly grainy compared to Model A.
Verdict: Both FLUX.2 [klein] 9B and Stable Diffusion 3.5 Large failed the specific negative constraint to have the horse on top of the astronaut (both produced the standard astronaut-riding-horse image). FLUX.2 [klein] 9B is the better image in terms of raw detail and clarity, despite the unwanted text residue, while Stable Diffusion 3.5 Large offers a more cohesive and cinematic atmosphere.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent adherence to the complex scene layout including the back seat passenger.
- + Highly realistic textures on the capybara fur and the woman's clothing.
- + Captures the requested 'bored' expression of the passenger perfectly.
- − The text on the driver's cap is slightly nonsensical ('NEW TALA').
- − Minor lighting inconsistency on the capybara's paws compared to the dashboard.
Stable Diffusion 3.5 Large
- + Vibrant colors and high contrast image quality.
- + Great detail on the capybara's whiskers and jacket.
- − Completely failed to include the businesswoman in the back seat.
- − The capybara's anatomy is distorted, with legs and jeans appearing human-like and awkward.
- − The composition is a side profile rather than the requested 'inside the taxi' view showing the passenger.
Verdict: FLUX.2 [klein] 9B followed the prompt instructions near-perfectly, successfully including both the capybara driver and the bored businesswoman in the background. Stable Diffusion 3.5 Large failed to include the passenger entirely and struggled with the internal composition of the car, resulting in a much simpler and less accurate image.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent typography with perfect spelling in a gothic font.
- + High-quality cinematic lighting with a clear central jack-o-lantern.
- + Precise adherence to the thorn and spiderweb border request.
- − The parchment texture is subtle, appearing more like a clean frame than weathered paper.
Stable Diffusion 3.5 Large
- + Strong vintage parchment aesthetic with torn edges.
- + Good use of vertical space for the scroll banner.
- + Detailed illustrative style for the trees and fence.
- − Failed to include the specific event details (Date, Time, Location) at the bottom.
- − The lighting is overly blown out around the moon, making the scene less moody.
- − Text rendering is inconsistent, with some artifacts in the bottom banner.
Verdict: FLUX.2 [klein] 9B followed every instruction perfectly, including all requested text and specific date/location details with flawless typography. Stable Diffusion 3.5 Large captured a great vintage parchment texture but failed to include several required lines of text and had significantly lower clarity in its rendering.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent typography rendering with clean, bold text precisely as requested.
- + Simple, elegant 3D miniature aesthetic with soft lighting.
- + Accurate isometric perspective and clean execution of a diorama base.
- − The flag icon is incorrect, resembling the flag of Yemen rather than Japan.
- − The sushi roll in the foreground has a confusing structural layout where the salmon sits atop a cut maki.
Stable Diffusion 3.5 Large
- + Correctly identifies the Japanese flag for the icon.
- + Highly detailed textures for the fish and rice that lean more towards 'realistic PBR'.
- + Good variety of sushi types on the diorama.
- − Failed to place the text directly on the background as a graphic element, placing it on a sign instead.
- − The scene is more cluttered than the 'minimal' request specified.
- − Minor artifacts on the chopsticks and garnish shapes.
Verdict: FLUX.2 [klein] 9B followed the graphic design instructions much better, correctly placing the text and adopting a cleaner, more 'minimal' 1-to-1 isometric style. Stable Diffusion 3.5 Large produced a more complex and visually rich 3D scene but failed the specific text placement instruction and rendered much more garnish than requested. FLUX.2 is the winner for its superior layout and adherence to the 'ultra-clean' aesthetic, despite the incorrect flag colorization.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent clarity and sharpness in the fur textures and butterfly details.
- + Rich, vibrant colors in the wildflower meadow.
- + Strong lighting effects with clear God rays and backlighting.
- − Failed to include the baby bunny requested in the prompt.
- − The kitten's eye anatomy is slightly distorted and unsettling.
Stable Diffusion 3.5 Large
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox kit.
- + Better sense of action and 'tumbling' motion as requested.
- + Dreamy, cinematic bokeh and lighting that fits the 'wholesome vibe'.
- − The fox's face lacks anatomical detail and looks slightly blurry.
- − Butterflies are less detailed and simplified compared to the other model.
Verdict: Stable Diffusion 3.5 Large wins this comparison because it followed the prompt instructions more accurately, including all four specific animals while FLUX.2 [klein] 9B missed the bunny. While FLUX.2 has superior technical sharpness and texture, Stable Diffusion 3.5 Large better captured the playful movement and the specific cast of characters requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent typography with correct spelling and accent placement.
- + Clean vector emblem style with balanced composition.
- + Effective use of subtle paper texture on the background.
- − The cloche design is slightly bulky compared to traditional minimalist logos.
Stable Diffusion 3.5 Large
- + Elegant vintage aesthetic with corner ornaments.
- + Good color palette adherence following the warm brown and cream tones.
- − Incorrect spelling of 'Caffè' as 'Cafféé'.
- − The cloche is detached from the base in a physically impossible way.
- − Text is slightly cramped within the banner.
Verdict: FLUX.2 [klein] 9B followed the prompt more accurately, specifically regarding the correct spelling of the name and the application of a clean vector style. Stable Diffusion 3.5 Large suffered from significant spelling errors and a disjointed icon design, despite having a pleasant vintage texture.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Successfully followed the instructional step sequence with specific icons for launch, orbit, and landing.
- + Very high text legibility for headers, even with minor spelling errors.
- + Clean vector aesthetic with the requested NASA-inspired color palette.
- − Includes several typos in the labels such as 'AOILLO', 'EARDHT', and 'LANDINING'.
- − The flow between the steps is a bit disconnected/cluttered compared to a professional infographic.
Stable Diffusion 3.5 Large
- + Strong composition and artistic design that feels like an authentic retro poster.
- + Captures the high-contrast aesthetic and color palette perfectly.
- − Failed to provide the requested 6-step logical sequence.
- − Includes non-Apollo elements like a Space Shuttle shaped rocket and a ringed planet.
- − Text is completely illegible and nonsensical.
Verdict: FLUX.2 [klein] 9B followed the prompt's structural instructions much more effectively, providing the specific 6-step sequence and readable (if slightly misspelled) text. Stable Diffusion 3.5 Large produced a more visually striking poster, but it failed on almost every specific technical requirement, including the icon types and logical progression of the Apollo 11 mission.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency