Distilled version of Black Forest Labs' FLUX.2 [dev] outperforming it at a cheaper price. Developed by fal.ai.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [dev] Turbo
#3 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#28 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [dev] Turbo
70.8%
win rate
Ties
12.5%
Stable Diffusion 3.5 Large
16.7%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Perfect adherence to spatial instructions with the book clearly on top and the sphere inside.
- + Highly realistic texture on the glass cube, including dust and fingerprints.
- + Excellent lighting and depth of field that creates a cohesive, photographic feel.
- − The plant is inside/behind the cube in a slightly confusing way regarding the cube's back face.
Stable Diffusion 3.5 Large
- + Clean, sharp rendering of the glass and sphere.
- + Good lighting highlights on the surface of the table.
- − Failed to place the book on top of the cube, instead placing the cube on the book.
- − The plant is barely visible and does not clearly interact with the glass as requested.
- − Physics issues where the sphere appears to be floating inside the cube.
Verdict: FLUX.2 [dev] Turbo followed every spatial instruction in the prompt perfectly, placing the red book on top of the glass cube and the blue sphere inside it. In contrast, Stable Diffusion 3.5 Large failed the spatial challenge by placing the cube on top of the book and rendering a floating sphere with no clear ground plane inside the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent adherence to all prompt details including motion blur from passing cars.
- + Highly realistic skin texture and facial details.
- + Accurate representation of the bicycle and repair tools on the ground.
- − The transition between the man's knees and the wet pavement is slightly awkward/clipped.
Stable Diffusion 3.5 Large
- + Strong cinematic mood with effective lighting and reflections.
- + Good depiction of light rain and wet pavement.
- − Missed the 'motion blur from passing cars' instruction as the background car is static.
- − Noticeable anatomy issues with the hands appearing distorted and fused.
- − The man seems to be standing/hovering strangely over the center of the bike frames.
Verdict: FLUX.2 [dev] Turbo followed the prompt much more accurately, successfully incorporating specific elements like motion blur and a 50mm lens look. Stable Diffusion 3.5 Large delivered a moody image but failed on the car motion and had significant structural issues with the man's hands and his physical relationship to the bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent adherence to the 'beads' instruction in the braids
- + Realistic skin texture with natural-looking scars and dirt
- + Superb metal texture with convincing torchlight reflections and scratches
- − The torch in the background is a bit distracting in its placement
Stable Diffusion 3.5 Large
- + Very intricate and beautiful engraving on the plate armor
- + Strong cinematic lighting and composition
- + Good preservation of the 'battle-worn' aesthetic
- − Completely missed the 'small beads' instruction in the hair
- − The background soldiers have a slightly 'melted' look common in older AI generations
Verdict: FLUX.2 [dev] Turbo followed the prompt much more accurately, including specific details like the beads in the hair that Stable Diffusion 3.5 Large missed. While Stable Diffusion 3.5 Large produced beautiful armor engravings, FLUX.2 offered better overall realism in the skin textures and lighting consistency.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent adherence to the grid layout requirement for food photos.
- + Highly legible main headers and body text with professional typography.
- + Very clean, modern layout that feels like a functional, finished graphic design piece.
- − Sub-headers and secondary text contain significant gibberish characters.
- − Layout goes slightly off-grid with the oversized 'Mains' photo compared to the top section.
Stable Diffusion 3.5 Large
- + Vibrant food photography that pops well against the white background.
- + Unique composition that places the text menu in a central column flanked by images.
- − Much poorer text rendering, with almost all menu items being illegible or nonsensical.
- − The layout feels more like a background pattern than a functional restaurant menu.
- − The categories (Appetizrs, Maimaes) include heavy spelling errors.
Verdict: FLUX.2 [dev] Turbo produces a much more realistic and usable menu design with clear sections and a professional grid layout. While Stable Diffusion 3.5 Large offers vibrant food images, its chaotic text rendering and disjointed composition make it feel less like a finished design product.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent text rendering with impressive fiery glow effects
- + Perfectly captures the 'exploded' view with all components floating
- + Photorealistic food textures and dynamic sauce splashes
- − The sauce droplets look slightly stylized compared to the very realistic patty
Stable Diffusion 3.5 Large
- + Strong fiery lighting and atmospheric embers
- + Good texture on the charred meat and melting cheese
- − Completely failed to include the requested text
- − Ignored the 'exploded' instruction, providing a mostly stacked burger
- − The burger is floating but lacks the dynamic separation of ingredients requested
Verdict: FLUX.2 [dev] Turbo followed every instruction, including complex text rendering with a fiery effect and the specific 'exploded' layout. Stable Diffusion 3.5 Large delivered a high-quality visual with good lighting but failed entirely on the text prompt and structural layout of the burger.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent text rendering with no spelling errors.
- + Authentic chalk texture and natural handwriting variations.
- + Precise adherence to the requested date and menu prices.
Stable Diffusion 3.5 Large
- + Cozy cafe atmosphere with good interior composition.
- + Correct date in the subtitle despite other errors.
- − Numerous spelling errors including 'TODAAY' and 'Ottpups'.
- − The text looks more digital than handwritten in several locations.
- − Failed to follow the requested menu layout and content.
Verdict: FLUX.2 [dev] Turbo followed the prompt with near-perfect accuracy, rendering the complex menu items and specific date exactly as requested with a realistic chalk texture. Stable Diffusion 3.5 Large struggled significantly with text legibility and spelling, and failed to maintain the requested handwritten aesthetic throughout.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent photorealistic texture on the space suit and horse hair.
- + Balanced cinematic composition with clear background details.
- + Accurate rendering of complex tackle and horse anatomy.
- − Failed the negative constraint; the astronaut is on top of the horse, not vice versa.
Stable Diffusion 3.5 Large
- + Dynamic sense of movement with the cosmic dust and lighting.
- + High contrast and ethereal aesthetic.
- + Cinematic lighting that blends the subjects into the environment.
- − Failed the negative constraint; the astronaut is on top of the horse.
- − Minor distortion in the horse's facial structure and front legs.
Verdict: Both FLUX.2 [dev] Turbo and Stable Diffusion 3.5 Large failed the specific spatial logic constraint 'horse on top, not vice versa,' interpreting the prompt as a standard horse-riding astronaut. FLUX.2 produced a much sharper image with superior detail in the suit and horse, while Stable Diffusion 3.5 Large offered a more dreamlike, painterly quality but with less anatomical precision.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent adherence to the 'bored' businesswoman in the back seat prompt.
- + Very realistic lighting and taxi interior texture.
- + Captured both paws on the steering wheel as requested.
- − The capybara's hands look slightly more like primate fingers than capybara paws.
- − The taxi sign on top of the car is visible from the inside, which is physically incorrect.
Stable Diffusion 3.5 Large
- + High detail on the capybara's fur and whiskers.
- + Vibrant colors and sharp focus on the main character.
- − Completely failed to include the businesswoman in the back seat.
- − The capybara has human legs and jeans, which was not requested.
- − The capybara is not holding the steering wheel with its paws.
Verdict: FLUX.2 [dev] Turbo followed the complex prompt requirements much better than Stable Diffusion 3.5 Large, correctly including the businesswoman, the bored expression, and the specific driver positioning. Stable Diffusion 3.5 Large missed several key instructions, including the passenger and the requested physical interaction with the steering wheel, despite having high-quality textures.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent typography with perfect spelling of all text elements including date and location.
- + High-quality gothic aesthetic with cohesive lighting and color palette.
- + Superior composition that balances the central jack-o'-lantern with the border and text.
- − The parchment texture is slightly less distinct than the high-contrast paper edges in the competitor.
Stable Diffusion 3.5 Large
- + Strong 'vintage parchment' texture with realistic frayed and burnt edges.
- + Dramatic lighting and high-contrast silhouette elements.
- − Completely failed to include the specific event details (date, time, location).
- − Typography is a bit disjointed and varies significantly in style.
- − The format is slightly vertical despite the 'square format' instruction.
Verdict: FLUX.2 [dev] Turbo is the clear winner as it followed every instruction, including the specific event details and the requested square format. Stable Diffusion 3.5 Large failed to include the date, time, and location, and its overall design feels less like a polished invitation and more like a generic horror poster.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Perfectly follows text instructions with clean, floating typography at top-center.
- + Exceptional realism in the sushi textures and wood grain of the diorama base.
- + Highly clean and professional composition that feels like modern digital art.
- − The 45-degree angle is slightly flattened, leaning more towards a side view than a true top-down isometric view.
- − Includes ginger and wasabi but misses the specific requested 'flag icon' separate from the text.
Stable Diffusion 3.5 Large
- + Excellent adherence to the '3D cartoon' and 'isometric' style with a clear diorama feel.
- + Creative interpretation of the sushi varieties and miniature environment.
- + Includes all requested elements including the flag icon and the diorama base.
- − Failed the text placement instruction, putting the text on a sign within the scene instead of top-center.
- − The scene is significantly more cluttered than the 'minimal garnish' requested.
- − Noticeable artifacts around the flags and some lighting inconsistencies.
Verdict: FLUX.2 [dev] Turbo produced a much cleaner and more professional-looking image that strictly followed the typography instructions. While Stable Diffusion 3.5 Large captured the 'cartoon' and 'isometric' look more vibrantly, its failure to place the text correctly and its cluttered composition made it less successful in meeting the specific prompt constraints.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent anatomical detail and realism across all four animals.
- + Includes all requested species with clear individual characteristics (tabby stripes, fox markings).
- + Superior lighting effects with visible god rays and dew sparkles as requested.
- − The scale of the butterflies is slightly large relative to the animals.
Stable Diffusion 3.5 Large
- + Captures a strong sense of motion and 'chasing' as requested.
- + Whimsical lighting and soft bokeh create a dreamlike atmosphere.
- − The kitten lacks distinct tabby markings and looks more like a generic ginger cat.
- − Significant anatomical artifacts, particularly the butterfly merged with the bunny's ear.
- − Lower overall sharpness and detail compared to Model A.
Verdict: FLUX.2 [dev] Turbo significantly outperforms Stable Diffusion 3.5 Large by delivering high-fidelity textures and accurate anatomy for all four requested animals. While Stable Diffusion 3.5 Large captures the 'chasing' motion well, it suffers from several AI artifacts and fails to produce the specific tabby kitten requested, whereas FLUX.2 provides a polished, 8K masterpiece that perfectly follows the lighting and prompt details.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Perfect adherence to the requested text 'Caffè Florian' with the correct grave accent.
- + Excellent vintage aesthetic with authentic-looking grain and texture.
- + Clean, professional vector-style layout with a well-integrated cloche and banner.
- − The steam is a bit simplified compared to the rest of the detailed illustration.
Stable Diffusion 3.5 Large
- + Good use of the 'Est. 1720' date and classic decorative flourishes.
- + The cloche design includes a nice 'opening' effect to reveal steam.
- − Spelling error in the main name: 'Cafféé' instead of 'Caffè'.
- − The composition feels slightly disjointed with the large gap between the cloche and the banner.
- − The texture is mostly confined to the edges rather than integrated into the design.
Verdict: FLUX.2 [dev] Turbo is the clear winner as it followed the typography instructions perfectly, including the specific accent in 'Caffè'. Stable Diffusion 3.5 Large failed the text requirement by adding an extra 'e' and using the wrong accent, and its overall composition felt less cohesive as a professional logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [dev] Turbo
- + Excellent text rendering with almost perfect spelling of step names and crew members.
- + High prompt adherence, following all sequential steps and technical icons requested.
- + Clean, modern layout that functions well as an educational infographic.
Stable Diffusion 3.5 Large
- + Strong aesthetic appeal with a high-quality flat vector art style.
- + Effective use of the requested NASA-inspired color palette.
- + Good use of negative space and balance in the composition.
- − Failed to follow the requested sequential steps or include correct text.
- − Icons include non-existent planets/rings that don't relate to the Apollo 11 mission.
- − Text is mostly illegible or nonsensical 'lorem ipsum' style.
Verdict: FLUX.2 [dev] Turbo is the clear winner as it successfully created a functional, legible infographic that followed all six requested steps and rendered technical labels accurately. Stable Diffusion 3.5 Large produced a visually pleasing piece of art, but it failed completely on the information architecture, featuring garbled text and irrelevant celestial bodies.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency