Black Forest Labs' state-of-the-art image generation model with maximum quality and speed, supporting text-to-image and multi-reference image editing with up to 4MP output
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [pro]
#8 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [pro]
62.5%
win rate
Ties
0.0%
Stable Diffusion 3.5 Large
37.5%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [pro]
- + Perfect prompt adherence with spatial relationships correctly placed.
- + Excellent lighting and depth of field, creating a very realistic photographic look.
- + Material properties are convincing, with accurate reflections and transparency in the glass.
- − The plant in the background is quite blurred, though still clearly identifiable.
Stable Diffusion 3.5 Large
- + High level of detail in the glass texture, including realistic smudges and scratches.
- + Vibrant colors and sharp rendering of objects.
- − Failed the prompt's spatial instructions by placing the book under the sphere instead of on top of the cube.
- − Logical inconsistency: the book appears to be both inside and outside the glass base simultaneously.
Verdict: FLUX.2 [pro] followed every spatial instruction perfectly, correctly placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Large failed the spatial arrangement, placing the book at the base and erroneously putting the sphere on top of the book, while also exhibiting clipping issues where the book meets the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [pro]
- + Exceptional photographic realism with natural skin textures and visible pores.
- + Realistic mechanical details on the bicycle and authentic rain droplets on surfaces.
- + Superb handling of lighting and reflections on the wet pavement.
- − The motion blur on the background car is subtle rather than pronounced.
Stable Diffusion 3.5 Large
- + Good composition that captures the 'candid' and 'imperfect framing' requested.
- + Successfully incorporates motion blur on moving vehicles in the background.
- + Strong atmosphere with visible rain streaks.
- − The man's hands have significant anatomical distortions (merged fingers and strange shapes).
- − The bicycle's mechanical structure is nonsensical in several places.
- − Lower overall sharpness and texture quality compared to the other model.
Verdict: FLUX.2 [pro] produces a significantly more realistic and technically sound image, with incredible attention to skin texture, bicycle mechanics, and the physics of water. Stable Diffusion 3.5 Large succeeds in creating a more dynamic sense of motion and 'imperfect' street photography framing, but it fails significantly on anatomical details and mechanical coherence.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent adherence to hair beads and lighting prompts.
- + Superior texture on leather straps and clothing underlayers.
- + Stronger 'close portrait' composition with cinematic bokeh sparks.
- − The scars look a bit like digital paint or face paint rather than deep physical wounds.
Stable Diffusion 3.5 Large
- + Very intricate engraving details on the plate armor.
- + Complex hair braiding style.
- + Good depiction of dirt and weathering on the face.
- − Failed to include the requested beads in the hair.
- − Lighting feels flat and daylight-based rather than warm torchlight.
- − Less focus on the requested leather and cloth textures.
Verdict: FLUX.2 [pro] followed the prompt more comprehensively, successfully including the specific details of hair beads and warm torchlight reflections. While Stable Diffusion 3.5 Large produced beautiful armor engravings, it missed several key prompt elements and provided a broader shot rather than the requested close portrait.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent adherence to the grid layout with distinct sections for different food categories.
- + Clean, legible typography that mimics a real-world professional menu.
- + Highly realistic food photography that fits the restaurant aesthetic well.
- − Contains several spelling errors (e.g., 'MINS' instead of 'MAINS', 'MageFiza').
- − Logic errors in pricing and content pairing, such as 'Garlic Bread' appearing in every section including 'Mins' for $40.
Stable Diffusion 3.5 Large
- + Strong aesthetic appeal with a high-contrast minimalist look.
- + Effective use of a colorful photo grid bordering the text.
- + Bold, impactful sans-serif headline typography.
- − Text rendering is poor with significant gibberish throughout the body and subheaders.
- − The layout is less practical as a functional menu, feeling more like a poster than a list of items.
- − The cropping on the sides suggests a tiled pattern rather than a finished page design.
Verdict: FLUX.2 [pro] produced a far more professional and usable menu layout with realistic food images and a clear organizational structure, despite some minor typos in section headers. Stable Diffusion 3.5 Large created a visually interesting artistic composition, but failed significantly on legibility and the practical requirements of a menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [pro]
- + Perfect adherence to all text requirements including price and secondary message
- + Excellent 'exploded' view with clearly suspended, distinct ingredients
- + Clean professional layout that looks like a finished advertising piece
- − The lighting on the top bun is slightly flat compared to the fiery surroundings
Stable Diffusion 3.5 Large
- + Highly vibrant fire and ember effects across the entire image
- + Appetizing textures on the burger patties and melting cheese
- − Completely failed to include any of the requested text
- − Missing the 'exploded' view; the burger is mostly assembled rather than suspended in components
- − The composition feels cluttered with the fire overlapping the food
Verdict: FLUX.2 [pro] followed the prompt instructions perfectly, successfully integrating all three requested text elements and the 'exploded' visual style. Stable Diffusion 3.5 Large failed to include any text or the exploded effect, resulting in a standard burger image that does not meet the user's specific advertising requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent text accuracy, rendering all requested menu items perfectly.
- + The chalk texture is highly realistic with dust and smear effects.
- + Followed the elegant cursive style and specific date accurately.
- − The 'chalkboard' looks more like a piece of slate than a traditional framed board.
- − The background is very blurred, sacrificing some cafe context.
Stable Diffusion 3.5 Large
- + Provides a wider environmental shot showing more of the cozy café setting.
- + Good layout with decorative chalk frames.
- − The text is full of spelling errors and gibberish (e.g., 'TODAAY', 'Ottpups', 'Cholcalte').
- − Failed to render the specific date requested (2024 instead of 2026).
- − The handwriting looks more like a generic digital font than natural chalk.
Verdict: FLUX.2 [pro] followed the prompt instructions perfectly, rendering the text with 100% accuracy and a very convincing chalk texture. Stable Diffusion 3.5 Large struggled significantly with the text rendering, introducing numerous misspellings and failing to provide the requested date or the specific cursive handwriting style.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellently follows the difficult spatial constraint of the horse being on top.
- + High level of detail in the astronaut suit and cinematic lighting.
- + Clean, sharp composition with no visible blurring artifacts.
- − The anatomy where the horse meets the astronaut is slightly confusing and fragmented.
Stable Diffusion 3.5 Large
- + Great sense of movement and 'cosmic dust' atmosphere.
- + High-quality rendering of the Earth's surface and clouds.
- − Failed the negative constraint entirely by placing the astronaut on top of the horse.
- − The astronaut's hands and the horse's reins are poorly integrated and look messy.
Verdict: FLUX.2 [pro] followed the specific, surreal instruction to put the horse on top of the astronaut, which is a significant win for prompt adherence. Stable Diffusion 3.5 Large defaulted to the more common trope of an astronaut riding a horse, failing the core challenge of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent photorealism with cinematic lighting and realistic textures.
- + Perfect adherence to all prompt elements, including the passenger and expression.
- + High-quality rendering of the car interior and raindrops on the window.
- − The capybara's hands look slightly more like human hands wearing gloves than animal paws.
Stable Diffusion 3.5 Large
- + Bright, vibrant colors and sharp focus on the capybara.
- + Clear depiction of the jacket and hat textures.
- − Completely failed to include the businesswoman in the back seat.
- − The capybara's anatomy is distorted with limbs emerging awkwardly from its chest.
- − The scale of the capybara relative to the car seat feels unnatural.
Verdict: FLUX.2 [pro] followed the prompt instructions perfectly, capturing the surreal scene with high photorealism and including both the capybara driver and the bored passenger. In contrast, Stable Diffusion 3.5 Large omitted the passenger entirely and suffered from anatomical glitches regarding the animal's limbs.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [pro]
- + Perfect text rendering for all requested details including date, time, and location.
- + High-quality cinematic lighting with a realistic 3D jack-o-lantern.
- + Excellent adherence to the border description with webs and thorns.
Stable Diffusion 3.5 Large
- + Strong vintage parchment aesthetic for the background.
- + Good use of negative space in the center for the moon and bats.
- − Failed to include the specific event details (Date, Time, Location) at the bottom.
- − Text rendering on the banner and title is inconsistent and blocky.
- − Jack-o-lantern is small and off-center, contrary to the 'central' prompt requirement.
Verdict: FLUX.2 [pro] followed every part of the prompt, including the complex text instructions and centralized composition. Stable Diffusion 3.5 Large missed the majority of the event details and struggled with overall text legibility and layout balance.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [pro]
- + Perfectly adheres to the text placement and style requirements.
- + Captures the clean, 3D cartoon, isometric aesthetic flawlessly.
- + Uses a minimalist raised diorama base as requested.
- − The stylized rice grains are quite chunky and abstract compared to the 'realistic PBR' instruction.
Stable Diffusion 3.5 Large
- + Excellent realistic PBR textures, especially on the salmon and rice.
- + High level of detail in the food rendering.
- − Failed to place text at the 'top-center' as requested, opting for a small sign instead.
- − Included extra elements like chopsticks and side bowls not specified in the prompt.
- − The isometric angle is less precise than Model A.
Verdict: FLUX.2 [pro] followed the complex layout instructions much better than Stable Diffusion 3.5 Large, accurately placing the bold text at the top and maintaining a clean isometric 3D cartoon style. While Stable Diffusion 3.5 Large produced more realistic textures for the food, it ignored several negative and positioning constraints, resulting in a cluttered composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent fur detail and texture across all three subjects
- + Beautiful lighting with well-defined dew sparkles and god rays
- + Dynamic and engaging composition with animals interacting directly
- − Missed the baby bunny requested in the prompt
- − Large dog paws appear slightly anatomically odd in their reach
Stable Diffusion 3.5 Large
- + Included all four requested animals (puppy, kitten, bunny, fox)
- + Strong sense of movement and 'chasing' as requested
- + Bright, joyful color palette
- − The tabby kitten's appearance is closer to a generic ginger/brown kitten than a distinct tabby
- − Lower level of fine detail in the fur and background compared to Model A
- − Blurry artifacts on the butterfly wings
Verdict: Stable Diffusion 3.5 Large followed the prompt more accurately by including all four animals, whereas FLUX.2 [pro] missed the baby bunny entirely. However, FLUX.2 [pro] produced a much higher quality image with superior textures, more realistic lighting, and more expressive character designs, making it the more visually impressive output despite the missing element.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [pro]
- + Perfect adherence to text and accent marks in 'Caffè Florian'.
- + Clean vector emblem style with sophisticated cross-hatching detail.
- + Professional composition that looks like a real minimalist logo.
- − The 'Est. 1720' banner is slightly small relative to the main text.
Stable Diffusion 3.5 Large
- + Good use of vintage parchment-style texture on the background.
- + Followed the instruction for the 'Est. 1720' text.
- − Included an extra letter in the name, spelling it 'Cafféé'.
- − The cloche illustration is disjointed with a strange pipe/arm detail appearing under the dome.
- − The composition feels cluttered and lacks the minimalist vector aesthetic requested.
Verdict: FLUX.2 [pro] produced a high-quality, professional logo that perfectly followed all text and stylistic requirements. In contrast, Stable Diffusion 3.5 Large suffered from a spelling error in the main brand name and a messy, nonsensical illustration of the cloche dome.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [pro]
- + Excellent layout that follows the chronological steps requested
- + Clean vector aesthetic with readable main headings
- + Good adherence to the requested NASA-inspired color palette
- − Nonsense filler text for the minor descriptions
- − The 'Launch' icon resembles a generic shuttle rather than a Saturn V
Stable Diffusion 3.5 Large
- + Includes a high level of visual detail and technical-looking callouts
- + Captures the NASA aesthetic through colors and vintage-style diagramming
- − Failed to provide a clear step-by-step infographic layout
- − Text is largely illegible or misspelled (e.g., 'Lannch')
- − Composition is cluttered and chaotic compared to the clean layout requested
Verdict: FLUX.2 [pro] followed the prompt's structural requirements much better, creating a clear vertical progression from launch to landing with distinct icons. Stable Diffusion 3.5 Large produced a more cluttered, non-linear composition that ignored the specific six-step instruction and suffered from significant text distortions.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency