Black Forest Labs' 12-billion parameter flow transformer for high-quality text-to-image generation, suitable for personal and commercial use with streaming support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [dev]
#16 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#57 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [dev]
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photorealistic rendering of glass and light reflections
- + Accurate perspective and object placement
- + High detail in the texture of the book and wooden table
- − The glass structure is more of a hollow pedestal than a simple solid cube
Stable Diffusion 3.5 Medium
- + Accurately places the plant behind the glass as requested in the prompt
- + Correctly identifies all prompt elements
- − Noticeable artifacts on the glass edges and floating sphere
- − Lighting is somewhat flat and overexposed in the background
- − The sphere lacks high-quality refraction and texture compared to the other model
Verdict: FLUX.1 [dev] produced a much more visually compelling and realistic image with superior glass physics and lighting, although it interpreted the cube as a glass frame. Stable Diffusion 3.5 Medium followed the spatial instruction of placing the plant specifically behind the glass better, but the overall image quality is significantly lower with more digital artifacts and less convincing textures.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent anatomical realism in the man's face and hands.
- + High-quality rendering of textures, especially the leather jacket and wet pavement.
- + Accurate interpretation of shallow depth of field and soft bokeh.
- − The motion blur on the passing car is quite subtle compared to the request.
- − The man appears to be holding the bike rather than actively repairing it.
Stable Diffusion 3.5 Medium
- + Stronger 'candid' feel with a more dynamic, imperfect composition.
- + Colors and lighting feel very cinematic, especially the red reflections on the wet street.
- + Captures the atmosphere of light rain and street activity effectively.
- − Noticeable anatomical distortion in the hands and how they grip the bike.
- − The bicycle's geometry is physically inconsistent, especially near the handlebars and seat.
- − Lack of fine detail in the man's facial features compared to Model A.
Verdict: FLUX.1 [dev] is the clear winner due to its superior anatomical accuracy and technical execution, producing a believable repairman with realistic skin and clothing textures. While Stable Diffusion 3.5 Medium captured the 'candid' and 'cinematic' mood well, it failed on basic structural details, particularly with the man's hands and the bicycle's frame.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent skin texture with realistic pores and subtle freckles.
- + Intense, lifelike eye rendering with clear reflections.
- + Clean composition with professional-looking bokeh.
- − Failed to include 'small beads' in the hair braids.
- − Armor lacks the 'ornate engraved' level of detail requested, appearing more like plain hammered metal.
- − Skin appears too clean and modeled for a 'battle-worn' character.
Stable Diffusion 3.5 Medium
- + Successfully incorporated small beads into the complex braided hair.
- + Superior adherence to 'battle-worn' prompt with visible dirt and faint scars.
- + Highly detailed ornate engraving on the plate armor and visible leather/cloth textures.
- − Skin texture on the forehead is slightly blotchy and looks more like digital noise than natural dirt.
- − The warm torchlight is a bit overexposed on the cheek, losing some facial detail.
Verdict: Stable Diffusion 3.5 Medium is the winner for its superior prompt adherence, successfully capturing specific details like the beads in the hair, the engraved armor, and the gritty 'battle-worn' aesthetic. While FLUX.1 [dev] produced a very clean and attractive portrait, it missed several key descriptive elements and felt too polished for the requested concept.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent typographic hierarchy and legibility
- + Very clean and professional white space management
- + High-quality realistic food images that feel consistent
- − Missed the 'grid' requirement for food photos
- − Did not include a specific 'pizza' section label
Stable Diffusion 3.5 Medium
- + Successfully implemented a grid layout for food photos
- + Includes multiple food items as requested
- − Text is largely illegible and uses stylized, inconsistent fonts
- − Layout feels cluttered and less 'minimalist' than requested
- − Pricing and text alignment are messy and unrealistic
Verdict: FLUX.1 [dev] produces a much more professional and aesthetically pleasing menu that looks like a real-world design, despite missing the 'grid' instruction. Stable Diffusion 3.5 Medium adheres better to the grid request but fails significantly on text clarity, whitespace management, and overall minimalist branding.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent 'exploded' view with clear separation of all burger elements.
- + Accurate rendering of the secondary text and price.
- + Clean, professional composition with a magical, starry atmosphere.
- − Completely failed to include the primary title 'MAGIC BURGER'.
- − Lacks the requested fiery starburst for the price.
- − The background is more starry than fiery as requested.
Stable Diffusion 3.5 Medium
- + Successfully integrated all requested text strings including 'MAGIC BURGER'.
- + Includes a starburst graphic for the price as requested.
- + Vibrant fiery background matches the requested aesthetic well.
- − Failed to provide an 'exploded' burger, showing a mostly intact levitating burger instead.
- − Slight text distortion on 'ONLY'.
- − Lower realism on the burger textures compared to Model A.
Verdict: FLUX.1 [dev] produced a much better visual interpretation of the 'exploded' concept and higher image fidelity, but it completely missed the main title of the ad. Stable Diffusion 3.5 Medium followed the text and background instructions far better, though it failed to properly separate the burger components in mid-air. Stable Diffusion 3.5 Medium is the likely winner for better adhering to the specific text and element requirements of the advertising prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent text rendering with almost 100% accuracy to the prompt's requested menu items.
- + The chalk texture looks authentic with realistic dust and varying stroke pressure.
- + Strong composition that maintains a clean, readable café aesthetic.
- − The writing style is a bit too uniform, appearing slightly more like a digital font than natural handwriting in some sections.
- − Minor spelling artifacts in the bottom disclaimer text like 'drish' instead of 'fresh'.
Stable Diffusion 3.5 Medium
- + High level of chalk realism with smudge marks and textured eraser lines.
- + Captured the 'cursive' element of the prompt more aggressively than its competitor.
- − Complete failure in text legibility, with most words being gibberish or heavily misspelled.
- − Did not follow the requested date or price format specified in the prompt.
- − The layout is messy and cluttered, making it look like a rough sketch rather than a professional menu board.
Verdict: FLUX.1 [dev] is the clear winner as it successfully rendered almost all the specific text requested in a readable, attractive format. Stable Diffusion 3.5 Medium struggled significantly with the text-to-image challenge, producing illegible 'hallucinated' words and failing to follow the content instructions in the prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent anatomical rendering of both the horse and the astronaut.
- + Strong cinematic lighting with a soft atmospheric glow from the planet.
- + Clean, professional composition with smooth textures and clear details.
- − Prompt followed incorrectly; the astronaut is riding the horse, but the prompt requested the horse riding the astronaut.
Stable Diffusion 3.5 Medium
- + Dynamic composition with a wide-angle perspective of the planet below.
- + Included some interesting patches of light on the horse's coat.
- − Prompt followed incorrectly; failed the 'horse on top' requirement.
- − Anatomy is poor, with the horse having too many legs or malformed limb structures.
- − Astronaut's legs are floating disconnected from their body.
Verdict: Both models failed the negative constraint and the specific instruction 'horse on top', instead defaulting to the cliché of an astronaut riding a horse. FLUX.1 [dev] is the clear winner as it produced a high-quality, anatomically correct, and cinematic image, whereas Stable Diffusion 3.5 Medium suffered from severe anatomical glitches and distorted limbs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photorealism in the human subject and phone details.
- + Correctly places the woman in the passenger area while the capybara drives.
- + Strong skin textures and realistic lighting from city lights.
- − The woman appears to be in the front passenger seat rather than the back seat.
- − The capybara's paws are not placed logically on the steering wheel.
Stable Diffusion 3.5 Medium
- + Successfully places the passenger in the back seat as requested.
- + The capybara's expression and clothing are very detailed.
- + Correctly centers the capybara as the driver with paws near the wheel.
- − The passenger is not looking at a phone.
- − The passenger is blurry and lacks the 'bored' professional expression requested.
- − The composition feels slightly more like a collage than a single photorealistic shot.
Verdict: FLUX.1 [dev] produces a much higher level of photorealism, especially with the human subject and the tech interaction, but fails the spatial requirement of putting her in the back seat. Stable Diffusion 3.5 Medium follows the spatial layout better by placing the woman in the back, but misses key prompt elements like the phone and high-fidelity textures. FLUX.1 [dev] is the preferred choice for its superior image quality and character rendering.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Strong prompt adherence for specific text strings
- + Clean, professional graphic design aesthetic with consistent lighting
- + Near-perfect spelling of the date, time, and location details
- − The main title text is corrupted into 'Lalloween Rantcl'
- − Lacks the requested parchment texture, appearing more like a modern digital poster
Stable Diffusion 3.5 Medium
- + Excellent 'vintage parchment' texture and gothic atmosphere
- + Includes both webs and thorns as requested in the border
- − Significant spelling errors throughout all text fields
- − Poor composition with text overlapping the parchment edges
- − The year is incorrectly rendered as '226' instead of '2026'
Verdict: FLUX.1 [dev] followed the layout and text instructions much more accurately, producing a usable invitation design despite a small error in the main title. Stable Diffusion 3.5 Medium captured the vintage parchment aesthetic better but failed significantly on text rendering, spelling, and character spacing.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent adherence to the isometric perspective and diorama base request.
- + Includes the requested flag icon and Japanese characters.
- + High-quality soft clay-like textures with clean lighting.
- − Text rendering is slightly messy with 'SUSH CATON' instead of just 'SUSHI'.
- − The diorama base is slightly cropped at the bottom.
Stable Diffusion 3.5 Medium
- + Precise and bold text rendering for 'JAPAN' and 'SUSHI'.
- + Very vibrant colors and clear, high-resolution textures.
- + Perfectly centered square composition.
- − Missed the request for a raised diorama base.
- − Failed to include the flag icon.
- − The perspective is more of a side-angle than a 45-degree top-down isometric view.
Verdict: FLUX.1 [dev] followed the stylistic and compositional instructions much better, correctly including the diorama base, flag icon, and isometric perspective. Stable Diffusion 3.5 Medium produced cleaner text and more vibrant food textures but ignored several key elements of the prompt like the diorama and flag.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent whimsical composition with clear butterfly interaction
- + Clean character designs with soft, appealing lighting
- − Failed to include a tabby kitten, instead showing two puppies and two rabbit-like hybrids
- − Stylistically leaned toward 3D cartoon/animation rather than the requested hyper-photorealism
Stable Diffusion 3.5 Medium
- + Followed the animal variety better by including a kitten and a distinct fox
- + Excellent fur texture and lighting on the animals
- + Colors are vibrant and fit the 'lush wildflower' prompt well
- − Failed to include the baby bunny from the prompt
- − There is a slight anatomical smudging on the kitten's paws
Verdict: While both models failed to include all four requested animals, Stable Diffusion 3.5 Medium came much closer to the requested 'hyper-photorealistic' style with high-quality fur textures. FLUX.1 [dev] produced a charming image, but it resembles a Pixar-style animation more than a realistic scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean vector aesthetic suitable for a logo
- + Excellent alignment and balance
- + Legible 'Est. 1720' text
- − Significant spelling errors in the main name ('Flarilaan') and bottom text ('Reseaurant')
- − Addition of random numbers (11011, 1941) not in prompt
Stable Diffusion 3.5 Medium
- + Strong vintage illustration style with good texture
- + Accurate representation of a cloche dome integrated into the design
- + Closer spelling of 'Caffee' and 'Florrian'
- − Incorrect date text ('Est 170' instead of '1720')
- − Text on the lower banner is garbled and unreadable
- − Cluttered composition compared to the 'minimalist' request
Verdict: While both models failed to provide perfect spelling, FLUX.1 [dev] produced a much more usable 'minimalist' logo with clean vector lines and professional spacing. Stable Diffusion 3.5 Medium captured the vintage texture and cloche motif effectively but failed on the specific date and the minimalist requirement, resulting in a design that is too busy for a standard logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [dev]
- + Follows the requested NASA-inspired color palette perfectly with a professional navy and muted red.
- + Excellent layout balance with a clear central illustrative element and supporting icons at the bottom.
- + Crisp, clean vector lines that accurately reflect the requested modern flat-vector style.
- − The text labels are largely nonsensical gibberish despite the prompt's clear step titles.
- − Includes extraneous celestial bodies like Saturn that were not part of the Apollo 11 mission scope.
Stable Diffusion 3.5 Medium
- + Attempts to follow the sequential list of 6 steps more literally than the comparative image.
- + Uses a clean sans-serif typeface that is very legible.
- − Composition is cluttered and lacks a clear focal point or logical flow.
- − The iconography is inconsistent, featuring wireframe spheres mixed with silhouettes and rough shapes.
- − Fails to capture the 'modern vector' aesthetic, appearing more like a low-fidelity scan of a poster.
Verdict: FLUX.1 [dev] is the clear winner for its superior aesthetic quality and adherence to the requested graphic design style. While both models struggled with generating the exact text requested for the steps, FLUX.1 [dev] produced a balanced, professional-looking infographic with a cohesive color palette, whereas Stable Diffusion 3.5 Medium produced a disjointed and visually messy layout with inconsistent iconography.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding