Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Perfectly follows all spatial instructions, including the book on top and sphere inside.
- + High photographic realism with convincing reflections and soft window lighting.
- + Excellent depth of field with the plant correctly positioned behind the glass.
- − The glass cube has a mirrored base which wasn't specifically requested, though it adds to the aesthetic.
Stable Diffusion 3.5 Large
- + Good realistic textures on the glass and wooden table.
- + Accurately captures the lighting direction from the left.
- − Failed the spatial reasoning of the prompt by placing the sphere on top of the book and the cube over both.
- − The book is inside the cube rather than the cube sitting on the table with the book on top.
- − Composition feels slightly cluttered with the background elements.
Verdict: FLUX.1 Kontext [dev] followed every detail of the prompt perfectly, correctly placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Large struggled with the spatial relationships, placing the book and sphere inside the cube instead of the arrangement requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent high-resolution detail in the man's face and hair.
- + Very clean reflections on the wet asphalt.
- + Realistic bicycle structure and components.
- − The subject is just standing with the bike rather than 'repairing' it.
- − Failed to include the requested motion blur for passing cars.
- − The lighting on the man feels slightly artificial compared to the background.
Stable Diffusion 3.5 Large
- + Stronger adherence to the 'repairing' action and 'candid' feel.
- + Successful motion blur on the background vehicle as requested.
- + Realistic hunched posture and skin texture on the arms.
- − Significant anatomy issues with the man's hands and feet.
- − The bicycle's pedals and chain area are physically incoherent.
- − The rain effect appears as static white streaks rather than integrated moisture.
Verdict: Stable Diffusion 3.5 Large followed the prompt's stylistic instructions much better, capturing the candid action, motion blur, and cinematic framing, but suffered from poor anatomical and mechanical coherence. FLUX.1 Kontext [dev] produced a much cleaner, higher-quality image, but it failed to include the requested motion blur and the subject is posing rather than repairing the bike.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent depiction of warm torchlight highlights on the face and armor
- + Highly detailed engraving patterns on the plate armor
- + Very lifelike skin texture and realistic eye reflection
- − Missed the request for braided hair
- − Lack of visible beads in the hair
- − The 'war-worn' aspect is limited to a couple of small scratches
Stable Diffusion 3.5 Large
- + Successfully incorporated complex braided hair as requested
- + Strong depiction of dirt and grime on the skin for a battle-worn look
- + Excellent texture on the cloth and chainmail underlayers
- − The 'beads' in the hair are not clearly defined or visible
- − The background lighting feels slightly less like a singular 'warm torch' and more like general daylight with fire
- − Skin texture is slightly over-sharpened compared to Model A
Verdict: Stable Diffusion 3.5 Large is the winner as it followed more of the specific prompt instructions, particularly the braided hair which FLUX.1 Kontext [dev] ignored completely. While FLUX.1 Kontext [dev] has slightly more natural skin lighting, Stable Diffusion 3.5 Large captures the 'battle-worn' aesthetic with superior textures on the face and under-armor layers.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent grid composition that integrates images and text seamlessly
- + High-quality, appetizing food photography with vibrant colors
- + Strong use of bold, modern typography that fits the minimalist theme
- − Nonsense text for headers and descriptions
- − Minimal logical structure for specific menu sections requested
Stable Diffusion 3.5 Large
- + Closer adherence to the formal structure of a menu with price points and dividers
- + Very clean, professional white space usage
- + Better clarity on category keywords like 'Appetizrs' and 'Maimaes' (Mains)
- − The grid overflows to the edges, feeling more like a wallpaper than a single menu page
- − Text becomes very compressed and illegible at the bottom
- − Food images are repetitive, mostly showing different versions of pizza
Verdict: FLUX.1 Kontext [dev] creates a more visually compelling and aesthetically modern layout that feels like a premium brand identity, though its text is mostly gibberish. Stable Diffusion 3.5 Large follows the 'menu' prompt more literally by including lists and prices, but its composition is cluttered and the image grid is less integrated. FLUX.1 Kontext [dev] is the winner for its superior visual quality and professional design ethics.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent text integration and legibility
- + Accurate inclusion of price in starburst
- + Clean and professional graphic design layout
- − The burger is not exploded/suspended as requested; it is fully assembled
- − Text has a minor spelling error: 'LNHLY' instead of 'ONLY'
- − Lower photorealistic detail in the burger texture
Stable Diffusion 3.5 Large
- + Incredible photorealistic detail and texture in the food
- + Dynamic lighting and sense of heat from the flames
- + Creative use of fire flowing through the layers
- − Completely failed to include any requested text (title, secondary message, price)
- − The burger is not 'exploded' or separated into mid-air components
- − Lacks the advertisement layout elements requested
Verdict: FLUX.1 Kontext [dev] followed the majority of the prompt instructions, including the specific text and pricing elements, though it failed the 'exploded' layout. Stable Diffusion 3.5 Large produced a much more visually impressive and detailed image of a burger, but entirely ignored the text and advertisement requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Captures a very realistic chalk texture and handwriting style.
- + Accurately renders the complex menu items mentioned in the prompt with high legibility.
- + Follows the request for 'elegant cursive' and natural variations in size.
- − Contains several spelling errors and repetitions like 'Mashroom' and 'with with'.
- − The date rendering is garbled and difficult to read.
Stable Diffusion 3.5 Large
- + Excellent contextual composition showing the cozy café environment.
- + Features a clean and aesthetically pleasing layout for the chalkboard.
- − Significant spelling errors throughout the board (e.g., 'Todaay', 'Cockies').
- − Fails to render the specific menu items correctly, opting for a grid layout not requested.
- − The font looks more like a digital brush than natural chalk handwriting.
Verdict: FLUX.1 Kontext [dev] followed the specific text instructions much better than Stable Diffusion 3.5 Large, capturing the requested menu items even though it suffered from minor typos and repetitions. FLUX.1 also delivered a much more convincing 'hand-drawn chalk' texture, whereas Stable Diffusion 3.5 Large provided a better background environment but failed on the primary task of rendering the specific text-rich menu accurately.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the specific 'horse on top' spatial instruction
- + High clarity and clean rendering of both the horse and the astronaut's face
- + Correct interpretation of the 'surreal' aspect by subverting physics.
- − The transition area where the horse meets the astronaut's back is a bit blurry
- − The background is somewhat plain with simple spheres instead of complex space nebulae.
Stable Diffusion 3.5 Large
- + Beautiful cinematic lighting and atmospheric effects
- + Highly detailed environment with nebulae and planet curvature
- − Failed the negative constraint; the astronaut is riding the horse, not vice versa
- − The horse's legs and the saddle area are messy and structurally incoherent
- − Clipped composition at the top of the helmet.
Verdict: FLUX.1 Kontext [dev] is the clear winner because it successfully followed the complex spatial instruction to place the horse on top of the astronaut. Stable Diffusion 3.5 Large produced a visually stunning scene but completely ignored the core prompt requirement, providing a standard 'astronaut riding a horse' image with significant blurring and anatomical artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the full prompt, including the woman in the back seat.
- + High photorealism with cinematic lighting and realistic textures.
- + Accurately places the capybara's paws on the steering wheel as requested.
- − One of the capybara's paws is resting on its lap rather than both being on the steering wheel.
- − The woman is sitting right next to the driver's seat rather than in the back seat.
Stable Diffusion 3.5 Large
- + Dynamic lighting and vibrant colors that suggest a city at night.
- + Good facial detail on the capybara driver.
- − Completely failed to include the human businesswoman in the back seat.
- − The capybara's paws are not placed on the steering wheel.
- − The scale of the capybara relative to the car seat looks slightly awkward.
Verdict: FLUX.1 Kontext [dev] followed the complex prompt much more effectively, including both the capybara driver and the bored passenger, whereas Stable Diffusion 3.5 Large ignored the passenger entirely. FLUX also achieved a higher level of photorealism and better interpreted the specific request for the capybara's pose and profession.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully rendered the date and time text correctly.
- + Excellent thorny border that follows the prompt.
- + Clear central glowing jack-o-lantern.
- − The smaller text on the scroll banner and location is garbled.
- − Missing the 'parchment' texture, opting for a flat black background.
Stable Diffusion 3.5 Large
- + Beautiful parchment texture and gothic banner illustration.
- + Captures the 'moody night sky' and 'twisted trees' with high artistic quality.
- + Legible text on the scroll banner.
- − Failed to include the specific date, time, and location details entirely.
- − The jack-o-lantern is off-center rather than central as requested.
- − Text rendering for 'Invitation' is very small compared to the header.
Verdict: FLUX.1 Kontext [dev] followed the text requirements more strictly by including the specific event details, though it struggled with the legibility of specific words. Stable Diffusion 3.5 Large produced a much more visually appealing and thematic 'vintage gothic' illustration, but it completely omitted the bottom half of the text prompt. FLUX.1 is better for a functional invitation, while Stable Diffusion 3.5 Large is better for an artistic poster but fails the prompt adherence test regarding text content.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering with clean, large bold text
- + Adheres perfectly to the solid light blue background and minimal garnish requirement
- + Achieves a high-clarity 3D cartoon/plastic aesthetic
- − The 'flag' icon is abstract and does not represent the Japanese flag
- − Interpretation of sushi is very simplistic and toy-like
Stable Diffusion 3.5 Large
- + Features a recognizable Japanese flag and high-quality PBR realistic materials
- + Includes complex textures and miniature details and captures the isometric angle well
- + Shows a variety of sushi items with impressive visual quality
- − Fails the text placement instruction by putting it on a small sign rather than 'top-center'
- − Ignored the 'minimal garnish' request, resulting in a cluttered composition
- − The background has unexpected noise/texture instead of being a solid color
Verdict: FLUX.1 Kontext [dev] followed the layout and aesthetic instructions much more accurately, particularly with the top-center text placement and solid background. While Stable Diffusion 3.5 Large produced more realistic and impressive materials, it failed several negative constraints like 'minimal garnish' and 'solid background', resulting in a cluttered image that missed the requested graphic design style. FLUX.1 is the winner for its superior composition and prompt adherence.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent fur texture and sharpness on the animals
- + Clear and vibrant lighting with a well-defined sun
- + Very clean composition with no major anatomical artifacts
- − Failed to include the bunny and the red fox kit
- − Included two kittens instead of a variety of species
- − Lacks the requested 'dew sparkles' and god rays
Stable Diffusion 3.5 Large
- + Followed the prompt perfectly including all four requested species
- + Beautiful atmosphere with 'dew sparkles' and shimmering light
- + Excellent sense of movement and 'tumbling' energy
- − Slightly lower anatomical clarity on the fox's facial features
- − A bit of AI-style glow/haze that reduces 'hyper-photorealistic' sharpness
Verdict: Stable Diffusion 3.5 Large is the clear winner for its superior prompt adherence, successfully rendering the puppy, kitten, bunny, and fox kit together, whereas FLUX.1 Kontext [dev] omitted half of the subjects. While FLUX.1 Kontext [dev] has slightly more realistic fur textures, Stable Diffusion 3.5 Large captured the magical atmosphere of dew and sunrise light much more effectively.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography with perfect spelling and accent marks
- + Clean vector aesthetic suitable for a professional logo
- + Well-centered and balanced composition
- − Missed the 'banner' requirement for the Est. 1720 text
- − Minimalist style is a bit too modern for a 'vintage' request
Stable Diffusion 3.5 Large
- + Successfully included the banner and vintage texture elements
- + Better atmospheric 'vintage' feel with subtle ornametation
- + Closer adherence to the various prompt descriptors like cloche steam
- − Spelling error in the main text adding an extra 'e' to Caffè
- − Vector execution is slightly messy with floating elements under the cloche
Verdict: FLUX.1 Kontext [dev] produced a much cleaner and professional-looking logo with perfect typography, but it simplified the prompt significantly by ignoring the requested banner and texture. Stable Diffusion 3.5 Large captured the requested 'vintage' atmosphere and banner details better, but failed on the text spelling ('Cafféé') and had less cohesive vector shapes. FLUX.1 is the winner for its functional design quality, despite being less detailed.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Follows the sequential step-by-step numbering (1-6) more closely
- + Better adheres to the 'clean, flat vector' style requested
- + Icons are distinct and separated for a clearer infographic layout
- − Significant spelling errors including the main title 'APOLO'
- − Iconography is very abstract and doesn't clearly represent a Saturn V or Lunar Module
- − Text blocks become garbled and unreadable
Stable Diffusion 3.5 Large
- + More sophisticated and aesthetically pleasing illustration style
- + Includes a recognizable Saturn V-style rocket and planetary bodies
- + Better use of the NASA-inspired color palette with more nuanced design
- − Fails the 'flat-vector' style by including complex textures and detailed globes
- − Does not follow the 6-step sequential instruction or specific icon requests
- − Includes a Space Shuttle-style orbiter which is historically inaccurate for Apollo 11
Verdict: FLUX.1 Kontext [dev] followed the structural requirements for a 6-step infographic, but failed significantly on spelling and icon clarity. Stable Diffusion 3.5 Large produced a much more professional-looking poster with better visual quality, but it ignored the specific step-by-step sequence and included a Space Shuttle which contradicts the Apollo 11 theme. Stable Diffusion 3.5 Large is the preferred choice for its superior composition and detail, despite the historical inaccuracies.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency