Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfect adherence to spatial instructions with the book on top and sphere inside.
- + Sophisticated lighting and caustic effects on the table.
- + Realistic glass refraction and depth of field.
- − The sphere texture looks slightly like foam or glitter rather than a smooth solid.
Stable Diffusion 3.5 Large
- + Clean, sharp edges on the glass cube.
- + Vibrant colors for the sphere and book.
- − Failed the spatial prompt; the book is inside/under the cube and the sphere is on the book.
- − Reflections inside the glass show books that are not present in the scene.
- − Lighting direction is inconsistent with the requested 'from the left' window light.
Verdict: FLUX.1 Kontext [max] followed the prompt instructions perfectly, correctly placing the red book on top of the cube and the blue sphere inside. Stable Diffusion 3.5 Large failed the spatial logic by placing the sphere on the book and the book underneath the cube, while also introducing halluncinated objects in the reflections.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of skin texture and hand details
- + Realistic depth of field and beautiful ground reflections
- + Comes closest to the requested 50mm lens look with a tight, intimate composition
- − The man's ethnicity is somewhat ambiguous compared to the specific prompt
- − The rain effect looks a bit like a vertical streak filter in some areas
Stable Diffusion 3.5 Large
- + Strong adherence to the 'Japanese elderly man' descriptor
- + Good capture of 'imperfect framing' and a wider street scene
- + Effective use of atmospheric lighting and wet pavement
- − The bicycle's geometry is broken (rear wheel/frame connection)
- − The man's hands and arms have anatomical warping and artifacts
- − Failed to incorporate the requested 'motion blur' for the car
Verdict: FLUX.1 Kontext [max] produces a much higher quality image with realistic textures and coherent physical details, particularly in the hands and the bicycle's mechanics. Stable Diffusion 3.5 Large does a better job of capturing the specific cultural look of the subject, but the image suffers from significant anatomical distortions and a physically impossible bicycle frame.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of warm torchlight reflecting off the skin and metal armor.
- + Superior skin texture featuring realistic sweat, dirt, and fine pores.
- + Detailed engraving on the plate armor that looks physically etched.
- − Missed the request for beads in the hair braids.
- − The lighting on the face is slightly oversaturated/reddened.
Stable Diffusion 3.5 Large
- + Stronger adherence to the specific braiding style requested.
- + Great attention to the 'battle-worn' aspect with more prominent facial scaring and grime.
- + Intricate engraving patterns on the chest plate and pauldrons.
- − The eyes look slightly glassier and less lifelike than Model A.
- − The bokeh sparks are less prominent and the overall lighting feels flatter than the requested torchlight effect.
Verdict: FLUX.1 Kontext [max] provides a more immersive atmosphere with superior lighting and texture, capturing the heat of the scene perfectly. Stable Diffusion 3.5 Large adheres more closely to character details like the hair braids but lacks the cinematic depth and skin realism found in the other model.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional layout that looks like a realistic printed menu.
- + High image quality for the food items with clean, uniform backgrounds.
- + Good logical structure with distinct sections and price points.
- − The text is mostly gibberish despite having a professional 'look'.
- − The sections don't clearly label 'Appetizers' or 'Mains' using recognizable English.
Stable Diffusion 3.5 Large
- + Successfully captured specific section headers like 'Appetizrs' and 'Maimaes' (albeit with typos).
- + Strong bold sans-serif typography that matches the prompt perfectly.
- + Vibrant colorful food photos and distinct accents.
- − The layout is cluttered and the food photos are cropped awkwardly at the edges.
- − The vertical column layout is less practical for a real restaurant menu than Image A.
Verdict: FLUX.1 Kontext [max] produces a much more realistic and professional-looking menu layout with higher visual quality in the food photography. While Stable Diffusion 3.5 Large followed the prompt's request for specific sections more closely and used bolder fonts, the overall composition is chaotic and less aesthetically pleasing than the clean, minimalist approach of FLUX.1 Kontext [max].
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to all text requirements with perfect spelling.
- + Achieves the 'exploded' and 'suspended' look requested for most components.
- + Clean, professional graphic design suitable for an advertisement.
- − The main burger stack remains mostly assembled rather than fully exploded.
- − The starburst element for the price is very small and lacks impact.
Stable Diffusion 3.5 Large
- + High visual energy with intense fire and ember effects.
- + Very photorealistic rendering of the meat patties and melted cheese.
- + Good vertical suspension of the burger above the coals.
- − Completely failed to include any of the requested text.
- − The burger is not 'exploded' or separated into its individual components.
- − Lacks the specific starburst element requested.
Verdict: FLUX.1 Kontext [max] is the winner as it successfully integrated all the complex text and layout requirements of the prompt, whereas Stable Diffusion 3.5 Large completely ignored the text instructions. While Stable Diffusion 3.5 Large produced a more visceral and fiery image, it failed to deliver the 'exploded' view and the specific ad copy required for the 'Magic Burger' campaign.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text rendering with perfect spelling for all specific menu items.
- + Realistic chalk texture and smudges on the blackboard surface.
- + Accurate adherence to the requested date of April 30, 2026.
- − The title is in print-style block letters rather than the 'elegant cursive' requested.
- − Composition is a tight framing of the board rather than showing the 'cozy café' interior.
Stable Diffusion 3.5 Large
- + Successfully captures a 'cozy café' atmosphere with seating and decor.
- + Layout is visually balanced and professional for a restaurant setting.
- − Significant spelling errors throughout the text (e.g., 'Todaay', 'Ottpups', 'Cholcalte').
- − Failed to include the correct year (2024 instead of 2026).
- − The complex items requested were heavily mangled or ignored in favor of gibberish text.
Verdict: FLUX.1 Kontext [max] far outperforms its competitor by accurately rendering the specific and lengthy text requirements with perfect spelling and realistic chalk aesthetics. While Stable Diffusion 3.5 Large provides a better sense of the café environment, its inability to follow the specific text prompts and the presence of numerous typos make it less useful for this prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + High resolution and clean textures on the space suit and horse fur.
- + Excellent anatomical accuracy for both the horse and the astronaut rider.
- + Cinematic lighting with a clear bokeh effect in the background.
- − Completely failed the negative constraint to put the horse on top.
- − Standard composition that lacks the requested 'surreal' quality.
Stable Diffusion 3.5 Large
- + Dynamic composition with stardust and cloud effects enhancing the cosmic theme.
- + Interesting cinematic atmosphere with a view of a city-lit planet below.
- − Completely failed the negative constraint to put the horse on top.
- − The horse's front legs exhibit anatomical clipping and distortion.
- − Lower overall clarity compared to the competitor.
Verdict: Both FLUX.1 Kontext and Stable Diffusion 3.5 Large failed to follow the specific spatial instruction 'horse on top, not vice versa,' instead producing the standard 'astronaut riding horse' trope. FLUX.1 Kontext is the superior image due to its significantly higher fidelity, better anatomical rendering, and cleaner details, whereas Stable Diffusion 3.5 Large suffered from anatomical artifacts in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent prompt adherence by including the passenger in the back seat as requested.
- + Higher level of photorealism with natural lighting and realistic fur textures.
- + Includes accurate 'TAXI' text on the driver's cap.
- − The capybara's paw looks slightly fused with the steering wheel.
- − The passenger is holding a phone to her ear rather than looking at it.
Stable Diffusion 3.5 Large
- + Clear depiction of both paws on the steering wheel as requested.
- + Vibrant colors and high-contrast lighting that emphasizes the night scene.
- − Completely failed to include the businesswoman in the back seat.
- − The capybara's anatomy appears somewhat distorted, especially the human-like legs wearing jeans.
- − The texture is overly smooth and looks more like a digital illustration than a photorealistic image.
Verdict: FLUX.1 Kontext is the clear winner because it followed the entire prompt, including the critical element of the bored businesswoman in the back seat. While Stable Diffusion 3.5 Large produced a visually striking image, it completely ignored the passenger and gave the capybara human legs and jeans, failing the photorealism requirement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with perfect rendering of all requested event details.
- + Highly atmospheric cinematic lighting and cohesive gothic art style.
- + Effective use of the border including thorns and webs as requested.
- − The parchment texture is less prominent than Model B.
- − The text is slightly repetitive at the bottom with the location listed twice.
Stable Diffusion 3.5 Large
- + Creative use of a tattered parchment 'paper' effect.
- + Highly detailed illustrative style for the trees and moon.
- + Good adherence to the request for a banner scroll.
- − Failed to include the specific event details (Date, Time, Location) at the bottom.
- − Text rendering on the banner includes garbled characters.
- − The composition feels a bit cluttered compared to a formal invitation.
Verdict: FLUX.1 Kontext [max] is the clear winner as it successfully rendered every piece of custom text requested in the prompt, including the specific date and location. While Stable Diffusion 3.5 Large has a lovely illustrative style and great parchment textures, it failed to follow the instructions regarding the event details.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent 3D cartoon style with soft, refined textures.
- + Clean, large, and perfectly rendered typography.
- + Simple and balanced composition that follows the 'minimal garnish' instruction.
- − Missed the small flag icon mentioned in the prompt.
Stable Diffusion 3.5 Large
- + Successfully included the small flag icon.
- + Higher variety of sushi types on the platter.
- + Realistic PBR materials on the sushi rice and toppings.
- − Failed the typography instruction by placing text on a small sign rather than 'at top-center'.
- − The scene feels cluttered and does not adhere to the 'minimal garnish' request.
- − Chopsticks are disproportionately small and poorly placed.
Verdict: FLUX.1 Kontext [max] captured the requested aesthetic much better, providing a clean, professional 3D graphic with excellent typography. While Stable Diffusion 3.5 Large managed to include the flag icon, it failed the layout and text placement instructions, resulting in a cluttered and less visually appealing composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to all requested species including a clear tabby kitten.
- + Beautiful rendering of god rays and warm golden hour lighting.
- + Very high detail in the fur texture and facial expressions.
- − The composition is a bit static with the animals sitting rather than 'chasing' or 'tumbling'.
- − The butterflies look somewhat flat and illustrative compared to the mammals.
Stable Diffusion 3.5 Large
- + Successfully captures the dynamic action of 'chasing and tumbling' requested in the prompt.
- + Atmospheric bokeh and dew sparkles are more pronounced.
- + Expressions are very joyful and match the 'wholesome' vibe.
- − Failed to include a 'tabby' kitten, opting for a solid ginger/brown kitten instead.
- − The anatomy of the kitten is slightly awkward with thin legs.
- − The fox kit in the background is a bit blurry and less distinct.
Verdict: FLUX.1 Kontext [max] delivered a more accurate representation of the specific species requested, particularly the tabby animal, and had superior lighting and fur detail. However, Stable Diffusion 3.5 Large captured the 'action' part of the prompt much better, showing the animals in motion rather than posing. FLUX.1 Kontext [max] is the winner for its overall technical polish and prompt adherence regarding the characters.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with correct spelling and accent placement.
- + Sophisticated woodcut-style texture that fits the vintage aesthetic.
- + Balanced, professional composition suitable for a real-world logo.
- − The steam is a bit small compared to the cloche.
Stable Diffusion 3.5 Large
- + Stronger emphasis on the steam elements.
- + Warm, aged paper background texture is well-executed.
- − Misspelled the name as 'Cafféé' with double vowels.
- − The composition is bottom-heavy with thin, fragile lines for the 'Est. 1720' section.
- − The cloche icon looks clunky and disconnected.
Verdict: FLUX.1 Kontext [max] creates a much more cohesive and professional logo with perfect text rendering and a high-quality vintage texture. Stable Diffusion 3.5 Large fails on the spelling of 'Caffè' and has a disjointed central icon that lacks the refined vector emblem style requested by the prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with clean, legible text including the names of the astronauts.
- + Strong adherence to the requested flat-vector style and NASA-inspired color palette.
- + Clear, logical visual flow that attempts to follow the numbered steps.
- − The sequence of steps is geographically/physically confusing (Translunar is placed below the moon).
- − The rocket design is a generic cartoon rocket rather than a Saturn V representative.
Stable Diffusion 3.5 Large
- + Includes more intricate technical-looking details that fit the 'infographic' theme.
- + Captures a high-contrast aesthetic that feels more cinematic.
- − Significant text corruption and illegible gibberish throughout the image.
- − Failed the flat-vector style request, opting for a more complex, noisy illustrative style.
- − Includes a Space Shuttle-style orbiter instead of the requested Saturn V or Lunar Module icons.
Verdict: FLUX.1 Kontext [max] is the clear winner as it successfully follows the stylistic instructions for a flat-vector infographic and renders legible, accurate text. While the flow of its diagram is slightly nonsensical, Stable Diffusion 3.5 Large fails significantly on prompt adherence by including an incorrect spacecraft type (a shuttle) and failing to produce any readable text.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency