Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#58 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photographic realism and lighting
- + Perfect adherence to spatial relationships like the plant behind the glass
- + Superior material textures for wood, glass, and paper
- − The sphere texture is slightly noisy or glittery compared to a smooth sphere
Stable Diffusion 3.5 Medium
- + Follows all core prompt instructions
- + Good color saturation
- − The sphere appears to be levitating unnaturally in the center of the cube
- − Overall image quality is noticeably less realistic and softer in focus
- − Object proportions are a bit clumsy
Verdict: FLUX.1 Kontext [max] produced a high-fidelity, photorealistic image with complex light refractions through the glass, whereas Stable Diffusion 3.5 Medium created a more basic, synthetic-looking scene. FLUX.1 handled the 'partially visible through the glass' instruction much better, making the scene feel grounded and physically accurate.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent anatomical detail in the man's hands and face
- + Superior technical execution of the bicycle mechanics and wet pavement reflections
- + Follows the shallow depth of field and texture prompts very effectively
- − The rain effect looks a bit like a digital overlay rather than a natural mist
- − The man's heritage appears more ambiguous than specifically Japanese
Stable Diffusion 3.5 Medium
- + Captures a more authentic 'candid street photo' atmosphere and lighting
- + Strong adherence to the 'imperfect framing' and 'motion blur' aspects of the prompt
- + Film-like grain and color palette feel natural to a 50mm lens shot
- − Very poor anatomical rendering of the hands
- − The man is standing/leaning rather than 'repairing' the bike
- − The bicycle geometry is warped and incoherent
Verdict: FLUX.1 Kontext [max] produces a much higher quality image in terms of detail and realism, particularly regarding the man's hands and the bicycle's construction. While Stable Diffusion 3.5 Medium does a better job of capturing the specific aesthetic of a gritty street photograph, it fails significantly on anatomical correctness and the 'repairing' action requested.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Sublime, high-frequency skin textures and pore details.
- + Excellent metal engraving detail and realistic torchlight reflections.
- + Superior rendering of hair filaments and lifelike eyes.
- − Missed the request for small beads in the hair braids.
- − The skin appears slightly too oily/sweaty in the highlight areas.
Stable Diffusion 3.5 Medium
- + Stronger adherence to the braided hair request.
- + Effective use of warm lighting against cool metallic shadows.
- + Good depiction of dirt and facial grime.
- − Lower overall resolution and slightly muddy textures on the skin.
- − Engravings on the armor appear more like painted patterns than depth-based metalwork.
- − Hair braids look somewhat plastic and lack individual strand realism.
Verdict: FLUX.1 Kontext [max] provides a significantly more realistic and high-fidelity output with superior texture work on both the skin and the ornate armor. While Stable Diffusion 3.5 Medium followed the braided hair instruction more closely, it suffered from lower clarity and less convincing materials compared to the lifelike quality of FLUX.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional layout with a logical hierarchy
- + High-quality, appetizing food photography
- + Clean typography and effective use of white space
- − Text is mostly gibberish despite looking visually correct
- − Uses script fonts for headers which slightly contradicts the 'bold sans-serif' prompt
Stable Diffusion 3.5 Medium
- + Includes a wider variety of food types matching the sections requested
- + Adheres better to the bold sans-serif font requirement
- − Poor image quality with significant artifacts in the food photos
- − Layout is cluttered and lacks professional alignment compared to Model A
- − Text is extremely distorted and messy
Verdict: FLUX.1 Kontext [max] produces a much more professional and realistic menu design that captures the 'casual dining' aesthetic perfectly, despite some minor text halluncinations. Stable Diffusion 3.5 Medium struggles significantly with visual clarity, presenting distorted food images and a cluttered, unappealing composition.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text integration with the requested glowing/fiery effect.
- + Highly detailed food textures and dynamic explosion effect.
- + Perfect adherence to the starburst and pricing placement.
- − The burger is not fully 'exploded' as much as it is surrounded by floating extra ingredients.
- − Multiple tomato and lettuce layers make the burger look slightly repetitive.
Stable Diffusion 3.5 Medium
- + Strong fiery background with vibrant colors.
- + Clean, modern starburst graphic for the price.
- + Good depth of field on the embers.
- − Failed to create an 'exploded' burger, showing a fully assembled floating burger instead.
- − Small artifacts in the text rendering, such as 'ONI_Y'.
- − Text lacks the requested 'fiery, glowing effect' and feels like a flat overlay.
Verdict: FLUX.1 Kontext [max] far outperformed the competition by accurately rendering all requested text elements with the specific fiery style and starburst requested. Stable Diffusion 3.5 Medium failed to follow the 'exploded' instruction for the burger and had minor typos in the secondary text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text rendering with perfect spelling for all requested items.
- + Convincing chalk texture with smudge marks on the board.
- + Consistent and realistic handwriting style throughout the entire image.
- − The title is in all-caps rather than the 'elegant cursive' specifically requested.
- − The font style, while looking handwritten, is slightly too uniform in some places.
Stable Diffusion 3.5 Medium
- + Successfully attempted a more artistic cursive style for the menu items.
- + Good visual presentation of a chalkboard with a wooden frame.
- − Failed significantly on text legibility and spelling, resulting in garbled words.
- − Incorrectly rendered the date and prices with overlapping symbols.
- − Included digital-looking outline effects rather than pure chalk texture.
Verdict: FLUX.1 Kontext is the clear winner as it perfectly captured the text content and displayed high-quality, realistic chalk textures. Stable Diffusion 3.5 Medium struggled with the prompt's specific text requirements, resulting in numerous spelling errors and incoherent lettering.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent anatomical rendering of the horse and lighting on the spacesuit.
- + Cinematic atmosphere with realistic lens flares and planetary backdrop.
- − Failed the core prompt instruction to have the horse on top of the astronaut.
Stable Diffusion 3.5 Medium
- + Successfully places the astronaut on top of a horse.
- + Correctly interprets the space setting with high-contrast lighting.
- − Failed the core prompt instruction to have the horse on top of the astronaut.
- − Noticeable anatomy errors with the horse's legs appearing elongated and disjointed.
Verdict: Both FLUX.1 Kontext and Stable Diffusion 3.5 Medium failed the spatial logic test requested in the prompt, which specifically asked for the horse to be on top of the astronaut. FLUX.1 Kontext is the better image overall due to its superior anatomical accuracy and cinematic lighting, whereas Stable Diffusion 3.5 Medium suffers from significant artifacts in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealism in taxi interior and car textures
- + High-quality fur rendering on the capybara
- + Correct depiction of both paws on the steering wheel
- − The passenger is on a phone call instead of just looking at it
- − The composition is a bit tight on the capybara's face
Stable Diffusion 3.5 Medium
- + Strong symmetrical composition providing a clear view of both subjects
- + The passenger's expression and pose perfectly match the requested bored look
- + Creative interpretation of the taxi driver emblem on the cap
- − The paws are placed awkwardly on the dashboard/wheel edge rather than gripping a wheel
- − Lighting on the capybara's face is slightly flat compared to the background
Verdict: FLUX.1 Kontext [max] wins on technical image quality and realistic lighting, capturing the textures of the taxi and capybara fur with impressive fidelity. While Stable Diffusion 3.5 Medium followed the character placement and bored expression of the passenger more accurately, its execution of the paws and steering wheel was less coherent.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with nearly perfect spelling and formatting
- + Moody, cinematic lighting that adheres well to the gothic aesthetic
- + Centered jack-o-lantern and thorn border are visually striking
- − Repeated the location text at the very bottom
- − Misinterpreted the scroll banner, placing the text on a dark ribbon instead of a traditional scroll
Stable Diffusion 3.5 Medium
- + Strong parchment texture aesthetic
- + Good incorporation of twisted trees and webs in the background
- − Severe spelling errors throughout the text (e.g., 'Halloweeen', 'Inviloween')
- − Messy text rendering for small details and dates
- − Poorly integrated jack-o-lanterns that look stuck on rather than glowing within the scene
Verdict: FLUX.1 Kontext [max] produced a professional-grade invitation with clean, legible text and a cohesive atmosphere. Stable Diffusion 3.5 Medium struggled significantly with the text requirements, resulting in numerous spelling errors and a cluttered layout.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography rendering with clean, stylish font.
- + Perfectly follows the 45-degree isometric camera angle on a diorama base.
- + High-quality 3D clay-like textures with soft, professional lighting.
- − Missed the small flag icon requested in the prompt.
Stable Diffusion 3.5 Medium
- + Clean, solid blue background as requested.
- + Good material rendering on the sushi roe (PBR style).
- − The text 'SUSHI' is oddly stylized with a white blocky shadow that reduces readability.
- − Failed to include the diorama base, showing just a plate floating.
- − Perspective is not a true isometric 45-degree angle.
Verdict: FLUX.1 Kontext [max] is the clear winner as it faithfully followed the composition instructions, including the isometric 45-degree angle and the diorama base. While both models missed the flag icon, FLUX.1 Kontext [max] produced much cleaner typography and a more cohesive 3D cartoon aesthetic compared to the slightly disjointed presentation of Stable Diffusion 3.5 Medium.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox.
- + Beautiful soft lighting with clear god rays and dew sparkles that match the prompt.
- + Superior fur texture and realistic eye reflections.
- − The animals are sitting still rather than 'playfully chasing' or 'tumbling' as requested.
Stable Diffusion 3.5 Medium
- + Strong, vibrant colors and high contrast.
- + Good quality butterfly rendering.
- + Expressive facial features on the animals.
- − Failed to include all creatures, missing the bunny entirely.
- − Fur looks somewhat digital and over-sharpened compared to the 'ultra-detailed soft fur' request.
- − The light source in the upper right is slightly blown out.
Verdict: FLUX.1 Kontext [max] is the clear winner as it followed the complex prompt by including all four distinct baby animals, whereas Stable Diffusion 3.5 Medium missed the bunny. Furthermore, FLUX.1 Kontext [max] achieved a much more realistic, soft '8K masterpiece' look with convincing lighting effects, while Stable Diffusion 3.5 Medium felt more like a digital illustration.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with perfect spelling and correct accent placement.
- + Follows the minimalist vector emblem style requested in the prompt.
- + Perfectly captures the requested 'Est. 1720' banner and cloche with steam.
- − The steam icon is very simple, though this fits the minimalist request.
Stable Diffusion 3.5 Medium
- + Intricate hand-drawn vintage illustrative style.
- + Good color palette adherence to warm brown and cream tones.
- − Fails on text accuracy with 'Florrian' and 'Est 170'.
- − Layout is cluttered and lacks the requested minimalist aesthetic.
- − The cloche dome is poorly defined and looks like a generic circular frame.
Verdict: FLUX.1 Kontext [max] produced a professional, clean logo that perfectly followed every instruction, including text rendering and specific vector style. Stable Diffusion 3.5 Medium struggled significantly with spelling and provided a cluttered illustration instead of a minimalist logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with clean, readable text for 'APOLLO 11' and the astronaut names.
- + High-quality vector aesthetic with clean lines and professional-looking silhouettes.
- + Good interpretation of the requested color palette.
- − Failed to include all 6 requested steps, merging and omitting several steps of the mission.
- − The rocket design is a generic cartoon rocket rather than a Saturn V icon.
- − The flow of the infographic is confusing and does not follow a clear chronological sequence.
Stable Diffusion 3.5 Medium
- + Successfully attempted all 6 steps of the mission in a logical vertical list.
- + Good use of the NASA-inspired color palette and flat-vector style.
- + Interesting radial composition for the background elements.
- − Text is largely nonsensical and suffers from significant spelling errors ('Lauan strip', 'Tralluunar', 'Trrrllity').
- − Iconography is inconsistent and does not match the specific descriptions requested (e.g., missing Saturn V icon).
- − The visual layout is cluttered and lacks the 'clean, modern' clarity requested.
Verdict: FLUX.1 Kontext [max] produces a much more polished and professional-looking graphic with superior text rendering, though it fails to follow the 6-step logical sequence of the prompt. Stable Diffusion 3.5 Medium attempts the full list of steps and follows the infographic structure better, but it is marred by severe text garbling and poor icon clarity. FLUX.1 Kontext [max] is the preferred choice for its higher visual quality and design coherence.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding