Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [dev] Black Forest Labs Stable Diffusion 3.5 Medium Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [dev]

17.1 arena score

#54 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Medium

16.8 arena score

#58 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [dev]

0%

win rate

Ties

0%

Stable Diffusion 3.5 Medium

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to spatial relationships and lighting direction.
  • + High-quality rendering of glass reflections and material textures.
  • + Realistic depiction of the green plant through the glass cube.
  • The sphere is quite large relative to the prompt 'small sphere'.

Stable Diffusion 3.5 Medium

  • + Successfully placed all requested elements in the scene.
  • + The sphere size better matches the 'small' descriptor.
  • The sphere appears to be levitating unnaturally in the center of the cube.
  • Lower overall sharpness and more visual artifacts compared to Model A.
  • Lighting is blown out and lacks the 'soft window light' quality.

Verdict: FLUX.1 Kontext [dev] produced a much more realistic and aesthetically pleasing image with convincing physics and lighting. While Stable Diffusion 3.5 Medium followed the prompt's instructions, it suffered from poor composition, levitating objects, and lower image clarity.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent high-resolution detail on the man's face and clothing
  • + Clear and vibrant rendering of the red bicycle
  • + Follows the shallow depth of field requirement well
  • The subject is posing with the bike rather than 'repairing' it
  • Lacks the requested 'motion blur' for passing cars
  • The composition is too centered and clean for a 'candid street photo' with 'imperfect framing'

Stable Diffusion 3.5 Medium

  • + Successfully captures the 'repairing' action and candid posture
  • + Excellent adherence to the 'imperfect framing' and 'candid street photo' aesthetic
  • + Realistic film-like texture and moody lighting
  • The bicycle geometry is distorted and physically nonsensical
  • Low visual fidelity on the face and hands
  • The rain looks more like static or noise than droplets

Verdict: Stable Diffusion 3.5 Medium much better captures the intended atmosphere, composition, and specific action (repairing) of the prompt, whereas FLUX.1 Kontext [dev] creates a generic, clean portrait. However, FLUX.1 Kontext [dev] is the winner because Stable Diffusion 3.5 Medium suffers from severe anatomical and mechanical distortions that make the image look broken upon closer inspection.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent anatomical realism in the face and eyes
  • + Highly detailed and readable engravings on the plate armor
  • + Subtle and realistic scarring on the face
  • Missed the request for braided hair
  • Torchlight effect is a bit flat across the face

Stable Diffusion 3.5 Medium

  • + Successfully incorporated braided hair with beads
  • + Strong, dynamic lighting from a specific torchlight source
  • + Excellent depiction of dirt and weathering on the skin
  • Armor engravings are messy and lack structural logic compared to Model A
  • Iris patterns in the eyes look slightly artificial or over-sharpened

Verdict: Stable Diffusion 3.5 Medium followed the specific details of the prompt better, specifically the braids and beads which FLUX.1 Kontext [dev] omitted. However, FLUX.1 Kontext [dev] produced a more polished and anatomically coherent portrait with superior armor craftsmanship. Stable Diffusion 3.5 Medium is the winner here for its better prompt adherence and more dynamic atmosphere.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong, high-quality food photography with vibrant colors.
  • + Excellent use of bold sans-serif typography that feels modern and professional.
  • + Clean, balanced grid layout that fits the minimalist aesthetic.
  • The text is nonsensical and has some artifacting/glitches over characters.
  • Lacks a clear list of actual food prices or a structure that resembles a functional menu page.

Stable Diffusion 3.5 Medium

  • + More realistic menu structure with headers, item descriptions, and pricing columns.
  • + Includes a wider variety of images that clearly represent different menu categories like pizza.
  • + Better overall organization as a practical multi-section document.
  • Image quality is low with significant muddy textures and artifacts.
  • The layout feels cramped and the typography is less impactful than requested.
  • Food photography appears dull and lacks the 'vibrant' quality requested.

Verdict: FLUX.1 Kontext [dev] excels in visual quality and professional graphic design, producing a stunning aesthetic with high-end photography, though it fails to create a functional text layout. Stable Diffusion 3.5 Medium follows the structural requirements of a menu more closely with sections and price points, but suffers from poor resolution and unappealing image clarity. FLUX.1 Kontext [dev] is the likely winner for its superior adherence to the 'modern minimalist' and 'vibrant' style prompts.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography rendering with the requested glowing effect.
  • + Good layout that feels like a professional advertisement.
  • + Clear, vibrant colors and sharp focus on the burger.
  • Failed the core 'exploded burger' prompt by showing a fully assembled burger.
  • Spelling error in 'ONLY' (rendered as 'LNHLY').

Stable Diffusion 3.5 Medium

  • + Very high photorealistic texture on the bun and patty.
  • + Dynamic background with realistic fire and embers.
  • + Correct spelling of the promotional text.
  • Failed the 'exploded burger' requirement, showing a mostly assembled burger.
  • The starburst element is stylized as lines rather than a solid starburst shape.
  • The text is small and difficult to read compared to the background.

Verdict: Both models failed the specific 'exploded burger' instruction, rendering whole burgers instead of suspended components. FLUX.1 Kontext [dev] produced a more effective advertisement layout with large, glowing text, though it suffered from a typo; Stable Diffusion 3.5 Medium achieved higher realism in the food textures but the text integration was less impactful for an ad.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent chalk-like texture on the lettering.
  • + Mostly legible text with few spelling errors.
  • + Clean and realistic composition that mimics a real cafe chalkboard.
  • Has repetitive text errors like 'Mushroom Mashroom' and 'with with'.
  • Date and some menu spelling ('Risoktso', 'Octpus') are garbled.
  • Text style is printing-heavy rather than the requested elegant cursive for the title.

Stable Diffusion 3.5 Medium

  • + Features more chalk smudges and realistic chalkboard degradation.
  • + Attempted artistic swirls and decorative accents in the lettering.
  • Text is largely illegible and contains numerous spelling failures.
  • The date and pricing are completely nonsensical.
  • The layout is cluttered and fails to follow the requested menu structure.

Verdict: FLUX.1 Kontext [dev] is the clear winner as it produces a largely readable menu that closely follows the prompt's structural requirements, despite some minor spelling repetition. Stable Diffusion 3.5 Medium fails significantly on prompt adherence, resulting in a jumbled mess of illegible text and confusing layout.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to the 'horse on top' instruction
  • + High-quality rendering of the astronaut's face and suit textures
  • + Clear, cinematic lighting with a minimalist background
  • The horse appears to be floating behind rather than clearly 'riding' the astronaut
  • The scale of the horse relative to the astronaut is slightly small for a riding pose

Stable Diffusion 3.5 Medium

  • + Beautiful background with stars and planet detail
  • + Dynamic composition with a sense of movement in the horse's mane
  • Completely failed the negative constraint/spatial instruction, showing the astronaut on top
  • Anatomical issues with the horse's legs, which appear elongated and tripled in sections
  • Lower facial detail due to the dark visor

Verdict: FLUX.1 Kontext [dev] followed the specific and difficult spatial instruction to place the horse on top of the astronaut, whereas Stable Diffusion 3.5 Medium defaulted to the standard trope of an astronaut riding a horse. FLUX.1 Kontext [dev] also produced much cleaner anatomy and higher overall image quality, avoiding the limb artifacts present in the other model.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent photorealism and texture on the capybara fur and the car interior.
  • + Successfully captures the specific requested expression of a bored businesswoman looking at her phone.
  • + Realistic lighting that matches a nighttime city environment.
  • The capybara only has one paw on the steering wheel instead of both.
  • The character looks more like a groundhog/beaver hybrid than a classic capybara.

Stable Diffusion 3.5 Medium

  • + Successfully placed both paws on the steering wheel as requested.
  • + The capybara anatomy is more accurate to the species.
  • + Good use of depth of field and vibrant city lights in the background.
  • Failed to include the phone or the specific 'bored' expression for the passenger, who is just blurred and looking forward.
  • The composition feels slightly cramped with the capybara's head appearing very large relative to the car.

Verdict: FLUX.1 Kontext [dev] followed the complex narrative instructions much better, correctly depicting the woman looking at her phone with a bored expression to complete the absurd scene. While Stable Diffusion 3.5 Medium captured the capybara's likeness and paw placement better, it failed the secondary requirement of the passenger's actions.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography for the main title
  • + High-quality central jack-o-lantern illustration
  • + Clean and symmetrical border design
  • Text in the scroll and the bottom location name are garbled
  • Lacks the 'parchment' texture requested, opting for a solid black background

Stable Diffusion 3.5 Medium

  • + Successfully captured the dark parchment and moody night sky textures
  • + Creative composition with twisted trees and spider webs
  • + Good interpretation of the 'vintage' aesthetic
  • Significant spelling errors throughout all text layers
  • Layout feels cluttered with the multiple jack-o-lanterns and misaligned text
  • Failed to include the specific 'scroll banner' design requested

Verdict: FLUX.1 Kontext [dev] produced a more polished and professional-looking poster with superior lighting, although it failed to render the parchment background. Stable Diffusion 3.5 Medium followed the thematic elements of the prompt more closely (parchment, sky, webs), but suffered from severe legibility issues and poor text rendering compared to the competitor.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography alignment and font clarity
  • + Perfectly captures the soft, 3D cartoon miniature aesthetic
  • + Very clean composition with a solid background and smooth textures
  • The flag icon is stylized but not a recognizable Japanese flag
  • The sushi anatomy is slightly abstract

Stable Diffusion 3.5 Medium

  • + More realistic texture on the nori and roe
  • + Good isometric perspective and lighting
  • Failed to include 'JAPAN' in large bold text as primary focus
  • Missing the requested flag icon entirely
  • Text rendering is slightly messy with artifacts on the letters

Verdict: FLUX.1 Kontext [dev] followed the layout and typography instructions much better, producing a clean, professional-looking graphic that matches the '3D cartoon' prompt perfectly. Stable Diffusion 3.5 Medium failed to prioritize the 'JAPAN' text and completely omitted the flag icon, while also struggling with text clarity.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent sense of motion and playfulness
  • + Realistic lighting and soft, believable fur textures
  • + High level of overall image clarity
  • Failed to include the baby bunny or red fox kit
  • Included two cats instead of the requested animals

Stable Diffusion 3.5 Medium

  • + Successfully included the red fox kit with distinct features
  • + Vibrant colors and beautiful backlighting
  • + Good facial expressions on the animals
  • Failed to include the baby bunny
  • Fur textures appear slightly over-sharpened and less 'fluffy' than Model A

Verdict: Both models failed to include all four requested animals, with both missing the baby bunny entirely. However, Stable Diffusion 3.5 Medium followed the prompt better by including the fox kit, whereas FLUX.1 Kontext [dev] produced a higher-quality, more realistic image despite only rendering a puppy and two kittens.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography and spelling of 'Caffè Florian'.
  • + Highly clean, minimalist vector aesthetic suitable for a modern logo.
  • + Accurate representation of the cloche dome and steam.
  • Missed the 'banner' element for the 'Est. 1720' text.
  • The typography is perhaps too modern/bold for a 'vintage' style.

Stable Diffusion 3.5 Medium

  • + Strong vintage/retro illustrative style with good texture.
  • + Successfully incorporated the requested banner element.
  • + Captures the 'warm brown and cream' color palette well.
  • Several spelling errors in the text, including 'Florrian' and 'Est 170'.
  • The composition is cluttered and loses the 'minimalist' requirement.
  • Illegible gibberish text on the bottom banner.

Verdict: FLUX.1 Kontext [dev] followed the text instructions perfectly and produced a professional, clean logo, though it leaned more modern than vintage. Stable Diffusion 3.5 Medium captured the requested vintage texture and banner elements better, but failed significantly on spelling and the 'minimalist' constraint. FLUX.1 is the winner for its functional design and perfect text rendering.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [dev]
Stable Diffusion 3.5 Medium

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Stronger flat vector aesthetic following the specified iconography style.
  • + Clean, modern typography that feels more like a deliberate poster.
  • + Better adherence to the 'navy, white, and muted red' palette.
  • Misspelled the primary subject title as 'APOLO'.
  • The icons are highly abstract and do not clearly represent the Saturn V or the Lunar Module as requested.
  • Text below headings is mostly garbled gibberish.

Stable Diffusion 3.5 Medium

  • + Uses a clear vertical steps layout that mimics actual mission phases.
  • + Clean line art for the planet/globe icons.
  • + Text is slightly more legible in specific labels like 'Launch' and 'Landing'.
  • Included a large red circle that feels out of place with the color palette and theme.
  • Iconography does not follow the requested stages (e.g., repeating globes for different steps).
  • Background is cluttered with noise/stars that detract from the 'clean, flat' vector requirement.

Verdict: FLUX.1 Kontext [dev] captured the requested modern vector aesthetic and color palette much better than Stable Diffusion 3.5 Medium, despite a spelling error in the title. Stable Diffusion 3.5 Medium followed the sequential layout better but failed significantly on the specific icon requests and the flat vector style, resulting in a cluttered composition.

Next steps

Explore each model