Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [pro] Black Forest Labs Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [pro]

20.3 arena score

#41 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.9 arena score

#29 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [pro]

0%

win rate

Ties

0%

Stable Diffusion 3.5 Large

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Perfect adherence to object placement instructions with the book on top and sphere inside.
  • + Extremely realistic wood texture and glass refraction.
  • + Photorealistic soft window lighting that accurately follows the leftward direction.
  • The blue sphere has a slightly fuzzy/felt texture rather than being a smooth geometric solid.

Stable Diffusion 3.5 Large

  • + High-quality glass material with realistic smudges and internal reflections.
  • + Bright, high-contrast lighting with clear shadows.
  • Failed the spatial reasoning of the prompt by placing the book inside the cube and the sphere on top of the book.
  • The plant is mostly above/beside the cube rather than being clearly visible through the glass.

Verdict: FLUX.1 Kontext [pro] followed every spatial instruction in the prompt perfectly, placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Large struggled with the spatial logic, swapping the positions of the sphere and the book, which resulted in a failure of prompt adherence despite its high visual quality.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent skin texture and facial details
  • + Strong sense of rain atmosphere with realistic water droplets
  • + Captures the candid feel and shallow depth of field perfectly
  • The subject is riding the bike rather than actively repairing it as requested
  • Minimal motion blur on the background vehicles

Stable Diffusion 3.5 Large

  • + Shows the subject actively engaged in repairing the bike
  • + Better use of the red color in the bicycle
  • + Good composition that uses the full street background
  • The man's proportions are slightly distorted (hunched posture looks unnatural)
  • Face and hair lack the fine detail and realism seen in Model A
  • Rain effect looks more like post-process streaks than an environmental interaction

Verdict: FLUX.1 Kontext [pro] produces a much more realistic and cinematic image with superior skin textures and lighting, though it missed the specific action of 'repairing'. Stable Diffusion 3.5 Large followed the 'repairing' action more accurately but failed to reach the same level of photorealism, particularly in the rendering of the human subject and the natural falloff of the background.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photorealistic skin texture and lifelike eyes
  • + Beautifully subtle lighting and bokeh effects
  • + High-quality material textures on the leather and fabric
  • The braids are very simple and lack the requested beads
  • Lower contrast on the scars and dirt makes them less prominent

Stable Diffusion 3.5 Large

  • + Complex braided hairstyle that captures the requested 'battle-worn' aesthetic well
  • + Very intricate engraving on the plate armor
  • + Dynamic lighting and clear presence of dirt and scars
  • Overall image has a slightly AI-sharpened, less organic look than Model A
  • Failed to include beads in the hair braids despite the prompt

Verdict: FLUX.1 Kontext [pro] produces a much more lifelike and cinematic portrait with superior skin rendering and soft lighting. However, Stable Diffusion 3.5 Large offers more intricate detail on the armor and a better interpretation of the 'battle-worn' look. FLUX.1 is the winner for its overall realism and superior compositional quality.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent font choice with bold sans-serif headers and consistent alignment.
  • + Matches the specific section requests for Appetizers, Pizza, and Mains perfectly.
  • + Clear, professional layout that feels like a functional menu.
  • The placeholder text for dishes is gibberish.
  • The food photos are slightly repetitive in content compared to the labels (e.g., pizza photo next to appetizers).

Stable Diffusion 3.5 Large

  • + Strong visual grid of food photos that frames the central content.
  • + High-quality, vibrant food photography with good variety.
  • + Captures the 'modern minimalist' aesthetic well.
  • Poor text legibility with significant artifacts in smaller fonts.
  • Mispelled section headers (Appetizrs, Maimaes) and incorrect section categories.
  • The layout is less practical for a real restaurant menu due to extreme verticality.

Verdict: FLUX.1 Kontext [pro] creates a much more functional and accurate menu layout that correctly follows the prompt's section requirements and typography style. Stable Diffusion 3.5 Large offers more vibrant food photography in a creative grid, but fails significantly on text rendering, spelling, and structural logic.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Perfectly renders all requested textual elements with the specified fiery glow.
  • + Excellent photorealistic textures on the brioche bun and melted cheese.
  • + Strong adherence to the exploded layout with clearly suspended components.
  • Repeats the price tag twice, creating some redundancy in the composition.
  • The 'starburst' for the price tag is more of a general glow than a graphic shape.

Stable Diffusion 3.5 Large

  • + Vibrant lighting with high-contrast flames and glowing embers at the base.
  • + Detailed rendering of the meat texture and fresh vegetable ingredients.
  • Completely failed to include any of the requested text elements.
  • The burger is mostly stacked rather than showing a 'dynamic, exploded' view.
  • Overall layout is less like a professional advertisement than Image A.

Verdict: FLUX.1 Kontext [pro] is the clear winner because it successfully integrated all complex text requirements and the 'exploded' compositional style. Stable Diffusion 3.5 Large produced a high-quality visual of a burger over fire but ignored the text and advertisement instructions entirely.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent text rendering with highly accurate spelling and adherence to the specific prompt items.
  • + Realistic chalk texture and natural-looking handwriting variations as requested.
  • + Perfectly captures the date requested (April 30, 2026).
  • The final lines at the bottom contain minor typos ('ous', 'tree' instead of 'free').
  • The cursive requested for the title is more of a print-italic than true elegant cursive.

Stable Diffusion 3.5 Large

  • + Strong composition showing a full cozy café environment with nice lighting and depth.
  • + Good chalk aesthetics with realistic dust and shading on the board.
  • Garbled and incorrect text rendering throughout the menu items.
  • Failed to follow the specific date (rendered 2024 instead of 2026).
  • The font style appears somewhat digital/uniform in the header rather than natural handwriting.

Verdict: FLUX.1 Kontext [pro] is the clear winner for its superior ability to render the specific, complex text requested in the prompt with high accuracy and a realistic chalk texture. While Stable Diffusion 3.5 Large provides a beautiful environmental shot, it fails significantly on the essential text-rendering requirements of the challenge, producing gibberish and incorrect details.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Perfectly adheres to the complex spatial instruction of having the horse on top of the astronaut.
  • + High degree of realism in the textures of the space suit and horse hide.
  • + Includes clever surreal details like the astronaut having hooves for feet.
  • The transition area where the horse is mounted to the astronaut's back is a bit anatomically messy.
  • The small astronaut-like figure on the horse's back adds confusion to the hierarchy.

Stable Diffusion 3.5 Large

  • + High visual appeal with cinematic lighting and a vibrant space background.
  • + Clear and coherent composition of an astronaut on a horse.
  • Completely failed the negative constraint/spatial instruction to have the horse on top.
  • The horse's muzzle has some minor distortion.

Verdict: FLUX.1 Kontext [pro] followed the difficult prompt instruction to flip the traditional dynamic, placing the horse on top of the astronaut, whereas Stable Diffusion 3.5 Large defaulted to the common 'astronaut on a horse' trope. FLUX.1 Kontext [pro] also leaned into the surrealism by giving the astronaut hooves, making it the clear winner for adherence and creativity.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photorealism with cinematic lighting and realistic fur textures.
  • + Captures the woman in the back seat as requested by the prompt.
  • + Consistent focus and depth of field that feels like a professional photograph.
  • The capybara only has one paw visible on the steering wheel instead of both.
  • The human character has a slight distortion on her hand/phone area.

Stable Diffusion 3.5 Large

  • + Vibrant colors and high-contrast night lighting.
  • + Creative clothing choice for the capybara.
  • + Detailed car interior with clear bokeh in the background.
  • Completely missed the 'human businesswoman in the back seat' part of the prompt.
  • The capybara's anatomy is less realistic, appearing more like an anthropomorphic caricature with human-like legs.
  • The 'paws on the steering wheel' are somewhat loosely rendered.

Verdict: FLUX.1 Kontext [pro] is the clear winner as it followed all prompt instructions, including the presence of the businesswoman in the back seat. While Stable Diffusion 3.5 Large produced a high-quality, colorful image, it failed to include the secondary character and created a more stylized, less photorealistic capybara.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent text rendering with correct spelling for the invitation title and specific event details.
  • + Sophisticated cinematic lighting on the jack-o-lantern and professional gothic typography.
  • + Cohesive and elegant composition that feels like a polished invitation.
  • Includes some unnecessary hallucinated text 'Your: Vorkleat: Iight & Spans' before the event details.

Stable Diffusion 3.5 Large

  • + Great 'dark parchment' texture and tattered border effect.
  • + Dynamic and spooky composition with prominent twisted trees.
  • + Legible banner text that matches the requested phrase almost perfectly.
  • Completely failed to include the specific event details (Date, Time, Location) requested.
  • The font for the title is basic sans-serif/serif and lacks the requested 'elegant gothic' style.
  • The jack-o-lantern is small and off-center rather than being the 'central' element.

Verdict: FLUX.1 Kontext [pro] is the clear winner for its superior ability to render the specific event details accurately and its professional-grade gothic typography. While Stable Diffusion 3.5 Large captured the 'parchment' texture very well, it failed to include half of the required text (Date, Time, Location), making it unsuccessful as a specific event invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography integrated naturally into the scene composition.
  • + High-quality 3D miniature Aesthetic with consistent lighting and PBR materials.
  • + Perfectly follows the 'top-center' text layout requirement.
  • The rice grains look more like white bubbles than realistic rice.

Stable Diffusion 3.5 Large

  • + High level of complex detail in the variety of sushi pieces.
  • + Good material texture on the wooden diorama base and chopsticks.
  • + Accurate sushi rice textures.
  • Failed to place text at 'top-center', instead putting it on a small sign.
  • Scene is cluttered and ignores the 'minimal garnish' instruction.
  • The flag icon is a physical prop rather than a graphic element as implied by the prompt.

Verdict: FLUX.1 Kontext [pro] followed the layout and typographic instructions perfectly, capturing the clean, miniature 3D aesthetic requested. Stable Diffusion 3.5 Large produced a more complex image but failed on the specific text placement and minimalism requirements, resulting in a cluttered composition that deviated from the prompt.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent fur texture and sharpness across all animals.
  • + Warm, cohesive color palette with beautiful rim lighting.
  • + Accurate counts for the specific animals requested.
  • The animals are sitting rather than 'playfully chasing and tumbling'.
  • The bunny looks more like a third kitten with long ears.

Stable Diffusion 3.5 Large

  • + Successfully captures the dynamic 'chasing' motion requested in the prompt.
  • + Each animal species is clearly distinct and identifiable.
  • + Dynamic composition with butterflies in the foreground and background.
  • The fox kit has a slightly distorted, flat face.
  • Noticeable artifacts around the retriever's front paw.
  • The lighting is somewhat hazy and lacks the sharp 8K detail seen in the competitor.

Verdict: Stable Diffusion 3.5 Large followed the behavioral instructions much better, capturing the action of chasing butterflies and tumbling. However, FLUX.1 Kontext [pro] produced a significantly higher quality image in terms of texture, lighting, and detail, even though the animals are mostly posing.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography with clean, vintage serif fonts
  • + Accurate spelling of the name and establishment date
  • + Strong minimalist vector aesthetic that perfectly matches the 'emblem style' prompt
  • Simple steam illustration is a bit basic compared to the cloche detail
  • Includes a minor typo in 'EEST.' for 'EST.'

Stable Diffusion 3.5 Large

  • + Elegant warm brown and cream color palette with nice textured background
  • + Creative use of decorative flourishes and multiple steam elements
  • Misspelled the name as 'Cafféé Florian'
  • Layout feels more like a generic postcard or flyer than a clean, minimalist logo
  • The 'Est. 1720' text is tiny and disconnected from the main banner element

Verdict: FLUX.1 Kontext [pro] produced a much more convincing professional logo with superior composition and typography, despite the minor typo in 'EST.'. Stable Diffusion 3.5 Large failed on spelling and created a cluttered design that strayed from the minimalist vector requirement.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [pro]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Legible and clean text rendering for the main title and most labels.
  • + Accurate adherence to the navy, white, and muted red color palette.
  • + Logical infographic layout with a clear numerical flow.
  • Nonsensical inclusion of Saturn-like rings on Earth and the Moon.
  • Incorrectly labels the 'Moon' icon as 'Earth' in the diagram sequence.

Stable Diffusion 3.5 Large

  • + Sophisticated composition that feels like a vintage NASA technical poster.
  • + High-quality vector aesthetic with clean lines and nice textures.
  • + Stronger 'small supporting details' representation as requested in the prompt.
  • Illegible, garbled text for all labels and small supporting data.
  • Poor prompt adherence regarding the specific 6-step sequence, instead showing a jumbled layout.
  • Incorrectly includes a space shuttle-style orbiter which was not part of Apollo 11.

Verdict: FLUX.1 Kontext [pro] creates a much more functional infographic with legible text and a coherent (though flawed) step-by-step logic. Stable Diffusion 3.5 Large has superior artistic appeal and texture, but fails completely on text readability and technical accuracy, including showing a space shuttle instead of the requested Saturn V. FLUX.1 Kontext [pro] is the winner for following the complex instructional nature of the prompt.

Next steps

Explore each model