Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [dev] Black Forest Labs Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [dev]

24.6 arena score

#16 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.9 arena score

#29 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [dev]

0%

win rate

Ties

0%

Stable Diffusion 3.5 Large

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent adherence to the spatial relationship of the book sitting on top of the cube.
  • + High photorealistic quality with believable light refraction and depth of field.
  • + Accurate interpretation of soft window lighting.
  • The glass structure is slightly thicker and more 'frame-like' than a standard hollow glass cube.
  • The blue sphere appears to be floating without a visible support mechanism.

Stable Diffusion 3.5 Large

  • + Successfully includes all prompt elements including the plant visible through the glass.
  • + Realistic textures on the wooden table and book cover.
  • Fails the spatial prompt: the red book is inside/under the cube rather than on top of it.
  • The glass cube looks more like an acrylic display case with visible seams.
  • The composition is a bit cluttered with the background furniture.

Verdict: FLUX.1 [dev] is the clear winner as it correctly followed the spatial instruction to place the book on top of the cube, whereas Stable Diffusion 3.5 Large placed the cube on top of the book. FLUX.1 [dev] also produced a much cleaner, more aesthetically pleasing image with superior lighting and professional-looking depth of field.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent handling of shallow depth of field and bokeh from background lights.
  • + Highly realistic skin textures and clothing materials.
  • + Effective use of atmospheric lighting and pavement reflections.
  • The car in the background lacks the requested motion blur, appearing mostly static.
  • The man seems to be just standing with the bike rather than actively 'repairing' it.

Stable Diffusion 3.5 Large

  • + Stronger sense of 'candid' street photography with a more realistic, less polished composition.
  • + The subject's pose and actions more closely represent the act of repairing a bicycle.
  • + Includes a variety of textures such as the wire basket and wet hair that feel authentic.
  • The rain particles are rendered as very long, unnatural streaks that feel slightly artificial.
  • The anatomy/proportions of the man's arms and hands appear slightly distorted upon closer inspection.

Verdict: FLUX.1 [dev] produces a more cinematic and aesthetically pleasing image with superior skin textures and lighting, but it feels somewhat staged. Stable Diffusion 3.5 Large captures the 'candid' and 'repairing' aspects of the prompt more accurately, though the rain effects are less convincing. FLUX.1 [dev] is the winner due to its significantly higher overall image quality and photographic realism.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent photographic realism in the eyes and skin texture.
  • + Beautiful shallow depth of field and bokeh lighting.
  • + Clean, professional-grade composition.
  • Missed the request for 'ornate engraved' plate armor, opting for simpler hammered metal.
  • Failed to include 'small beads' in the hair braids.
  • The character looks more like a fashion model with freckles than a 'battle-worn' paladin.

Stable Diffusion 3.5 Large

  • + Accurately captured the 'ornate engraved' armor and leather/cloth underlayer detail.
  • + Strong adherence to the 'battle-worn' and 'scars' requirements.
  • + Dynamic braids that match the gritty fantasy aesthetic.
  • Eyes lack the lifelike realism found in FLUX.1 [dev].
  • The lighting is a bit flat compared to the warm torchlight atmosphere requested.
  • Overall image composition is slightly more cluttered with the background army.

Verdict: While FLUX.1 [dev] produced a stunning, high-fidelity portrait, it failed several specific prompt requirements such as the engraving on the armor and beads in the hair. Stable Diffusion 3.5 Large followed the complex text instructions much more accurately, delivering a truly battle-worn warrior in intricately detailed plate armor, making it the better interpretation of the prompt despite slightly lower facial realism.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent layout that remains functional for a real-world restaurant
  • + Accurate typography with legible headers like 'Appetizers' and 'Mains'
  • + Clean minimalist aesthetic that feels professional and spacious
  • Missed the 'grid' request for food photos, grouping them in corners instead
  • Missing a dedicated 'Pizza' section as requested in the prompt

Stable Diffusion 3.5 Large

  • + Successfully implemented a grid layout for food photos
  • + Bold, impactful sans-serif typography for the main headers
  • + High-quality, vibrant food photography that covers variety
  • Text rendering is poor with many typos and nonsensical words like 'MAIMAES'
  • The design feels a bit crowded and lacks the 'minimalist' whitespace of Model A
  • Failed to properly organize sections into Appetizers, Pizza, and Mains as distinct lists

Verdict: FLUX.1 [dev] is the winner because it creates a functional, clean menu design that feels realistic and professional, whereas Stable Diffusion 3.5 Large struggles significantly with text legibility and logical grouping. While Stable Diffusion 3.5 Large followed the 'grid' instruction more closely, the resulting image is cluttered and the text is mostly gibberish compared to the relatively clean output of FLUX.1 [dev].

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent adherence to the 'exploded' request with clear separation of ingredients.
  • + Partial rendering of the specific requested text and pricing.
  • + Clean, professional composition suitable for a high-end food advertisement.
  • Missed the primary title 'MAGIC BURGER' entirely.
  • The 'starburst' for the price is missing, opting for plain text.
  • The background is more starry/magical than the requested fiery/ember-filled atmosphere.

Stable Diffusion 3.5 Large

  • + Stunning photorealistic textures on the burger patties and cheese.
  • + Captures the fiery, glowing ember atmosphere perfectly with intense lighting.
  • + Excellent visual quality and vibrant colors.
  • Failed to provide an 'exploded' view; the burger is mostly assembled.
  • Failed to include any of the requested text elements.
  • Does not meet the 'ad-style' requirements of the prompt.

Verdict: FLUX.1 [dev] followed the structural layout of the prompt much better by creating an exploded view and including some of the required text, though it missed the main title. Stable Diffusion 3.5 Large produced a more visually striking, high-fidelity image with superior lighting, but failed nearly all specific instruction-following regarding layout and text. FLUX.1 [dev] is the preferred choice for a functional advertisement mockup.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent text rendering with nearly perfect spelling of the entire prompt.
  • + Texture of the chalk and handwriting feel very realistic and consistent.
  • + Effective central composition that makes the menu items easy to read.

Stable Diffusion 3.5 Large

  • + Great environmental context, showing the menu within a beautifully rendered café setting.
  • + Aesthetically pleasing layout with greenery and nice lighting.
  • Numerous spelling errors including 'TODAAY' and 'Cholcalte Chip'.
  • The text looks more digital/vector than like actual hand-drawn chalk.
  • Incorrect year (2024 instead of 2026).

Verdict: FLUX.1 [dev] followed the text prompt with incredible precision, rendering almost all requested items and prices correctly in a convincing handwritten style. Stable Diffusion 3.5 Large provided a better overall scene composition but failed significantly on text accuracy, containing many typos and a digital-looking font style rather than the requested chalk texture.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Clean and cinematic lighting effect
  • + Excellent rendering of the spacesuit and horse's mane textures
  • + Composition is balanced and clear
  • Anatomical error with five horse legs visible
  • Failed the logical constraint of the prompt (horse on top)

Stable Diffusion 3.5 Large

  • + Dynamic sense of scale and atmosphere with space clouds
  • + Highly detailed background featuring a planetary curve and nebula effects
  • + Good integration of the rider and saddle
  • Failed the logical constraint of the prompt (horse on top)
  • Horse's face and bridal are slightly distorted
  • Perspective of the astronaut's legs and the horse's back is a bit messy

Verdict: Both models failed the specific spatial reasoning constraint 'horse on top, not vice versa,' instead producing the more common trope of an astronaut riding a horse. FLUX.1 [dev] produced a cleaner, more cinematic image but suffered from a jarring anatomical glitch with five legs, whereas Stable Diffusion 3.5 Large offered more complex background details but had slightly poorer coherence in the subject's anatomy.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Successfully includes the businesswoman in the back seat as requested
  • + Excellent photographic realism and cinematic lighting
  • + Precisely adheres to the instruction of 'both front paws on the steering wheel'
  • The 'TAXI' text on the cap is slightly mirrored or stylized incorrectly

Stable Diffusion 3.5 Large

  • + High detail on the capybara's fur and clothing textures
  • + Strong color saturation that emphasizes the yellow taxi theme
  • Completely fails to include the businesswoman in the back seat
  • The capybara's anatomy appears somewhat distorted with humanoid legs surfacing from the seat
  • Only one paw is near the steering wheel while the other is in its lap

Verdict: FLUX.1 [dev] followed every detail of the complex prompt, including the specific interaction between the driver and the passenger, which is central to the requested scene. Stable Diffusion 3.5 Large failed to include the passenger entirely and struggled with the physical placement of the capybara in the driver's seat.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent thorns and branch detail as requested in the border.
  • + Strong cinematic lighting with a vibrant, glowing jack-o-lantern.
  • + Includes the date and time details exactly as requested.
  • Major text errors including 'Lalloween Rantcl' and redundant '7pm, 7pm'.
  • Lacks the 'webs' mentioned in the border prompt.

Stable Diffusion 3.5 Large

  • + Much better parchment texture and 'vintage' aesthetic.
  • + Captures the spider webs and twisted trees with high fidelity.
  • + Nearly perfect text rendering for the title and the scroll banner phrase.
  • Failed to include the specific event details (Date, Time, Location) at the bottom.
  • The jack-o-lantern is small and not the central focus as requested.

Verdict: FLUX.1 [dev] followed the instruction to include the specific event details like date and location, whereas Stable Diffusion 3.5 Large completely omitted them. However, Stable Diffusion 3.5 Large produced a significantly more atmospheric and visually accurate 'vintage gothic' design with superior legibility for the text it did include, while FLUX.1 [dev] struggled with nonsensical spelling errors and secondary time repetitions.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent adherence to the 'cartoon scene' and 'miniature 3D' aesthetic
  • + Clean and soft PBR material rendering
  • + Accurate placement of text and flag as a graphic overlay
  • Text includes typos like 'SUSH CATON'
  • Sushi variety is limited to six identical pieces

Stable Diffusion 3.5 Large

  • + High variety of sushi types including rolls and nigiri
  • + Very detailed textures on the rice and fish
  • + Successful rendering of the flag icon and text elements
  • Failed the 'top-center text' instruction by placing it on a small sign
  • Perspective is slightly lower than the requested 45-degree isometric angle
  • Composition feels cluttered compared to the 'minimal' request

Verdict: FLUX.1 [dev] followed the stylistic and layout instructions much better, delivering a clean, isometric 3D miniature that looks like a professionally rendered icon, despite minor text errors. Stable Diffusion 3.5 Large produced a more detailed and realistic food scene, but missed the specific layout requirements for the text and the overall minimalist 'diorama' aesthetic.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent character consistency with vibrant colors
  • + Clean composition with a very clear sunrise backlight effect
  • Stylized cartoon/illustration look fails the 'hyper-photorealistic' requirement
  • Missing the kitten entirely, substituting with extra puppies and fox-like creatures
  • Anatomically simplified paws and faces

Stable Diffusion 3.5 Large

  • + Successfully achieves a hyper-photorealistic look with realistic fur texture and lighting
  • + Correctly includes the specific animals requested: puppy, kitten, bunny, and fox kit
  • + Dynamic poses that better reflect the 'chasing and tumbling' part of the prompt
  • Background bokeh is a bit cluttered with white sparkles
  • The fox in the background is slightly out of focus compared to the other animals

Verdict: Stable Diffusion 3.5 Large is the clear winner as it adhered to both the specific animal list and the requested 'hyper-photorealistic' style. FLUX.1 [dev] produced a high-quality image, but it appeared as a 3D animation/cartoon style and failed to include the tabby kitten.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent typography style that fits the vintage theme
  • + Crisp vector-like execution of the cloche and steam icon
  • + Well-balanced vertical composition
  • Significant spelling errors in the brand name ('Flariláan') and 'Restaurant'
  • Inclusion of extra, nonsensical numbers ('11011', '1941')

Stable Diffusion 3.5 Large

  • + Successfully rendered the text 'Caffé Florian' with minimal errors
  • + Excellent application of subtle paper texture in the background
  • + Captures the vintage aesthetic through ornamentation and banner style
  • The 'Est. 1720' text is outside of the banner, which contradicts the prompt
  • The cloche graphic is somewhat clunky with odd internal shapes below the dome

Verdict: Stable Diffusion 3.5 Large is the clear winner for its superior text accuracy, correctly rendering the restaurant name where FLUX.1 [dev] failed with multiple typos. While FLUX.1 [dev] produced a cleaner vector icon, its inability to spell the primary subject and the addition of random numbers makes it unusable as a logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [dev]
Stable Diffusion 3.5 Large

AI Judge Analysis

FLUX.1 [dev]

  • + Stronger adherence to the flat-vector style requested.
  • + Better iconography alignment with the specific mission steps.
  • + Much cleaner layout with more legible, though gibberish, typography.
  • The main rocket is a generic stylized vector rather than a Saturn V.
  • Includes irrelevant planetary bodies like Saturn which were not part of the mission.

Stable Diffusion 3.5 Large

  • + Includes a more complex, technical feel often associated with infographics.
  • + Good color palette adherence.
  • Fails the 'flat vector' style by including detailed, photographic textures on the Moon and Earth.
  • Shows a Space Shuttle or similar craft instead of a Saturn V rocket.
  • Extremely cluttered composition with chaotic line work that doesn't represent a coherent process.

Verdict: FLUX.1 [dev] is the clear winner as it successfully captured the 'flat-vector' aesthetic and the clean infographic layout requested in the prompt. While both models struggled with technical accuracy (Stable Diffusion 3.5 Large generated a space shuttle and FLUX.1 [dev] added extra planets), FLUX.1 [dev] produced a much more professional and visually appealing design that follows the sequential structure of the prompt.

Next steps

Explore each model