Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] Black Forest Labs FLUX.2 [klein] 4B Black Forest Labs

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell]

19.2 arena score

#44 of 62 in Text-to-Image

Skill signature · Text-to-Image

FLUX.2 [klein] 4B

22.3 arena score

#32 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell]

0%

win rate

Ties

0%

FLUX.2 [klein] 4B

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + High visual clarity and vibrant colors.
  • + Accurately represents the spatial arrangement of the plant behind the glass cube.
  • + Follows the lighting direction from the left well.
  • Included extra elements not in the prompt, specifically a second blue sphere on top of the book.
  • The sphere inside the cube appears to be floating unnaturally.

FLUX.2 [klein] 4B

  • + Strictly followed the quantity of objects mentioned in the prompt.
  • + Realistic interaction between the sphere and the bottom surface of the cube.
  • + Excellent texture work on the red book cover.
  • The glass refractive effects on the plant behind it are slightly less convincing than in Model A.
  • The cube structure is very thin, appearing more like a plastic container than heavy glass.

Verdict: FLUX.2 [klein] 4B followed the prompt more accurately by only including one blue sphere, whereas FLUX.1 [schnell] hallucinated an additional sphere on top of the book. While FLUX.1 [schnell] has slightly more artistic rendering of the glass and lighting, FLUX.2 [klein] 4B is the winner for its better adherence to the specific object counts and physics.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent shallow depth of field with beautiful bokeh
  • + High skin texture realism and natural facial details
  • + Strong colors with rich reflections on the wet pavement
  • The cars in the background are too sharp, lacking the requested motion blur
  • Hands have slightly confusing geometry where they meet the handlebars

FLUX.2 [klein] 4B

  • + Visible rain droplets add to the atmospheric realism requested
  • + Effective candid 'imperfect framing' that feels photographic
  • + Better representation of a typical Japanese street scene
  • The transition from the man to the bike seat is anatomically confusing
  • Overall image is a bit softer/less detailed than the competition
  • Lacks significant motion blur on the background vehicles

Verdict: Both models struggled with the 'motion blur from passing cars' instruction, but FLUX.1 [schnell] produced a much higher quality image with superior depth, color, and texture. FLUX.2 [klein] 4B captured the 'rain' and 'candid' feel better, but was let down by lower levels of detail and more noticeable structural issues with the character's posing.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Extremely detailed skin texture and lifelike eyes
  • + Dramatic lighting with intense highlights and shadows that convey the 'warm torchlight' well
  • + Superior rendering of fine facial hair and skin pores
  • The composition is a bit tight, losing almost all the plate armor detail requested
  • The beads in the hair are barely visible and look more like leather wraps
  • Overall look is a bit too hyper-processed/saturated

FLUX.2 [klein] 4B

  • + Excellent adherence to the 'ornate engraved plate armor' and 'beads' requirements
  • + Includes bokeh sparks and visible torchlight in the background
  • + Shows the character's full torso and equipment, including the cloth underlayer and leather straps
  • Faces and skin lack the 'lifelike' texture seen in the other model, appearing a bit flat
  • The scars look like surface-level paint rather than integrated skin damage
  • The lighting on the face is a bit too even, missing the dramatic contrast of torchlight

Verdict: FLUX.1 [schnell] creates a much more visceral and high-fidelity facial portrait but fails to showcase the detailed armor, which was a core part of the prompt. FLUX.2 [klein] 4B follows the prompt instructions much better by including the armor, various layers, and specifically requested beads, resulting in a more complete interpretation of the 'Battle-worn Paladin'.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong minimalist aesthetic with clean white space.
  • + Legible typography for main headers.
  • + Photos are integrated within a clear grid structure.
  • Nonsense text for one of the main category headers ('ORFEFUS' instead of Mains).
  • Limited use of vibrant accents as requested.
  • The body text is largely gibberish.

FLUX.2 [klein] 4B

  • + Excellent variety of colorful food photography.
  • + Stronger adherence to the 'vibrant accents' and 'casual dining' vibes.
  • + Includes pricing and more complex layout detail.
  • Several spelling errors in headers ('MDANS', 'MDAINS', 'FRAZZA').
  • The grid layout feels slightly cramped compared to the other model.
  • Font choices are a bit futuristic for a standard minimalist menu.

Verdict: Both models struggled with the specific text requirements, though FLUX.1 [schnell] followed the minimalist design prompt more closely with its use of negative space. FLUX.2 [klein] 4B provided much better food photography and more 'vibrant' elements, but the repeated spelling errors in the headers make it less functional as a design template.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent photorealistic texture on the patty and bun.
  • + Vibrant fiery background that matches the requested theme.
  • + Good sense of motion with the liquid cheese and flying ingredients.
  • Serious spelling error in the primary title ('AGIC' instead of 'MAGIC').
  • Confusing and repetitive price rendering with a nonsensical '€699' in the starburst.
  • The burger is mostly intact rather than 'exploded' into individual suspended components.

FLUX.2 [klein] 4B

  • + The price tag and starburst are rendered clearly and accurately.
  • + Atmospheric lighting on the burger creates a strong sense of a fiery environment.
  • + Text styling matches the requested glowing, fiery effect.
  • The main title is a chaotic jumble of letters ('MAAC AGID BUIRCGERR').
  • The burger is fully assembled and lacks the 'exploded' suspended component requirement.
  • The background effect is more like generic sparks than a fiery environment with embers.

Verdict: Both models failed the specific 'exploded burger' instruction, instead opting for a floating whole burger. FLUX.1 [schnell] produced much higher quality food textures and a better background, but failed on the text spelling and pricing logic. FLUX.2 [klein] 4B handled the price tag well, but the main title is unreadable gibberish, making it less effective as an advertisement.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + The text is sharp and highly legible in most sections.
  • + Achieves a clean chalk-on-blackboard aesthetic.
  • Failed significantly on text accuracy, outputting 'Pril' instead of 'April' and 'Taffle Mushmnctiomm'.
  • The handwriting looks more like a digital marker font than authentic chalk-textured cursive.

FLUX.2 [klein] 4B

  • + Excellent chalk texture including smudges and dust on the board.
  • + Followed the elegant cursive style instruction for the headers.
  • + Much closer adherence to the specific menu text requested, despite some spelling errors.
  • Several spelling errors like 'Trufffl', 'Musheram', and 'Ootrpous'.
  • The spacing of 'S TPECIALS' is slightly disconnected.

Verdict: FLUX.2 [klein] 4B is the clear winner for its superior atmospheric rendering of a real chalkboard, including smudges and authentic chalk grain. While it struggled with some character spelling, FLUX.1 [schnell] failed to follow the specific menu content and handwriting style instructions, providing a much more generic and less accurate result.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellently captures the difficult role-reversal prompt of a horse riding an astronaut.
  • + Highly cinematic lighting and composition with the planetary backdrop.
  • + Effective surrealist quality that challenges physics as requested.
  • The horse appears to have two heads or a double torso anomaly.
  • The astronaut's anatomy is slightly mangled where the horse is seated.

FLUX.2 [klein] 4B

  • + High visual clarity and sharp textures on the horse and spacesuit.
  • + Realistic rendering of the galaxy and planetary horizon.
  • Completely failed the semantic instruction for the horse to be on top.
  • Produces a generic 'astronaut on horse' image seen frequently in AI models.
  • The horse's legs show anatomical glitches near the hooves.

Verdict: FLUX.1 [schnell] is the clear winner because it successfully interpreted the challenging logic of the prompt ('horse on top, not vice versa'), which is a common failure point for LLMs and diffusion models. While FLUX.2 [klein] 4B has high resolution, it defaulted to the standard trope of an astronaut riding a horse, ignoring the specific directive for a role reversal.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent photographic lighting and depth of field
  • + Clear, high-quality text on the cap
  • + Highly realistic textures on the capybara fur
  • The capybara only has one paw near the wheel, and it is holding the side rather than having both front paws on the wheel as requested
  • The capybara is looking at the camera rather than the road
  • The car interior layout is slightly ambiguous

FLUX.2 [klein] 4B

  • + Perfectly adheres to the instruction of having both front paws on the steering wheel
  • + Captures the 'bored businesswoman' expression very accurately
  • + The wide-angle framing makes the interior and the scene feel more coherent
  • The cap on the capybara's head has garbled text/logo instead of a clear taxi label
  • The lighting is a bit flatter compared to the more cinematic Model A

Verdict: FLUX.2 [klein] 4B followed the specific behavioral instructions much better, correctly placing both paws on the wheel and capturing the exact mood of the passenger. While FLUX.1 [schnell] has superior lighting and texture detail, it failed the specific pose requirement for the paws and feels more like a portrait than a narrative scene.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong composition and layout for the graphic design
  • + Good color contrast and use of glowing elements
  • Terrible text accuracy with numerous spelling errors and repeated lines
  • The aesthetic feels a bit generic compared to the 'vintage gothic' request

FLUX.2 [klein] 4B

  • + Excellent atmosphere with a true vintage gothic parchment feel
  • + Much better rendering of the requested border with webs and thorns
  • + Detailed and cinematic lighting on the central jack-o-lantern
  • Significant spelling errors in the large title text
  • Missing the specific '7pm' time detail from the prompt

Verdict: FLUX.2 [klein] 4B better captures the vintage gothic aesthetic and cinematic lighting requested, whereas FLUX.1 [schnell] feels like a more modern, flat graphic. While both models struggled with text, FLUX.2 [klein] 4B followed the stylistic instructions for the border and background texture far more effectively.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the 'diorama base' instruction with a clean 3D render feel.
  • + Perfectly rendered flag of Japan.
  • + Clean, professional typography that is centered according to the prompt.
  • Missed the 'SUSHI' text entirely under the heading.
  • The sushi roll features slightly messy red toppings that look less professional than the rest of the image.

FLUX.2 [klein] 4B

  • + Beautiful soft refined textures on the salmon and rice grains.
  • + Attempts both 'JAPAN' and 'SUSHI' text as requested.
  • + Good 45-degree isometric composition.
  • Failed to render the flag of Japan correctly, showing a generic red-and-white striped icon instead.
  • Spelling error in the text ('SUSH' instead of 'SUSHI').
  • Ignored the request for a 'raised diorama base', placing the plate directly on the background.

Verdict: FLUX.1 [schnell] followed the structural elements of the prompt much better, successfully creating the raised diorama base and the correct flag, though it missed one word of text. FLUX.2 [klein] 4B had superior material textures but failed significantly on the flag, text spelling, and the base component. FLUX.1 [schnell] is the clearer winner for its professional layout and accuracy to the specific scene requirements.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent soft lighting and bokeh effect
  • + Vibrant colors that enhance the 'masterpiece' vibe
  • + Beautiful butterfly details and placement
  • Failed to include a rabbit, instead showing two felines
  • The creature between the kitten and fox is an anatomical hybrid that doesn't clearly represent a puppy, kitten, or rabbit

FLUX.2 [klein] 4B

  • + Successfully captured the active 'tumbling' and 'chasing' motion described
  • + Clearly shows the golden sunrise and distinct god rays
  • + Fur texture is highly detailed and distinct for each species
  • Failed to include the baby bunny, showing two kittens instead
  • The butterfly on the left has slightly disjointed wings

Verdict: Both models failed to include all four requested animals, with both omitting the baby bunny in favor of a second kitten. FLUX.2 [klein] 4B is the winner because it better captures the dynamic composition of 'tumbling' and 'chasing', and more accurately rendered the requested 'god rays' and sunrise atmosphere compared to the more static pose in FLUX.1 [schnell].

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully captured the vector emblem style with a circular frame.
  • + Excellent color palette consistency with the 'warm brown and cream' request.
  • + Clean, professional-looking typography despite spelling errors.
  • Significant spelling error ('FRAMILAN' instead of 'Florian').
  • Incorrect date '7720' instead of the requested '1720'.
  • Missing the 'steam' element from the cloche dome.

FLUX.2 [klein] 4B

  • + Included the steam element and the 'Est. 1720' date correctly.
  • + Good subtle texture on the background as requested.
  • + Better cloche dome illustration with highlights.
  • Spelling error in the name ('FLAXTION' instead of 'Florian').
  • Redundant text by repeating the establishment date twice.
  • The cloche handle and steam are slightly off-center.

Verdict: Both models failed to correctly spell the primary name 'Florian,' which is a common hurdle for text integration. FLUX.2 [klein] 4B followed the prompt details more closely by including the steam and the correct year, whereas FLUX.1 [schnell] produced a cleaner vector composition but hallucinated the date to be in the far future.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell]
FLUX.2 [klein] 4B

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the cohesive navy and white NASA-inspired palette.
  • + Sophisticated layout that uses orbit rings as a central design element.
  • + Achieves a sleek, professional vector aesthetic with consistent iconography.
  • The text is completely illegible gibberish.
  • The rocket icon looks more like a cartoon shuttle than a Saturn V.
  • Fails to clearly delineate the requested 6 distinct steps in chronological order.

FLUX.2 [klein] 4B

  • + Successfully follows the 6-step structure requested in the prompt.
  • + Contains partially legible text and labels that align with the mission stages.
  • + Clever inclusion of crew silhouettes at the bottom.
  • The color palette includes green and more varied tones that deviate from the specific NASA prompt.
  • The 'Saturn V' rocket design is inaccurately proportioned and stylized.
  • The layout is a bit cluttered compared to a modern infographic.

Verdict: FLUX.1 [schnell] creates a much more visually appealing and professional-looking poster, though its content is purely decorative and the text is unreadable. FLUX.2 [klein] 4B follows the complex instructions much more accurately, attempting all six specific steps with better text rendering, but fails to maintain the requested color palette and sophisticated layout of the former. FLUX.1 [schnell] is the winner for its superior 'modern vector' style which was the primary aesthetic goal.

Next steps

Explore each model