Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
FLUX.1 [schnell]
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfect adherence to spatial instructions with the sphere inside and book on top.
- + Excellent photorealistic lighting and caustic reflections on the table.
- + Realistic material textures for the wooden table and the book cover.
- − The sphere is slightly textured/fuzzy rather than smooth, though it fits the scene.
- − The text on the book spine is gibberish.
FLUX.1 [schnell]
- + Very clean, high-resolution aesthetic.
- + Vibrant colors and attractive plant foliage.
- − Failed the spatial prompt by placing a second sphere on top of the book and floating the sphere inside.
- − The blue sphere inside is disproportionately large compared to the description 'small'.
- − The glass cube physics look less like a solid object and more like a frame.
Verdict: FLUX.1 Kontext [max] significantly outperformed FLUX.1 [schnell] by following all spatial instructions perfectly, placing the sphere inside the cube and the book on top. FLUX.1 [schnell] struggled with the logic of the scene, adding an extra sphere on top of the book and creating a floating effect that wasn't requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of rain with visible streaks and atmospheric depth.
- + Captured the requested 50mm shallow depth of field perfectly.
- + The skin texture and hands look very realistic and age-appropriate.
- − The bike chain and spokes have some structural clipping and inconsistencies.
- − The rain streaks are very heavy, bordering on a 'filter' look in some areas.
FLUX.1 [schnell]
- + Natural composition and color balance with good reflections on the wet pavement.
- + The man's pose feels very candid and grounded in the environment.
- + Strong adherence to the 'elderly Japanese man' descriptor.
- − The cars in the background lack the motion blur requested in the prompt.
- − The 'light rain' is barely visible, appearing more as a post-rain environment.
- − Anatomical issues with the hands where they grip the handlebars.
Verdict: FLUX.1 Kontext [max] is the superior image as it much more accurately depicts the requested rain, 50mm depth of field, and motion blur effects. While FLUX.1 [schnell] captures a decent street scene, it misses several technical atmospheric prompts and has noticeable anatomical flaws in the hands.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of ornate engraved plate armor with realistic metallic reflections.
- + Very skin-like texture with believable pores and faint scars as requested.
- + Effective use of warm lighting and bokeh sparks that frame the character well.
- − Missed the detail of small beads in the hair braids.
- − The glow in the eyes borders on supernatural, whereas the prompt asked for lifelike.
FLUX.1 [schnell]
- + Stronger adherence to the 'hair braided with small beads' part of the prompt.
- + Intense character expression that fits the 'battle-worn' theme.
- + Clear depiction of the cloth underlayer and leather elements.
- − The armor looks more like studded leather than the requested 'ornate engraved plate armor'.
- − Skin texture appears slightly overprocessed or airbrushed in some areas compared to the other model.
- − The background lighting is a bit flat and lacks the specific 'bokeh sparks' requested.
Verdict: FLUX.1 Kontext [max] produced a superior image in terms of material quality, specifically the ornate plate armor and the realistic skin texture. While FLUX.1 [schnell] followed the minor instruction for beads in the hair more accurately, it failed to deliver the silver plate armor that defines a paladin, making the Kontext version the more impressive and prompt-accurate output overall.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional layout with clear typographic hierarchy
- + High-quality, distinct food photography that fills the grid nicely
- + Strong adherence to the minimalist, clean, and professional aesthetic
- − The text content is mostly gibberish despite the good layout
- − Lacks clear header sections for 'Appetizers' and 'Mains' as requested
FLUX.1 [schnell]
- + Successfully included specific headers for 'Appetizers', 'Pizza', and 'Mains'
- + Follows the white background and grid instruction well
- + Layout feels functional and realistic for a casual dining menu
- − The food images are repetitive and some look like low-resolution textures
- − Includes a strange non-existent word 'ORFEFUS' as a section header
- − Text alignment on the right side is somewhat messy
Verdict: FLUX.1 Kontext [max] produced a much more visually appealing and professional-looking design, though it failed to include the specific requested section headers. FLUX.1 [schnell] adhered closer to the prompt's specific content requirements (sections for pizza/mains), but the visual quality of the food photography and the overall layout were significantly lower than the other model.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI judge analysis unavailable for this challenge.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfect spelling and adherence to the requested text for all menu items.
- + Very realistic chalk texture with slight smudges and varying stroke weight.
- + Excellent composition that feels like a real café chalkboard with professional layout.
- − The title is more of a print-style block letter than the 'elegant cursive' requested.
- − The background around the board is slightly cluttered/cropped.
FLUX.1 [schnell]
- + Includes more of the background environment of the café.
- + Text has a casual handwritten appearance.
- − Frequent and severe spelling errors, such as 'Taffle Mushmnctiomm' and 'Octtoopus'.
- − Fails to render the date correctly, truncating April to 'Pril'.
- − The text quality degrades at the bottom with nonsensical 'glitch' text repetitions.
Verdict: FLUX.1 Kontext [max] far outperforms FLUX.1 [schnell] in this challenge, producing clear, perfectly spelled, and aesthetically pleasing text that aligns almost exactly with the prompt. FLUX.1 [schnell] struggled significantly with text coherence, resulting in numerous spelling errors and repetitive, nonsensical phrases at the bottom of the board.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + High visual clarity and cinematic lighting
- + Excellent texture detail on the astronaut suit and horse fur
- + Realistic spatial composition
- − Failed the specific logic of the prompt by placing the astronaut on top of the horse
FLUX.1 [schnell]
- + Successfully followed the complex instruction of placing the horse on top of the astronaut
- + Creative interpretation of the surreal requirement
- + Dynamic and interesting composition
- − Anatomical errors with the horse, specifically having two heads/chests merged
- − Lower overall detail in the background and planet texture compared to Model A
Verdict: FLUX.1 Kontext [max] produced a much higher quality image in terms of rendering and realism, but completely ignored the specific inversion requested in the prompt. FLUX.1 [schnell] followed the difficult instruction to put the horse on top of the astronaut, making it the superior choice for prompt adherence despite some anatomical clipping and lower fidelity.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent fur texture rendering and realistic lighting from the city street.
- + High fidelity in the taxi interior details, such as the steering wheel and seatbelt.
- − The passenger is holding a phone to her ear like a call instead of looking at it as requested.
- − The capybara only has one paw visible on the steering wheel.
FLUX.1 [schnell]
- + Perfectly captures the businesswoman looking at her phone with a bored expression.
- + Better interior composition that clearly shows both the driver and the passenger as requested.
- + Includes both of the capybara's paws on or near the wheel.
- − The capybara's face is slightly less photorealistic and more stylized compared to Model A.
- − The phone rendering in the woman's hands is a bit blurry/distorted.
Verdict: While FLUX.1 Kontext [max] has superior texture quality and lighting, FLUX.1 [schnell] followed the specific character instructions much better, particularly the 'looking at phone' and 'bored expression' details for the passenger. FLUX.1 [schnell] provides a more complete narrative composition that fits all elements of the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and adherence to almost all specific details.
- + Authentic vintage gothic aesthetic with a dark, parchment-like texture.
- + Very high artistic quality in the illustration of the pumpkin and twisted trees.
- − Repeated the location text twice at the bottom.
- − The scroll banner is somewhat blended into the background rather than being a distinct element.
FLUX.1 [schnell]
- + Good use of color hierarchy with high-contrast orange text.
- + Clearer representation of a scroll banner as requested.
- − Significant text errors including typos and repeated words (e.g., 'a a night', 'firiichts', 'Tive').
- − Composition feels less 'cinematic' and more like a flat graphic design.
- − Failed to correctly display the event details, adding nonsensical lines like 'Fate: butistigtion'.
Verdict: FLUX.1 Kontext [max] produced a much more professional and atmospheric image that truly feels like a gothic invitation, with nearly perfect text rendering. While it suffered from a minor repetition in the location field, it far outperformed FLUX.1 [schnell], which struggled significantly with spelling, syntax, and overall artistic coherence.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a friendly, cartoonish style.
- + Superior 3D modeling feel with soft lighting and pleasing textures.
- + Complete composition including dipping sauce, wasabi, and chopsticks.
- − Missed the request for a 'small flag icon'.
- − The font color is a bit dull compared to the vibrant food.
FLUX.1 [schnell]
- + Successfully included the Japanese flag icon.
- + Strict adherence to the 'isometric' perspective on a white diorama base.
- + High contrast text that is easy to read.
- − Missing the word 'SUSHI' requested in the prompt.
- − The 'JAPAN' text is poorly placed at the very top edge.
- − The sushi roll graphics look a bit flat and less appetising compared to the 3D nigiri.
Verdict: FLUX.1 Kontext [max] produced a much more appealing and cohesive 3D scene with better lighting and a more complete set of props, though it missed the flag icon. FLUX.1 [schnell] followed the technical layout and flag requirement better but failed to include the word 'SUSHI' and lacked the 'soft refined textures' requested for the food.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the prompt by including all four distinct animals: puppy, kitten, bunny, and fox.
- + Beautiful soft lighting with clear sunbeam effects and sparkling highlights.
- + Consistent and distinct fur textures for each different animal species.
- − The eyes of the bunny and fox look slightly artificial/uncanny.
- − Some butterflies are blurry or blend too much into the background glow.
FLUX.1 [schnell]
- + High quality fur rendering and cute, expressive facial features on the animals.
- + Subtle, realistic lighting on the grass and flower petals.
- − Failed to include the requested baby bunny, instead generating two kittens.
- − The fox and kittens share very similar facial structures, lacking species-specific anatomical distinction.
- − The butterflies look like flat overlays rather than being integrated into the 3D space.
Verdict: FLUX.1 Kontext [max] is the winner as it accurately followed the complex prompt by including all four specific animals (puppy, kitten, bunny, and fox), whereas FLUX.1 [schnell] missed the bunny entirely. Additionally, FLUX.1 Kontext [max] better captured the requested 'god rays' and 'whiteflower meadow' atmosphere, creating a more cohesive and magical scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography including the correct grave accent on 'Caffè'
- + High accuracy with prompt details like 'Est. 1720' and the steam element
- + Authentic vintage texture and woodblock-style illustration
- − The steam is a bit simplified compared to the rest of the illustration
FLUX.1 [schnell]
- + Clean vector-style composition
- + Warm and inviting cream color palette
- − Contains significant spelling errors: 'Framilan' and '7720'
- − Missing the steam element requested in the prompt
- − Cloche dome looks more like a building or a chess piece
Verdict: FLUX.1 Kontext [max] delivered a near-perfect execution of the prompt, capturing the exact requested text, date, and imagery with a beautiful vintage hand-drawn texture. In contrast, FLUX.1 [schnell] failed on several key prompt elements, including incorrect spelling of the brand name and wrong digits for the establishment date.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Features a more detailed and aesthetically pleasing illustration style.
- + Higher quality rendering of the lunar module on the surface.
- + Captures the sense of a classic educational poster with varied icons.
- − Terrible text accuracy with confusing labels like 'Moon' pointing to Earth and 'Saturn' appearing in a lunar mission.
- − Failed to follow the requested sequential 6-step infographic structure.
- − Chaotic arrow logic that doesn't clearly represent a mission path.
FLUX.1 [schnell]
- + Follows the 6-step infographic structure much more effectively than the competitor.
- + Cleaner, more consistent icons that match the 'vector' prompt better.
- + Uses a symmetrical layout that fits the modern infographic style.
- − The main rocket illustration is nonsensical with overlapping fuselages.
- − Text is completely illegible 'gebbierish' throughout the image.
- − The central composition is a bit empty and lacks the NASA-inspired charm of the first image.
Verdict: FLUX.1 [schnell] followed the structural logic of a 6-point infographic much better, though the actual illustrations and text are poor. FLUX.1 Kontext [max] produced much nicer individual illustrations and a better color palette, but the infographic itself is technically inaccurate (mislabeling Earth as Moon) and fails to organize the information into the requested steps. FLUX.1 Kontext [max] is the preferred choice for its higher visual polish, despite the logical errors.
Explore each model
Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps