Black Forest Labs' 12-billion parameter multimodal flow transformer for in-context image generation and editing with character consistency, typography handling, and commercial-ready quality
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [pro]
#41 of 62 in Text-to-Image
FLUX.2 [flex]
#14 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [pro]
0%
win rate
Ties
0%
FLUX.2 [flex]
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent photographic realism with natural window light and curtains.
- + Realistic wood grain texture and high-quality glass transparency.
- + The cube structure is thin and elegant, fitting the 'glass cube' description perfectly.
- − The plant is more in the background than 'behind' the cube, though still visible.
- − The blue sphere has a felt-like texture rather than being a standard smooth sphere.
FLUX.2 [flex]
- + Perfectly adheres to the spatial prompt with the plant clearly positioned behind the glass cube.
- + Smooth, high-quality rendering of the blue sphere and red book.
- + Good application of soft window light from the left side.
- − The glass cube is missing its top face, appearing more like a glass box or stand with the book resting on the side edges.
- − The composition feels slightly more digital and less like a real photograph compared to Model A.
Verdict: Both models followed the prompt instructions very well. FLUX.1 Kontext [pro] produced a more convincing photograph with superior lighting and textures, while FLUX.2 [flex] handled the spatial relationship of the plant behind the glass more effectively. FLUX.1 Kontext [pro] is the winner for creating a physically complete glass cube, whereas Model B's cube lacks a top surface.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent shallow depth of field with realistic background bokeh.
- + Natural skin texture and convincing rainfall effects.
- + Captures the 'candid' and 'imperfect framing' request perfectly.
- − The subject is holding the bike rather than repairing it.
- − The bicycle anatomy around the handlebars is slightly confusing.
FLUX.2 [flex]
- + Accurately depicts the act of 'repairing' the bicycle.
- + Includes strong motion blur from passing cars as requested.
- + Good reflection detail on the wet pavement.
- − The bicycle frame geometry is physically impossible (missing seat tube).
- − The 'imperfect framing' feels less natural and more cluttered with the large post on the left.
Verdict: FLUX.1 Kontext [pro] creates a much more convincing photographic aesthetic with superior skin textures and lighting, though it fails to show the man actually repairing the bike. FLUX.2 [flex] adheres better to the specific actions in the prompt like 'repairing' and 'motion blur', but the image suffers from significant structural errors in the bicycle's frame. FLUX.1 Kontext [pro] is the winner for its superior visual quality and realism.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Naturalistic lighting and shallow depth of field
- + Highly realistic skin textures and subtle, believable weathering
- + Excellent artistic composition with a cinematic feel
- − Missed the beads in the braids
- − The scars are extremely faint, almost invisible
FLUX.2 [flex]
- + Matches all prompt details including beads in the hair and prominent scars
- + Highly detailed engraving and textures on the armor
- + Clear depiction of warm torchlight and bokeh sparks
- − The scars look like fresh, clean cuts rather than old battle-worn scars
- − Lighting is a bit flat and looks more like a high-end CGI render than a photograph
Verdict: FLUX.1 Kontext [pro] creates a much more lifelike and photographically convincing portrait with superior lighting and depth, though it misses minor ornaments like beads. FLUX.2 [flex] adheres strictly to every detail of the prompt, including specific hair accessories and scarring, but the overall image feels more like a game character than a real person. FLUX.1 Kontext [pro] is the winner for its superior visual quality and realism.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Strong typography and effective use of white space for a minimalist look.
- + Food images are high quality and clearly distinguishable.
- + The layout neatly categorizes into the three requested sections.
- − The grid requested is represented more as a staggered list rather than a true photo grid.
- − Includes a pizza under the 'Mains' heading which causes categorical overlap.
- − Text contains more gibberish characters than Model B.
FLUX.2 [flex]
- + Adheres better to the 'grid' request for the food photos.
- + More professional-looking color blocks and UI-inspired design.
- + Better rendering of realistic currency and menu item structures.
- − The 'Appetizers' heading at the top is redundant as a title for the whole page.
- − Layout feels slightly cluttered compared to Model A's minimalist approach.
- − Text becomes illegible in the smaller descriptions.
Verdict: FLUX.2 [flex] produced a more professional, commercially viable layout that correctly implemented the requested photo grid. While FLUX.1 Kontext [pro] captured the 'minimalist' aesthetic well with cleaner spacing, its categorical logic was flawed, whereas FLUX.2 [flex] felt like a more complete design concept for a casual dining menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography with a glowing red and gold border.
- + High photorealistic detail in the texture of the bun and melted cheese.
- + Dynamic atmosphere with realistic embers and smoke.
- − Repeated the price text '€6.99' twice in different sizes.
- − Missed the specific 'starburst' request for the price element.
- − The burger layers are not fully 'exploded' or separated, appearing mostly stacked.
FLUX.2 [flex]
- + Perfectly followed the instruction to place the price in a starburst.
- + Better separation of ingredients conveying the 'exploded' motion requested.
- + Creative fiery effect applied to the internal letters of the main title.
- − The 'LIMITED TIME ONLY' text is slightly less legible than Image A.
- − The bottom bun seems to be melting or drooping unrealistically into the sauce.
Verdict: FLUX.2 [flex] is the superior choice as it adhered more closely to the specific structural requirements of the prompt, including the starburst for the price and the dynamic separation of the burger layers. While FLUX.1 Kontext [pro] produced a high-quality image, it failed to generate the starburst and redundantly included the price twice, cluttering the composition.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent chalk texture and grainy detail on the letters
- + Perfect adherence to the requested date and menu item prices
- + Realistic wood frame and board surface
- − Text at the bottom has spelling errors ('ous' instead of 'us', 'tree' instead of 'free')
- − Title is print-style rather than the requested 'elegant cursive'
FLUX.2 [flex]
- + Beautiful cursive handwriting style for the title and menu items as requested
- + Atmospheric lighting and composition within a café setting
- + Accurate spelling throughout the entire text
- − Chalk texture is slightly too smooth, leaning towards a 'chalk marker' look
- − Less natural variation in the thickness of the chalk strokes compared to Model A
Verdict: Both models followed the prompt instructions well, but FLUX.2 [flex] is the superior choice because it correctly interpreted the 'elegant cursive' requirement and maintained perfect spelling. While FLUX.1 Kontext [pro] had superior chalk texture, its significant spelling errors and use of block letters for the title made it less successful overall.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent adherence to the 'horse riding astronaut' specific prompt by placing the horse in a seating position on the astronaut's back.
- + High cinematic quality with realistic space lighting and textures.
- + Highly detailed and clean rendering of the space suit and horse fur.
- − Anatomical weirdness where the astronaut's legs end in hooves.
- − A small, unexplained secondary astronaut figure is fused into the horse's back.
FLUX.2 [flex]
- + Vibrant, surreal color palette with a celestial nebula background.
- + Clearer distinction Between the horse's hooves and the astronaut's hands/boots.
- − Fails the specific prompt logic; the horse is jumping over/touching the astronaut rather than 'riding' them.
- − The horse's front legs are awkwardly merged into the astronaut's shoulders/hands area.
Verdict: FLUX.1 Kontext [pro] followed the complex spatial instructions much better by actually depicting the horse in a riding posture atop the astronaut, whereas FLUX.2 [flex] produced a more generic 'beside each other' composition. Although FLUX.1 Kontext [pro] has a strange glitch with hoof-feet on the astronaut, its adherence to the surreal 'horse on top' request makes it the clear winner.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent photorealistic texture on the capybara fur
- + Natural, cinematic lighting that feels professional
- + Accurate depiction of a dark jacket as requested
- − The passenger is on a phone call rather than looking at her phone as requested
- − The composition is quite tight, showing less of the Manhattan environment
FLUX.2 [flex]
- + Perfect adherence to the passenger's action and expression
- + Stronger sense of place with the recognizable Manhattan backdrop through the windows
- + Clearer view of the 'taxi driver cap' and the full cabin interior
- − The hands of the capybara look a bit more like human hands than paws
- − The lighting on the capybara's face is a bit flatter than in the other image
Verdict: Both models followed the complex prompt well, but FLUX.2 [flex] achieved better prompt adherence by correctly depicting the bored businesswoman looking down at her phone. FLUX.2 [flex] also provided a better sense of scale and environment, whereas FLUX.1 Kontext [pro] felt more like a close-up portrait with slightly superior fur textures.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography style that matches the gothic aesthetic
- + Detailed spooky border with webs
- + Dramatic, high-contrast cinematic lighting
- − Includes nonsense text 'Your: Vorkleat: Iight & Spans' that was not in the prompt
FLUX.2 [flex]
- + Perfect text accuracy for all required fields
- + Strong representation of the dark parchment texture with torn edges
- + Clean, readable layout
- − Lighting is a bit flatter compared to the other model
- − The outer white border slightly breaks the 'poster' feel
Verdict: FLUX.2 [flex] is the superior choice because it followed the text instructions perfectly, whereas FLUX.1 Kontext [pro] hallucinated an extra line of gibberish. While FLUX.1 Kontext [pro] had more dramatic lighting and border detail, FLUX.2 [flex] delivered a more accurate and usable invitation with better parchment texture.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent full head of hair that matches the beard color and texture.
- + Seamless blend between the new hair and the existing sideburns/beard.
- − Slightly alters the facial structure, making the man look younger and changing the eye shape.
- − Changes the frame style of the glasses to a solid black plastic.
FLUX.2 [flex]
- + Perfect preservation of the facial features, skin texture, and original glasses.
- + Maintains the exact lighting and background from the source image.
- − The hair is quite short and thin, not meeting the 'full, thick' part of the prompt.
- − The hairline near the temples looks a bit unnatural and abrupt.
Verdict: FLUX.1 Kontext [pro] creates a much more convincing head of thick hair that satisfies the prompt's aesthetic requirements, but it fails to perfectly preserve the subject's face and glasses. FLUX.2 [flex] excels at source preservation, keeping the identity and accessories identical to the original, but the hair it adds is too short and lacks the requested density. FLUX.1 Kontext [pro] is the winner for better fulfilling the core creative request despite the minor drift in facial likeness.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent 3D cartoon art style with charming, rounded fonts
- + Very clean composition with high-quality soft lighting and textures
- + Accurate 45-degree isometric perspective
- − The flag icon is stylized as an envelope/tag rather than a traditional flag shape
FLUX.2 [flex]
- + Features a variety of sushi types including nigiri and maki
- + Very clean, flat vector-style flag and text rendering
- + Adheres well to the requested diorama base and solid background
- − The text is plain and lacks the '3D cartoon' aesthetic requested in the prompt
- − Shadow on the diorama base is strangely cut off on the left side
Verdict: FLUX.1 Kontext [pro] captured the aesthetic of a '3D cartoon scene' much better with its rounded, playful typography and soft, cohesive lighting. While FLUX.2 [flex] provided more variety in the sushi itself, the overall composition felt flatter and the text was less integrated into the artistic style.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent preservation of the subject's clothing and facial structure in a stylized format.
- + High-quality comic book art style with clean lines.
- − Completely ignored the prompt instructions regarding the job (TV anchor), dogs, and hockey.
- − Failed the editing task by only changing the style without adding the requested elements.
FLUX.2 [flex]
- + Successfully integrated all requested elements including the TV news desk, multiple dogs, and hockey gear.
- + Exaggerated caricature style that fits the 'humorous' request perfectly.
- + Good source preservation of the subject's outfit and facial features within the new style.
- − Small text artifacts on the news desk ('INEIANAES').
- − The hockey element is slightly understated compared to the dogs, appearing mostly as jerseys and one stick.
Verdict: FLUX.1 Kontext [pro] failed the image editing task by only converting the source image into a cartoon style while ignoring all specific content requests like the TV anchor job, dogs, and hockey. FLUX.2 [flex] successfully followed all instructions, creating a detailed caricature scene that incorporates every requested theme while maintaining the likeness and clothing of the person in the source image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent soft, golden lighting that matches the requested sunrise vibe.
- + High-quality fur texture and expressive, cute facial features on all animals.
- + Vibrant color palette with a dense, beautiful wildflower meadow.
- − Failed to include a distinct baby bunny, instead generating two cat-like creatures in the center.
- − The composition is static with animals sitting rather than 'playfully chasing and tumbling'.
FLUX.2 [flex]
- + Successfully included all four distinct animal types: puppy, kitten, bunny, and fox kit.
- + Captured the 'playfully chasing and tumbling' action much better with dynamic posing.
- + Includes very clear 'god rays' and dew sparkles as requested in the prompt.
- − The fox kit has slightly unusual black limb coloring that looks a bit unnatural.
- − The butterflies are somewhat large and distracting compared to the more subtle ones in Model A.
Verdict: While FLUX.1 Kontext [pro] has a slightly more masterfully painted aesthetic and lighting, it failed to follow the core prompt instructions by omitting the bunny. FLUX.2 [flex] adhered to all prompt requirements, including all species and the active tumbling/chasing motion, making it the more accurate and successful output.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent preservation of the iconic character expressions and poses.
- + Captures the hand-painted watercolor texture of Ghibli films perfectly.
- + Maintains the distinct color palette of the original photo while stylizing it.
FLUX.2 [flex]
- + Achieves a very soft, nostalgic, and dreamy pastel aesthetic.
- + High-quality linework that adheres well to a classic anime style.
- + Good background detail with stylized signage.
- − Fails to preserve the character emotions, specifically making the girlfriend look happy rather than jealous.
- − The man's stubble and facial structure look slightly less like the original subject compared to Model A.
Verdict: FLUX.1 Kontext [pro] is the superior edit because it successfully translates the specific narrative tension of the original 'distracted boyfriend' meme into the Ghibli art style, keeping the jealous facial expression intact. While FLUX.2 [flex] has a beautiful, soft aesthetic, it fails as an image edit by changing the woman's expression from angry to smiling, losing the original meaning of the scene.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent preservation of the subject's face and body features.
- + The hair motion looks very natural and integrated with the character.
- + Subtle and realistic addition of floating leaves.
- − The leash handle has been slightly simplified causing a minor loss of detail.
FLUX.2 [flex]
- + Very dramatic and high-energy hair motion.
- + High quantity of leaves creates a very lively feel as requested.
- − Noticeable distortion of the subject's right leg (left side of image) making it look unnaturally wide.
- − Leaves appear somewhat pasted-on and lack realistic blurring for the speed implied by the hair.
Verdict: Both models followed the instructions well, but FLUX.1 Kontext [pro] is the clear winner due to its superior anatomical preservation. While FLUX.2 [flex] achieved a more energetic hair effect, it significantly distorted the subject's leg and hip area, whereas FLUX.1 Kontext [pro] maintained the original person's proportions while adding the requested motion.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent texture that mimics paper or woodblock printing
- + Clear and bold typography that fits the vintage aesthetic
- + Accurate rendering of all prompt elements including the banner and steam
- − Spelling error in the banner text ('EEST' instead of 'EST')
- − The cloche illustration is slightly blocky compared to a clean vector style
FLUX.2 [flex]
- + Elegant arched typography with correct accent on the name
- + Perfect spelling of all requested text
- + Clean vector lines and more balanced composition
- − The 'subtle texture' requested in the prompt is almost non-existent
- − The steam effect is a bit faint for a logo
Verdict: FLUX.2 [flex] is the likely winner because it followed the text instructions perfectly, whereas FLUX.1 Kontext [pro] included a spelling error ('EEST'). While FLUX.1 Kontext [pro] captured the vintage texture and grit better, the clean composition and accuracy of FLUX.2 [flex] make it a more professional logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent NASA-inspired color palette.
- + Artistic and integrated vector style with a strong sense of composition.
- − Confused scientific imagery, including a bizarre Saturn-like ring on the rocket and multiple moons/planets.
- − Text labels are cluttered and contain inaccuracies like labeling the Moon as 'Earth' and including 'Collins' randomly.
- − Fails to clearly delineate the specific steps 1 through 6.
FLUX.2 [flex]
- + Follows the infographic structure perfectly with clear, labeled steps.
- + Clean, modern flat-vector icons that match the prompt's request for crisp lines.
- + Highly legible and accurate typography with correct spelling.
- − Included only 5 distinct steps instead of the 6 requested in the prompt.
- − The layout is a bit more rigid and standard compared to a creative poster design.
Verdict: FLUX.2 [flex] is much more successful because it understands the logic of an infographic, providing clear icons and accurate text labels for nearly all the requested steps. FLUX.1 Kontext [pro] creates a more visually interesting piece of art, but it fails completely as an infographic due to nonsensical labels (labeling the Moon as Earth) and the inclusion of random planetary rings on the Saturn V rocket.
Explore each model
Black Forest Labs' precision image generation model with maximum control, reliable text rendering, and complete creative control supporting up to 4MP output