Black Forest Labs' 12-billion parameter multimodal flow transformer for in-context image generation and editing with character consistency, typography handling, and commercial-ready quality
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [pro]
#43 of 62 in Text-to-Image
FLUX.2 [klein] 4B
#32 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [pro]
0%
win rate
Ties
0%
FLUX.2 [klein] 4B
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent adherence to lighting instructions with clear soft light from the left.
- + Superior composition with the plant clearly visible through the glass as requested.
- + Realistic materials, particularly the texture of the red book and the wooden table grain.
- − The blue sphere appears slightly fuzzy or felt-like, which may not be the intended texture for a geometric primitive.
FLUX.2 [klein] 4B
- + Accurately places all requested elements in the scene.
- + Good reflection of the blue sphere on the base of the glass cube.
- + Clean, minimalist aesthetic.
- − The plant is positioned more to the side rather than 'behind the cube', making it less visible 'through the glass'.
- − The book seems slightly thin and less detailed compared to Model A.
Verdict: FLUX.1 Kontext [pro] is the winner as it followed the spatial instructions more accurately, specifically placing the plant behind the glass cube so it is visible through it. Both models handled the complex task of stacking objects and placing items inside glass well, but FLUX.1 Kontext [pro] had superior lighting and texture details.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent skin texture and realistic facial detail
- + Strong cinematic atmosphere with effective shallow depth of field
- + Accurate depiction of rain droplets and dampened clothing textures
- − The man appears to be riding or holding the bike rather than repairing it
- − Missing the explicit motion blur from cars mentioned in the prompt
FLUX.2 [klein] 4B
- + The subject is actively engaged in a repair/inspection posture
- + Effective use of reflections on the wet pavement
- + Good composition that shows the full red bicycle
- − Anatomical errors in the hands and legs
- − Structural issues with the bicycle frame and chain alignment
- − Lacks the requested motion blur for passing cars
Verdict: FLUX.1 Kontext [pro] produces a much higher quality portrait with superior skin textures and lighting, capturing the cinematic feel requested, though it misses the 'repairing' action. FLUX.2 [klein] 4B follows the 'repair' instruction better but suffers from significant anatomical distortions and a less realistic finish. FLUX.1 Kontext [pro] is the clear winner for its photographic fidelity.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent photorealistic skin and hair textures
- + Highly detailed and authentic-looking engraved plate armor
- + Atmospheric lighting that feels natural and moody
- − Missed the request for small beads in the braids
- − The 'close portrait' framing obscures some requested details like leather straps
FLUX.2 [klein] 4B
- + Successfully included beads in the braided hair
- + Shows more of the armor and leather strap details
- + Good interpretation of battle-worn skin with more prominent scars
- − The lighting feels artificial and slightly flat for a torchlit scene
- − Armor engravings look somewhat digital and repetitive
- − The fire/torch in the background has a slightly lower graphical fidelity
Verdict: FLUX.1 Kontext [pro] produces a significantly more lifelike and cinematic portrait with superior skin and metal Rendering, though it missed the specific detail of beads in the hair. FLUX.2 [klein] 4B adhered better to the 'beads' and 'leather straps' prompts, but the overall image quality feels more like a video game render than a realistic photograph.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography rendering with readable sans-serif fonts
- + Clean and orderly layout that mimics a real professional menu
- + Accurate representation of the requested sections (Appetizers, Pizza, Mains)
- − The placement of food items is somewhat inconsistent, like a pizza in the 'Mains' section
- − Limited to only three images rather than a full grid
FLUX.2 [klein] 4B
- + Successfully creates a more complex grid of food photos as requested
- + High visual quality and color saturation in the food photography
- + Includes branding elements like a logo and header
- − Significant text artifacts and 'gibberish' across the menu items
- − Failed to properly label sections, using 'Frazza' and 'Mdalns' instead of requested titles
- − Layout feels a bit cluttered compared to Model A
Verdict: FLUX.1 Kontext [pro] is the superior choice because it generates legible, professional-looking typography and a layout that functions as a real menu. While FLUX.2 [klein] 4B handles the 'grid' request better with more photos, its inability to spell requested section names and its messy text rendering make it unuseable for a design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent text rendering with no spelling errors.
- + Realistic lighting and texture on the burger ingredients.
- + High-quality 'fiery' aesthetic that matches the prompt perfectly.
- − Missed the 'starburst' request for the price.
- − Included the price twice, which was not requested.
FLUX.2 [klein] 4B
- + Successfully included the price in a starburst graphic.
- + Good sense of motion with the radial blur on embers.
- − Severe spelling errors in the main title ('MAAC AGIR BUIRCGRE').
- − The burger is slightly separated but lacks the true 'exploded' look of individual suspended components.
- − Overall image quality and text integration are messy.
Verdict: FLUX.1 Kontext [pro] is the clear winner due to its superior text legibility and photorealistic rendering of the burger, whereas FLUX.2 [klein] 4B failed significantly on the primary typography. Although FLUX.2 [klein] 4B followed the starburst instruction, the jumbled letters and lower overall visual fidelity make it unusable as an advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent spelling accuracy for all mandatory menu items.
- + Very realistic chalk texture with slight graininess in the strokes.
- + Good composition and readability.
- − The 'cursive' request for the title was not fully followed, appearing more as print caps.
- − Minor spelling errors in the footer text at the bottom ('ous' instead of 'us', 'tree' instead of 'free').
FLUX.2 [klein] 4B
- + Successfully used elegant cursive handwriting for the menu items.
- + Excellent chalk smears and board texture that add to the realism of a hand-wiped chalkboard.
- − Significant spelling errors across almost every word (e.g., 'Truffel Musheram', 'Ootrpous', 'Brawn Buter').
- − Poor spacing in the title with a large gap inside 'SPECIALS'.
Verdict: FLUX.1 Kontext [pro] is the clear winner as it accurately spelled all the specific menu items requested in the prompt, whereas FLUX.2 [klein] 4B suffered from numerous spelling hallucinations. While FLUX.2 [klein] 4B had a more convincing hand-wiped chalkboard texture, the lack of legibility and text accuracy makes it less useful.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Strictly followed the difficult reversal prompt with the horse physically on top of the astronaut.
- + Higher level of surrealism and creativity in the composition.
- + Excellent cinematic lighting and textures on both the spacesuit and horse fur.
- − The astronaut's legs transform into hooves at the bottom, which is a bit of a logical hallucination.
- − The small secondary astronaut riding the horse creates some confusion in the scene logic.
FLUX.2 [klein] 4B
- + High visual clarity and sharp detail on the horse and stars.
- + Balanced composition with a beautiful starry background.
- − Failed the prompt instruction 'horse on top, not vice versa'.
- − Presents a standard trope image rather than the requested surreal reversal.
Verdict: FLUX.1 Kontext [pro] successfully followed the specific prompt instruction to place the horse on top of the astronaut, resulting in a truly surreal image. FLUX.2 [klein] 4B ignored the core constraint of the prompt, providing a standard 'astronaut riding a horse' image despite the 'not vice versa' warning.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent fur texture rendering on the capybara
- + Photorealistic lighting on the car exterior reflections
- + Accurate focus and shallow depth of field
- − The passenger is on a phone call instead of looking at her phone as requested
- − Only one paw is visible on the steering wheel
FLUX.2 [klein] 4B
- + Perfect adherence to the instruction for the passenger's action and expression
- + Both front paws are correctly placed on the steering wheel
- + Composition captures the 'inside the taxi' feeling much better
- − The capybara's paws look slightly like human hands covered in fur
- − The interior light is a bit harsh for a night scene
Verdict: FLUX.2 [klein] 4B followed the prompt much more accurately, correctly depicting the passenger looking at her phone and the capybara having both paws on the wheel. While FLUX.1 Kontext [pro] produced a more high-end cinematic texture, it failed to capture the specific multi-subject actions requested.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent text rendering with clean, legible gothic fonts
- + Impressive details on the jack-o-lantern and cobweb border
- + High-contrast cinematic lighting that fits the spooky theme
- − Includes an extra line of hallucinated text between the date and time
FLUX.2 [klein] 4B
- + Strong composition with twisted trees framing the image
- + Accurately represents the thron and cobweb border combination
- + Good atmosphere in the moody night sky
- − Significant spelling errors in the title text and location
- − Incorrect year in the date and missing the time entirely
- − Jack-o-lantern rendering is slightly generic compared to Model A
Verdict: FLUX.1 Kontext [pro] is the clear winner as it successfully rendered most of the complex text instructions with high legibility and beautiful typography. While FLUX.2 [klein] 4B captured the requested aesthetic elements like thorns and twisted trees well, it failed significantly on the text requirements, producing garbled words and missing key information.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Successfully added thick hair with a natural hairline.
- + Preserved the lighting and shadows of the original scene well.
- − Significantly altered the facial features, making the person look like a different individual.
- − The glasses frame thickness and shape were noticeably changed.
FLUX.2 [klein] 4B
- + Excellent source preservation, keeping the glasses and facial features nearly identical to the original.
- + The hair texture and style feel very natural and integrated with the existing beard.
- − Slight blurring where the hair meets the background on the far left.
Verdict: FLUX.2 [klein] 4B is the clear winner because it successfully applied the edit while perfectly preserving the identity of the person in the source image. In contrast, FLUX.1 Kontext [pro] generated a full head of hair but fundamentally changed the underlying facial structures and eyewear, failing the 'preserve facial features' part of the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent text rendering with accurate spelling and stylization.
- + Beautiful 3D render with soft lighting and a clear diorama style base.
- + Followed all elements of the prompt including the specific flag icon placement.
- − The rice grains are slightly oversized, looking more like pearls than rice.
FLUX.2 [klein] 4B
- + Natural look to the sushi textures and garnishing.
- + Clean isometric perspective.
- − Spelling error in the text, writing 'SUSH' instead of 'SUSHI'.
- − The flag icon is incorrect, showing a red/white striped flag instead of the Japanese flag.
- − Lacks the 'diorama base' requested, featuring only a standard plate.
Verdict: FLUX.1 Kontext [pro] followed every specific detail of the prompt, including the text, icon, and the diorama base concept. FLUX.2 [klein] 4B failed significantly on the text rendering, misspelling the main subject and using an incorrect flag icon.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent preservation of the source image's clothing and pose.
- + Successfully captures the subject's likeness in a stylized caricature format.
- − Completely ignored all thematic requests including the job, dogs, and hockey.
- − The background is blank and lacks the requested context.
FLUX.2 [klein] 4B
- + Successfully incorporates the TV anchor profession and the love for dogs.
- + Features an exaggerated caricature style that fits the 'humorous' prompt.
- − Failed to include any hockey-related elements.
- − The likeness to the original subject is weaker than the other model, especially the face shape.
Verdict: FLUX.1 Kontext [pro] created a great caricature of the subject's face and clothing but failed to include any of the requested thematic elements like the job, dogs, or hockey. FLUX.2 [klein] 4B followed most of the prompt instructions by placing the character in a news studio with dogs, even though it missed the hockey detail and altered the clothing and pose significantly. FLUX.2 [klein] 4B is the winner because it actually attempted the complex creative task requested, whereas FLUX.1 Kontext [pro] only performed a style transfer on the person.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent soft lighting and bokeh effect.
- + High-quality fur texture rendering.
- − Failed to include the fox kit, instead generating two cats/rabbit hybrids.
- − The animals are sitting still rather than 'playfully chasing butterflies' as requested.
FLUX.2 [klein] 4B
- + Successfully included the retriever puppy, kittens, and fox kit.
- + Better captures the active 'chasing' and 'tumbling' motion from the prompt.
- + Distinct, detailed butterfly renderings.
- − Anatomy errors, particularly the kitten in the center having two tails.
- − Missing the baby bunny requested in the prompt.
Verdict: FLUX.2 [klein] 4B is the winner because it better captures the action and variety requested, including the distinct fox kit that FLUX.1 Kontext [pro] missed. While both models struggled to include all four specific animals correctly, FLUX.2 [klein] 4B provided a more dynamic scene that matched the joyful, chasing vibe described.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Captures the iconic Ghibli watercolor texture and line art perfectly
- + Preserves the exact facial expressions and poses of all three characters from the original meme
- + Maintains the vibrant red and blue color palette while translating it into an illustrative style
- − The lighting is a bit flat compared to the 'dreamy' request in the prompt
FLUX.2 [klein] 4B
- + Successfully applies a soft, pastel color palette and dreamy lighting
- + Excellent hand-painted watercolor aesthetic for the background and clothing
- − Fails to preserve the emotional context, turning the girlfriend's angry expression into a pleasant smile
- − The central man's expression is too neutral compared to the comical 'duck face' in the source image
Verdict: FLUX.1 Kontext [pro] is the clear winner because it successfully converts the photo into a Ghibli illustration while perfectly preserving the character's unique expressions that make the original meme recognizable. FLUX.2 [klein] 4B provides a beautiful 'dreamy' aesthetic, but it fails the editing task by changing the core narrative of the image, making the angry girlfriend look happy and disinterested.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent conservation of the subject's face and original features.
- + Highly realistic integration of the wind effect on the hair.
- − The number of flying leaves is very minimal, barely fulfilling that part of the prompt.
- − Change in the subject's posture and hand position creates some inconsistencies with the original.
FLUX.2 [klein] 4B
- + Successfully adds a large volume of flying leaves to create a very energetic feel.
- + Maintains the original scale and positioning of the subject and dog more accurately than Model A.
- − The face of the woman is noticeably altered from the source image.
- − Applying motion blurs to the foreground leaves occasionally looks like a post-process overlay rather than a seamless edit.
Verdict: FLUX.1 Kontext [pro] creates a much more natural-looking hair blowing effect and preserves the face better, but it fails to deliver on the 'flying leaves' request significantly. FLUX.2 [klein] 4B follows the instruction for leaves and energy much better, though it slightly changes the woman's facial features and some leaves look a bit artificial. FLUX.2 is the better choice for this specific prompt as it adheres to all parts of the motion request.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography rendering for 'Caffè Florian'
- + Clean and balanced banner layout
- + Appropriately subtle paper texture
- − Minor spelling error in the banner with 'EEST' instead of 'EST'
- − The cloche button and steam lines are slightly offset
FLUX.2 [klein] 4B
- + Classic serif typography choice
- + Good use of negative space in the cloche illustration
- + Better adherence to the 'subtle texture' on the background
- − Incorrectly spelled name as 'FLAXTION'
- − Redundant date text by including it twice
- − The banner shape has awkward, sharp geometric wings that clash with the vintage style
Verdict: FLUX.1 Kontext [pro] is the clear winner because it correctly spells the primary brand name 'Caffè Florian' and maintains a professional emblem style, despite a minor typo in the 'Est.' banner. FLUX.2 [klein] 4B fails significantly on text adherence by misspelling the brand name and repeating the foundation date unnecessarily.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography for the main title and most labels
- + Dynamic composition using a trajectory map
- + Clean vector aesthetic that adheres well to the NASA-inspired color palette
- − Confuses planets, showing a ringed planet (Saturn) near the rocket and top right instead of Earth/Moon
- − Steps are numbered and labeled incorrectly, with '4. Earth' and '2. Collins' making no logical sense
FLUX.2 [klein] 4B
- + Attempts a grid-based infographic layout that is easy to read
- + Includes specific requested icons like the Lunar Module and astronaut silhouettes
- + Correctly identifies components like the Lunar Module and provides a sense of the mission steps
- − Severe spelling errors in almost all text including the main header 'APLLO 1UT 1J'
- − The rocket design is distorted and lacks the crisp lines of a professional vector style
- − Failed to follow the specific 6-step sequence outlined in the prompt
Verdict: FLUX.1 Kontext [pro] creates a much more visually appealing and professional-looking poster, though its data accuracy is poor, confusing Saturn for Earth and creating nonsensical labels. FLUX.2 [klein] 4B followed the request for specific icons like the Lunar Module better, but the catastrophic text rendering and lack of vector cleanliess make it unusable as an infographic. FLUX.1 Kontext [pro] is the winner for its superior aesthetic quality and coherent text rendering, even though the content is factually confused.
Explore each model
Black Forest Labs' compact, open-source image generation model with sub-second inference, optimized for production and near real-time applications with multi-reference support