Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of caustics and light refraction through glass
- + Strong cinematic lighting and depth of field
- + Clean, modern aesthetic
- − The plant is highly blurred and less 'partially visible through glass' than 'behind the scene'
- − The sphere has a textured/glittery appearance rather than a smooth finish
Wan 2.7
- + Perfect adherence to the spatial requirement of the plant seen through the glass
- + Realistic wooden table texture and natural window light
- + Higher degree of realism in the book and sphere textures
- − The glass cube has some structural inconsistencies in its internal edges
- − The reflection of the sphere on the left face of the cube looks slightly misplaced
Verdict: Both models followed the prompt instructions very well, but Model B (Wan 2.7) provided a more literal and successful interpretation of the plant being visible through the glass cube. FLUX.1 Kontext [max] produced a more aesthetically polished and artistic image with superior lighting, but Model B felt more grounded in reality and better captured the specific layout of the scene.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent shallow depth of field and bokeh effects
- + Strong photographic quality with realistic rain and reflections
- + Detailed interaction between the subject's hands and the bicycle chain
- − The rain looks slightly like long vertical streaks rather than droplets
- − The man's ethnicity appears somewhat ambiguous compared to the specific request
Wan 2.7
- + Perfectly captures the 'imperfect framing' and 'candid' aesthetic requested
- + Highly realistic skin textures and authentic Japanese urban setting
- + Natural clothing textures and convincing wet surfaces
- − Lacks the 'motion blur from passing cars' specified in the prompt
- − Depth of field is slightly deeper than a 50mm lens wide open would typically produce
Verdict: Wan 2.7 produces a much more authentic 'candid' street photo that feels like a real moment captured in Japan, effectively hitting the skin texture and 'no stylization' requirements. While FLUX.1 Kontext [max] has superior bokeh and lighting, it feels slightly more staged and less characteristic of a spontaneous street photograph. Wan 2.7 is preferred for its superior adherence to the 'candid' and 'natural' aspects of the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Ornate engraved plate armor is rendered with high detail and realistic reflections.
- + Excellent facial hair texture and intense eye detail.
- + Composition feels very intimate and meets the 'close portrait' requirement well.
- − Missed the request for small beads in the braided hair.
- − Scars are very faint or absent compared to the prompt's 'battle-worn' request.
Wan 2.7
- + Excellent adherence to the 'hair braided with small beads' and 'faint scars' prompts.
- + Excellent texture on leather straps and the cloth underlayer as requested.
- + Lighting accurately reflects warm torchlight across the armor and skin.
- − The sparks look a bit like static dots rather than dynamic bokeh elements.
- − The engraving on the armor is somewhat less intricate/ornate than the other model.
Verdict: Wan 2.7 is the clear winner as it successfully incorporated every specific detail of the prompt, including the beads in the hair, the scars, and the textures of the leather and cloth. While FLUX.1 Kontext [max] produced a high-quality cinematic image, it failed to include the beads and the battle-worn facial details requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Features high-quality, professional food photography that feels consistent.
- + Accurately represents the request for a minimalism with a clean white background.
- + Includes decorative script accents that add a touch of personality.
- − The text is largely illegible gibberish.
- − The grid is repetitive, featuring mostly identical-looking pizzas rather than a variety of food types.
Wan 2.7
- + Highly legible headers and realistic, believable prices and descriptions.
- + Features a varied grid of items including soup, salad, burger, pasta, and dessert.
- + Excellent use of bold sans-serif fonts and functional section dividers.
- − Includes some unnecessary background elements (pen, rosemary, olive oil) that weren't specifically requested.
- − Minor spelling errors in smaller text descriptions.
Verdict: Wan 2.7 is the superior choice because it generates a functional, highly readable menu with diverse food categories and clear typography. FLUX.1 Kontext [max] produces a visually pleasing aesthetic but fails significantly on text legibility and variety of dishes.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the meat patty and bun
- + Clean and legible typography
- + Realistic lighting that integrates the burger with the environment
- − The burger is not truly 'exploded' but remains mostly stacked
- − Misses the starburst requirement for the price tag
Wan 2.7
- + Perfect interpretation of the 'exploded' request with suspended components
- + Excellent fiery text effects on the title
- + Correctly applies all requested elements including the starburst for the price
- − The ingredients look slightly more illustrative/rendered than photorealistic
- − The price text is a bit small within its starburst container
Verdict: Wan 2.7 is the clear winner as it followed every instruction in the prompt, specifically the 'exploded' burger layout and the inclusion of the price within a starburst. While FLUX.1 Kontext [max] has more realistic food textures, it failed to separate the burger components or include the starburst, resulting in a more static image.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent chalk texture including authentic smudges and dust on the board.
- + Strong adherence to the 'handwritten' request with natural variations in letter size and spacing.
- + Accurate text rendering for all requested items.
- − The 'elegant cursive' requirement for the title was not fully met, favoring a print/script mix instead.
- − Lighting is a bit harsh/flat compared to the cozy cafe atmosphere requested.
Wan 2.7
- + Beautiful 'cozy café' aesthetic with warm lighting and wooden background.
- + Text is exceptionally clear and legible.
- + Successfully incorporated the cursive flourishes in the title text.
- − The text appears too perfect and uniform, resembling a digital font rather than natural chalk handwriting.
- − The chalk texture on the letters is very subtle, looking more like a graphic overlay than physical chalk.
Verdict: FLUX.1 Kontext [max] is the winner because it successfully captured the authentic feel of a handwritten chalkboard, including the physical imperfections and chalk dust requested in the prompt. While Wan 2.7 created a more visually pleasing 'cozy' environment, its text looks too much like a digital font, failing the primary requirement for a realistic handwritten style.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent cinematic lighting and texture on the horse's coat.
- + Strong anatomical consistency with a dynamic, jumping-like pose.
- − Failed the negative constraint; the astronaut is riding the horse, not the horse on top of the astronaut.
Wan 2.7
- + Rich background detail with galaxies and planets.
- + Clear and crisp textures on the spacesuit and horse tack.
- − Failed the negative constraint completely, providing a standard 'astronaut riding horse' image.
- − The legs of the horse are unnaturally long and thin compared to the body.
Verdict: Both FLUX.1 Kontext [max] and Wan 2.7 failed the specific spatial logic prompt to have the horse on top of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. FLUX.1 Kontext [max] is the preferred image because it features much more cinematic lighting and a more natural physical form for the horse compared to the awkward leg anatomy in Wan 2.7.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent fur texture and photographic lighting.
- + High-quality blurred background showing Manhattan lights.
- + Realistic rendering of the capybara's paw on the steering wheel.
- − The passenger appears to be in the front seat next to the driver, not the back seat.
- − The passenger is holding a phone to her ear rather than looking at it.
Wan 2.7
- + Successfully places the human in the back seat as requested.
- + The capybara has two paws on the steering wheel.
- + The passenger is looking at her phone with a bored expression.
- − The fur texture is much lower quality and looks more like an illustration than a photo.
- − The capybara's paws look more like bird talons or claws than capybara paws.
- − The taxi sign on top of the car is unnaturally large and distorted.
Verdict: Wan 2.7 adhered much better to the complex spatial prompt by placing the woman in the back seat and showing her looking at a phone, whereas FLUX.1 Kontext [max] put her in the front passenger seat. However, FLUX.1 Kontext [max] produced a significantly more realistic and aesthetically pleasing image with superior textures and lighting, while Wan 2.7 has noticeable anatomical and structural artifacts.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent atmospheric lighting and moody color palette.
- + Perfect text rendering with high clarity.
- + Effective integration of the jack-o-lantern within the background.
- − Redundant location text at the bottom.
- − The layout is a bit centered and heavy compared to a traditional flyer.
Wan 2.7
- + Ornate and intricate border design that adheres well to the 'webs and thorns' request.
- + Superior layout for a physical invitation with clear sections.
- + Creative inclusion of extra thematic elements like the cauldron and books.
- − Illustration style is a bit car-toonish rather than 'cinematic'.
- − Minor artifacts in the text rendering (e.g., 'Arches' is slightly cramped).
Verdict: FLUX.1 Kontext [max] excels in mood and cinematic lighting, creating a high-quality poster feel with perfect text, though it suffers from slight repetition in the event details. Wan 2.7 provides a more authentic 'invitation' layout with a beautiful border and more visual variety, but the overall style is less cinematic and more illustrative. FLUX.1 wins on overall visual impact and atmospheric quality.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully added a large volume of hair with high detail.
- + Maintains the rugged aesthetic of the original character.
- − Significantly altered facial features, making the person look younger and different from the source.
- − The hair texture appears slightly metallic or over-sharpened compared to the beard.
Wan 2.7
- + Excellent preservation of original facial features and identity.
- + The added hair integrates perfectly with the existing beard and lighting.
- + Highly realistic, natural-looking hairline and texture.
- − None identified for this specific task.
Verdict: Wan 2.7 is the clear winner as it successfully added the requested hair while perfectly preserving the person's identity and the image's original lighting. FLUX.1 Kontext [max] added the hair but fundamentally changed the person's face, failing the source preservation requirement of the editing task.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent 3D toy-like texture and material rendering
- + Perfectly clean and centered composition
- + Clear bold text that follows the requested style
- − Missed the request for a small flag icon
- − The diorama base is a bit plain compared to the subject
Wan 2.7
- + Successfully included the Japanese flag icon mentioned in the prompt
- + More variety in the sushi types which makes for a more interesting diorama
- + High level of detail in the food textures and miniature props
- − Text is slightly off-center and the layout feels more cramped
- − Minor artifacts present on the shrimp piece and the smaller garnishes
Verdict: Both models followed the prompt well, but Wan 2.7 is the likely winner for including all requested elements, including the small flag icon which FLUX.1 Kontext [max] omitted. While FLUX.1 Kontext [max] has slightly cleaner aesthetics, the variety and completeness of the scene in Wan 2.7 better capture the 'isometric miniature diorama' concept.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Captures the 'TV news anchor' vibe very clearly with the desk and microphone setup.
- + Maintains high character consistency by keeping the denim jacket from the source image.
- + Includes all elements: dogs, hockey (stick in top right), and profession.
- − The hockey element is cut off at the edge of the frame.
- − The addition of glasses deviates significantly from the original subject's appearance.
Wan 2.7
- + Excellent caricature style that remains recognizable as the person from the source image.
- + Very creative implementation of the hockey theme with a miniature rink and a dog in a helmet.
- + Clear text rendering in speech bubbles adds to the humorous 'caricature' feel.
- − The 'Dogs Rolee!' text contains a small spelling error.
- − The composition is a bit crowded with many floating elements.
Verdict: Both models followed the instructions well, but Wan 2.7 produced a superior caricature that balanced humor and detail much better than FLUX.1 Kontext. Wan 2.7's interpretation included creative touches like the dog in a hockey helmet and a miniature rink, while FLUX.1 Kontext added unnecessary glasses and cut off the hockey stick, though it did a better job of preserving the subject's original clothing.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent soft lighting and dreamy glow that matches the 'wholesome vibe'.
- + High degree of fur detail and realistic textures for all four animals.
- + Strong composition with animals looking toward the light source and butterflies.
- − The rabbit has a somewhat creepy, human-like facial structure.
- − The animals are quite static and posed rather than 'playfully chasing'.
Wan 2.7
- + Successfully captures dynamic movement and action with the animals in mid-stride.
- + Excellent dew sparkles on the grass which was specifically requested in the prompt.
- + The fox and kitten have very realistic, species-appropriate facial features.
- − The golden retriever puppy's front paw has a slightly awkward anatomical bend.
- − The lighting feels a bit more digital and less natural than Model A.
Verdict: Wan 2.7 is the superior choice for this prompt as it truly captured the 'playfully chasing' and 'tumbling' aspect of the request, whereas FLUX.1 Kontext [max] produced a more static, posed portrait. Wan 2.7 also accurately included the dew sparkles in the grass, providing a more comprehensive adherence to the specific details requested.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfectly captures the Studio Ghibli cel-shaded character design and aesthetic.
- + Maintains the composition and iconic poses of the original meme flawlessly.
- + ExCELLENT color rendering with soft, painterly textures on the clothing and backgrounds.
- − The facial expressions on the man and the girlfriend are slightly softened/muted compared to the original intensity.
Wan 2.7
- + Strong watercolor texture that feels very hand-painted.
- + Preserves the specific facial features and expressions of the actors better than Model A.
- − The style feels more like a standard Western watercolor sketch rather than the specific 'Studio Ghibli' anime aesthetic requested.
- − Line work is a bit scratchy and lacks the clean line-art style of Ghibli films.
Verdict: FLUX.1 Kontext [max] is the clear winner as it accurately interprets the 'Studio Ghibli' instruction, producing an image that looks like a legitimate still from an anime. While Wan 2.7 does a good job of preserving the source actors' likenesses in a painted style, it fails to capture the specific stylistic hallmarks of the requested animation studio.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Strong prompt adherence with wind-blown hair and visible green leaves.
- + Added a dynamic walking pose for the woman and dog which increases the lively feel.
- + Improved facial expression and hair texture.
- − Significant change to the original composition, including the dog's position and the woman's arm posture.
- − Occasional anatomical oddities like the extended left hand fingers near the dog's head.
Wan 2.7
- + Excellent source preservation, keeping the main figures and background almost identical.
- + Effective use of falling leaves and wind-tousled hair while maintaining image integrity.
- + Natural integration of the motion effects into the existing scene.
- − Motion effect on the hair is slightly less dramatic than in the other model.
- − The leaves appear somewhat static rather than 'flying' with motion blur.
Verdict: FLUX.1 Kontext [max] creates a much more 'energetic' and 'lively' scene by re-posing the characters to look like they are in mid-stride, though it fails at source preservation by changing the image too much. Wan 2.1 is the superior image editor, as it successfully adds the wind and leaves while perfectly preserving the identity, pose, and background of the original photo.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography including the requested accent on 'Caffè'
- + Perfect implementation of a minimalist vintage aesthetic
- + Highly professional woodblock/stamp texture that adds character
- − The steam element is very simple and could be more graceful
Wan 2.7
- + Ornate emblem style with good balance of extra elements like stars and leaves
- + Accurate color palette according to the warm brown and cream request
- − Misspelled name as 'Florion' instead of 'Florian'
- − The cloche is transparent, which looks more like a modern display case than a classic metal cloche
- − Composition is a bit cluttered for a 'minimalist' request
Verdict: FLUX.1 Kontext [max] produced a superior, professional-grade logo that perfectly followed the minimalist vector emblem style and correctly spelled the name. Wan 2.7 failed on the typography with a spelling error and missed the minimalist mark by creating an over-detailed, circular badge.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Strong artistic vector style with clean linework.
- + Adheres well to the requested color palette.
- + Good spatial composition and layout.
- − Nonsensical text labels such as 'UNONY MODULLE' and 'TRANSLUNAR LAARTH'.
- − Confuses celestial bodies, labeling the Earth as 'MOON'.
- − Includes a random ringed planet (Saturn) which wasn't requested and isn't part of the mission path.
Wan 2.7
- + Excellent adherence to the sequential 1-6 step infographic structure.
- + Highly accurate and readable text including dates, times, and velocities.
- + Very clean, professional NASA-inspired aesthetic with consistent iconography.
- − Small spelling error in 'DESCRIPT' and 'TRANQUILIRY'.
- − The icons are very small relative to the total canvas size.
- − Minor misalignment in the vertical dotted line connecting the icons.
Verdict: Wan 2.7 is the clear winner as it successfully created a logical, sequential infographic that follows all 6 requested steps with relevant data and high text legibility. In contrast, FLUX.1 Kontext [max] produced a visually pleasing illustration but failed significantly as an infographic, mislabeling the Earth as the Moon and generating gibberish text.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits