Black Forest Labs' aesthetically-tuned 12-billion parameter flow transformer optimized for high-quality images with incredible aesthetics, suitable for personal and commercial use
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Krea [dev]
#47 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Krea [dev]
0.0%
win rate
Ties
0.0%
Wan 2.6
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent photorealistic rendering of glass and reflections
- + Very clean, minimal composition
- + Accurately represents soft window light from the left side
- − The green plant is more of a background element and less visible 'through' the glass relative to Model B
Wan 2.6
- + Stronger adherence to the 'plant behind the cube visible through the glass' instruction
- + Realistic texture on the red book and wooden table
- + Accurate spatial placement of all objects
- − The lighting is a bit harsh compared to the requested 'soft' light
- − The glass cube has some slight geometric warping on the right edge
Verdict: Both models followed the spatial instructions perfectly. FLUX.1 Krea [dev] produced a cleaner, more aesthetically pleasing photographic image with superior lighting, while Wan 2.6 did a slightly better job of showing the plant directly through the glass as requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Successfully captures motion blur from passing cars
- + Genuinely feels like a candid 50mm shot with slightly imperfect framing
- + Natural skin textures on the man
- − The man is pushing or holding the bike rather than repairing it
- − The bicycle geometry is slightly mangled around the pedals and chain area
- − Lack of wetness on the man's clothing despite the rain
Wan 2.6
- + Excellent adherence to the 'repairing' action
- + Highly realistic skin textures and wetness details on clothing and objects
- + Superior handling of reflections and lighting from city storefronts
- − Misses the 'motion blur' requested for passing cars
- − The rain droplets on the jacket look a bit overly uniform/spherical
Verdict: Wan 2.6 is the clear winner for its superior storytelling and texture work, capturing the 'repairing' action with high fidelity and realistic wet surfaces. While FLUX.1 Krea [dev] followed the motion blur and framing instructions more closely, it failed to depict the primary action of the prompt and had significant anatomical issues with the bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Symmetry in composition and lighting is well-balanced.
- + Excellent engraving detail on the plate armor.
- + Clean, professional bokeh and lighting effects.
- − The 'scars and dirt' look more like blood splatters than integrated skin textures.
- − Hair braids are thin and less prominent compared to the prompt's focus.
Wan 2.6
- + Superior rendering of grit, dirt, and realistic skin texture for a 'battle-worn' look.
- + Prominent and beautifully detailed beads in the hair braids.
- + Exceptional texture work on frayed fabric and leather straps.
- − The background fire is slightly distracting from the figure.
- − The armor engraving is slightly less intricate than Model A.
Verdict: While both models followed the prompt well, Wan 2.6 provided a more authentic 'battle-worn' aesthetic with realistic skin textures and more distinct details on the leather and cloth layers. FLUX.1 Krea produced a very clean and noble portrait, but the blood-like splatters felt less like the requested dirt and scars compared to the gritty realism of Wan 2.6.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent grid layout with clear food photography.
- + Highly professional typographic hierarchy and clean sans-serif fonts.
- + Consistent and vibrant accent colors used for category headers.
- − Includes a duplicate 'Appetizers' section instead of a distinct 'Pizza' section.
- − The text is largely unreadable gibberish.
Wan 2.6
- + Adheres better to the section requirements including Pizza and Mains.
- + Successful use of vibrant, multi-colored geometric accents as requested.
- + Layout feels more expansive and modern for a casual dining setting.
- − Large amount of dead space in the 'Pizza' photo grid area.
- − Text contains significant spelling errors and incoherent characters.
Verdict: FLUX.1 Krea (dev) produces a more realistic and professionally balanced layout that looks like a real physical menu, though it fails to include the specific 'Pizza' section header. Wan 2.6 follows the prompt's structural requirements more closely by including all requested sections and vibrant accents, but the composition is slightly clunkier with more wasted space.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Clean and legible typography
- + Solid photographic lighting on the burger ingredients
- + High resolution for the main food subject
- − Missed the 'fiery, glowing effect' for the text entirely
- − The 'starburst' for the price is missing (it is in a separate logo at the top)
- − The 'exploded' effect is very static and lacks motion
Wan 2.6
- + Excellent adherence to the 'fiery, glowing' text style requested
- + Dynamic 'exploded' composition with a true sense of motion and dripping sauce
- + Perfect placement and rendering of the price starburst and text components
- − Slightly lower clarity on the fine textures of the lettuce and patty compared to the other model
- − The fiery background is a bit busy, making the bottom text slightly harder to read
Verdict: Wan 2.6 is the clear winner as it followed every stylistic instruction, particularly the fiery glowing text and the starburst price tag. While FLUX.1 Krea produced a clean image, it failed to apply the requested 'fiery' effects to the typography and the burger composition felt very static rather than energetic.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent spelling on the additional items at the bottom.
- + Clean, legible composition with centered framing.
- + Uses distinct cursive for the menu items as requested.
- − Text rendering style looks mechanical and digital rather than chalky.
- − Frequent spelling errors in the listed items, such as 'Gruffle' and 'tcoman'.
- − The title 'TODAY'S SPECIALS' is in block caps rather than the requested elegant cursive.
Wan 2.6
- + Superior chalk texture with realistic dusty residue and varying opacity.
- + Flawless adherence to all text content and prices in the prompt.
- + Atmospheric café lighting and realistic depth of field.
- − The frame of the chalkboard is slightly cut off at the edges.
- − Chalk dust at bottom might be a bit heavy for some tastes.
Verdict: Wan 2.6 is the clear winner as it followed every textual instruction perfectly with zero spelling errors. FLUX.1 Krea struggled with the specific menu items, hallucinating words like 'Gruffle' and failing to provide the realistic chalk texture requested, appearing more like a digital font.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Successfully followed the difficult logic instruction of placing the horse on top of the astronaut.
- + High cinematic quality with realistic planetary lighting.
- + The mechanical rig on the horse's belly adds to the surreal theme.
- − The horse's front legs and pose are slightly awkward anatomically.
- − Minimal color palette compared to the other model.
Wan 2.6
- + Beautiful composition with vibrant nebulae and lighting effects.
- + Highly detailed texture on the horse's coat and equipment.
- + Good interpretation of the surreal 'space' environment.
- − Failed the negative/switch logic; the astronaut is riding the horse.
- − Physics-wise, it places the scene on a dusty ground rather than in space.
Verdict: FLUX.1 Krea [dev] is the clear winner because it successfully interpreted the specific instruction 'horse on top, not vice versa,' creating a surreal and unique image. Wan 2.6 ignored the logic switch and produced a standard, though visually pretty, astronaut-on-horse image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Features a very high-quality, professional-looking taxi driver cap that fits the prompt's theme.
- + Excellent fur texture and lighting on the capybara.
- + The passenger in the back looks truly bored as requested.
- − Includes an extra person in the passenger seat not mentioned in the prompt.
- − The positioning of the capybara's paws on the wheel is less clear and anatomical than Model B.
Wan 2.6
- + Excellent composition that shows both subjects and the city lights more clearly.
- + Highly realistic raindrops on the windshield and vibrant Manhattan neon lights.
- + Great adherence to the request for both paws to be on the steering wheel.
- − The passenger looks slightly distressed rather than 'completely normal' or 'bored'.
- − The cap looks more like a police hat than a standard taxi driver cap.
Verdict: Both models struggled with the 'inside' perspective, opting for exterior shots through the window, but Wan 2.6 captured the atmosphere of New York at night much better with its rain and lighting effects. FLUX.1 Krea (dev) produced a very sharp image but hallucinated an extra person in the front seat, which cluttered the composition.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Elegant border with integrated spiderwebs
- + High contrast lighting in the jack-o-lantern
- + Clean graphic design aesthetic
- − Significant text errors like 'Pasty' instead of 'Party' and duplicated words
- − Font used in the bottom scroll is difficult to read
- − The trees lack texture and appear as flat silhouettes
Wan 2.6
- + Excellent text accuracy across the entire invitation
- + Highly detailed twisted trees with realistic texture
- + Great atmospheric lighting and moody blue night sky
- − The golden glitter effect on the main text may feel slightly less 'gothic' than requested
- − Some overlapping of thorny vine and cobweb borders looks slightly cluttered
Verdict: Wan 2.6 is the clear winner as it followed every text requirement perfectly with no spelling errors, whereas FLUX.1 Krea [dev] failed on the main title and multiple small details. Additionally, Wan 2.6 delivered much better texture and depth in the background elements like the twisted trees.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent typography with a friendly, rounded aesthetic that matches the 3D cartoon theme.
- + Very clean, high-clarity rendering of the salmon nigiri with PBR-like light interactions.
- + Solid adherence to the minimal garnish and diorama base constraints.
- − The flag icon is placed as a physical 3D prop rather than part of the text header as requested.
- − The diorama base is slightly less centered vertically compared to the other model.
Wan 2.6
- + Follows the layout instructions perfectly by placing the flag icon within the text section.
- + Provides a more varied 'sushi' scene with multiple types of nigiri and traditional garnishes.
- + Stronger 45° isometric perspective with a professional, clean lighting setup.
- − The text has minor kerning and alignment inconsistencies (e.g., the 'S' in SUSHI is slightly offset).
- − The garnish is slightly more than 'minimal' compared to the other model.
Verdict: Both models followed the prompt exceptionally well, but Wan 2.6 is the winner for its superior layout accuracy, placing the flag icon exactly where requested alongside the text. While FLUX.1 Krea produced a very pleasing 3D cartoon style, Wan 2.6 offered a more complete interpretation of the subject matter with better isometric framing.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent fur texture and lighting on the animals.
- + High clarity and sharp focus across the subjects.
- + Clean, artistic composition with a dreamlike quality.
- − Missed the baby bunny requested in the prompt.
- − Included two kittens instead of one kitten and one bunny.
Wan 2.6
- + Successfully included all requested animals: puppy, kitten, bunny, and fox.
- + Strong adherence to 'god rays' and 'dew sparkles' descriptors.
- + Dynamic and playful interaction between all four creatures.
- − The fox's front right leg is anatomically distorted.
- − Slightly more cluttered composition with some floating artifacts/seeds.
Verdict: Wan 2.6 is the clear winner for prompt adherence as it successfully included the baby bunny, whereas FLUX.1 Krea [dev] substituted it with a second kitten. Although FLUX.1 Krea [dev] has slightly more polished fur rendering, Wan 2.6 better captured the requested environmental effects like the dew and god rays while maintaining a cohesive, playful scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Successfully incorporates the banner and 'Est. 1720' text.
- + Captures an intricate woodcut/etching style that feels authentic to a vintage brand.
- + Includes the steam element above the cloche dome as requested.
- − Misspells the main name as 'Cafe FLANDRIN' instead of 'Caffè Florian'.
- − The composition is a bit cluttered with overlapping elements.
Wan 2.6
- + Correctly spells the restaurant name 'Caffè Florian'.
- + Features a clean, minimalist vector design that aligns with the 'logo' request.
- + Accurately represents the warm brown and cream tones with a clear cloche dome.
- − The 'Est. 1720' banner is small and placed awkwardly off to the side.
- − The typography is somewhat plain compared to the requested 'classic' style.
Verdict: Both models followed the stylistic cues well, but Wan 2.6 is the clear winner for its perfect spelling and professional logo composition. While FLUX.1 Krea produced a more artistic vintage illustration, the significant misspelling and cluttered typography make it less effective as a logo compared to the clean layout of Wan 2.6.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Captures the requested vector infographic style with icons.
- + Strict adherence to the requested NASA-inspired color palette.
- + Includes a sequence of numbered steps as requested in the prompt.
- − Icons do not accurately match the specific descriptions (e.g., Saturn V, lunar module).
- − Text labels are largely nonsensical or garbled.
- − Logical flow of the infographic is confusing and spatially inconsistent.
Wan 2.6
- + Text rendering for the titles and names is clean and accurate.
- + Uses a simple, modern design aesthetic.
- − Fails entirely to provide the 6-step infographic requested.
- − Missing the required icons (Saturn V, orbit rings, trajectory, lunar module).
- − Composition is empty and lacks the complexity of an informational poster.
Verdict: FLUX.1 Krea [dev] followed the complex instructions much better than Wan 2.6, attempting the specific 6-step infographic structure and iconography requested. While the icons in FLUX.1 Krea [dev] are somewhat generic and the text is messy, Wan 2.6 failed significantly on instruction following, omitting nearly all requested visual elements and steps.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English