FP8 quantized variant of Black Forest Labs' FLUX.1 [schnell] model, offering ~2x faster inference with reduced precision while maintaining high-quality image generation in 4 steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell] FP8
#47 of 62 in Text-to-Image
Wan 2.7
#38 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell] FP8
50.0%
win rate
Ties
25.0%
Wan 2.7
25.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent photographic quality and lighting.
- + Clean, minimalist aesthetic that feels modern.
- + Accurately represents the transparency of the glass cube.
- − The cube geometry is slightly tall (rectangular prism) rather than a perfect cube.
- − The internal blue sphere is floating on a floating shelf rather than sitting on the base.
Wan 2.7
- + Highly realistic textures on the wooden table and red book.
- + Internal reflections and refractions of the blue sphere are very convincing.
- + Perfectly follows all spatial instructions in the prompt.
- − The glass cube appears to have thicker, mirrored edges which creates a double image of the sphere.
- − The plant in the background is less 'visible through the glass' compared to image A.
Verdict: Both models followed the complex prompt instructions perfectly. Wan 2.7 produced a more realistic scene with impressive material textures on the book and table, while FLUX.1 [schnell] FP8 opted for a cleaner, more stylized commercial photograph look. Wan 2.7 is the winner due to the superior handling of realistic textures and the correct placement of the sphere within the cube's volume.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent use of color with the vibrant red bicycle and warm reflections.
- + Cinematic lighting and composition that draws the eye to the subject.
- + Highly detailed facial texture and skin realism.
- − Failed to include motion blur for passing cars.
- − The bicycle geometry is slightly warped near the hands and seat post.
Wan 2.7
- + Successfully captured visible rain drops and more realistic wet fabric textures.
- + The bicycle design is highly realistic for a Japanese context.
- + Accurate street atmosphere that feels authentically like a Japanese city.
- − Faces in the background are strangely obscured or blurred in an unnatural way.
- − Missing the motion blur for passing cars as requested in the prompt.
Verdict: Both models failed to deliver the motion blur on passing cars, but they excelled in different artistic directions. FLUX.1 [schnell] FP8 produced a more cinematic, high-contrast image with superior skin details, while Wan 2.7 delivered a more authentic 'candid' street photography feel with better environmental textures like wet clothing and falling rain.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong cinematic lighting with vibrant warm tones.
- + Intense character expression and lifelike eye detail.
- + Clean, modern digital painting aesthetic.
- − Missed the request for 'braided hair with small beads' almost entirely.
- − Overly smooth skin texture lacks the requested 'faint scars and dirt'.
- − Armor detail is visible but less 'engraved' than Model B.
Wan 2.7
- + Excellent adherence to the 'braided hair with beads' and 'scars' prompt.
- + Highly detailed engraved plate armor that feels authentic and battle-worn.
- + Great mechanical detail on leather straps and buckles.
- − The lighting is a bit flat compared to the dramatic 'torchlight' in Model A.
- − Facial expression is somewhat static and less intense.
Verdict: Wan 2.7 is the clear winner for prompt adherence, accurately depicting the specific braided hair, beads, and scars that FLUX.1 [schnell] FP8 ignored. While FLUX.1 [schnell] FP8 offers a more dramatic and colorful lighting setup, Wan 2.7 provides a much more convincing historical fantasy aesthetic with superior texture on the armor and leather.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Features a true white background as requested
- + Clean and simple layout that mimics a physical bi-fold menu
- − Text is largely gibberish and poorly rendered
- − Food photos lack diversity in color and presentation
- − Significant layout errors like heading overlaps and spelling mistakes like 'PIZZAL'
Wan 2.7
- + Excellent text legibility and realistic layout
- + High-quality, vibrant food photography in a clear grid
- + Strong adherence to the 'vibrant accents' and 'professional' prompt requirements
- − Includes extraneous physical objects like a pen and olive oil which were not part of the design prompt
- − Small spelling errors in dish names like 'Bisge' and 'Calanfri'
Verdict: Wan 2.7 produced a much more professional and usable design that strictly followed the grid and font requirements, whereas FLUX.1 [schnell] FP8 struggled with text rendering and basic spelling. While Wan 2.7 included unwanted background props, the core menu design is vastly superior in both visual quality and prompt adherence.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong photographic rendering of the burger texture.
- + Effective use of real fire at the bottom of the frame.
- − Numerous spelling errors including 'LIIMITED' and 'LIMIED TIME NEEY'.
- − Incorrect price displayed as €69 and €0.99 instead of €6.99.
- − The burger is largely intact rather than being a fully exploded view of all components.
Wan 2.7
- + Perfect text rendering for all requested copy including the correct price.
- + Excellent adherence to the 'exploded' request with clear separation of ingredients.
- + Superior interpretation of the fiery glowing effect on the typography.
- − Some seeds on the bun appear slightly disconnected or floating unnaturally.
- − The lettuce and sauce have a slightly more illustrative feel rather than strictly photorealistic.
Verdict: Wan 2.7 is the clear winner as it followed every instruction in the prompt, including complex text rendering and the specific 'exploded' layout. FLUX.1 [schnell] FP8 failed significantly on the text, producing major typos and the wrong price, while also failing to properly separate the burger components.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent chalk texture on the board surface and lettering
- + Realistic variation in handwriting style that feels human
- − Numerous spelling errors including 'Mushnroom', 'Octopuss', and 'malk cokies'
- − Text is repetitive and cluttered with nonsensical price duplicates
Wan 2.7
- + Perfect spelling adherence for all menu items and the date
- + Consistent and attractive cursive-style font
- + Clean and professional composition with warm lighting
- − The font looks digital and lacks the requested 'natural variations in letter size' and 'chalk texture'
- − Lettering is too perfectly uniform, making it look like a graphic overlay rather than real chalk
Verdict: Wan 2.7 produced a much more usable image with perfect spelling and a clean layout, although it ignored the request for realistic chalk texture and natural handwriting variations. FLUX.1 [schnell] FP8 captured the appearance of real chalk on a board much more effectively but failed significantly on text accuracy and spelling.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent adherence to the 'horse on top' spatial instruction
- + Cinematic lighting and high level of detail on the horse and machinery
- + Strong surreal atmosphere with the horse riding a life-support-style backpack
- − Confusing anatomy with multiple horse segments or a second head emerging from the back
Wan 2.7
- + Clean, high-resolution rendering
- + Good composition with various celestial elements
- + Accurate depiction of an astronaut suit and horse gear
- − Failed the core prompt instruction of 'horse on top, not vice versa'
- − Less 'surreal' than requested, opting for a standard fantasy trope
Verdict: This was a logic-test prompt where the model was asked to invert the typical 'astronaut riding horse' scene. FLUX.1 [schnell] FP8 correctly interpreted the instruction by placing the horse on top of the astronaut's equipment, whereas Wan 2.7 defaulted to the common trope and failed the primary constraint.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent close-up detail on the capybara's fur and expression.
- + High-quality lighting that creates a cinematic nocturnal atmosphere.
- + Clean text rendering on the hat and background taxi.
- − The woman is holding two phones simultaneously, which was not requested.
- − The perspective makes the capybara look unusually large compared to the interior.
- − Paws are not clearly on the steering wheel as requested.
Wan 2.7
- + Accurately depicts the capybara with both front paws on the steering wheel.
- + More realistic wide-angle perspective of a car interior.
- + The woman's 'bored' expression better matches the requested tone.
- − The woman is sitting in the passenger seat instead of the back seat.
- − The capybara's head has a slightly 'pasted on' look with sharp edges against the jacket.
- − Lower overall sharpness and texture detail compared to the other model.
Verdict: FLUX.1 [schnell] FP8 produces a more visually stunning and detailed image with better lighting, though it fails on subtle logic like the number of phones and paw placement. Wan 2.7 follows the mechanical instructions of the prompt more closely (paws on wheel, bored expression) but fails the spatial requirement of having the passenger in the back seat. Ultimately, FLUX.1 [schnell] FP8 is preferred for its superior photorealism and artistic quality despite the second phone.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Atmospheric cinematic lighting
- + Central glowing jack-o-lantern fits the prompt well
- − Numerous spelling errors in text
- − Date and time details are garbled
- − Border looks distorted like torn fabric rather than a designed frame
Wan 2.7
- + Excellent text rendering with perfect spelling
- + Highly detailed gothic border with thorns, webs, and skulls
- + Clearer composition that balances the vintage parchment feel with a cinematic illustration
- − Illustration style leans more towards a storybook than a 'dark parchment poster'
Verdict: Wan 2.7 is the clear winner as it followed the complex text instructions perfectly, whereas FLUX.1 [schnell] FP8 struggled with spelling and legibility. Wan 2.7 also executed the specific border and banner requests with much higher artistic detail and layout coherence.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent soft-shading and 3D cartoon surface quality
- + Good use of isometric perspective
- + Clean, minimal composition
- − Text rendering is broken and fails to follow the prompt instructions
- − Sushi models are generic and lack varied detail
Wan 2.7
- + Perfect text rendering following exactly what was requested
- + Higher level of detail and variety in sushi types (shrimp, tuna, salmon)
- + Includes thoughtful additions like soy sauce and ginger while remaining clean
- − Chopsticks are slightly merged into the plate geometry
Verdict: While FLUX.1 [schnell] FP8 captures the 'soft cartoon' lighting very well, it fails completely on the text and flag components. Wan 2.7 provides a much more accurate interpretation of the prompt, delivering perfect text, a variety of detailed sushi models, and a more polished diorama feel.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent vibrant golden lighting and warm atmosphere
- + Consistent stylistic approach across all animals
- + Very high emotional appeal with expressive eyes
- − Missed the baby bunny entirely, replacing it with extra kittens/foxes
- − Animals are oddly fused together in a cluster
- − Anisotropic distortion on some of the butterflies
Wan 2.7
- + Successfully included all four requested species: puppy, kitten, bunny, and fox
- + Strong adherence to the 'playfully chasing' and 'tumbling' action
- + Superior rendering of 'dew sparkles' and distinct blades of grass
- − The fox's face looks slightly more mature than a 'kit'
- − The transition between the kitten and bunny's paws is a bit awkward
- − Slightly less 'masterpiece' lighting compared to the dramatic warmth of the other model
Verdict: Wan 2.7 is the clear winner as it successfully followed the complex prompt requirements, including all specific animal types (golden retriever, tabby kitten, bunny, and fox), whereas FLUX.1 [schnell] FP8 failed to include the bunny. Wan 2.7 also better captured the requested movement of animals chasing butterflies in a realistic meadow with visible dew sparkles.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong minimalist vector aesthetic
- + Good adherence to the requested warm brown and cream color palette
- − Significant spelling errors in the brand name ('AFe FLAMILAN')
- − The central icon resembles a building dome rather than a food cloche
- − The typography layout is unbalanced and messy
Wan 2.7
- + Excellent rendering of the cloche dome as requested
- + Clean, professional emblem composition
- + Much more accurate text rendering despite a small vowel error
- − Spelling error in the name ('Florion' instead of 'Florian')
- − Less 'vintage' feel than the prompt suggests due to very clean lines
Verdict: Wan 2.7 is the clear winner as it correctly depicted the requested cloche dome with steam and maintained a high level of visual clarity and composition. FLUX.1 [schnell] FP8 failed significantly on the text rendering and the central icon, which looks more like a cathedral dome than a restaurant cloche.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Features a very clean, professional navy-based color palette.
- + Excellent vector aesthetic with crisp, modern icons and layout.
- − Text is largely gibberish with frequent misspellings.
- − Icons don't strictly follow the requested technical sequence (labels like 'Saturn Vicon' are literal interpretations).
Wan 2.7
- + Significantly better text rendering with legible and accurate information.
- + Includes the full six-step sequence requested in a clear, vertical hierarchy.
- + Successfully incorporates additional details like the three astronauts and NASA logo.
- − Suffers from a few minor typos like 'Descript' for Descent and 'Tranquiliry'.
- − Layout is a bit more crowded compared to the spacious feel of Model A.
Verdict: Wan 2.7 is the clear winner because it successfully carries out the complex instructional nature of the prompt, providing legible text and the correct chronological sequence of the mission. While FLUX.1 [schnell] FP8 has a slightly more sophisticated graphic design feel, its inability to produce readable or accurate text makes it less effective as an infographic.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits