FP8 quantized variant of Black Forest Labs' FLUX.1 [schnell] model, offering ~2x faster inference with reduced precision while maintaining high-quality image generation in 4 steps
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell] FP8
#47 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell] FP8
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent photographic realism and lighting
- + Detailed blue sphere with internal textures
- + Clean and modern composition
- − The glass object is a rectangular prism, not a cube
- − Includes a strange horizontal glass shelf inside the container that wasn't requested
Qwen Image Max
- + Accurately depicts the glass object as a cube
- + Excellent adherence to spatial relationships, showing the plant through the glass
- + Realistic lighting and shadow play on the wooden table
- − The blue sphere appears to be floating without visible support
- − Black pattern on the red book was not specified in the prompt
Verdict: Qwen Image Max followed the spatial instructions more accurately, providing a true cube whereas FLUX.1 [schnell] FP8 generated a tall rectangular prism. Qwen Image Max also handled the transparency effect of the plant behind the glass more effectively, though FLUX.1 [schnell] FP8 had slightly more vibrant and aesthetically pleasing material textures.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Extremely high skin texture and pore detail
- + Dramatic and accurate lighting contrast
- + Intense, lifelike eye rendering
- − Missed the request for 'braided hair with small beads'
- − The armor is very dark and lacks the requested ornate engravings
- − Composition is perhaps too close, losing context of the leather and cloth layers
Qwen Image Max
- + Perfect adherence to all prompt elements including hair beads and scars
- + Excellent rendering of ornate engraved plate armor
- + Great inclusion of bokeh sparks and torchlight source
- − Skin texture is slightly over-sharpened compared to a natural photograph
- − Depth of field is slightly deeper than requested
- − Focus is slightly softer on the eyes compared to the armor
Verdict: Qwen Image Max is the superior model for this prompt as it followed every instruction, including the specific hair style and detailed armor engravings that FLUX.1 [schnell] FP8 ignored. While FLUX.1 [schnell] FP8 has more realistic skin and lighting, it failed to represent the character's gear and attributes accurately.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + High resolution and clean rendering of the central burger.
- + Effective use of negative space for ad layout.
- − Significant text errors including typos and incorrect price.
- − The burger is mostly assembled rather than 'exploded' into individual components.
- − Text lacks the requested 'fiery, glowing effect'.
Qwen Image Max
- + Perfect text adherence for all requested phrases including the price.
- + Successfully applied the fiery glowing effect to all text elements.
- + Better sense of motion with flying ingredients and embers.
- − The 'starburst' for the price is more of a glowing flare than a traditional starburst shape.
Verdict: Qwen Image Max is the clear winner as it followed every instruction, including complex text rendering, fiery effects, and the specific price of €6.99. In contrast, FLUX.1 [schnell] FP8 failed significantly on the text, producing several spelling errors and an incorrect price of €69.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent text rendering on the taxi signs.
- + Well-lit and vibrant colors consistent with Manhattan at night.
- + Clearer view of the passenger holding her phone.
- − The passenger is holding two phones simultaneously, which is an odd artifact.
- − The taxi driver cap looks like a plastic toy rather than a professional uniform hat.
Qwen Image Max
- + Superb texture on the capybara's fur and jacket.
- + Much more realistic taxi interior with wear and tear on the ceiling.
- + The capybara's hands look more natural on the steering wheel.
- − The passenger's face is slightly blurry compared to the driver.
- − The capybara's head is disproportionately large for the body.
Verdict: Qwen Image Max is the winner due to its superior photorealistic textures and the gritty, realistic atmosphere of the taxi interior. While FLUX.1 [schnell] FP8 handles the lighting and text well, it suffers from a major logical artifact where the passenger is holding two phones.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong cinematic lighting on the central pumpkin
- + Moody, atmospheric background with subtle silhouettes
- − Significant text rendering errors and misspellings throughout the image
- − Banner design is split into two disjointed pieces
- − Lacks the requested thorns and webs in the border decoration
Qwen Image Max
- + Perfect text rendering for all requested details including the address and date
- + Accurate adherence to all design elements like webs, thorns, and twisted trees
- + Excellent vintage gothic aesthetic with clear, readable typography
- − The parchment texture is a bit clean, leaning more toward 2D illustration than cinematic
Verdict: Qwen Image Max is the clear winner as it followed every instruction perfectly, including complex text rendering with zero typos. FLUX.1 [schnell] FP8 struggled significantly with the text, producing several nonsensical words and failing to include the specific border elements requested.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent soft 3D stylized aesthetic.
- + Clean isometric composition with a nice layered base.
- + High-quality lighting and shaders.
- − Text rendering is broken, displaying 'JAPAN' and 'JAPAN' with a dot.
- − Missing the word 'SUSHI'.
- − The rice texture is slightly amorphous.
Qwen Image Max
- + Perfect adherence to text instructions, including 'JAPAN', 'SUSHI', and the flag icon.
- + Excellent realistic PBR textures on the fish and rice.
- + Very clean glass-like diorama base.
- − Dropped the 'minimal garnish' instruction by adding many green onions.
- − The transition between the fish and avocado looks slightly artificial.
Verdict: While FLUX.1 [schnell] FP8 captured the 'cartoon scene' style more effectively with its soft shapes, it failed significantly on the text requirements. Qwen Image Max followed every specific instruction including complex text placement and the flag icon, while providing much more realistic food textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent soft lighting and bokeh effect
- + Vibrant, warm color palette that fits the 'joyful' vibe
- + High-quality fur textures on the animals
- − Failed to include a rabbit
- − Includes multiple kittens instead of a variety of animals
- − Anatomical issues with the fox/kitten hybrid in the bottom right
Qwen Image Max
- + Includes all requested animals: retriever, kitten, fox, and the additional puppy for the 'tumbling' action
- + Dynamic composition that captures the 'tumbling together' instruction perfectly
- + Beautiful rendering of god rays and wildflower meadow variety
- − The bunny is missing from the group
- − Slightly less 'fluffy' texture on the fox compared to the retriever
Verdict: While both models failed to include the bunny, Qwen Image Max is the superior image as it correctly interpreted the 'tumbling together' action and included a much better variety of the requested animals. FLUX.1 [schnell] FP8 produced a beautiful but repetitive group of kittens with anatomical errors, whereas Qwen Image Max captured the dynamic playfulness and specific species requested in the prompt with greater accuracy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean vector aesthetic
- + Includes the established date correctly
- − Serious spelling errors in the brand name ('AFe FLAMILAN')
- − The icon resembles a dome/cupola more than a food cloche
- − The typography is messy and overlapping
Qwen Image Max
- + Perfect spelling of 'Caffè Florian'
- + Accurate representation of a retro cloche dome with stylized steam
- + Professional composition and high-quality shading
- − The 'Est. 1720' text is slightly off-center on its banner
- − Textured background is a bit heavy-handed compared to the minimalist request
Verdict: Qwen Image Max successfully executed almost all aspects of the prompt, most importantly spelling the brand name correctly and providing an accurate 'cloche' icon. FLUX.1 [schnell] FP8 failed significantly on the text rendering, producing a nonsensical name, and interpreted the cloche as more of an architectural dome.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.