Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0.0%
win rate
Ties
0.0%
Z-Image Turbo
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent sharp resolution and vibrant colors.
- + Captures the soft window light and reflections on the wood surfaces well.
- + Highly detailed plant foliage.
- − Failed the prompt by adding an extra blue sphere on top of the book.
- − The sphere inside appears to be floating rather than resting.
Z-Image Turbo
- + Followed the spatial prompt instructions perfectly with all objects in the right place.
- + Realistic glass behavior and reflection on the bottom surface.
- + Natural, believable indoor lighting.
- − The plant is very blurry and lacks detail compared to the other model.
- − Slightly less crisp texture on the book and table.
Verdict: Z-Image Turbo is the clear winner for prompt adherence, correctly placing each object as described, whereas FLUX.1 [schnell] hallucinated an additional sphere on top of the book. While FLUX.1 [schnell] has superior image sharpness and background detail, it failed the fundamental task of representational accuracy.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent shallow depth of field and bokeh realism.
- + Superior cinematic lighting and reflection quality on the pavement.
- + Strong adherence to the 'imperfect framing' request with a candid feel.
- − The motion blur on the car feels slightly static and lacks directional streaks.
- − Minor anatomical confusion where the hands meet the handlebars.
Z-Image Turbo
- + Natural and detailed skin texture on the subject's arms.
- + Visible rain streaks providing clear atmospheric context.
- + Authentic bicycle geometry for a 'mamachari' style bike.
- − Lacks the requested motion blur for passing cars, which appear sharp.
- − The composition is a bit flat and lacks the 'cinematic' depth requested.
Verdict: FLUX.1 [schnell] captures the request for a cinematic look with much better depth of field and beautiful lighting, though it struggles slightly with the bicycle's mechanical details. Z-Image Turbo provides a very realistic subject and captures the rain well, but it fails to incorporate the requested motion blur and feels more like a standard snapshot than a cinematic street photo. FLUX.1 [schnell] is preferred for its superior atmosphere and adherence to the technical photography prompts.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Intense, sharp facial details and micro-textures on skin
- + High-contrast lighting creates a dramatic atmosphere
- + Accurate shallow depth of field on the background
- − Fails to show silver 'engraved plate armor' clearly, appearing more like dark studded leather
- − The hair beads are oversized and look more like leather binding
Z-Image Turbo
- + Perfect adherence to 'engraved plate armor' with clear decorative filigree
- + Multiple braids clearly decorated with small silver beads
- + Includes visible bokeh sparks and the torch light source explicitly
- − The facial skin texture is slightly softer compared to the other model
- − Composition is a medium shot rather than the requested 'close portrait'
Verdict: Z-Image Turbo is the clear winner for prompt adherence, accurately depicting every specific detail from the hair beads to the ornate silver plate armor. While FLUX.1 [schnell] provides a more intense high-resolution close-up, it fails to deliver the requested armor type and fine decorative details.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean, minimalist layout consistent with professional menu design
- + Correctly included headers for Pizza and Appetizers
- + Good management of white space
- − Text rendering for menu items is illegible and distorted
- − Misspelled the 'Mains' category as 'Orfefus'
- − The grid of food photos lacks consistency in lighting and style
Z-Image Turbo
- + Excellent high-resolution food photography that looks appetizing
- + Stronger adherence to the 'grid' and 'vibrant accents' request with the orange bars
- + Cleaner rendering of prices and bold sans-serif fonts
- − Major spelling error in a central header ('PIZZA MANS')
- − Included a nonsensical section header ('SE TIIION')
- − The layout is a bit cramped compared to the minimalist request
Verdict: FLUX.1 [schnell] captures the 'minimalist' aesthetic more accurately with its elegant use of white space, but it fails significantly on text legibility and category naming. Z-Image Turbo produces much higher quality food imagery and a more vibrant layout, though it suffers from a glaring typo in the main header. Z-Image Turbo is the likely winner for its superior visual impact and professional-grade food photography.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic texture on the meat and bun
- + Very high-quality embers and fire effects
- + Dynamic movement with debris
- − Failed the primary text prompt with 'AGIC BURGER'
- − Repeated messy price text at bottom right
- − Burger is mostly assembled rather than 'exploded' with suspended components
Z-Image Turbo
- + Perfect text rendering for all requested strings
- + Accurately applied the fiery, glowing effect to the text
- + Clean starburst implementation for the price
- − The burger is not at all 'exploded' or separated
- − Visual quality has a slightly more plastic, AI-generated look compared to Model A
- − The background is less detailed
Verdict: Z-Image Turbo is the clear winner for this task because it correctly rendered all pieces of requested text and applied the 'fiery' style to them as requested. FLUX.1 [schnell] produced a more photorealistic burger with better fire effects, but failed significantly on the typography with misspellings and redundant numbers.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features a realistic cafe background setting
- + Good chalk texture on the individual letters
- − Terrible text spelling with many nonsensical words
- − Incorrect date formatting ('Pril' instead of 'April')
- − Included multiple $ symbols and incorrect pricing numbers
Z-Image Turbo
- + Excellent text accuracy with perfectly readable menu items
- + Superior chalk texture and realistic smudging on the board surface
- + Followed the date and price instructions precisely
- − Text is very orderly and lacks the 'elegant cursive' requested for the title
- − Slight spelling error in 'Mustroom' for Mushroom
Verdict: Z-Image Turbo is the clear winner as it produces coherent, readable text that follows the prompt's specific menu items and pricing accurately. FLUX.1 [schnell] failed significantly on the text generation task, producing gibberish words and failing to spell the month correctly.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully followed the difficult spatial constraint of the horse being on top of the astronaut
- + Excellent cinematic lighting and composition with the planetary background
- + High level of detail on the space suit and saddle
- − The horse has two heads/torsos appearing to merge into one tail
- − The astronaut's posture is somewhat distorted to fit the horse
Z-Image Turbo
- + Natural and realistic textures on the horse's coat and mane
- + Clean, high-resolution rendering of the space suit
- − Completely failed the negative constraint/instruction for the horse to be on top
- − Lacks the 'surreal' quality requested in the prompt
- − Traditional subject matter that ignores the specific user request
Verdict: FLUX.1 [schnell] is the clear winner as it successfully interpreted the challenging 'horse on top' spatial instruction, creating a truly surreal image. While the horse in FLUX.1 [schnell] has anatomical artifacts (two heads), Z-Image Turbo failed the core prompt requirement by placing the astronaut on top of the horse.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent interior taxi lighting and texture realism.
- + Accurately represents the 'bored' expression of the passenger.
- + High detail in the capybara's fur and clothing.
- − One paw is missing from the steering wheel, failing the 'both front paws' prompt.
- − The hat is more of a beanie/cloche style rather than a traditional driver cap.
Z-Image Turbo
- + Successfully placed both paws on the steering wheel as requested.
- + The capybara's head shape and profile feel very authentic.
- + The hat design perfectly matches a traditional taxi driver uniform.
- − The passenger is sitting in the front passenger seat instead of the back seat.
- − The hands on the steering wheel look more like human/monkey hands than capybara paws.
- − The lighting on the capybara's neck appears a bit inconsistent with the car interior.
Verdict: While FLUX.1 [schnell] captures the atmospheric lighting and the passenger's 'bored' attitude in the back seat better, it fails on technical prompt details like paw placement. Z-Image Turbo captures the driver attributes more accurately but places the passenger in the front seat, which changes the dynamic of the scene. FLUX.1 [schnell] is slightly preferred for its superior photorealism and correct spatial arrangement of the characters.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong cinematic lighting and atmosphere
- + Crisp illustrative style consistently applied
- + Includes all requested artistic elements like twisted trees and bats
- − Significant text errors including repetitions and typos like 'firiichts'
- − Layout of the text is cluttered at the bottom
- − Fails to portray the 'dark parchment' texture effectively
Z-Image Turbo
- + Excellent adherence to the 'dark parchment' and gothic aesthetics
- + Highly accurate text rendering with minimal errors
- + Superior composition with a clear central jack-o-lantern and thorn/web border
- − Minor typo in location name ('Archves' instead of 'Arches')
- − The scroll banner is split into small pieces rather than being one continuous element
Verdict: Z-Image Turbo is the clear winner as it successfully captured the requested 'dark parchment' look and managed nearly all the text prompts with high accuracy. FLUX.1 [schnell] struggled significantly with the textual elements, producing several typos and redundant lines of text that ruined the utility of the invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the isometric perspective and dioramic presentation.
- + High-quality rendering of textures, especially the salmon and the rice grains.
- + Correctly followed the flag icon instruction with the Japanese flag.
- − Missed the secondary 'SUSHI' text instruction.
- − The red sauce decoration on the salmon looks slightly unnatural and messy.
Z-Image Turbo
- + Followed the text instructions more thoroughly including both 'JAPAN' and 'SUSHI'.
- + Pleasant, soft 3D cartoon aesthetic that fits the 'miniature' prompt.
- + Good use of shadows to anchor the diorama base.
- − Included the flag of China instead of the flag of Japan.
- − The sushi roll geometry is a bit simplified compared to the realistic PBR request.
- − Text placement is slightly crowded toward the top edge.
Verdict: FLUX.1 [schnell] produced a much higher quality render with superior textures and the correct national flag, though it failed to include the word 'SUSHI'. Z-Image Turbo followed the text commands more literally but failed a major logic test by using a Chinese flag for a Japanese dish, while also offering a simpler visual style. FLUX.1 [schnell] is the winner for its professional finish and accuracy to the cultural theme.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Exquisite lighting with vibrant orange/golden tones
- + High level of detail on the puppy's fur and the surrounding flowers
- + Clear representation of a red fox kit with distinct features
- − Failed to include a rabbit, instead showing two kitten-like creatures alongside the puppy and fox
- − The second kitten has anatomy that borders on a fennec fox or hybrid creature
Z-Image Turbo
- + Excellent adherence to the prompt by including all four distinct animals (golden retriever, kitten, bunny, and fox)
- + Better sense of action and 'tumbling' as described in the prompt
- + Noticeable dew sparkles on the grass which was requested
- − The fox's eyes look somewhat artificial or glass-like
- − The butterfly anatomy is slightly distorted compared to Model A
Verdict: While FLUX.1 [schnell] (Model A) provides superior artistic lighting and textures, it failed to generate the baby bunny requested in the prompt. Z-Image Turbo (Model B) followed the instructions much more accurately, including all four specific animals with a dynamic composition that better reflects the 'tumbling' and 'chasing' aspect of the prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector-style emblem construction
- + Professional banner layout
- + Appropriate color palette
- − Major spelling errors in the brand name ('CAFEÉ FRAMILAN')
- − Incorrect date ('7720' instead of '1720')
- − Missing the 'steam' element requested
Z-Image Turbo
- + Perfect text rendering of the brand name and date
- + Includes the steam element accurately
- + Strong minimalist aesthetic
- − Minimal texture on background
- − Slightly less 'vintage' character than model a
Verdict: Z-Image Turbo is the clear winner as it followed all textual instructions perfectly, including the exact name 'Caffè Florian' and the correct historical date '1720'. In contrast, FLUX.1 [schnell] hallucinated several characters, resulting in a misspelled brand name and an impossible date, despite having a slightly more elaborate emblem design.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the color palette and dark navy background
- + Strong aesthetic composition with a central focus and orbital diagrams
- + Captured the 'flat-vector' and 'subtle gradient' style perfectly
- − Text is largely illegible gibberish
- − The rocket design is stylized but lacks the specific Saturn V profile requested
- − Failed to clearly number or separate the six requested steps
Z-Image Turbo
- + Text is much more legible with recognizable terms like 'Earth Orbit' and 'Tranquility'
- + Icons for Earth and Moon are very clear and use the requested orbit ring iconography
- + Layout effectively highlights key mission phases with readable typography
- − Misspelled the primary header as 'APOLIO E 11'
- − Included orange and yellow colors that were not part of the requested NASA-inspired palette
- − Missed a few of the requested six steps, stopping before showing a clear surface landing icon
Verdict: FLUX.1 [schnell] captures the requested aesthetic and color palette much better, appearing like a professional infographic, though its text is unreadable. Z-Image Turbo has superior text rendering and clearer iconography, but it fails to follow the color constraints and has a typo in the main header. FLUX.1 [schnell] is the likely winner for its superior composition and adherence to the specific stylistic artistic guidelines.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering