Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Z-Image Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light and reflections, including caustic patterns on the table.
- + High level of detail in the textures of the wood, book, and sphere.
- + Very accurate spatial reasoning with the plant correctly refracted through the glass.
- − The text on the book spine is nonsensical gibberish.
- − The blue sphere has an unusually rough, glittery texture instead of a smooth finish.
Z-Image Turbo
- + Clean, simple composition that matches all primary prompt elements.
- + The blue sphere has a realistic, smooth glass-like quality.
- + Accurate lighting from the left following the prompt's direction.
- − Lower overall resolution and slightly blurry textures on the plant and background.
- − The plant is simple and lacks the detail and realistic leaf structure shown in Model A.
- − Refractions through the glass are less complex and photorealistic.
Verdict: FLUX.1 Kontext [max] produced a significantly more high-resolution and atmospheric image with stunning light effects and realistic refraction of the plant through the glass. Z-Image Turbo followed the prompt accurately but yielded a softer, less detailed image with a much simpler representation of the background elements. FLUX.1 Kontext [max] is the winner due to its superior visual quality and technical execution of glass physics.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'candid' and 'imperfect framing' prompt with a low-angle perspective.
- + Superior rendering of rain droplets and light reflections on the pavement.
- + Captures the cinematic atmosphere and shallow depth of field perfectly.
- − The bike chain and spoke details are physically impossible (entwined incorrectly).
Z-Image Turbo
- + Good rendition of natural skin texture on the man's arms and face.
- + The subject is clearly Japanese and fits the 'elderly' description well.
- − Failed to include 'motion blur from passing cars' as the background cars are sharp.
- − Missed the 'repairing' action, as the man appears to be simply standing with or holding the bike.
- − Lacks the cinematic lighting and wet pavement reflections requested.
Verdict: FLUX.1 Kontext [max] far better captured the requested atmosphere, including the cinematic lighting, motion blur, and wet surfaces, resulting in a cohesive street photography aesthetic. While Z-Image Turbo produced a clear image, it failed several key stylistic prompts like motion blur and the specific act of repairing the bicycle, feeling more like a standard snapshot.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Exceptional macro detail on the engraved plate armor and skin texture
- + Strong adherence to the 'close portrait' instruction with a powerful composition
- + High-quality lighting consistency with golden torchlight reflecting across the face and metal
- − Missed the 'hair braided with small beads' detail, showing hanya thick braids
- − Skin texture appears slightly over-sharpened and excessively sweaty
Z-Image Turbo
- + Includes the 'small beads' in the hair braids as specifically requested
- + Good variety of textures including chainmail, leather, and engraved plate
- + Atmospheric composition with a visible torch source and bokeh
- − The 'close portrait' is more of a medium shot, losing some facial detail
- − The fire source appears slightly detached from the hand/torch handle
Verdict: FLUX.1 Kontext [max] produces a much more impactful and high-resolution close-up with incredible detail on the armor engravings, though it missed the specific mention of beads in the hair. Z-Image Turbo followed the 'beads' prompt better and included more varied armor types, but the overall image quality and lighting realism are lower than the competitor.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Clean, professional layout with a focus on whitespace
- + Accurate representation of a full menu sheet including pricing and logo
- + High-quality, appetizing food photography
- − Text is mostly illegible gibberish
- − The grid layout isn't as distinct for the food photos
Z-Image Turbo
- + Strong adherence to the 'grid' request for photos
- + Better text rendering with clear headers like 'Appetizers' and 'Pizza'
- + Vibrant orange accents create a cohesive visual style
- − Layout is a bit cluttered and lacks formal menu hierarchy
- − Misspelling of 'Mains' as 'Mans'
- − Image proportion is square rather than a standard vertical menu
Verdict: Z-Image Turbo followed the layout instructions more closely by providing a clear grid of photos and visible category headers like 'Appetizers'. While FLUX.1 Kontext [max] produced a more realistic-looking piece of stationery, it failed to provide the distinct sections and bold font headers requested in the prompt.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the meat and buns
- + Strong sense of motion with exploding ingredients
- + Highly legible and accurately integrated text
- − The burger components are partially assembled rather than fully exploded horizontally or vertically
- − The starburst for the price is very small and lacks impact
Z-Image Turbo
- + Perfect adherence to the starburst requirement for the price
- + Vibrant glowing effect on the text and assets
- + Clean and professional graphic design layout
- − Failed the 'exploded burger' requirement, showing a mostly assembled burger
- − Less photorealistic texture compared to the competitor
- − Background feels more like a studio set than a fiery void
Verdict: FLUX.1 Kontext [max] captures the 'exploded' and 'photorealistic' aspects of the prompt much better than the competitor, presenting a more dynamic image. While Z-Image Turbo followed the 'starburst' instruction more accurately and has clean typography, it failed to separate the burger components in mid-air as requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text accuracy with no spelling errors
- + Includes realistic chalk smudges and texture on the board
- + Successfully completed the truncated prompt text with 'Cookies'
- − The heading is in sans-serif print rather than cursive as requested
- − Text style is a bit too uniform, bordering on marker style rather than dusty chalk
Z-Image Turbo
- + Beautiful chalk texture with realistic variation in thickness and opacity
- + Captures a hand-drawn feeling with better character spacing and slant
- − Includes a spelling error with 'Mustroom' instead of 'Mushroom'
- − Failed to include 'elegant cursive' for the title as requested
Verdict: FLUX.1 Kontext produced a much more accurate image in terms of spelling and completing the partial prompt, whereas Z-Image Turbo suffered from a significant spelling mistake ('Mustroom'). While Z-Image Turbo had a more authentic chalk texture, FLUX.1 Kontext is the clear winner for its reliability in rendering complex text and the surrounding cafe environment.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photographic lighting and cinematic atmosphere
- + High detail in the horse's fur and the astronaut's suit textures
- + Beautiful composition with the curve of the planet in the background
- − Failed the specific spatial instruction for the horse to be on top of the astronaut
Z-Image Turbo
- + Clean, sharp image quality
- + Good anatomical rendering of the horse
- + Consistent lighting on the subjects
- − Failed the specific spatial instruction for the horse to be on top of the astronaut
- − Background is relatively flat and less cinematic compared to Model A
Verdict: Both models failed the specific logic test of placing the horse on top of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. FLUX.1 Kontext [max] is the superior image due to its cinematic lighting, detailed textures, and much more immersive space environment.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent fur texture rendering and realistic lighting on the capybara.
- + Captures a very cinematic, high-quality photorealistic aesthetic.
- + The passenger is correctly placed in the background with a phone.
- − Only one paw is clearly visible on the steering wheel instead of both requested.
- − The passenger appears to be making a call rather than just looking at the screen.
Z-Image Turbo
- + Better adherence to the 'both front paws on the steering wheel' instruction.
- + The passenger's expression and pose perfectly match the 'bored businesswoman looking at phone' prompt.
- + Realistic integration of the animal's limbs with the steering wheel.
- − The lighting is a bit flat compared to the cinematic quality of the competitor.
- − The capybara's fur looks slightly less detailed and soft than in the other image.
Verdict: Both models followed the complex prompt very well, but Z-Image Turbo captures the specific details of the prompt more accurately, such as the businesswoman's bored expression and both paws on the wheel. While FLUX.1 Kontext [max] has superior lighting and textures, it struggles slightly with the specific positioning of the paws and the passenger's action.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a dark, polished gothic aesthetic.
- + Perfect adherence to the dark parchment and cinematic lighting request.
- + Accurate text rendering for all lines including the secondary banner.
- − Repetitive location text at the bottom.
- − The layout is a bit cluttered with the text overlapping the dark background elements.
Z-Image Turbo
- + Strong composition with a clear, readable scroll-based design.
- + Good inclusion of all requested elements like thorns, webs, and twisted trees.
- + Vibrant glow on the jack-o-lantern.
- − Spelling error in the location ('The Archves' instead of 'The Arches').
- − The parchment looks a bit too bright and clean for a 'dark gothic' theme.
- − Secondary text at the top is very small and plain.
Verdict: FLUX.1 Kontext [max] captured the dark, moody atmosphere of the prompt much better than the competitor, delivering a truly 'gothic' aesthetic. While Z-Image Turbo had a very clear layout, it suffered from a spelling error and a slightly more generic 'clipart' feel. FLUX.1 Kontext [max] is the winner for its superior text accuracy and atmospheric lighting.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'full, thick head of hair' request.
- + Maintains very high visual consistency with the original background and lighting.
- + The texture and volume of the hair look natural and integrated.
- − The facial features were slightly altered, making the person look like a younger/different version of the source.
Z-Image Turbo
- + The facial identity is preserved slightly better in terms of bone structure.
- − Failed the primary instruction to add a full head of hair, providing only a slight buzz-cut shadow.
- − Significant change to the background scenery (added grass and mountains).
- − Removed the subject's glasses entirely.
Verdict: FLUX.1 Kontext [max] successfully added a full, thick head of hair as requested and maintained the background and lighting of the original image, despite some slight facial softening. Z-Image Turbo completely failed the main task, altered the background, and removed the subject's glasses, resulting in a poor edit.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography style that matches the 3D aesthetic.
- + High-quality 3D renders with realistic textures on the fish and rice.
- + Captures the diorama feel well with the addition of chopsticks and wasabi.
- − Missing the requested flag icon.
- − The perspective is slightly lower than a true 45-degree isometric view.
Z-Image Turbo
- + Successfully included a flag icon as requested.
- + Clean, minimalist composition that adheres strictly to the 'small diorama base' request.
- + Accurate isometric perspective.
- − Used the flag of China instead of Japan, which is a major hallucination/error given the text 'JAPAN'.
- − The single piece of sushi is less visually interesting than the pair in the other model.
Verdict: FLUX.1 Kontext [max] produced a much higher quality render with better textures and more appealing 3D typography, though it missed the flag icon. Z-Image Turbo followed the prompt's structural instructions but made a significant contextual error by displaying the Chinese flag next to the word 'JAPAN' and the Japanese dish sushi.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully creative and fun caricature style consistently applied.
- + Includes all requested elements like the hockey stick, dog, and news desk.
- + Maintains character likeness through hair and outfit color.
- − The hockey stick is positioned awkwardly in the top corner.
- − Complete departure from the original photograph's realistic style.
Z-Image Turbo
- + Preserves the source face and photographic style with high fidelity.
- + Subtly adds a small dog in the background area.
- − Fails to apply the core caricature request.
- − Misses the television anchor and hockey themes entirely.
- − The added dog is very small and lacks detail.
Verdict: FLUX.1 Kontext [max] followed the complex instructions to create a caricature incorporating news, hockey, and dog elements, even though it completely changed the art style. Z-Image Turbo largely ignored the edit instructions, failing to produce a caricature or the specific career and hobby themes requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of warm golden light and volumetric god rays
- + High consistency in the 'lush meadow' environment with vibrant wildflowers
- + Strong adherence to the 'butterflies' prompt with multiple elements integrated into the scenery
- − The animals are sitting relatively still rather than 'playfully chasing' as requested
- − The bunny has a slightly artificial, cartoonish expression compared to the others
Z-Image Turbo
- + Better captures the 'playfully chasing' and 'tumbling together' action requested in the prompt
- + Superior fur texture and realism on the golden retriever and fox
- + Naturalistic dew sparkles on the grass in the foreground
- − Lighting is less dramatic, lacking the requested 'god rays' effect seen in the other model
- − Butterflies appear a bit static and disconnected from the animals' eyelines
Verdict: While FLUX.1 Kontext [max] creates a more magical atmosphere with beautiful lighting and god rays, Z-Image Turbo captures the requested movement and interactions between the animals much more effectively. Z-Image Turbo's animals look more photorealistic and active, though it misses the dramatic lighting effects that FLUX.1 Kontext [max] excelled at.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfectly captures the Studio Ghibli art style with hand-painted watercolor textures.
- + Excellent preservation of the source image's composition and the 'distracted boyfriend' meme logic.
- + Successfully translates the character expressions and poses into a nostalgic anime aesthetic.
- − The woman in the red dress is now in sharp focus, which differs from the bokeh effect in the original image.
Z-Image Turbo
- + Maintains the realistic depth of field from the source image.
- − Completely failed to apply the Ghibli illustration style, resulting in a slightly filtered photo rather than art.
- − The character faces are altered to new people rather than artistic interpretations of the originals.
- − Lacks the requested soft pastel colors and hand-painted textures.
Verdict: FLUX.1 Kontext [max] is the clear winner as it successfully reimagined the famous meme image in a beautiful, recognizable Studio Ghibli art style while keeping the scene's iconic composition intact. Z-Image Turbo failed the prompt entirely, providing a slightly desaturated photo that does not meet any of the stylistic requirements for an illustration or hand-painted look.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'windy hair' prompt with natural-looking flow.
- + Successfully added flying leaves throughout the scene.
- + Strong preservation of the original subjects' faces and clothing style.
- − Minor anatomic issues with the woman's left hand hovering unnaturally near the dog's head.
- − The leash has become detached and floats in a loop.
Z-Image Turbo
- + Effectively added falling and flying leaves to the composition.
- + Maintains a high level of visual quality and consistency with the source lighting.
- − The hair shows very little motion compared to Model A, failing that part of the prompt.
- − Noticeable change to the woman's facial features, losing the likeness of the source image.
- − The woman's left hand is poorly rendered and appears to be missing most fingers.
Verdict: FLUX.1 Kontext [max] is the winner as it much more effectively captures the 'dynamic motion' requested, particularly in the hair, which Z-Image Turbo largely ignored. While FLUX.1 Kontext [max] struggled with the leash and hand placement, it preserved the subject's identity much better than Z-Image Turbo, which significantly altered the woman's face and generated a distorted hand.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with proper accented characters and vintage texture
- + Accurately follows the banner request for the 'Est. 1720' text
- + Provides a genuine hand-drawn vintage feel with consistent woodblock-style texture
- − Slightly less 'minimalist' than a modern corporate logo, though fitting for the vintage brief
Z-Image Turbo
- + Clean vector-style execution
- + Good color palette adherence
- + Legible typography
- − Missed the 'banner' element for the date, using horizontal lines instead
- − The cloche icon is very generic and lacks the requested vintage detail
- − Overall appearance looks like a modern digital imitation rather than an authentic vintage emblem
Verdict: FLUX.1 Kontext [max] is the clear winner as it followed every part of the prompt, including the specific request for a banner and subtle texture. While Z-Image Turbo produced a clean logo, it missed the banner element and lacked the character and artistic depth of the vintage style requested.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography and spelling throughout the graphic.
- + Sophisticated composition with a clean, professional NASA-inspired color palette.
- + Included creative additions like the correctly identified crew member silhouettes.
- − Failed to include exactly 6 steps as requested in the sequence.
- − The rocket icon looks more like a generic cartoon rocket than a Saturn V.
- − Logic flow of the arrows and labels is slightly confusing.
Z-Image Turbo
- + Successfully captured the requested flat-vector style with crisp icons.
- + Followed the color palette requirements accurately.
- + The rocket design more closely resembles the staging of a heavy-lift launch vehicle.
- − Multiple spelling errors including 'APOLIO E 11', 'Translurian', and 'Descenty'.
- − Poorly aligned text and floating elements create a messy composition.
- − Missing some of the requested steps in the infographic sequence.
Verdict: FLUX.1 Kontext [max] produced a much higher quality, professional-looking poster with correct spelling and beautiful typography, though it simplified the sequence of steps. Z-Image Turbo followed the vector style well but was plagued by significant spelling errors and a disorganized layout. FLUX.1 Kontext [max] is the preferred choice for its clear communication and superior visual appeal.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering