Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#58 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0.0%
win rate
Ties
0.0%
Z-Image Turbo
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent crispness and detail in the texture of the red book and the wooden table grain.
- + Very high visual quality with beautiful reflections of the window in the blue sphere.
- + Correct spatial placement of all elements, including the plant being partially visible through the glass.
- − The plant sits quite close to the cube, making the composition feel a bit crowded.
Z-Image Turbo
- + Good adherence to all prompt elements including the glass cube, sphere, book, and plant.
- + Effective use of depth of field, keeping the focus strictly on the glass cube.
- + Accurate lighting from the left consistent with the prompt.
- − The red book has slightly blurred edges and less realistic texture compared to the other model.
- − The sphere looks slightly small relative to the volume of the cube.
Verdict: Both models followed the complex spatial prompt perfectly. FLUX.1 Kontext [dev] is the winner due to its superior texture rendering, particularly the realistic paper edges of the book and the fine wood grain, whereas Z-Image Turbo appears slightly softer and less detailed.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent city bokeh and lighting/reflections
- + High facial detail and realism
- + Captures a cinematic mood with shallow depth of field
- − The subject is sitting on the bike rather than repairing it
- − Rain effects look a bit like a vertical line overlay
- − Cars in background are static and not blurred by motion
Z-Image Turbo
- + Matches the 'imperfect framing' request with a candid street feel
- + Better adherence to the 'repairing' action as the man is leaning over the bike
- + The background car shows more believable motion/candid positioning
- − Low resolution and lacks fine detail in the face and hands
- − Anatomical issues with the man's right leg disappearing into the bike frame
- − Lighting is flat compared to the cinematic request
Verdict: FLUX.1 Kontext [dev] produced a much higher quality image with beautiful cinematic lighting and sharp details, but it failed to show the man 'repairing' the bike, instead showing him riding it. Z-Image Turbo followed the composition and 'repair' instruction better but suffered from poor technical quality, including anatomical errors and low resolution. FLUX.1 Kontext [dev] is the winner for its professional aesthetic and realistic skin textures, despite the slight prompt deviation regarding the specific action.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent high-contrast lighting and warm atmosphere
- + Intricate engraved details on the plate armor
- + Clean, sharp facial features and intense gaze
- − Missed the request for braided hair with beads
- − The face appears too 'clean' for a battle-worn character despite a small mark
- − Cloth and leather textures are less varied than requested
Z-Image Turbo
- + Perfect adherence to specific details like braided hair with beads
- + Realistic application of dirt and scars on the skin
- + Complex layering of mail, cloth, and leather under the armor
- − The torch flame looks a bit artificial and flat compared to the rest of the image
- − Slightly less 'high-art' polish in the skin texture compared to Model A
Verdict: Model B (Z-Image Turbo) is the clear winner as it followed every specific detail of the prompt, including the braided hair with beads and the layered clothing, which Model A missed. While FLUX.1 Kontext [dev] produced a high-quality cinematic portrait, it ignored several key descriptive elements that were central to the character's design.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Stronger visual grid composition for high-end casual dining
- + High-quality, appetizing food photography with vibrant colors
- − Text consists of nonsensical gibberish
- − The layout feels more like a magazine spread than a functional menu
Z-Image Turbo
- + Follows the specific section requests for appetizers, pizza, and mains
- + Readable text and pricing structure mimic a real menu layout
- + Adds orange vibrant accents as requested
- − Food photography is slightly less refined with some odd rendering in the pasta
- − Text has minor spelling errors like 'SE TIIION'
Verdict: Z-Image Turbo is the clear winner here as it successfully followed the structural requirements of the prompt, including specific sections for appetizers and pizza with pricing. FLUX.1 Kontext [dev] produced higher quality imagery, but failed to create a functional menu layout, instead opting for decorative gibberish and a less logical structure.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong thematic background with coals and flames
- + Bold, readable main title text
- − Failed to provide an 'exploded' view of components
- − Spelling error in the secondary message
Z-Image Turbo
- + Excellent typography with better glowing effects
- + Higher level of textural detail in the food components
- − Failed to provide the requested 'exploded' mid-air suspension
- − Starbust design is slightly cluttered
Verdict: Both models failed the specific 'exploded' composition instruction, instead rendering standard stacked burgers. Z-Image Turbo wins due to superior text accuracy and more realistic food textures, whereas FLUX.1 Kontext [dev] had a significant spelling error and less polished typography.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent chalk texture and realistic handwriting variation
- + Captures a cozy cafe aesthetic with the frame and lighting
- − Significant spelling errors throughout the menu items
- − Garbled text in the date line and duplicated words
Z-Image Turbo
- + Nearly perfect spelling and legibility
- + Accurately follows the specific text instructions for the date and menu items
- − Texture looks somewhat cleaner/digital rather than rough chalk
- − Minor typo 'Mustroom' in the first menu item
Verdict: Z-Image Turbo is the clear winner as it successfully rendered almost all the complex text requested in the prompt with high legibility. While FLUX.1 Kontext [dev] had a more convincing chalk texture, it failed significantly on prompt adherence by misspelling words and creating garbled characters in the date.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully follows the difficult logic-inverting instruction of having the horse on top.
- + High resolution with realistic textures on the spacesuit and horse fur.
- + Cinematic lighting with a clear sense of depth and scale against the planet below.
- − The horse's hind legs appear somewhat detached or anatomically confused near the astronaut's back.
Z-Image Turbo
- + Good dynamic pose for the horse.
- + Crisp lighting on the astronaut's gear.
- − Completely failed the primary prompt instruction by placing the astronaut on top of the horse.
- − The horse's anatomy is distorted, particularly with a floating front-left hoof and a missing leg.
- − Generic composition that ignores the 'surreal' aspect of the prompt's structural requirement.
Verdict: FLUX.1 Kontext [dev] is the clear winner as it successfully interpreted the challenging 'horse on top' spatial instruction, creating a surreal and high-quality image. Z-Image Turbo defaulted to a standard astronaut on a horse, failing the logical constraint of the prompt while also suffering from significant anatomical artifacts in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent photorealism with a high degree of detail in the fur and clothing
- + Accurately captures the 'bored' expression of the passenger
- + Creates a convincing atmospheric night lighting consistent with a city environment
- − The capybara only has one paw on the steering wheel instead of both as requested
- − The capybara's anatomy is slightly more humanoid/broad-shouldered than a real capybara
Z-Image Turbo
- + Followed the instruction to have both paws on the steering wheel
- + Very accurate anatomical representation of a capybara's head and face
- + The taxi driver cap looks more official and traditional
- − The passenger is in the front seat next to the driver, not the back seat as requested
- − The lighting feels a bit too bright and flat for a night scene inside a car
- − Less emotional depth in the passenger's face compared to Model A
Verdict: FLUX.1 Kontext [dev] produced a much more cinematic and atmospheric image that perfectly captured the requested mood and the bored expression of the passenger in the back seat. While Z-Image Turbo followed the 'both paws' instruction better, it failed the spatial positioning by placing the passenger in the front seat, which undermines the prompt's narrative.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong, legible main title rendering
- + Includes all requested date and time details clearly
- − The secondary banner text is largely gibberish
- − The spelling of 'The Arches' is corrupted into 'The Argiiah's'
- − Missing the 'parchment' texture requested, opting for a flat black background
Z-Image Turbo
- + Excellent adherence to the 'dark parchment' and 'moody sky' aesthetic
- + Highly accurate text rendering for the title and event details
- + Beautiful composition with high-quality atmospheric lighting and gothic elements
- − Very minor typo in the location ('The Archves' instead of 'The Arches')
- − The specific requested quote was placed at the top instead of a scroll banner
Verdict: Z-Image Turbo is the clear winner for its superior visual quality and adherence to the requested 'vintage gothic parchment' aesthetic. While FLUX.1 Kontext [dev] struggled with background textures and experienced significant text corruption in the smaller fonts, Z-Image Turbo delivered a polished, cinematic, and functional invitation design.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully added a full head of hair as requested.
- + Maintained the overall color palette and background composition.
- + The hair texture matches the beard density well.
- − Failed to preserve original facial features, significantly altering the man's face.
- − The glasses were replaced with a different style and the lighting on the skin became smoother and less realistic.
Z-Image Turbo
- + Preserved the original facial structure, expression, and skin texture much better than the competitor.
- + Kept the original lighting and background intact.
- − Completely failed the primary edit instruction by only adding a slight buzz cut/stubble instead of a 'full, thick head of hair'.
- − Modified the glasses and some background elements unnecessarily.
Verdict: This is a trade-off between edit adherence and source preservation. FLUX.1 Kontext [dev] followed the prompt to provide a full head of hair, but in doing so, it completely changed the man's face, making him unrecognizable from the source. Z-Image Turbo preserved the man's identity much better, but failed to provide the 'thick' hair requested, opted instead for a very short buzz cut.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering with clean edges
- + Perfectly solid light blue background as requested
- + Clean, minimalist 3D rendering that matches the 'cartoon' aesthetic
- − The flag icon is an abstract geometric shape rather than a specific flag
- − The sushi design is overly simplified, appearing more like a toy than food
Z-Image Turbo
- + Beautiful PBR textures on the fish and rice grains
- + Multi-layered diorama base adds depth and visual interest
- + Professional 3D lighting and soft shadows
- − Displays the flag of China instead of the flag of Japan
- − Text rendering is slightly less sharp than the competitor
Verdict: FLUX.1 Kontext [dev] followed the technical formatting and typography requirements with high precision, but failed to create a recognizable flag. Z-Image Turbo produced a much more visually appealing 3D model with superior textures, however, it made a significant factual error by placing a Chinese flag in a scene labeled 'JAPAN'.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully converts the image into a stylized caricature aesthetic.
- + Incorporates the TV anchor profession with a TV screen and the dog interest with a cartoon dog.
- + Captures the essence of the person's features while maintaining the denim jacket outfit.
- − Completely misses the 'hockey' requirement.
- − The text 'JOB' and other elements are a bit nonsensical and clutter the background.
Z-Image Turbo
- + Excellent preservation of the original person's face and the high-resolution photo quality.
- + Subtly adds a small dog in the background.
- − Fails to follow the 'caricature' and 'exaggerated' style instructions entirely.
- − Fails to incorporate the 'tv show anchor' or 'hockey' profession/hobby elements.
- − The change is extremely minimal and does not fulfill the creative brief.
Verdict: FLUX.1 Kontext [dev] followed the core of the prompt by attempting a caricature style and including two of the three requested thematic elements (TV and dogs), even though it missed hockey. Z-Image Turbo failed almost every part of the prompt, providing a nearly identical photo to the source with only a tiny dog added. FLUX.1 Kontext [dev] is the clear winner for actually attempting the creative transformation requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Dynamic sense of movement and playful energy
- + Consistent warm lighting and golden hour glow
- + High-quality fur texture and sharp focus on the central characters
- − Failed to include the fox and the bunny as requested in the prompt
- − Characters on the right appear to be two kittens instead of the diverse species requested
Z-Image Turbo
- + Successfully included all four requested animals: golden retriever, tabby kitten, bunny, and fox kit
- + Excellent interpretation of 'tumbling together' with overlapping characters
- + Features realistic dew sparkles and soft god rays
- − The butterfly on the right has a slightly anatomical mutation in the wings
- − The puppy's front paw resting on the bunny has slightly merged fur textures
Verdict: Z-Image Turbo is the clear winner as it successfully rendered all four specific animals requested in the prompt, whereas FLUX.1 Kontext [dev] only generated a puppy and two kittens. Z-Image Turbo also better captured the 'tumbling together' aspect of the prompt and included the fine details like dew sparkles more effectively.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the 'Studio Ghibli' art style with clean line work and anime aesthetics.
- + Perfectly preserves the composition and silhouettes of the original meme characters.
- + Captures the requested mood with vibrant but gentle colors.
- − The art style leans slightly more 'modern anime' than the classic hand-painted Ghibli texture.
Z-Image Turbo
- + Successfully applies a softer, warmer color grade to the photo.
- − Completely fails the primary instruction to transform the photo into an illustration.
- − The faces of the characters are slightly distorted compared to the original, losing the specific expressions of the meme.
- − Maintains a photographic style instead of the requested hand-painted or dreamy background art.
Verdict: FLUX.1 Kontext [dev] successfully transformed the image into a high-quality anime illustration that perfectly mirrors the source material while following the stylistic prompt. Z-Image Turbo failed most of the core instructions, only applying a slight color filter while remaining in a photographic style. FLUX.1 Kontext [dev] is the clear winner for its thorough execution of the artistic transformation.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent preservation of the original subjects' faces and clothing
- + Subtle but effective hair movement that looks natural
- + Maintains the bridge and background layout perfectly
- − Very few leaves added, making the 'dynamic motion' feel somewhat weak
- − The few leaves that are present look static and lack motion blur
Z-Image Turbo
- + Successfully adds many flying leaves across the scene to enhance the atmosphere
- + Noticeable hair movement that fits the prompt well
- − Substantially changes the woman's face, losing her original likeness
- − Alters background elements like the bridge architecture and foliage significantly
- − Introduces anatomical issues such as a missing dog leash
Verdict: FLUX.1 Kontext [dev] is the clear winner for image editing as it perfectly preserves the identity of the woman and the details of the scene while applying subtle motion. In contrast, Z-Image Turbo fails as an editor because it recreates the entire image, leading to a loss of the subject's likeness and the removal of the dog's leash.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering with correct accent marks
- + Clean vector style that matches the minimalist request
- + Perfect alignment and color palette adherence
- − Missed the 'banner' element for the date
- − The cloche illustration is slightly abstract/segmented
Z-Image Turbo
- + Uses a more 'classic' serif typography that feels vintage
- + The cloche shape is more traditional and recognizable
- + Good use of subtle texture in the background
- − Missed the 'banner' element for the date
- − Line weight is a bit inconsistent between the cloche and the text accents
Verdict: Both models followed the prompt well, but FLUX.1 Kontext [dev] produced a cleaner, more professional vector emblem with superior text rendering and better-defined edges. While Z-Image Turbo captured a slightly better 'vintage' font style, it lacked the precision and modern polish of the FLUX.1 output.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully captures a cohesive NASA-inspired dark navy color palette.
- + Includes a higher number of infographic steps to match the prompt's sequence.
- − Text rendering is mostly gibberish or heavily distorted.
- − Icons are overly abstract and messy, failing the request for crisp, clean vector lines.
- − Spelling error in the main title ('Apolo').
Z-Image Turbo
- + Clean, modern vector aesthetic with sharp lines and high clarity.
- + Text is much more legible and mostly spelled correctly.
- + Icons for the Earth, Moon, and Saturn V are high quality and recognizable.
- − Missing several stages of the requested 6-step sequence.
- − Includes an unnecessary 'E' in the main title and minor typos like 'Descenty'.
- − Layout is a bit sparse compared to a full infographic.
Verdict: Z-Image Turbo is the winner because it actually delivers the requested clean vector style with high visual quality and legible text. While FLUX.1 Kontext [dev] attempts more of the requested steps, its execution is marred by illegible gibberish text and messy, disorganized icons that fail the 'clean and modern' requirement.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering