Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#21 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light and shadows from the left.
- + Highly realistic glass refraction and caustics on the table.
- + Perfect adherence to the plant being behind the cube.
- − The blue sphere has a slightly textured, non-solid appearance.
Qwen Image 2.0
- + Smooth, clean rendering of the blue sphere.
- + Sharp focus on the objects.
- − The plant appears to be physically intersecting with the back of the cube or even inside it, rather than just behind it.
- − The glass cube lacks realistic thickness and logical refraction compared to Model A.
- − Shadows do not realistically reflect light coming from a specific window source.
Verdict: FLUX.1 Kontext [max] produced a superior, photorealistic image with a deep understanding of physics, particularly in how light passes through glass to create caustics on the wood. While Qwen Image 2.0 followed the basic prompt instructions, it struggled with spatial depth, making the plant look like it was merging with the cube.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent atmosphere with realistic falling rain and wet pavement reflections.
- + Captures the requested 'imperfect framing' and 'motion blur' from passing cars very effectively.
- + Sophisticated lighting that feels cinematic and natural.
- − The bike chain and mechanical details are physically nonsensical.
- − The man's hands appear slightly distorted and lack anatomical precision.
Qwen Image 2.0
- + Exceptional skin texture and facial details that feel very 'candid photo'.
- + Superior mechanical realism on the bicycle, including the chain and pedal structure.
- + Perfectly captures the 'shallow depth of field' requested.
- − Does not show much 'motion blur' from the background cars.
- − The presence of a partial second figure on the right is a bit distracting.
Verdict: FLUX.1 Kontext [max] creates a much more atmospheric and cinematic scene that perfectly adheres to the 'motion blur' and 'rain' elements of the prompt. However, Qwen Image 2.0 provides vastly better skin textures and anatomical/mechanical correctness, making it feel more like a real photograph despite missing some of the environmental effects. Qwen Image 2.0 is the winner for its superior realism and detail in the subject.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Expert skin texture with realistic pores and sweat sheen
- + Incredible detail on the floral engravings of the plate armor
- + Beautiful shallow depth of field with realistic bokeh sparks
- − The 'small beads' in the hair are largely missing or indistinct
- − Composition is very tight, clipping the top of the head
Qwen Image 2.0
- + Accurately included the small colorful beads in the braids
- + Clearer depiction of battlefield scarring and dirt
- + Excellent representation of the leather and cloth textures requested
- − Skin texture appears slightly over-sharpened or gritty compared to Model A
- − Hand anatomy on the sword hilt is slightly awkward
Verdict: FLUX.1 Kontext [max] produced a more cinematic and high-fidelity image with superior lighting and metal reflections, while Qwen Image 2.0 followed the specific prompt details like the beads in the braids and the leather straps more accurately. FLUX.1 Kontext [max] is the overall winner due to its lifelike eyes and professional-grade skin and metal rendering.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent full-page layout composition that feels like a real usable physical menu.
- + Clean white background with professional use of negative space.
- + Professional-looking brand logo and accent colors.
- − The food photos lack diversity, consisting almost entirely of pizza.
- − Some fonts used for headers are script-based rather than the requested bold sans-serif.
- − Text contains significant gibberish and artifacts.
Qwen Image 2.0
- + Strong adherence to the requested bold sans-serif header fonts.
- + Strict adherence to the 'grid' layout and category requirements (Appetizers/Pizza/Mains).
- + High-quality, diverse food photography that accurately represents different dish types.
- − The layout is slightly repetitive with text overlapping or clipping at the bottom.
- − The price points are all '100', making it feel more like a template than a finished design.
- − The tight margins make the composition feel a bit cramped compared to Model A.
Verdict: Both models followed the prompt well, but Qwen Image 2.0 provided better content diversity for the food photos and followed the typography instructions for bold sans-serif fonts more accurately. While FLUX.1 Kontext [max] created a more aesthetically pleasing full-page layout, its failure to include varied food items (mostly showing only pizza) makes Qwen Image 2.0 the more successful interpretation of a restaurant menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and clean font choices.
- + High resolution with vibrant colors that pop against the background.
- + Good integration of glowing effects on the text and embers.
- − The burger is not truly 'exploded' as requested, with most layers still stacked.
- − The floating objects on the sides look like rolls rather than internal burger components.
- − Missing the starburst for the price element.
Qwen Image 2.0
- + Captures the 'fiery, glowing effect' on the text much more effectively.
- + Successfully includes the requested starburst for the price tag.
- + Better sense of vertical separation ('exploded' effect) between the top bun and the patty.
- − The price text is slightly misaligned compared to the starburst graphic.
- − The 'LIMITED TIME ONLY' text is smaller and less impactful than in Model A.
- − Some components like the lettuce and bottom bun are still relatively compressed together.
Verdict: Qwen Image 2.0 adhered better to the specific stylistic elements of the prompt, successfully including the fiery text effect and the price starburst. While FLUX.1 Kontext [max] produced a cleaner, more advertising-ready aesthetic with sharper text, it failed to separate the burger components or include the starburst, making Qwen the more accurate interpretation of the prompt's details.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and spelling accuracy.
- + The chalk smudges on the blackboard look very realistic.
- + Captures the full list of items perfectly, including a creative completion of the truncated prompt.
- − The handwriting style is a bit too uniform, bordering on looking like a font.
- − The title font is not strictly 'elegant cursive' as requested.
Qwen Image 2.0
- + The handwriting has more natural, varied 'slant' and 'size' as requested in the prompt.
- + Atmospheric lighting and composition create a stronger 'cozy café' feel.
- + The chalk texture is very dusty and authentic.
- − Some text alignment is slightly messy, making it harder to read.
- − Minor artifacting where some letters blend into the chalk smudges.
Verdict: Both models followed the prompt exceptionally well, particularly regarding the specific date and menu items. FLUX.1 Kontext [max] produced much cleaner, more legible text and successfully completed the truncated 'Brown But...' prompt as 'Brown Butter Chocolate Chip Cookies'. However, Qwen Image 2.0 captured the requested handwriting variations and 'elegant cursive' title more accurately, feeling more like a real handwritten sign rather than a digital recreation.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + High visual fidelity and cinematic lighting across the scene.
- + Clean and coherent horse anatomy with detailed fur and mane textures.
- + Excellent professional rendering of the astronaut suit.
- − Failed to follow the specific spatial instruction (horse on top of astronaut).
- − Standard interpretation lacking the requested surrealism.
Qwen Image 2.0
- + Dynamic composition with interesting foreground elements like water droplets.
- + Creative scales on the horse adding a unique, surreal element.
- + Good depth of field and color contrast.
- − Failed to follow the specific spatial instruction (horse on top of astronaut).
- − Anatomical issues in the horse's legs and hooves, particularly where they merge with background elements.
Verdict: Both models completely failed the negative constraint and logic reversal requested in the prompt, which asked for the horse to be on top of the astronaut; instead, both provided the standard 'astronaut riding a horse' trope. FLUX.1 Kontext [max] is the superior image in terms of technical execution and realistic textures, while Qwen Image 2.0 offers a more creative interpretation of the horse's appearance but suffers from anatomical glitches.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the capybara's fur and the taxi interior.
- + Well-integrated seatbelt and lighting that matches the night scene.
- − The passenger is holding a phone to her ear like a call instead of 'looking at her phone' as requested.
- − Only one paw is clearly visible on the steering wheel.
Qwen Image 2.0
- + Perfect adherence to the instruction of the passenger looking at her phone.
- + Clearly shows both paws on the steering wheel.
- + The capybara's professional, forward-facing expression is spot on.
- − The paws look somewhat distorted and merge awkwardly with the steering wheel texture.
- − The overall image quality has a slightly higher contrast/digital look compared to the softness of Model A.
Verdict: Both models followed the complex prompt well, but Qwen Image 2.0 adhered more closely to specific behavioral details, such as the passenger looking at her phone and the capybara having both paws on the wheel. FLUX.1 Kontext [max] produced a more convincing photorealistic texture, particularly for the fur and leather, but failed on the specific passenger interaction.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a modern gothic aesthetic
- + Detailed and atmospheric border design with webs and thorns
- + Accurate rendering of almost all text elements
- − Included redundant text by repeating 'The Arches' and 'NYC'
- − The central jack-o-lantern is quite dark, losing some requested glow
Qwen Image 2.0
- + Beautiful composition with clear central jack-o-lantern and twisted trees
- + Captured the 'parchment' aesthetic more convincingly across the whole background
- + Correctly interpreted all text requirements without redundancy
- − Slightly less 'gothic' and more 'spooky illustration' style
- − The scroll banner text is small and a bit cramped
Verdict: Both models followed the prompt exceptionally well. FLUX.1 Kontext [max] has superior graphic design elements particularly in the border and title font, though it stumbled on text logic by repeating the location twice. Qwen Image 2.0 provided a more balanced illustration with better lighting on the pumpkin and a truer vintage parchment feel, making it the more visually appealing invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent 3D cartoon art style consistent with the 'miniature' prompt.
- + Smooth, refined textures and lighting create a high-quality 3D render look.
- + Clean typography that matches the aesthetic of the scene.
- − Omitted the requested flag icon.
- − The text 'JAPAN' has a slightly unconventional font choice for the subject matter.
Qwen Image 2.0
- + Successfully included the requested flag icon.
- + Higher variety of sushi types on the plate.
- + Good wood grain texture on the diorama base.
- − Failed to capture the 'cartoon' style, opting for a more realistic photographic look.
- − Text layout is less integrated with the overall composition.
- − Lighting feels flat compared to the 3D render style requested.
Verdict: FLUX.1 Kontext [max] followed the stylistic instructions much better, delivering a beautiful 3D isometric cartoon scene with refined textures, whereas Qwen Image 2.0 produced a more standard photographic composite. Although Qwen included the flag icon that FLUX.1 missed, FLUX.1's superior composition and adherence to the 'cartoon' and 'refined texture' prompts make it the more visually appealing result.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Features a very warm, dreamy, and cohesive aesthetic that fits the 'wholesome' vibe.
- + Consistent lighting across all animals with beautiful god rays and bokeh.
- − The animals are static and sitting in a row, lacking the 'tumbling' and 'chasing' action requested.
- − The bunny has a slightly stylized, almost cartoonish appearance compared to the others.
Qwen Image 2.0
- + Excellent adherence to the 'tumbling' and 'playfully chasing' action in the prompt.
- + Superior anatomical realism and texture detail in the fur and eyes.
- + Dynamic composition with clear interaction between all four distinct animals.
- − The butterflies appear a bit flat/superimposed compared to the rest of the 3D scene.
Verdict: While FLUX.1 Kontext [max] creates a beautiful, glowy atmosphere, it fails to capture the requested action, showing the animals sitting still. Qwen Image 2.0 perfectly follows the prompt's instructions for a playful, tumbling scene while maintaining a higher level of photorealistic detail in the animals' fur and features.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with correct Italian diacritics
- + Perfectly captures the vintage minimalist vector emblem style
- + Subtle paper texture across the background and elements is very high quality
- − The steam element is very small compared to the cloche
Qwen Image 2.0
- + Includes all prompted elements in a clear arrangement
- + Smooth gradients and consistent warm color palette
- − Text is somewhat generic and lacks the requested 'classic' or 'vintage' feel
- − Steam is inconsistently placed inside the dome surface rather than rising from it
- − Less 'minimalist' than requested due to heavy shading and gradients
Verdict: FLUX.1 Kontext [max] perfectly executes the vintage minimalist aesthetic with superior typography and a authentic woodcut-style texture. Qwen Image 2.0 provides a decent interpretation but feels more like a modern clip-art illustration than a classic 18th-century heritage logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography and inclusion of correctly named astronauts.
- + Clean modern aesthetic with professional-looking illustrations.
- + High visual quality and resolution of individual elements.
- − Failed to follow the requested chronological sequence of icons.
- − Text placement for 'Translunar' and 'Lunar Orbit' is confusingly swapped or misplaced.
- − Rocket icon is generic/cartoonish rather than a Saturn V.
Qwen Image 2.0
- + Followed the specific 6-step sequence and labels perfectly.
- + Much better representation of the Saturn V and Lunar Module shapes.
- + Composition flows logically from top to bottom as an infographic.
- − One minor typo 'Translunjar' in a label.
- − Text rendering on the white border at the bottom is slightly less crisp than Model A.
- − Relies on a very dark background which makes it feel less 'airy' than a modern vector poster.
Verdict: Qwen Image 2.0 is the clear winner because it actually followed the instructions to create a 6-step infographic, whereas FLUX.1 Kontext [max] produced a jumbled layout with incorrect labels. Qwen Image 2.0 also captured the technical details of the Apollo hardware much more accurately than FLUX.1 Kontext [max], despite a single small typo.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request