Black Forest Labs' precision image generation model with maximum control, reliable text rendering, and complete creative control supporting up to 4MP output
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [flex]
#14 of 62 in Text-to-Image
Qwen Image
#40 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [flex]
0%
win rate
Ties
0%
Qwen Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent photographic quality with realistic soft lighting.
- + Accurate glass refraction and reflections including the plant through the cube.
- + Strong composition with high-fidelity textures on the book and table.
- − The sphere is quite large relative to the cube, pushing the definition of 'small' in the prompt.
Qwen Image
- + Successfully captured all prompt elements including the small sphere size.
- + Realistic wooden table texture and grain.
- + Correct lighting direction from the window on the left.
- − The glass cube has strange internal vertical lines that look like structural supports rather than a solid glass object.
- − The red book looks slightly more artificial compared to the surrounding environment.
Verdict: Both models followed the prompt instructions perfectly, including the specific spatial relationships between objects. FLUX.2 [flex] produced a more aesthetically pleasing and physically coherent image, whereas Qwen Image suffered from odd artifacts inside the glass cube that resembled internal pillars.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent adherence to the 'motion blur from passing cars' prompt requirement
- + Highly realistic skin textures and believable lighting
- + Strong adherence to the 50mm shallow depth of field aesthetic
- − The bicycle frame geometry is slightly warped/nonsensical near the seat post
Qwen Image
- + Good reflections on the wet pavement
- + Natural pose of the elderly man
- − Failed to include motion blur on the background cars
- − The rain effect looks like static vertical lines rather than integrated into the scene
- − Lacks the cinematic lighting and depth of the competitor
Verdict: FLUX.2 [flex] much better captures the 'cinematic but realistic' atmosphere requested, particularly through the successful use of motion blur on the background traffic. Qwen Image feels flatter, more like a standard digital photo, and fails to incorporate the specific motion blur and 50mm lens look as effectively.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent high-resolution skin texture and realistic scarring.
- + Superb detail on the engraved armor plating and leather buckles.
- + Realistic lighting and bokeh integration with the torchlight.
- − The beads in the hair are a bit plain and blend into the hair color.
- − The composition is a very standard frontal passport-style shot.
Qwen Image
- + Dynamic lighting and composition with a strong sense of mood.
- + Great variety in bead colors and detail in the braided hair.
- + Excellent texture on the chainmail and leather straps.
- − The 'bokeh sparks' look like stylized digital vector art rather than realistic light artifacts.
- − The scarring looks more like face paint or blood smears than actual healed or fresh wounds.
- − The light source (torch) has some digital artifacts and clipping.
Verdict: FLUX.2 [flex] wins on pure technical realism, particularly regarding face texture, authentic-looking scars, and the physical believability of the armor. While Qwen Image offers a more dynamic composition and colorful details in the hair, it is let down by the cartoonish 'spark' effects and less realistic skin rendering.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent professional layout that looks like a real-world menu.
- + High-quality, appetizing food photography with consistent lighting.
- + Clear hierarchy with headings, item names, and prices.
- − The main header 'APPETIZERS' is confusing as it oversees a section containing pizza and mains.
- − Small body text and 'RESTAURANT' subtitle have minor spelling/artifacting issues.
Qwen Image
- + Highly vibrant and playful color palette that fits the 'vibrant accents' prompt.
- + Creative grid layout for the food imagery.
- + Bold, legible sans-serif typography.
- − The food photos look slightly artificial and less realistic than Model A.
- − Text rendering (e.g., 'Pizzaurant', 'Pippeeieeraisuk') is significantly more garbled.
- − Layout feels a bit cluttered with large blocks of color.
Verdict: FLUX.2 [flex] produces a much more professional and realistic menu design that adheres to the 'clean professional layout' requirement. While Qwen Image has good vibrant energy, its text rendering and food quality are inferior to FLUX.2 [flex], which feels like a usable template.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent photorealistic texture on the meat and bun
- + Highly effective fiery/glowing text rendering
- + Clean and professional starburst design for the price
- − The 'exploded' effect is more of a stack than a dynamic dispersal
- − Lower bun is slightly melting/distorted at the bottom
Qwen Image
- + Excellent dynamic motion with ingredients flying outwards
- + Higher degree of creativity in the exploded view
- + Accurate rendering of all requested text elements
- − Less photorealistic, appearing more like a 3D render/illustration
- − Lighting on the burger feels a bit flat compared to the intense background
Verdict: Both models followed the prompt instructions very well, correctly rendering the text and price. FLUX.2 [flex] produced a much more photorealistic and appetizing product, making it superior for an actual advertisement, whereas Qwen Image captured the 'dynamic' and 'exploded' aspect of the prompt with more energy and movement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [flex]
- + Perfect spelling of all menu items including the cut-off prompt portion.
- + Highly realistic chalk texture with smudges and natural variations.
- + Elegant and consistent cursive handwriting that adheres to the prompt's style.
- − The lighting creates a slight glare at the top, making the title a bit harder to read compared to the center.
Qwen Image
- + Clear and legible handwriting with good layout balance.
- + Bright, well-lit composition within the cafe setting.
- − Contains a significant spelling error in 'Risotto' (spelled as 'Risoto').
- − Incorrectly rendered the date as '20026' instead of '2026'.
- − The lettering lacks the gritty, authentic chalk texture found in the other image.
Verdict: FLUX.2 [flex] delivered a superior result with perfect spelling, much more realistic chalk textures, and a convincing elegant cursive style. Qwen Image struggled with basic spelling and misinterpreted the year as 20026, while also providing a cleaner, more digital-looking font that lacked the requested chalk texture.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [flex]
- + Successfully followed the specific instruction for the horse to be on top of the astronaut.
- + High spatial detail in the galaxy background and asteroid field.
- + Cinematic lighting and dynamic composition.
- − The astronaut's hands are awkwardly holding the horse's front legs.
- − The 'riding' concept is slightly more like 'carrying' due to the pose.
Qwen Image
- + Clean, clear render with good lighting on the spacesuit.
- + Realistic textures on the horse's coat.
- − Failed the negative constraint/specific instruction for the horse to be on top.
- − Anatomical issues with the horse's back legs which appear fused or incorrectly positioned.
Verdict: FLUX.2 [flex] is the clear winner as it successfully interpreted the difficult prompt instruction to have the horse on top of the astronaut, creating a surreal and cinematic image. Qwen Image defaulted to the standard 'astronaut on horse' trope, failing the prompt's central requirement and exhibiting anatomical flaws in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent adherence to lighting and background atmosphere for night-time Manhattan.
- + The capybara's hands/paws on the wheel are surprisingly well-rendered for a hybrid animal concept.
- + The passenger is perfectly positioned in the back seat with a boredom that matches the prompt.
- − The capybara's fur texture appears slightly over-sharpened compared to the rest of the scene.
Qwen Image
- + Clean, professional aesthetic for the capybara's uniform and cap.
- + The passenger has a very realistic and fitting bored expression.
- + Good use of bokeh for the city lights through the windshield.
- − The passenger appears to be sitting in the front passenger seat or an awkwardly positioned middle seat rather than the back seat.
- − The capybara's hands look more like monkey hands than capybara paws.
- − The taxi sign on the roof says 'YOXI' unintentionally.
Verdict: FLUX.2 [flex] is the clear winner for its superior composition, correctly placing the passenger in the back seat to create the requested taxi dynamic. While Qwen Image provides a clean and crisp render, the anatomical choice for the paws and the incorrect seating arrangement make it less faithful to the prompt's narrative.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent text rendering with no spelling errors.
- + Sophisticated parchment edge effect with integrated skulls.
- + Strong cinematic lighting and atmospheric fog.
- − The parchment takes up the whole image area rather than being a poster on a background.
Qwen Image
- + Good inclusion of the thorns and web border requested.
- + Strong color contrast with a deep blue night sky.
- − Significant spelling error in the title ('Halle Party').
- − Text rendering on the scroll is partially broken and disconnected.
- − Visual quality of the bats is somewhat cartoonish compared to the rest of the scene.
Verdict: FLUX.2 [flex] produced a professional-grade invitation with perfect spelling and a cohesive vintage aesthetic. Qwen Image captured the specific border elements well but failed on the critical 'Halloween' title text and had minor rendering issues on the scroll. FLUX.2 [flex] is much more usable as an actual graphic design piece.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent text rendering with clean, bold typography.
- + Superior material textures that realistically blend 3D cartoon style with PBR properties.
- + Accurate sushi models including recognizable nigiri, maki, ginger, and wasabi.
- − The flag icon is floating in space rather than being integrated naturally into the scene.
Qwen Image
- + Strong isometric composition with a nice multi-layered diorama base.
- + Creative inclusion of a physical flag and chopsticks to enhance the miniature scene.
- + Captures the requested soft, playful aesthetic well.
- − Text layout is slightly crowded at the top.
- − Small flag icon near the text is distorted and barely recognizable.
- − The sushi rice texture looks more like large beads than rice grains.
Verdict: FLUX.2 [flex] produced a much cleaner and more professional-looking graphic with superior text rendering and highly detailed PBR materials on the sushi itself. While Qwen Image did a great job with the diorama base and added charming details like chopsticks, its text handling and specific material textures were not as refined as FLUX.2 [flex].
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent dynamic motion with the puppy and kitten appearing to jump through the field.
- + Highly detailed lighting with distinct god rays and dew sparkles on the flowers.
- + Strong adherence to all animal types mentioned in the prompt.
- − The fox kit has a slightly unnatural leg anatomy during its stride.
- − The background tree feels a bit stock-photo-like compared to the foreground.
Qwen Image
- + Beautiful composition with the puppy centered and larger, creating a clear focal point.
- + Soft, pleasing bokeh and high-quality fur textures on the kitten and bunny.
- + Very expressive and clean facial features on all four animals.
- − Less sense of movement compared to Model A, with animals appearing more posed.
- − The butterflies are somewhat oversized relative to the animals.
Verdict: Both models followed the prompt exceptionally well, capturing all four specific animals and the golden hour atmosphere. FLUX.2 [flex] succeeded better at the 'tumbling' and 'chasing' aspect of the prompt with more energetic poses, while Qwen Image produced a more balanced and aesthetically centered composition that leans into the 'wholesome' vibe.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent typography including the correct grave accent on 'Caffè'.
- + Clean vector aesthetic with perfect symmetry and balance.
- + Accurate adherence to all prompt elements, including the banner and steam style.
- − The 'Est. 1720' text is slightly small within the banner.
Qwen Image
- + Strong 'vintage' feel with thicker lines and bolder colors.
- + Good use of texture on the background.
- − Major typographical errors with overlapping and garbled letters in 'Florian'.
- − Incorrect accent mark on 'CAFFÉ' (using an acute accent instead of the requested grave accent).
- − The 'Est. 1720' text uses a modern sans-serif font that clashes with the vintage theme.
Verdict: FLUX.2 [flex] produced a professional-grade vector logo with perfect spelling and elegant composition. Qwen Image struggled significantly with typography, creating an unreadable mess of letters in the brand name and failing to match the classic aesthetic requested.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [flex]
- + Excellent text rendering with no spelling errors.
- + Clean, professional vector aesthetic that perfectly matches the requested NASA-inspired palette.
- + Logical flow and distinct iconography for each mission step.
- − Missing the final 6th step (Landing) on the surface.
- − Small 'floating' artifacts present in the central divider line.
Qwen Image
- + Includes a moon surface for the final landing step.
- + Captures the NASA logo style and astronaut silhouettes well.
- − Numerous spelling errors including 'ApolL', 'Sarth Orbit', and 'Arnnstrong'.
- − Included prompt instructions like '(Stop at landing)' as literal text in the image.
- − Cluttered composition with overlapping elements and inconsistent numbering.
Verdict: FLUX.2 [flex] produced a significantly more professional and polished infographic with perfect typography and a sophisticated flat-vector style. While Qwen Image attempted more of the specific prompt details like the landing surface, its output was marred by severe spelling errors and a failure to filter out meta-instructions from the text. FLUX.2 is the clear winner for its functional design and clarity.
Explore each model
Alibaba's Qwen image model