Black Forest Labs' compact, open-source image generation model with sub-second inference, optimized for production and near real-time applications with multi-reference support
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
FLUX.2 [klein] 4B
#32 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [klein] 4B
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent adherence to lighting instructions with realistic soft window light.
- + Features a very clean, high-quality glass cube with realistic reflections.
- + Strong photographic realism and natural composition.
- − The plant is significantly out of focus compared to the rest of the scene.
Qwen Image Max
- + Successfully includes all elements of the prompt including the plant and sphere.
- + Good lighting on the wooden table surface.
- + Provides a more detailed view of the plant behind the objects.
- − The blue sphere appears to be floating unnaturally in the center of the cube.
- − The perspective of the cube’s bottom edges against the table is slightly distorted.
Verdict: FLUX.2 [klein] 4B produces a more photorealistic image with better physics, as the sphere sits naturally on the bottom of the glass cube. While Qwen Image Max includes more detail in the background plant, its sphere is floating in a way that breaks immersion, making FLUX.2 the preferred choice for visual quality.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent photographic quality and skin texture realism
- + Clear and intricate floral-style engraving on the plate armor
- + Subtle and realistic implementation of scars and dirt
- − The braids and beads are a bit simple and sparse compared to the prompt
- − Lighting on the face is a bit flat despite the bright torches in the background
Qwen Image Max
- + Exceptional detail on the braids and beads, matching the 'small beads' request perfectly
- + Highly 'battle-worn' appearance with deep scars and heavily weathered leather
- + Very dynamic and atmospheric use of torchlight and bokeh sparks
- − The face has a slightly over-processed or 'gritty' digital look compared to the naturalism of Model A
- − The armor engraving is a bit busy and less distinct in some areas
Verdict: While FLUX.2 [klein] 4B produces a much more realistic, high-fidelity photographic portrait, Qwen Image Max captures the specific details of the prompt better, especially regarding the 'battle-worn' aesthetic and the complexity of the beads and braids. Model A feels like a clean studio photo of an actor, whereas Model B feels more like a lived-in character in a fantasy world.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent photorealistic texture on the burger bun and meat patties.
- + Clear and well-centered price starburst element.
- + Secondary text 'LIMITED TIME ONLY' is spelled perfectly.
- − Major spelling errors in the primary title 'MAGIC BURGER'.
- − The burger is not 'exploded' or separated; it is a standard stacked burger.
- − Background embers feel static and less dynamic.
Qwen Image Max
- + Perfect text rendering for all requested strings including 'MAGIC BURGER'.
- + Captures the 'exploded' and 'mid-air motion' prompt requirements much more effectively.
- + Dynamic background with realistic smoke and fiery effects.
- − The €6.99 text slightly overlaps the bottom bun.
- − The lettuce texture looks slightly more digital than the meat.
Verdict: Qwen Image Max is the clear winner as it followed every instruction, including the difficult 'exploded' layout and complex text rendering. FLUX.2 [klein] 4B failed significantly on the primary title spelling and provided a standard burger stack rather than the requested exploded view.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent character placement and framing of both the capybara and the passenger.
- + Highly realistic fur texture and cap integration on the capybara.
- + The passenger's expression perfectly captures the 'bored' instruction.
- − The paws on the steering wheel look slightly detached from the body.
- − The car interior feels a bit small and cramped for a full-size sedan.
Qwen Image Max
- + Strong perspective that highlights the driver's task and the city exterior.
- + Excellent rendering of the car's dashboard and steering wheel.
- + The capybara's professional facial expression is very well executed.
- − The passenger in the back looks distorted and is partially cut off by the edge.
- − The capybara has primate-like hands instead of paws gripping the wheel.
- − Significant cracking texture on the car's interior ceiling is distracting.
Verdict: FLUX.2 [klein] 4B is the clear winner as it successfully renders both characters in a cohesive, cinematic composition that perfectly matches the requested mood. While Qwen Image Max has impressive details on the car's dashboard, the passenger is poorly rendered and the capybara's hands are anatomically incorrect for the species.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Atmospheric cinematic lighting on the pumpkin
- + Strong integration of background elements like the shadowy trees
- − Significant spelling errors throughout almost all text fields
- − Year is missing a digit (206 instead of 2026)
- − Title text is largely illegible
Qwen Image Max
- + Excellent text rendering with no spelling mistakes
- + Followed all layout instructions including the scroll banner and footer details
- + Intricate border detail combining thorns and webs perfectly
- − The transition between the central vignette and the parchment background is slightly harsh
Verdict: Qwen Image Max is the clear winner as it successfully rendered every piece of requested text with perfect spelling and placement. FLUX.2 [klein] 4B failed significantly on the text-to-image aspect, producing gibberish for the main title and missing numbers in the date.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent 3D miniature 'toy' aesthetic
- + Clean and minimal composition
- + Accurate 45-degree isometric perspective
- − Typos in text ('SUSH' instead of 'SUSHI')
- − Incorrect flag icon (looks like Austria or a stylized bar)
- − The plate lacks the requested 'diorama base' look
Qwen Image Max
- + Perfect text rendering for 'JAPAN' and 'SUSHI'
- + Includes a correct Japanese flag icon
- + Accurately interprets the 'diorama base' and 'miniature' concepts with the glass stand
- − Texture on the tamago (egg) sushi is slightly less refined than the salmon
- − The shadows under the text are a bit heavy compared to the soft lighting of the scene
Verdict: Qwen Image Max followed every instruction perfectly, including accurate spelling of 'SUSHI' and the correct Japanese flag, which FLUX.2 [klein] 4B failed to do. Qwen Image Max also better captured the 'miniature 3D diorama' feel with a more complex set of sushi and a clear glass base, whereas FLUX.2 felt like a standard product photo.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent fur texture rendering and soft lighting.
- + Captures the 'wholesome' and 'joyful' vibe very effectively with expressive eyes.
- + Clearly defined species including the red fox kit and tabby kitten.
- − Failed to include a rabbit in the scene.
- − Included two kittens instead of the requested combination of animals.
Qwen Image Max
- + Dynamic 'tumbling' interaction between the golden retriever puppies and the fox.
- + Beautiful use of god rays and wildflower variety.
- + High micro-detail in the fur and whiskers.
- − Completely missed the baby bunny requirement.
- − Replaced the kitten with a second puppy in the middle, failing to provide one of each requested species.
Verdict: Both models failed to include all four requested animals, with FLUX.2 [klein] 4B missing the rabbit and providing two kittens, while Qwen Image Max also missed the rabbit and provided two puppies. FLUX.2 [klein] 4B is slightly preferred for its superior facial expressions and 'cuteness' factor that better aligns with the prompt's request for expressive eyes and a joyful vibe.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Clean vector aesthetic suitable for a modern minimalist logo.
- + Subtle, high-quality paper texture background.
- + Good use of negative space in the cloche icon.
- − Misspelled the primary text as 'CAFFÈ FLAXTION'.
- − Redundant 'Est. 1720' text appears twice, cluttered at the bottom.
- − The banner design is broken/disconnected.
Qwen Image Max
- + Perfect text rendering for both the name and the date.
- + Superior artistic detailing on the cloche and steam.
- + Excellent composition with a professional, cohesive banner design.
- − Background texture is a bit heavy/distressed rather than 'subtle'.
- − Slight overlap between the steam and the top edge of the frame.
Verdict: Qwen Image Max is the clear winner as it followed all textual instructions perfectly, whereas FLUX.2 [klein] 4B failed to spell the restaurant name correctly and included redundant text. Qwen Image Max also provided a much more sophisticated vector illustration style that feels premium and historically appropriate.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.