Black Forest Labs' aesthetically-tuned 12-billion parameter flow transformer optimized for high-quality images with incredible aesthetics, suitable for personal and commercial use
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
FLUX.1 Krea [dev]
#48 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Krea [dev]
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent photographic realism and soft lighting in line with the prompt.
- + Highly convincing glass textures and refractions within the cube.
- + The blue sphere has realistic internal reflections and light play.
- − The plant is mostly behind the cube but less visible through the glass itself compared to model b.
Qwen Image Max
- + Successfully shows the plant partially visible through the glass as requested.
- + Bright, clear lighting and sharp wooden grain details.
- + Accurate placement of all prompt elements.
- − The blue sphere appears to be floating unnaturally in the center of the cube.
- − The perspective of the book on the cube feels slightly skewed.
Verdict: FLUX.1 Krea [dev] produces a more aesthetically pleasing, cinematic photograph with superior glass material rendering, though the sphere rests on the bottom rather than floating. Qwen Image Max follows the transparency requirement for the plant more literally, but the floating placement of the sphere makes the composition feel less grounded and realistic.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent high-contrast cinematic lighting
- + Intricate floral-style engraving on the metal armor
- + Very sharp and lifelike eye details
- − The blood spatter looks a bit like paint rather than battle-worn dirt and scars
- − The braided hair is quite simple with very few beads
Qwen Image Max
- + Superb texture on leather straps and tattered cloth underlayer
- + Detailed scars and weathered skin texture create a gritty aesthetic
- + Colorful beads in the hair add a nice level of detail requested
- − The sparks have a slightly streaky, digital look
- − The torch in the background is a bit distracting from the subject
Verdict: Both models followed the prompt exceptionally well, but Qwen Image Max stands out for its superior textural work on the leather, cloth, and weathered skin. While FLUX.1 Krea [dev] has a very clean, high-fashion cinematic look, Qwen Image Max captures the 'battle-worn' aspect with more realistic grime and damage.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Clean and legible typography for the main title.
- + Good photorealistic texture on the bun and patty.
- + Dynamic presence of real flames at the bottom of the frame.
- − Failed to place the price in a starburst as requested.
- − The burger is less 'exploded' and more of a static stacked vertical assembly.
- − Includes several artifacts and gibberish text at the bottom and right side.
Qwen Image Max
- + Excellent adherence to text styles, including the fiery glow and starburst for the price.
- + Strong sense of motion with suspended, angled components and flying embers.
- + High level of detail in the food textures, particularly the tomato and grilled patty.
- − The background is slightly more cluttered with sparks compared to the clean fire in Model A.
- − The 'Limited Time Only' text style is slightly less professional than the main header.
Verdict: Qwen Image Max is the clear winner as it followed every specific instruction in the prompt, including the complex text effects and the starburst for the price. While FLUX.1 Krea produced a high-quality image, it failed to incorporate the requested starburst and resulted in a more static composition with unwanted text artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Successfully captured the bokehed New York skyline through the window.
- + Clearly depicted the businesswoman in the back seat using her phone as requested.
- + The 'TAXI' text on the cap is rendered correctly.
- − Included an extra person in the front passenger seat who was not in the prompt.
- − The perspective is more from the outside looking in, rather than 'inside' the taxi as specified.
Qwen Image Max
- + Excellent interior perspective that feels immersive and realistic.
- + Highly detailed fur texture and realistic lighting on the capybara.
- + Successfully conveys the 'bored' expression of the businesswoman in the background.
- − The capybara's hands look more like primate hands than capybara paws.
- − The taxi interior shows some minor structural inconsistencies near the ceiling.
Verdict: Qwen Image Max is the winner as it perfectly captured the 'inside' perspective requested by the prompt, whereas FLUX.1 Krea [dev] took a side-view approach from outside the vehicle. Qwen Image Max also followed the character count more accurately, while FLUX.1 Krea added an extra person in the front seat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Strong composition with a nicely illustrated jack-o-lantern
- + Captured the thorny border and cobwebs effectively
- − Numerous text errors including 'Pasty Halloween' and 'You are are invited'
- − Incorrect event details rendering with overlapping text in the bottom banner
- − Date is missing the day and month properly
Qwen Image Max
- + Perfect text rendering of the title, banner, and event details
- + Highly detailed background with a moody sky, crescent moon, and atmospheric trees
- + Complete adherence to every prompt element including the specific date and location
- − The thorn border is a bit dense, making the edges look slightly cluttered
Verdict: Qwen Image Max is the clear winner as it followed every instruction, including rendering complex text strings without a single spelling error. While FLUX.1 Krea [dev] produced a nice illustration, it failed significantly on the text elements, resulting in typos and garbled details.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent adherence to the 'minimal' prompt with a single high-quality salmon nigiri.
- + Features clean, rounded typography that matches the 3D cartoon aesthetic.
- + Very lighting and shadow work on the pedestal and plate.
- − The transition from the rice to the fish texture looks slightly jagged upon close inspection.
Qwen Image Max
- + Features a variety of sushi types that creates a more visually interesting scene.
- + Sophisticated use of materials, particularly the transparency of the glass diorama base.
- + Extremely sharp rendering of textures on the sushi toppings.
- − Includes more garnish than the 'minimal' request specified.
- − The placement of the flag icon is slightly awkward, partially overlapping the text area.
Verdict: FLUX.1 Krea [dev] adhered more closely to the 'minimal' and 'cartoon' style requests, producing a very clean and professional graphic. Qwen Image Max produced a more complex and detailed set of 3D models with superior material rendering, though it was less strictly 'minimal' in its composition. FLUX.1 Krea is the likely winner for its superior balance of typography and image space.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent fur texture rendering and soft lighting.
- + Clear individual subjects with expressive, cute features.
- + Good bokeh effect and depth of field in the meadow.
- − Missed the baby bunny requirement entirely.
- − Includes two kittens instead of one kitten and one bunny.
Qwen Image Max
- + Better captures the 'tumbling together' playful interaction requested.
- + Includes very atmospheric god rays and sparkling dew effects.
- + More vibrant variety of wildflowers.
- − Missed the baby bunny requirement, substituting a second golden retriever puppy.
- − Anatomical issues with the fox's front leg placement.
- − One butterfly appears to be floating without a body.
Verdict: Both models failed to include all four requested animals, specifically missing the baby bunny. FLUX.1 Krea [dev] produced a cleaner, more realistic image with superior fur textures, while Qwen Image Max excelled at the environmental effects and the playful interaction between the animals, despite some anatomical errors.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Elegant etching style that feels authentically vintage.
- + Good use of textured borders for a paper-like feel.
- − Severely botched spelling of 'Florian' as 'FLANOR'IN'.
- − Composition is a bit cluttered where the text overlaps the banner and dome.
Qwen Image Max
- + Perfect text rendering of the brand name 'Caffè Florian'.
- + Cleaner 'vector emblem' aesthetic that matches the logo prompt.
- + Accurate interpretation of the 'Est. 1720' banner.
- − The steam is a bit more illustrative than minimalist.
Verdict: Qwen Image Max is the clear winner because it followed the text requirements perfectly, rendering 'Caffè Florian' and 'Est. 1720' with zero errors. FLUX.1 Krea produced a beautiful vintage illustration style, but it failed on the primary task of a logo: correct spelling of the brand name.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.