Black Forest Labs' 12-billion parameter flow transformer for high-quality text-to-image generation, suitable for personal and commercial use with streaming support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [dev]
#16 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [dev]
0%
win rate
Ties
0%
Qwen Image 2512
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent lighting and photorealistic textures on the book and sphere.
- + Creative glass refraction and internal reflections.
- + High aesthetic appeal with a clean, modern look.
- − The glass structure is more of a hollow frame/stand than a solid closed cube.
- − The sphere appears to be floating mid-air inside the structure rather than resting.
Qwen Image 2512
- + Accurately depicts a standard closed glass cube.
- + Follows the positioning of the green plant behind the glass more clearly.
- + Excellent wood grain texture and light source consistency.
- − The sphere's reflection on the bottom glass panel is slightly misaligned with its physical position.
- − The blue sphere has a slightly flatter, more 'plastic' matte texture compared to the environment.
Verdict: Both models followed the prompt instructions very well. FLUX.1 [dev] produced a more artistically striking image with superior lighting and materials, whereas Qwen Image 2512 was more literal in its interpretation of a 'glass cube'. Qwen Image 2512 is the winner because it better represents the spatial relationship between the sphere, the cube, and the plant behind it.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent handling of shallow depth of field and bokeh
- + The bicycle design is sleek and the colors are vibrant
- + Correct execution of rain droplets and pavement reflections
- − The subject is holding the handlebars rather than repairing the bike
- − The background cars feel static rather than having motion blur
- − The framing is very centered and deliberate, missing the 'imperfect' request
Qwen Image 2512
- + Successfully captures an 'imperfect' candid framing and pose
- + The man appears more engaged with the bike, fitting the repair theme better
- + Highly realistic skin textures and believable lighting
- − Anatomy issues with the left hand resting on the seat
- − The bicycle geometry is slightly warped in the rear frame
- − The depth of field is a bit busy compared to the requested 50mm look
Verdict: FLUX.1 [dev] produces a much cleaner, more aesthetically pleasing image with superior technical rendering of rain and light, but it fails to capture the 'repairing' action or the 'imperfect framing'. Qwen Image 2512 feels more like a genuine candid street photo and follows the 'imperfect' and 'candid' keywords better, though it suffers from some structural distortions in the hand and bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [dev]
- + Features very lifelike and expressive eyes.
- + Excellent depth of field with soft, pleasing bokeh sparks.
- − Missed the request for beads in the hair.
- − Skin appears too clean and smooth for a 'battle-worn' character with only minor freckling instead of scars.
- − The armor lacks the requested ornate engravings.
Qwen Image 2512
- + Perfect adherence to all prompt details including beads, scars, dirt, and ornate engravings.
- + Outstanding texture on the leather straps and metal hardware.
- + Captures the 'battle-worn' aesthetic much more convincingly.
- − The torch in the background is a bit distracting and less 'bokeh' than requested.
Verdict: Qwen Image 2512 followed every specific detail of the prompt, including complex elements like beads in the hair, ornate engravings on the armor, and realistic leather straps. FLUX.1 [dev] produced a high-quality portrait but failed to include the beads and engravings, resulting in a character that looked more like a clean fashion model than a battle-hardened paladin.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean, professional typographic layout that mimics a real restaurant menu.
- + High-quality, isolated food photography that integrates well with the white space.
- + Accurate sans-serif font usage and clear sectional hierarchy.
- − Missed the specific 'grid' layout requested for the food photos.
- − Included a strange phrase at the bottom ('Chin, you clone!') instead of standard menu info.
- − Layout is slightly too sparse in the center.
Qwen Image 2512
- + Successfully incorporated the grid of food photos requested in the prompt.
- + Vibrant color accents and clear bold headings.
- + More diverse food representation within the layout.
- − The text is highly distorted and contains many nonsensical characters.
- − Pricing and text alignment are inconsistent and messy.
- − The 'grid' layout feels a bit cramped compared to the minimalist request.
Verdict: FLUX.1 [dev] produced a more professional and realistic menu design with superior legibility and high-end aesthetics, although it failed to follow the grid layout for photos. Qwen Image 2512 followed the grid requirement better, but the text rendering is poor and the overall composition feels cluttered and less 'minimalist' than requested. FLUX.1 [dev] is the winner for its commercial-grade visual quality and cleaner execution.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean, symmetrical composition with a nice vertical stack.
- + Photorealistic texture on the burger buns and patties.
- − Completely missing the main 'MAGIC BURGER' title text.
- − Lacks the requested 'starburst' for the price tag.
- − The background is more of a generic flame at the bottom rather than a fiery atmosphere.
Qwen Image 2512
- + Excellent adherence to all text requirements, including the fiery 'MAGIC BURGER' title.
- + Features a very dynamic 'exploded' look with flying ingredients and embers.
- + Correctly integrates the price within a starburst graphic as requested.
- − The 'LIMITED TIME ONLY' text is missing the 'TIME' keyword, appearing as 'LIMITED ONLY'.
Verdict: Qwen Image 2512 is the clear winner as it followed almost all prompt instructions, including complex text rendering, the starburst price tag, and the specific fiery atmosphere. FLUX.1 [dev] failed to include the primary title and several stylistic elements, resulting in a much simpler image that did not meet the advertisement requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent text legibility and alignment.
- + Accurately rendered all requested menu items with correct pricing.
- + Clean, modern chalkboard aesthetic.
- − The text looks more like a digital font than hand-drawn chalk.
- − Lowercase 'sh' in the bottom text is messy and contains spelling errors like 'drish' and 'aur'.
Qwen Image 2512
- + Text has a much more convincing chalk texture with smudges and dust.
- + Stronger adherence to the 'elegant cursive' style for the lettering.
- + Better captures the natural variations and slant requested in the prompt.
- − Spelling error in 'Risitto' (instead of Risotto).
- − Some letters are slightly cut off by the frame on the left.
Verdict: While FLUX.1 [dev] produced very legible text, it failed to capture the 'handwritten chalk' feel, resulting in something that looks like a digital overlay. Qwen Image 2512 much better understood the texture and artistic style of a chalkboard, providing a more authentic atmosphere despite a single spelling error in the word 'Risitto'.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent cinematic lighting and atmosphere.
- + Good adherence to the requested horse riding astronaut concept.
- + Clean, artistic composition with a dreamy feel.
- − The horse's anatomy is distorted, particularly with five legs visible.
- − Fails the negative constraint 'horse on top, not vice versa' by placing the astronaut on top.
Qwen Image 2512
- + High level of technical detail in the space suit and saddle.
- + Realistic horse anatomy and textures.
- + Clear facial features visible through the helmet.
- − Fails the negative constraint 'horse on top, not vice versa' by placing the astronaut on top.
- − The composition is a bit cramped compared to the other model.
Verdict: Both FLUX.1 [dev] and Qwen Image 2512 failed the tricky negative constraint to have the horse on top of the astronaut, instead providing the standard astronaut-riding-horse trope. FLUX.1 [dev] delivers a much more cinematic and surreal atmosphere, but it has significant anatomical errors in the horse's legs, while Qwen Image 2512 is technically superior regarding realistic details and anatomy.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photographic lighting and shallow depth of field.
- + Very realistic capybara fur texture and human skin tones.
- + Good adherence to the 'bored expression' prompt for the passenger.
- − The passenger appears to be in the front seat rather than the back as requested.
- − The capybara only has one paw clearly on the steering wheel instead of both.
Qwen Image 2512
- + Perfectly follows the spatial layout with the passenger in the back seat.
- + Accurately depicts both front paws on the steering wheel.
- + The capybara's hat looks more like a formal driver/service cap.
- − The passenger's expression looks more like a pout or sadness than 'bored/normal'.
- − Slightly less 'cinematic' lighting compared to Model A.
Verdict: FLUX.1 [dev] produced a more visually striking and realistic image in terms of texture and lighting, but it failed to place the passenger in the back seat. Qwen Image 2512 followed all parts of the complex prompt perfectly, including the specific positions of the paws and the passenger's seating location, making it the more accurate tool for this specific request.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Strong composition with a central glowing Jack-o-lantern
- + Clean thorn border that frames the central elements well
- − Significant text errors including 'Halloween Party' written as 'Lalloween Ranty'
- − Repetitive time data and character artifacts in the location text
- − Missing the spider web elements requested in the prompt
Qwen Image 2512
- + Excellent adherence to all prompt elements including thorns and spider webs
- + Accurate rendering of the secondary scroll banner text
- + Superior atmospheric lighting and more detailed 'twisted trees'
- − Minor spelling error in the main title ('Hallowern' instead of 'Halloween')
- − The layout is a bit crowded compared to Model A
Verdict: Qwen Image 2512 is the clear winner as it followed almost every specific instruction, including the inclusion of spider webs and the specific placement of text on a scroll. FLUX.1 [dev] failed significantly on text legibility and accuracy, while Qwen Image 2512 produced a much more polished and atmospheric invitation despite a small typo in the header.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent soft, clay-like textures as requested
- + Clean and minimal diorama base
- + High-quality lighting and depth of field
- − Text rendering is messy with extra characters and misspellings (SUSH CATON)
- − The fish texture on the sushi looks slightly repetitive
Qwen Image 2512
- + Perfect adherence to text instructions with bold, stylized fonts
- + Stronger variety in the sushi models (nigiri and maki)
- + Excellent isometric composition and diorama detailing
- − Slightly busier than requested for 'minimal garnish'
- − The flag icon is placed next to the text rather than below 'SUSHI'
Verdict: Qwen Image 2512 is the clear winner as it perfectly followed the complex text instructions while maintaining the desired isometric 3D cartoon aesthetic. While FLUX.1 [dev] produced a very beautiful, soft image, it failed to render the text correctly and lacked the variety in sushi types shown by Qwen.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent depiction of golden hour lighting and bokeh
- + Consistent artistic style across all characters
- − Failed to include a kitten and instead generated two puppies
- − Characters look more like Pixar-style 3D models than photorealistic animals
- − Anatomy of the ears and paws on the fox/bunny hybrids is unrealistic
Qwen Image 2512
- + Successfully included all four specific animals: puppy, kitten, bunny, and fox kit
- + High level of texture detail in the fur and butterfly wings
- + Follows the 'photorealistic' requirement much better than the competing image
- − The scale of the animals is slightly off, with the bunny and kitten being very small compared to the puppy
- − The paws of the puppy are fused or anatomically unclear beneath the bunny
Verdict: Qwen Image 2512 is the clear winner as it adhered to all parts of the prompt, including the specific list of animals which FLUX.1 [dev] failed by omitting the kitten. While FLUX.1 [dev] produced a very cute, stylized image, Qwen Image 2512 met the 'hyper-photorealistic' and animal-specific requirements with much higher accuracy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean vector aesthetic suitable for a real logo
- + Accurate color palette of warm brown and cream
- + Correctly included the 'Est. 1720' text
- − Major spelling errors in the brand name ('Flariláan') and descriptors ('RESEAURANT')
- − Introduction of random, contradictory dates ('11011' and '1941')
Qwen Image 2512
- + Perfect text rendering of 'Caffè Florian'
- + Excellent execution of the requested steam effect
- + Consistent vintage texture on the background and illustrations
- − The cloche is overly detailed and illustrative rather than a 'minimalist' logo
- − Slightly less formal composition for an emblem
Verdict: Qwen Image 2512 is the clear winner because it correctly spelled the requested brand name 'Caffè Florian', whereas FLUX.1 [dev] produced significant typos and added confusing extra numbers. While Qwen's style is more illustrative and less minimalist than requested, its overall quality and accuracy to the text prompt are far superior.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent adherence to the requested NASA-inspired color palette.
- + Superior composition with a balanced, orbital flow.
- + Clean, professional vector aesthetic that looks like actual graphic design.
- − Text consists of nonsensical 'gibberish' placeholders despite some legible characters.
- − Includes irrelevant celestial bodies like ringed planets not requested in the prompt.
Qwen Image 2512
- + Successfully renders legible English text for most labels.
- + Includes specific requested icons like the lunar module and crew names.
- + Strictly follows the numbered steps requested in the prompt.
- − Confusing layout with repetitive numbering (two step 2s, two step 3s).
- − Visual style is less 'flat vector' and more illustrative with complex shading and textures.
- − Vertical composition feels cluttered compared to a standard infographic.
Verdict: FLUX.1 [dev] produced a far more aesthetically pleasing and professionally composed vector infographic, though its text is illegible. Qwen Image 2512 followed the logical instructions more literal by including the numbered steps and names, but the layout is disorganized with redundant labels and a less consistent visual style. FLUX.1 [dev] is preferred for its superior design and color palette, which better matches the 'modern vector' request.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.