Black Forest Labs' aesthetically-tuned 12-billion parameter flow transformer optimized for high-quality images with incredible aesthetics, suitable for personal and commercial use
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Krea [dev]
#47 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Krea [dev]
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent photorealistic rendering of glass and reflections
- + Coherent physics with the sphere resting on the base of the cube
- + Strong adherence to the lighting prompt with soft light coming from the left
- − The sphere is quite dark and closer to navy than a vibrant blue
Qwen Image 2.0
- + Bright and clear colors and textures
- + Highly accurate placement of the plant behind the cube visible through the glass
- − The blue sphere is awkwardly floating in the center of the cube, lacking physics
- − The glass cube has strange internal reflections that look like multiple spheres
- − The book looks slightly disconnected from the top of the cube
Verdict: FLUX.1 Krea [dev] produces a much more realistic image with convincing lighting and physical interaction between the objects. While Qwen Image 2.0 captures the colors and plant visibility well, it fails the logic test by making the sphere float unnaturally and creating confusing reflections in the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent handling of motion blur on passing cars as requested.
- + Captures a wider cinematic view with effective street reflections.
- + Consistent lighting and moody atmosphere that fits the cinematic description.
- − The subject is not actually 'repairing' the bike; he is just holding it.
- − The bicycle's structure is slightly nonsensical with the pedal/chain assembly.
Qwen Image 2.0
- + Strong adherence to the 'repairing' aspect of the prompt with a functional pose.
- + Incredibly high skin texture detail and realistic facial features.
- + Effective 'imperfect framing' that heightens the candid photography feel.
- − Missing the 'motion blur' requested for the passing cars.
- − Rain is less visible compared to the other model.
Verdict: Qwen Image 2.0 provides a much more convincing character study with superior skin textures and a pose that actually depicts 'repairing,' though it misses the motion blur requirement. FLUX.1 Krea [dev] captures the environmental atmosphere and motion blur better, but the subject is passive and the bicycle anatomy is poor. Qwen Image 2.0 is the preferred choice for its realism and prompt adherence regarding the primary action.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent realization of torchlight reflecting on metal surfaces
- + High-quality textures on the engraved armor and fabric
- + Natural integration of bokeh sparks within a dark, moody environment
- − The 'scars' look more like dried blood or surface smears than physical tissue damage
- − Beads in the hair are very small and easy to miss
Qwen Image 2.0
- + Stronger adherence to the 'battle-worn' descriptor with visible physical scars
- + Detailed beads in hair that match the colorful aesthetic
- + Effective use of shallow depth of field against a more active background
- − The eyes look somewhat supernatural/glowing rather than just lifelike
- − The left hand is anatomically distorted with awkward finger placement
- − The armor engraving and metal texture feels slightly flatter compared to Image A
Verdict: FLUX.1 Krea [dev] produces a more cinematic and technically cohesive image with superior lighting and metal textures. While Qwen Image 2.0 captures the specific character traits like the colored beads and prominent scars better, it suffers from anatomical issues in the hand and less convincing material renders.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent structure that closely resembles a real vertical flyer or menu page
- + Professional use of white space and typographic hierarchy
- + Clean, logical separation between text descriptions and the photo grid
- − Text is largely nonsensical scribble
- − Repeats the 'Appetizers' section heading twice instead of including 'Pizza'
Qwen Image 2.0
- + Successfully includes all three requested sections: Appetizers, Pizza, and Mains
- + Extremely high food photography quality with vibrant colors
- + Legible bold sans-serif headers
- − The layout is a simple grid of images rather than a functional menu with descriptions
- − Lack of vertical page context makes it look more like a website category page than a printed menu
Verdict: FLUX.1 Krea [dev] produces a much more realistic menu layout that feels like a professional design for a casual dining establishment, though it fails to include the 'Pizza' header. Qwen Image 2.0 has superior food photography and adheres better to the header requirements, but its composition is a basic grid that lacks the text-heavy utility of a real menu. FLUX is preferred for its superior grasp of graphic design and composition.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent food photography quality with highly realistic textures.
- + All requested text elements are present and readable.
- + Clean, professional ad layout.
- − Failed the 'fiery, glowing effect' request for the primary text.
- − The €6.99 price is just floating rather than in a starburst.
- − Includes garbled AI gibberish text at the bottom.
Qwen Image 2.0
- + Perfectly captured the fiery, glowing effect on the headline text.
- + Strong sense of motion with embers and steam.
- + Correctly placed the price inside a fiery starburst as requested.
- − The 'LIMITED TIME ONLY' text is smaller and lacks the fiery effect specified.
- − Some slight clipping on the bottom of the starburst.
Verdict: Qwen Image 2.0 followed the stylistic instructions much more closely, successfully rendering the fiery text and the price starburst. While FLUX.1 Krea [dev] produced a very clean food image, it missed the specific formatting cues for the typography and included distracting artifacts at the bottom of the frame.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent chalk texture and board smears
- + Clean, centered composition with a realistic wooden frame
- − Significant text errors including 'Gruffle' and 'Clvrucuations'
- − Failed to follow the menu item list as requested, adding extra incorrect lines
Qwen Image 2.0
- + Perfect text accuracy for all requested items and dates
- + Very realistic 'cozy café' atmosphere with depth of field
- + Natural variations in chalk stroke thickness and handwriting slant
- − The layout is slightly cut off on the right edge
- − Minor chalk dust artifacts are heavy in the center
Verdict: Qwen Image 2.0 followed the prompt instructions perfectly, rendering every word of the requested text correctly and with a beautiful, natural chalk aesthetic. FLUX.1 Krea [dev] failed on text legibility and accuracy, hallucinating gibberish words and repeating items despite the high-quality visual texture of the board itself.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent cinematic lighting and composition
- + High quality rendering of the earth and space background
- + Interesting mechanical/sci-fi integration with the horse's saddle
- − Prompt adherence failure: the astronaut is riding the horse, not the other way around
- − Minor anatomical issues where the horse's front legs merge with the suit/equipment
Qwen Image 2.0
- + Sharp details on the spacesuit and horse's coat texture
- + Vibrant colors and clear background details
- + Creative addition of floating water droplets in zero-g
- − Prompt adherence failure: the astronaut is riding the horse, not the other way around
- − The horse's front right leg has an awkward, unnatural joint structure
Verdict: Both models failed the specific spatial logic test of 'horse on top, not vice versa,' instead producing the common trope of an astronaut riding a horse. FLUX.1 Krea [dev] is preferred for its superior cinematic lighting and cleaner composition, whereas Qwen Image 2.0 has slight anatomical distortions in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent photorealism and cinematic lighting
- + Clean, professional taxi driver cap and jacket rendering
- + Accurate depiction of a bored businesswoman looking at a phone
- − Includes an extra passenger in the front seat not mentioned in the prompt
- − Steering wheel position is slightly awkward for the capybara's reach
Qwen Image 2.0
- + Perfect adherence to the single passenger requirement
- + Dynamic angle that clearly shows the capybara's paws on the steering wheel
- + Very realistic capybara fur texture
- − The passenger is sitting in the middle/shadowy area rather than clearly the 'back seat'
- − Composition is slightly cluttered with the car window reflections
Verdict: Both models captured the surrealism of the scene well, but FLUX.1 Krea [dev] provided a more cinematic, high-quality image despite hallucinating an extra passenger. Qwen Image 2.0 followed the prompt details more accurately regarding the number of characters and the placement of the paws, but FLUX.1 Krea [dev] achieved a superior aesthetic that looked like a cohesive movie still.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Features a very intricate border with high-contrast cobwebs.
- + The jack-o-lantern has a vibrant, warm glow that creates good focal interest.
- − Several severe typos in the text including 'Pasty Halloween' and 'to a nights Tright'.
- − The event details are cluttered and poorly formatted at the bottom.
Qwen Image 2.0
- + Excellent text rendering with near-perfect spelling and elegant gothic font selection.
- + Superior atmospheric depth with misty twisted trees and a clouded moon.
- + Layout is cleaner and more professionally balanced for an invitation.
- − The thorns on the border are a bit repetitive and less ornate than Model A.
Verdict: Qwen Image 2.0 is the clear winner as it successfully follows the complex text instructions with almost perfect accuracy, whereas FLUX.1 Krea produced significant spelling errors and incoherent sentences. Qwen Image 2.0 also achieved a much better 'vintage gothic' mood through its use of lighting, mist, and parchment textures.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent adherence to the 'miniature 3D cartoon' and 'isometric diorama' style
- + Text is beautifully stylized and perfectly centered
- + Material textures on the sushi and base are clean and refined
- − The flag is placed in the scene rather than at the top-center with the text
Qwen Image 2.0
- + Followed instructions for text and flag placement at the top-center
- + Good variety of sushi types on the plate
- + Realistic textures on the wooden base and fish
- − Failed to capture the requested '3D cartoon scene' aesthetic, opting for realism instead
- − The text composition feels a bit cramped at the top
- − The garnish is more than 'minimal' compared to typical diorama styles
Verdict: FLUX.1 Krea (dev) followed the aesthetic and layout instructions much more effectively, producing a high-quality isometric diorama that perfectly matches the requested artistic style. While Qwen Image 2.0 followed the text/flag placement more literally, it failed to deliver the 3D cartoon/miniature look requested in the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent fur texture and lighting on the animals
- + Very clean composition with a cinematic shallow depth of field
- + Good rendering of the puppies, kittens, and fox
- − Failed to include the requested baby bunny
- − Includes two kittens instead of one kitten and one bunny
- − Butterflies appear a bit flat and pasted on
Qwen Image 2.0
- + Successfully included all four requested animals (dog, cat, bunny, fox)
- + Captured the 'tumbling together' action much better than Model A
- + Stronger adherence to 'god rays' and 'dew sparkles' in the background
- − The fox's neck and body connection looks anatomically awkward while rolling
- − Visual quality is slightly less 'clean' than Model A's sharp focus
Verdict: Qwen Image 2.0 is the clear winner for prompt adherence as it correctly included the golden retriever, kitten, fox, and bunny, whereas FLUX.1 Krea [dev] substituted the bunny with a second kitten. While FLUX.1 Krea [dev] has slightly more consistent lighting and polish, Qwen Image 2.0 better captured the playful, tumbling interaction and all specific environmental effects requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Excellent vintage engraving style with refined cross-hatching
- + Good use of subtle paper texture for a retro feel
- + Sophisticated typography layout
- − Spelling error in the main name ('FLANOR`IN' instead of 'Florian')
- − The 'Est. 1720' is in a small badge rather than a primary banner
Qwen Image 2.0
- + Perfect spelling of 'Caffè Florian'
- + Clean and readable vector logo aesthetic
- + Accurate placement of 'Est. 1720' on the banner as requested
- − Simplistic shading lacks the sophisticated 'vintage' feel of a heritage brand
- − Steam is depicted inside/on the cloche rather than rising from it
- − Composition is a bit generic
Verdict: FLUX.1 Krea (dev) captures a much more authentic vintage aesthetic with superior texture and line work, but it fails significantly on the spelling of the core brand name. Qwen Image 2.0 follows all text instructions perfectly and provides a clean, usable logo, though it lacks the artistic depth and 'classic' feel of its competitor. Qwen Image 2.0 is the winner primarily due to text accuracy, which is critical for a logo challenge.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Krea [dev]
- + Strong vector aesthetic with a clean, high-contrast style.
- + Accurate NASA-inspired color palette for the background and typography.
- + Crisp text for the main header.
- − Fails significantly on step-by-step logic, including nonsensical icons like a planet with rings for step 1.
- − Nonsensical text and number ordering (1, 2, 3, 4, 0, 6).
- − Icons do not match the requested subject matter (e.g., generic people instead of a lunar module).
Qwen Image 2.0
- + Excellent adherence to the content of all 6 requested steps.
- + Very clean typography with correctly spelled labels for almost every stage.
- + Accurate iconography for the Saturn V, Lunar Module, and orbital trajectories.
- − Minor spelling error in 'Translunjar'.
- − The vertical composition is slightly cramped toward the bottom.
Verdict: Qwen Image 2.0 followed the complex multi-step instructions almost perfectly, providing accurate icons and text for each phase of the mission. FLUX.1 Krea [dev] failed to provide relevant iconography, showing generic space assets and a broken numerical sequence that disregarded the mission steps. Qwen's ability to render legible, mostly correct text and appropriate vector symbols makes it the superior choice for an infographic.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request