Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#58 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0%
win rate
Ties
0%
Qwen Image 2512
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to lighting instructions with a clear soft glow from the left.
- + Highly realistic textures on the wood grain and glass surfaces.
- + Superior compositional balance with the plant well-integrated behind the object.
- − The plant is slightly more obscured than requested, though still visible through the glass.
Qwen Image 2512
- + Successfully placed all requested elements in the correct spatial relationship.
- + The glass has a realistic green tint commonly found in thick-paned glass.
- − The glass refractive logic is inconsistent, showing an odd duplication of the sphere to the right.
- − The lighting is somewhat flat and does not clearly emphasize the 'from the left' instruction as strongly as Model A.
Verdict: Both models followed the complex spatial prompt perfectly. FLUX.1 Kontext [dev] is the winner due to its superior lighting, more realistic material rendering, and lack of visual artifacts compared to the glass refraction errors and sphere duplication seen in Qwen Image 2512.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent wet pavement reflections and rain-streaked atmosphere
- + Clear, high-resolution rendering of the subject and bicycle
- − The man is just standing with the bike rather than 'repairing' it
- − The framing is very centered and clean, lacking the requested 'imperfect framing' and 'motion blur'
Qwen Image 2512
- + Successfully captures the requested 'imperfect framing' for a candid feel
- + Better representation of repair work with the subject's squatting posture
- + Skin texture and lighting feel significantly more realistic and less digital
- − Anatomy issues with several extra fingers on the man's hands
- − The bike's physical structure is slightly warped in the middle
Verdict: While FLUX.1 Kontext [dev] produces a much cleaner and technically sharp image, it fails to incorporate the specific 'candid' and 'motion blur' instructions, resulting in a posed look. Qwen Image 2512 captures the photographic style and mood of the prompt much more accurately, though it suffers from significant AI artifacts in the hands.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong composition with dramatic warm lighting
- + Very detailed and clean engraving on the plate armor
- + Good adherence to the bokeh sparks and shallow depth of field
- − Failed to include the hair braids and beads requested in the prompt
- − Skin looks slightly too smooth and clean for a 'battle-worn' character
- − The facial scars look painted on rather than physical texture
Qwen Image 2512
- + Excellent adherence to specific details like braided hair with beads
- + Realistic skin texture including dirt, pores, and physical scarring
- + Higher level of detail on the leather straps and metal weathering
- − The torch in the background is a bit distracting in the composition
- − The sparks look slightly more artificial than the lighting in Model A
Verdict: Qwen Image 2512 is the clear winner as it successfully incorporated every element of the prompt, including the specific hair braids and bead details that FLUX.1 Kontext [dev] ignored. While both models produced high-quality armor textures, Qwen Image 2512 captured the 'battle-worn' aesthetic much more convincingly through realistic skin imperfections and weathered equipment.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Features high-quality, professional food photography with realistic textures.
- + Uses a modern, asymmetrical layout that feels high-end.
- + Excellent use of whitespace and clean sans-serif typography.
- − The text content is largely gibberish.
- − The layout is somewhat fragmented and doesn't clearly separate the specific requested sections.
Qwen Image 2512
- + Excellently organized layout with clear sections for Appetizers, Pizza, and Mains as requested.
- + Good use of vibrant color accents in the UI elements.
- + The grid system for food photos is very clean and structured.
- − Food photography looks slightly more artificial and repetitive compared to the other model.
- − Font rendering is somewhat distorted and inconsistent in weight.
Verdict: Qwen Image 2512 followed the layout requirements much more effectively, providing the specific sections requested with a clear hierarchy that looks like a real menu. FLUX.1 Kontext [dev] produced significantly higher quality food photography and a more artistic composition, but failed to create a logical menu structure, feeling more like a magazine spread.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering for the main title
- + Clean and professional graphic design layout
- + Vibrant colors and high-contrast lighting
- − Failed the 'exploded burger' requirement, showing a mostly intact burger
- − Made a spelling error in 'ONLY' (rendered as LNHLY)
- − The starburst element looks like a flat vector graphic rather than a fiery effect
Qwen Image 2512
- + Brave interpretation of the 'exploded' requirement with dynamic flying ingredients
- + Text effectively integrates the requested fiery and glowing effects
- + Higher photorealistic detail in the food textures such as the patty and tomatoes
- − Text omits the word 'TIME' from the 'LIMITED TIME ONLY' phrase
- − The main title text is slightly less crisp than Model A
- − Composition is a bit more cluttered due to the debris
Verdict: Qwen Image 2512 is the preferred result because it much more accurately captured the 'exploded' and 'fiery' aesthetic of the prompt, whereas FLUX.1 Kontext [dev] presented a static burger and struggled with basic spelling. While Qwen missed one word in the secondary text, its superior adherence to the complex visual instructions and better textured food rendering makes it a more effective ad.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully rendered most words requested in the prompt.
- + Captures a authentic 'home-made' chalkboard look with natural imperfections.
- − Significant spelling errors and repetitions like 'Mashroom/Risoktso' and 'with with'.
- − The date rendering is garbled and unreadable.
- − The handwriting style is clunky rather than elegant cursive.
Qwen Image 2512
- + Excellent typography with a much more elegant cursive style as requested.
- + High level of background detail and realistic chalk textures, including eraser smudges.
- + Virtually perfect spelling and layout of all requested text.
- − One minor spelling error in 'Risitto' (should be Risotto).
Verdict: Qwen Image 2512 is the clear winner for its superior typography and adherence to the 'elegant cursive' requirement. While FLUX.1 Kontext [dev] struggled with basic spelling and layout, Qwen Image 2512 produced a professional-looking menu with realistic chalk textures and nearly perfect text accuracy.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully followed the difficult spatial prompt of the horse being on top
- + High cinematic quality and clean rendering of the astronaut's suit
- − The 'riding' aspect is a bit ambiguous as the horse is floating behind/above rather than sitting on the astronaut
- − Proportions of the horse's legs are slightly unusual
Qwen Image 2512
- + Excellent visual detail on the space suit and saddle equipment
- + Dynamic lighting and high resolution image
- − Failed the negative constraint entirely by placing the astronaut on top
- − A common trope interpretation that ignores the specific user request
Verdict: FLUX.1 Kontext [dev] successfully navigated the challenging prompt of placing the horse on top of the astronaut, showing much better prompt adherence despite a slightly less dynamic composition. Qwen Image 2512 ignored the specific instruction to reverse the positions, providing a standard 'astronaut riding a horse' image which fails the logic test of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent photographic lighting and skin textures on the passenger
- + High-quality rendering of fur and fabric
- − The capybara only has one hand on the steering wheel, failing the specific prompt instruction
- − The composition feels slightly cramped and cuts off the driver's perspective
Qwen Image 2512
- + Follows the prompt instruction for both paws on the steering wheel perfectly
- + The passenger's expression more accurately reflects the requested 'bored' look
- + Better framing that shows more of the taxi environment and the relationship between driver and passenger
- − The hands on the steering wheel are slightly distorts and look a bit more like a blend of human and animal anatomy
- − The passenger's hair and phone details are slightly softer than in Model A
Verdict: Qwen Image 2512 is the clear winner as it adhered much better to the specific constraints of the prompt, including the positioning of both paws on the wheel and the specific expression of the passenger. While FLUX.1 Kontext [dev] has slightly superior photorealistic textures, it failed on the motor control details requested in the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering for the main title and date footer.
- + Good adherence to the requested thorns and bats elements.
- + Strong contrast with the dark parchment background.
- − The text in the small banner is garbled and unreadable.
- − Visual style is a bit flat and resembles clip-art rather than 'cinematic lighting'.
- − The brand name 'The Arches' is misspelled as 'The Argiiah's'.
Qwen Image 2512
- + Beautiful cinematic lighting and atmosphere with a moody night sky.
- + Highly legible text across all areas, including the small banner.
- + Intricate border design that perfectly captures the webs and thorns requested.
- − Small spelling error in the word 'Hallowern' in the main title.
- − The bats are a bit small and less prominent than in the other model.
Verdict: Qwen Image 2512 produces a much more atmospheric and professional-looking invitation with superior lighting and composition, despite a minor typo in the main header. FLUX.1 Kontext [dev] has better spelling on the word 'Halloween', but falls apart on the banner text and lacks the requested cinematic quality.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Extremely clean and minimalist aesthetic.
- + Perfectly legible text rendering.
- − The flag icon is stylized and incorrect.
- − The sushi design is oversimplified and lacks realistic material texture.
Qwen Image 2512
- + Excellent 3D diorama composition with rich textures.
- + Accurately represents the Japanese flag icon.
- + Beautiful isometric perspective and detail level.
- − The word 'SUSHI' is slightly off-center compared to the top text.
- − Contains slightly more garnish than the 'minimal' request suggested.
Verdict: Qwen Image 2512 is the clear winner as it perfectly captures the '3D miniature diorama' and 'realistic PBR materials' requested, resulting in a much more professional look. While FLUX.1 Kontext [dev] is very clean, its interpretation is too abstract and fails to deliver the detailed miniature effect intended by the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Natural sense of movement and 'tumbling' as requested
- + Excellent bokeh and realistic lighting depth
- + Captures the 'joyful wholesome vibe' effectively
- − Failed to include all four requested animals, missing the fox and bunny
- − The kitten species are not as diverse as requested (two kittens/cats instead of variety)
- − Butterflies look a bit artificial compared to the environment
Qwen Image 2512
- + Successfully included all four specific animals: golden retriever, tabby kitten, bunny, and fox kit
- + Beautiful rendering of god rays and dew sparkles on the flowers
- + Highly detailed fur textures and correct anatomy for all animals
- − Static composition; animals are posing rather than 'playfully chasing and tumbling'
- − The butterflies appearing to land on the animals makes it feel more like a portrait than an action scene
Verdict: Qwen Image 2512 followed the complex prompt much more accurately by including all four distinct animals (puppy, kitten, bunny, and fox), whereas FLUX.1 Kontext [dev] only generated a puppy and two small cats/kittens. While FLUX.1 Kontext [dev] captured the requested movement and action better, Qwen Image 2512's superior adherence to the subject list and excellent lighting effects make it the winner.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the minimalist prompt requirement
- + Perfect typography rendering for all text elements
- + Clean vector-style execution with high contrast
- − Missed the request for a banner element
- − Minimalist style might be considered too basic for a 'vintage' aesthetic
Qwen Image 2512
- + Successfully included all elements including the cloche, steam, and banner
- + Beautiful vintage texture and cross-hatching detail
- + Classic typography fits the historical coffee house theme perfectly
- − Less 'minimalist' than Model A, opting for a more illustrative style
- − Small artifacts in the steam rendering
Verdict: While FLUX.1 Kontext [dev] produced a very clean and modern minimalist logo, Qwen Image 2512 much more accurately followed the complex prompt by including the 'Est. 1720' banner and the requested steam. Qwen's style better captures the 'vintage' and 'warm' atmosphere requested, whereas FLUX missed the banner entirely.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully uses the requested navy and muted red color palette.
- + Achieves a clean, flat-vector aesthetic that looks like professional graphic design.
- − The text is largely unreadable or nonsensical scrambled characters.
- − The icons are highly abstract and do not clearly represent the specific mission steps requested.
- − Includes a misspelling of 'Apollo' in the main header.
Qwen Image 2512
- + Excellent adherence to the iconography requested, including the Saturn V and Lunar Module.
- + Text is highly legible with accurate names like Armstrong and Collins included.
- + Layout is logical and follows the narrative arc of the mission steps.
- − Included 'Steps stop at landing:' text literally from the prompt instructions inside the image.
- − Has some redundant numbering and slight spelling errors in secondary labels like 'Desceeint'.
Verdict: Qwen Image 2512 is the clear winner as it successfully translated the complex infographic requirements into a logical, readable poster with accurate icons for each stage. While FLUX.1 Kontext [dev] has a nice graphic style, its complete failure to render legible text or recognizable space hardware makes it ineffective as an infographic.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.