OpenAI's previous generation image model with higher quality than DALL-E 2 and support for larger resolutions
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 3
#40 of 62 in Text-to-Image
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
Where the votes landed
DALL-E 3
0%
win rate
Ties
0%
FLUX.1 Kontext [dev]
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 3
- + High artistic detail with intricate textures on the wood and sphere.
- + Excellent lighting and atmosphere.
- + Correct placement of the plant behind the cube.
- − Failed the primary spatial instruction by putting the book inside the cube and the sphere on top of the book.
- − Conceptualized the cube with a bulky wooden frame rather than a simple glass cube.
FLUX.1 Kontext [dev]
- + Perfect adherence to spatial instructions including the book on top and sphere inside.
- + Realistic textures and reflections on the glass and wooden surface.
- + Correct interpretation of 'soft window light from the left' including the visible window.
- − The sphere is slightly large for the described 'small sphere' proportion.
- − The cube has a mirror-like bottom surface not specifically requested.
Verdict: DALL-E 3 failed significantly on spatial reasoning, placing the book inside the cube and the sphere on top of it, despite the prompt's clear instructions. FLUX.1 Kontext [dev] followed every instruction perfectly, accurately placing the sphere inside the cube and the book on top while maintaining high visual realism.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 3
- + Captures the 'repairing' action perfectly with the man crouching and using tools.
- + Excellent adherence to the 'imperfect framing' and 'shallow depth of field' with foreground bokeh elements.
- + Very realistic wet pavement reflections and cinematic atmospheric lighting.
- − The man's skin texture on his back looks slightly unnatural/overly aged.
- − Anatomy of the feet is a bit distorted.
FLUX.1 Kontext [dev]
- + High clarity and sharp rendering of the subject's face.
- + Explicitly shows raindrops throughout the frame.
- − Completely failed the 'repairing' prompt; the man is standing with/mounting the bike.
- − Composition is too centered and clean, missing the 'imperfect' and 'candid' street photography aesthetic.
- − Lacks the requested 'motion blur from passing cars'.
Verdict: DALL-E 3 followed the complex prompt requirements much more effectively, accurately depicting the specific action of 'repairing' the bike and utilizing creative foreground elements to achieve an 'imperfect' candid look. FLUX.1 Kontext [dev] failed to show the repair action and produced a more generic, posed-looking portrait that ignored several key technical descriptors like motion blur and framing style.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 3
- + Excellent close-up composition highlighting high-fidelity skin and armor textures.
- + Strong adherence to the 'engraved plate armor' and 'warm torchlight' prompts.
- + Superior rendering of fine details like the beard hair and the fabric grain on the collar.
- − Failed to include the requested 'braided hair with small beads'.
- − The scars appear a bit stylized/digital rather than naturally aged skin damage.
FLUX.1 Kontext [dev]
- + Captures a more intense, brooding emotion suitable for a battle-worn character.
- + Features subtle braiding in the hair, moving closer to that specific prompt requirement.
- + Good use of bokeh sparks in the background.
- − The armor engraving is less 'ornate' and looks more like cast molding compared to Image A.
- − The lighting is somewhat flat and lacks the dramatic warm reflections on the metal requested.
- − The texture on the leather straps is less distinct and detailed than the competitor.
Verdict: DALL-E 3 (Image A) produces a much more detailed and visually striking portrait that excels in texture and lighting, though it misses the specific detail of the hair braids. FLUX.1 Kontext [dev] (Image B) captures the character's mood and hair requirements better but falls behind in image clarity and the specific rendering of 'highly detailed' materials like leather and cloth.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 3
- + Provides a complete multi-page layout view suitable for a presentation.
- + Accurately includes several sections labeled for pizza, appetizers, and mains.
- + Uses vibrant color accents that enhance the professional layout.
- − Text rendering is very poor and distorted, even for large headings.
- − The grid layout feels a bit cluttered compared to a professional menu.
FLUX.1 Kontext [dev]
- + Excellent typography with clean, bold sans-serif fonts that look modern.
- + High-quality, appetizing food photography with vibrant colors.
- + Strong grid composition that follows contemporary design trends.
- − Text is largely gibberish with several misspelling artifacts.
- − Fails to clearly distinguish between 'appetizers', 'pizza', and 'mains' as requested.
Verdict: FLUX.1 Kontext [dev] creates a much more visually striking and modern design that accurately follows the 'bold sans-serif' and 'grid' requirements, despite the gibberish text. DALL-E 3 succeeds in following the specific section labels (Pizza, Mains) and presenting a full-page layout, but the overall image quality and text clarity are significantly lower than its competitor.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 3
- + Excellent 'exploded' effect with all components suspended separately.
- + Superior photorealistic detail in the food textures.
- + Dynamic and creative lighting that enhances the sense of motion.
- − Multiple spelling errors in the text including 'MAGIC BURGR' and 'Limiited'.
- − The price is not in a starburst as requested.
FLUX.1 Kontext [dev]
- + Perfect text rendering for 'MAGIC BURGER' and '€6.99'.
- + Includes the starburst element for the price as requested.
- + Clean composition with legible typography.
- − Failed to provide an 'exploded' view; the burger is mostly assembled and static.
- − Spelling error in 'ONLY' which appears as 'LNHLY'.
- − The visual style is more illustrative and less photorealistic than requested.
Verdict: DALL-E 3 captures the complex 'exploded burger' concept with much higher fidelity and realism, though it fails significantly on text accuracy. FLUX.1 Kontext [dev] handles the typography and specific layout features like the starburst better, but it fails to deliver the core 'dynamic, exploded' motion required by the prompt. DALL-E 3 is the preferred choice for the quality of the primary subject matter.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 3
- + Excellent chalk texture throughout the image
- + Atmospheric lighting and café environment details
- + Artistic flourishes and illustrations enhance the board aesthetic
- − Significant spelling errors on almost every line
- − Incorrect prices for items
- − The layout is cluttered and difficult to read as a functional menu
FLUX.1 Kontext [dev]
- + High legibility for the handwriting
- + Followed the menu items more accurately than the competitor
- + Clean composition that resembles a real café chalkboard
- − Repetitive words in text such as 'Mushroom Mashroom' and 'with with'
- − The date is garbled and unreadable
- − The chalk texture is very clean and lacks the natural dustiness of real chalk
Verdict: FLUX.1 Kontext [dev] is the winner because it provides a much more legible and accurate representation of the requested menu items, despite some repetitive word glitches. DALL-E 3 produces a more visually impressive and textured image, but fails significantly on word accuracy and price transcription.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 3
- + Excellent cinematic atmosphere and background detail
- + Stunning use of light and color in the nebulae and clouds
- + High level of polish on the astronaut and horse silhouettes
- − Failed the negative constraint to have the horse on top of the astronaut
- − Standard interpretation of the prompt without surreal reversal
FLUX.1 Kontext [dev]
- + Successfully followed the difficult spatial instruction of horse 'on top'
- + Creative and literal interpretation of the surreal request
- + Great facial detail inside the helmet
- − The composition is somewhat static and less cinematic than the competitor
- − Background is relatively plain compared to the vibrant space scene in Image A
Verdict: While DALL-E 3 produced a far more beautiful and cinematic image, it completely ignored the specific spatial instruction for the horse to be on top. FLUX.1 Kontext [dev] successfully followed the 'horse on top' constraint, creating a truly surreal image as requested, even if the overall environment is less visually striking.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 3
- + Excellent detailed rendering of capybara fur and whiskers.
- + Cinematic lighting consistent with a night scene in a city.
- + Clear taxi driver aesthetic with the uniformed cap and dashboard details.
- − Completely failed to include the human businesswoman in the back seat.
- − The cap is black/navy rather than the requested yellow.
FLUX.1 Kontext [dev]
- + Successfully included all prompt elements including the businesswoman and the cell phone.
- + The yellow cap color matches the prompt perfectly.
- + Superior composition that allows the viewer to see both the driver and the passenger clearly.
- − The capybara's face is slightly less 'photorealistic' and more anthropomorphized than Model A.
- − One of the front paws is resting on its lap rather than both being on the steering wheel.
Verdict: FLUX.1 Kontext [dev] is the clear winner as it successfully incorporated the entire narrative of the prompt, including the bored businesswoman in the backseat which DALL-E 3 omitted entirely. While DALL-E 3 had slightly better textures on the animal's fur, its failure to follow the character requirements makes it an incomplete response to the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent gothic aesthetic with intricate 3D-like textures and layered composition.
- + Highly cinematic lighting that makes the elements feel integrated and atmospheric.
- + Rich detail in the border elements including webs, thorns, and twisted trees.
- − Text rendering is almost entirely illegible gibberish, failing the technical prompt requirements.
- − Layout is overly cluttered, making it difficult to read as a functional invitation.
FLUX.1 Kontext [dev]
- + Most text is legible and follows the prompt instructions for specific dates and locations.
- + Clean, high-contrast layout that functions well as a graphic design piece.
- + Follows the specific request for a banner scroll and thorns border accurately.
- − The 'dark parchment' texture requested in the prompt is missing in favor of a plain black background.
- − The scroll text contains minor spelling errors and artifacts despite the main headings being clear.
Verdict: While DALL-E 3 produces a much more visually stunning and atmospheric 'cinematic' image, it completely fails to render the requested text accurately. FLUX.1 Kontext [dev] is the winner because it successfully carries out the primary function of an invitation by providing legible details and follows the specific text instructions, despite having a simpler visual style.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent execution of the isometric 3D diorama style.
- + High level of detail with multiple sushi types and realistic lighting.
- + Clean and visually appealing composition.
- − Completely ignored the request for large bold text at the top-center.
- − Failed to place 'JAPAN' and 'SUSHI' text in the image.
FLUX.1 Kontext [dev]
- + Successfully followed text instructions for 'JAPAN' and 'SUSHI'.
- + Accurate 45-degree isometric perspective.
- + Clean, minimalist interpretation of the cartoon scene.
- − The flag icon is abstract and does not represent the Japanese flag well.
- − Very simplistic 3D modeling compared to the level of detail requested in 'refined textures'.
- − The diorama base is extremely basic.
Verdict: FLUX.1 Kontext [dev] is the winner because it successfully incorporated the specific text requirements, which DALL-E 3 completely ignored. While DALL-E 3 produced a much more beautiful and detailed isometric scene, its total failure on prompt adherence regarding text makes it less successful for this specific challenge.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 3
- + Excellent adherence to the list of specific animals (puppy, kitten, bunny, fox).
- + Beautiful lighting effects including god rays and dew sparkles as requested.
- + Very high level of detail in the fur texture and background elements.
- − The butterflies have weirdly mammalian 'hamster-like' heads which is anatomically confusing.
- − Style leans more towards digital illustration/CGI than 'hyper-photorealistic'.
FLUX.1 Kontext [dev]
- + Natural, realistic lighting and photographic depth of field.
- + Dynamic poses that accurately capture the 'playfully chasing' and 'tumbling' prompt requirements.
- + Anatomically correct butterflies and more realistic fur textures.
- − Failed to include all requested animals; missing the bunny and the red fox kit.
- − The three animals included look very similar in color and style, reducing the variety requested.
Verdict: DALL-E 3 followed the complex prompt much better by including every specific animal requested, though it produced strange hybrid insect-mammal butterflies and a more 'artificial' look. FLUX.1 Kontext [dev] captured the photographic 'hyper-photorealistic' aesthetic and motion perfectly, but failed significantly on prompt adherence by omitting the bunny and fox. DALL-E 3 is the winner for successfully managing the dense subject list.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 3
- + Excellent execution of the 'vintage vector emblem' style with complex textures
- + Rich, warm color palette that fits the requested theme perfectly
- + High-quality typography and balanced composition
- − Failed to include the specific name 'Caffè Florian', replacing it with generic text
- − The banner does not contain the 'Est. 1720' text as requested, placing it at the bottom instead
FLUX.1 Kontext [dev]
- + Perfect adherence to the specific text requested including 'Caffè Florian'
- + Clean, minimalist design that follows the prompt's layout requirements
- + Accurate depiction of a styled cloche dome and steam icon
- − Lack of the 'banner' element specified in the prompt
- − Almost zero texture, appearing more like a modern flat icon than a vintage emblem
- − The font choice is a bit too modern for a 'vintage' and 'classic' request
Verdict: DALL-E 3 produced a much more visually appealing and stylistically accurate vintage logo, but completely failed to follow the specific text instructions. FLUX.1 Kontext [dev] followed the text requirements perfectly, but the resulting image is overly simplistic and lacks the requested texture and banner element. FLUX.1 is the winner for following the core identity of the prompt, as a logo without the correct name is non-functional.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 3
- + Excellent adherence to the NASA-inspired color palette.
- + Sophisticated, professional layout that feels like a real museum-quality infographic.
- + Highly detailed icons and complex vector-style shading.
- − Generated three variants in one image instead of a single poster.
- − Incorrectly includes Space Shuttle-style icons which are anachronistic to Apollo 11.
- − Text is largely illegible filler content.
FLUX.1 Kontext [dev]
- + Follows the specific 6-step structure requested in the prompt more clearly.
- + Minimalist flat-vector style is consistent throughout the icons.
- + Correctly attempts to label the steps as 'Launch' and 'Descent'.
- − Significant spelling errors ('APOLO', 'DESCNT', 'TRANGULIITY').
- − Text rendering is messy with overlapping characters and gibberish strings.
- − Composition is quite basic and lacks the 'modern infographic' aesthetic.
Verdict: DALL-E 3 produced a much more visually impressive and polished design that perfectly matches the requested color palette, though it mistakenly included Space Shuttle iconography. FLUX.1 Kontext [dev] followed the structural logic of the 6 steps more accurately but failed significantly on text rendering and overall graphic design quality. DALL-E 3 is the preferred choice for its professional aesthetic and adherence to the stylistic 'NASA' vibe.
Explore each model
Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities