Black Forest Labs' 12-billion parameter flow transformer for high-quality text-to-image generation, suitable for personal and commercial use with streaming support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [dev]
#16 of 62 in Text-to-Image
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [dev]
0%
win rate
Ties
0%
GPT Image 2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photorealism with convincing caustics and light refraction.
- + The sphere has a beautiful glass-like quality that fits the scene's aesthetic.
- + High-quality soft lighting that matches the requested window light direction.
- − The 'glass cube' appears more like a solid block of glass or a display case with thick pillars rather than a simple hollow cube.
- − The perspective of the blue sphere makes it look like it's floating or embedded rather than sitting on the bottom.
GPT Image 2
- + Perfect adherence to the object 'hollow glass cube' geometry.
- + Excellent spatial arrangement with all objects clearly defined and correctly placed.
- + High texture detail on the book cover and wooden table grain.
- − The blue sphere's texture is a bit flat/matte compared to the high-gloss environment.
- − The reflection of the sphere on the bottom panel of the cube is slightly simplified.
Verdict: Both models followed the prompt perfectly, but GPT Image 2 is the winner for its superior structural accuracy, rendering a clear, hollow glass cube that feels more physically grounded. While FLUX.1 [dev] produced a more artistically 'expensive' looking image with beautiful light play, its interpretation of the cube was more abstract and solid, making it less representative of the literal prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent handling of the rainy atmosphere and pavement reflections.
- + Accurately depicts the requested motion blur of background traffic.
- + Cinematic lighting with a clear 50mm shallow depth of field effect.
- − The subject is standing and holding the bike rather than actively 'repairing' it.
- − The bicycle's front wheel alignment and handlebars appear physically distorted.
GPT Image 2
- + Better adherence to the 'repairing' action, showing the man squatting with a toolbox.
- + The 'imperfect framing' prompt is well-interpreted with the foreground post and sign.
- + Contains legible Japanese text which adds to the requested realism and location.
- − Fails to depict the 'light rain' requested in the prompt, appearing mostly dry.
- − Motion blur on the background car is less pronounced than requested.
Verdict: FLUX.1 [dev] produced a much more atmospheric and cinematic image that perfectly captured the rain and lighting requested, though the subject was not actively repairing the bike. GPT Image 2 captured the specific action and local Japanese details much better, but failed to include the rainy weather and motion blur. FLUX.1 [dev] is the winner for its superior visual quality and adherence to the technical photography style requested.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photographic clarity and lighting
- + Captures the bokeh sparks requested in the prompt
- + High-quality skin texture and lifelike eyes
- − Fails to include the 'beads' in the hair
- − Missing the requested 'battle-worn' scars and dirt
- − The subject looks too clean and polished for the prompt requirements
GPT Image 2
- + Successfully includes small beads in the braided hair
- + Accurately depicts the 'battle-worn' state with dirt and faint scars
- + Ornate engravings on the plate armor are extremely detailed and clearly visible
- − The torchlight is slightly less dramatic than in the other version
- − Armor composition is a bit cluttered around the neck
Verdict: GPT Image 2 is the superior choice as it adhered to nearly every specific prompt detail, including the beads in the hair and the battle-worn skin textures that FLUX.1 [dev] omitted. While FLUX.1 [dev] produced a stunning, clean portrait, it lacked the narrative depth and specific character elements like the ornate engraving and scars requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean, professional minimalist layout with excellent use of negative space.
- + Accurate high-quality rendering of food assets with realistic soft lighting.
- + Successful execution of the modern airy aesthetic.
- − Text is mostly gibberish or placeholder characters.
- − Missing a dedicated 'Pizza' section as requested in the prompt.
- − Food images are not in a grid as specified.
GPT Image 2
- + Perfect adherence to all prompt elements including grid layout and specific sections for pizza/mains.
- + Fully legible and relevant text for all menu items and descriptions.
- + Excellent use of vibrant accents and bold sans-serif fonts.
- − Slightly crowded composition compared to a true minimalist style.
- − Some minor artifacts on small icons in the footer.
Verdict: GPT Image 2 is the superior choice because it fully delivers on the logical requirements of the prompt, providing distinct sections for Appetizers, Pizza, and Mains with legible English text. While FLUX.1 [dev] captures a more authentic 'high-end minimalist' aesthetic, it fails to include the requested grid layout and produces nonsensical text, making it less functional as a menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean layout with clear separation of burger components.
- + Accurate rendering of basic price and secondary text.
- − Failed to include the main 'MAGIC BURGER' title text.
- − Ignored the 'fiery' text effect and 'starburst' requirement for the price.
- − Lacks the sense of motion and dynamic energy requested in the prompt.
GPT Image 2
- + Perfect adherence to text requirements including 'MAGIC BURGER', secondary message, and price in a starburst.
- + Excellent fiery text effects and dynamic 'exploded' composition with sauce splashes.
- + Highly detailed textures on the meat patty, vegetables, and bun.
- − The composition is a bit crowded with several overlapping elements.
Verdict: GPT Image 2 followed every instruction in the prompt, including the specific text content and the complex fiery visual effects. While FLUX.1 [dev] produced a clean image, it failed to include the primary brand name and ignored all the requested text styling like the starburst and ember glow.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent text legibility and accuracy.
- + Good center alignment and overall layout cleanlineess.
- − The text looks more like a digital marker or vector font than actual chalk on a board.
- − Several spelling errors in the bottom two lines such as 'drish' and 'ufor aur'.
GPT Image 2
- + Perfectly depicts the requested chalk texture with realistic grain and pressure variations.
- + Followed all text prompts correctly with no spelling errors.
- + The handwriting style is much more authentic to a real cafe chalkboard.
- − The slant of the text is slightly inconsistent across lines, though this fits the 'natural variation' prompt.
Verdict: While both models followed the text instructions well, GPT Image 2 is the clear winner because it perfectly captured the requested chalk texture and handwriting style. FLUX.1 [dev] produced text that looks too much like a digital font and suffered from several typos in the footer text.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean, cinematic lighting and color grading
- + High visual quality with smooth textures
- + Good spatial background depicting the curvature of the Earth
- − Failed to follow the core instruction of the horse being on top
- − Anatomical issues with the horse's legs appearing rubbery and lacks joints
GPT Image 2
- + Perfect adherence to the complex logic of the prompt with a horse on top of an astronaut
- + Excellent textures on the spacesuit and lunar surface
- + Cleverly includes a NASA logo to add to the surreal realism
- − The horse's front legs merging into stirrups is slightly confusing but fits the surreal theme
Verdict: While FLUX.1 [dev] produced a high-quality cinematic image, it completely failed to follow the specific spatial instruction to place the horse on top. GPT Image 2 successfully followed the difficult 'horse riding astronaut' prompt with high detail and a perfectly surrreal execution.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent sharpness and texture on the capybara's fur
- + Accurately places the passenger in the seat beside or just behind the driver with a clear phone in hand
- − The passenger is sitting in the front passenger seat rather than the requested 'back seat'
- − Includes a 'TAXI' sign floating strangely inside the top of the car's interior
GPT Image 2
- + Perfectly depicts the businesswoman in the back seat as requested
- + Highly realistic textures on the jacket and the capybara's paws on the steering wheel
- + Excellent composition that feels more cinematic and professionally framed
- − The passenger is slightly out of focus, though this fits the depth of field
- − The capybara's hat is a bit oversized
Verdict: GPT Image 2 is the superior image because it accurately followed the instruction to place the passenger in the back seat, whereas FLUX.1 [dev] placed her in the front. GPT Image 2 also achieved a more realistic lighting and cinematic feel that closely aligns with the 'photorealistic' requirement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean vector-style illustration with high clarity
- + Bright, cinematic lighting on the jack-o-lantern
- + Composition is balanced and easy to read
- − Several text errors including 'Falloween Rantcy' and 'You Tre'
- − Fails to meet the 'vintage gothic' aesthetic, looking more like a modern digital sticker
- − Missing the requested scroll banner for the invitation text
GPT Image 2
- + Perfect adherence to the 'vintage gothic' and 'parchment' aesthetic
- + Flawless text rendering for both titles and details
- + Highly detailed background containing the requested bridge ('The Arches') and NYC skyline
- − The jack-o-lantern is slightly smaller in the frame compared to model a
- − The thorny border is very dense, which might feel cluttered to some
Verdict: GPT Image 2 followed the prompt's stylistic and textual instructions perfectly, delivering a sophisticated vintage aesthetic with 100% accurate text. In contrast, FLUX.1 [dev] produced a modern, simplified illustration with multiple typos and failed to capture the 'gothic parchment' feel.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent 45° isometric perspective and diorama base
- + High-quality soft lighting and refined textures
- + Accurate miniature toy aesthetic
- − Text rendering is poor with 'SUSH CATON' and a random kanji character
- − Sushi variety is limited to six identical pieces
- − The flag icon is stylized as an oval rather than a standard flag
GPT Image 2
- + Perfect text rendering for 'JAPAN' and 'SUSHI'
- + Highly detailed and diverse sushi selection with realistic PBR materials
- + Stronger composition with more visual depth
- − Slightly ignores the 'minimal' garnish instruction
- − Text placement is very close to the top edge
- − Perspective is slightly flatter than the requested 45° top-down
Verdict: GPT Image 2 is the clear winner as it perfectly rendered the requested text and flag icon, which FLUX.1 [dev] failed to do. While GPT Image 2 included more garnish than requested, its variety of sushi and superior adherence to the branding elements make it much more successful for the intended prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [dev]
- + Features a very clean and symmetrical composition.
- + Captures the 'wholesome' and 'soft' aesthetic with high polish.
- − Failed the realism aspect, resulting in a 3D-animation/Pixar style rather than hyper-photorealistic.
- − Missing the tabby kitten entirely, showing two dogs and two red-fox/rabbit hybrids instead.
- − Animals are standing still and posing rather than 'playfully chasing' or 'tumbling'.
GPT Image 2
- + Excellent adherence to the 'hyper-photorealistic' instruction with realistic fur textures and anatomy.
- + Accurately includes all four requested animal types: golden retriever, tabby kitten, bunny, and fox kit.
- + Dynamic composition that truly shows the animals 'playfully chasing' and 'tumbling'.
- − The fox's front right paw has some anatomical warping/blurring.
- − Some butterflies in the background are slightly distorted or simplified.
Verdict: GPT Image 2 followed the prompt with significantly higher accuracy, successfully including all four specific animal types in a realistic style. In contrast, FLUX.1 [dev] produced a non-photorealistic, stylized cartoon image and failed to include the tabby kitten, instead repeating species.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [dev]
- + Clean vector-style execution
- + Correct inclusion of all requested text elements
- − Major spelling errors in the primary name (Flariláan) and category (Reseaurant)
- − Random numbers '11011' and '1941' added without prompt instruction
GPT Image 2
- + Perfect text rendering for name and established date
- + Excellent vintage texture and cross-hatching detail
- + Stronger composition with a professional frame
- − Includes excessive scrollwork not explicitly mentioned, though it fits the theme
Verdict: GPT Image 2 is the clear winner as it correctly spells 'Caffè Florian' and 'Est. 1720' while capturing a superior vintage aesthetic with authentic textures. FLUX.1 [dev] fails significantly on typography, introducing several distracting misspellings and random numeric artifacts.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [dev]
- + Features a very clean, minimalist flat-vector aesthetic that matches the 'modern vector' prompt.
- + Utilizes an elegant circular flow layout for the infographic elements.
- − Text consists of garbled, nonsensical characters.
- − Fails to follow the specific requested steps and icons accurately, including several planets that aren't relevant.
GPT Image 2
- + Excellent adherence to all six requested steps with accurate iconography.
- + Superior text rendering with legible and correct names for the mission phases and crew.
- + Perfectly captures the NASA-inspired color palette and high-quality vector style.
- − The transition to 3D craters on the bottom moon surface slightly clashes with the flat-vector style of the icons above.
Verdict: GPT Image 2 is the clear winner for its exceptional prompt adherence, correctly depicting all six requested mission stages with legible text and icons. FLUX.1 [dev] produced a visually pleasing abstract layout, but failed to follow the logical steps and produced entirely illegible text.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following