Black Forest Labs' enhanced 12-billion parameter flow transformer with 6x faster generation than FLUX.1 [pro], delivering superior composition, detail, and artistic fidelity
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX1.1 [pro]
#50 of 62 in Text-to-Image
GPT Image 1
#32 of 62 in Text-to-Image
Where the votes landed
FLUX1.1 [pro]
0%
win rate
Ties
0%
GPT Image 1
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX1.1 [pro]
- + Exquisite rendering of glass refractions and reflections
- + Very high textural detail on the book and table
- + Atmospheric handling of background plant through glass
- − The glass container is a tall rectangular prism rather than a cube
- − The blue sphere appears to be floating in the center rather than resting on the bottom
GPT Image 1
- + The glass container is shaped much closer to a true cube
- + The blue sphere is resting naturally on the surface
- + Excellent spatial composition and lighting
- − The glass thickness is a bit inconsistent at the seams
- − Slightly less 'photorealistic' texture on the blue sphere compared to Model A
Verdict: Both models followed the complex spatial instructions well. FLUX1.1 [pro] produced a more visually stunning image with superior glass physics and lighting, but GPT Image 1 followed the 'cube' instruction more accurately while maintaining a naturalistic placement of the sphere.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent handling of wet pavement reflections and rainy atmosphere.
- + Captures the requested motion blur from passing cars well.
- + Realistic skin texture and age-appropriate lighting.
- − The bicycle geometry is broken, with a missing rear wheel and frame segment behind the man.
- − The man's posture is slightly awkward, hovering above the ground rather than standing or kneeling.
GPT Image 1
- + Stronger adherence to the action of 'repairing' the bicycle.
- + Superior skin texture and fine detail in the man's face and hands.
- + Better preservation of the bicycle's structural logic.
- − Missing the requested motion blur from passing cars.
- − Fails to show reflections on the pavement as clearly as the other model.
Verdict: FLUX1.1 [pro] better captures the environmental requirements like motion blur and wet pavement reflections, but suffers from significant structural errors in the bicycle and the man's anatomy. GPT Image 1 provides a much more intimate, realistic, and coherent subject with better skin textures and a logical pose, despite missing the motion blur. GPT Image 1 is preferred for its overall realism and lack of distracting artifacts.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX1.1 [pro]
- + Extremely high skin texture detail and realistic eyes.
- + Excellent use of shallow depth of field for a modern cinematic look.
- + Dynamic lighting with vibrant bokeh sparks.
- − Braided hair with beads is less prominent than requested.
- − Framing is very tight, cutting off much of the ornate armor details.
GPT Image 1
- + Excellent depiction of ornate engraved plate armor as requested.
- + Strong adherence to the braid and bead detail in the hair.
- + Warm torchlight atmospheric lighting is very convincing.
- − Skin textures and eyes look slightly more painterly/artificial compared to Model A.
- − The bokeh sparks are less integrated into the depth of the scene.
Verdict: While FLUX1.1 [pro] offers superior photorealism in the skin and eyes, GPT Image 1 follows the prompt's armor and hair requirements much more effectively. GPT Image 1 captures the 'ornate engraved plate' and 'braids with beads' more clearly, making it the better choice for this specific character design.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent presentation of the menu in a realistic lifestyle setting
- + Includes a high number of food images in a grid as requested
- + Sophisticated use of white space and typography hierarchy
- − The text is mostly illegible gibberish
- − The 'grid' of photos is somewhat cluttered and lacks clear alignment with the text sections
GPT Image 1
- + Highly legible bold sans-serif text
- + Perfectly clean and balanced grid layout
- + Clear categorization with vibrant color-coded accents
- − Minimalist to a fault, feeling slightly like a template rather than a finished design
- − Contains minor spelling errors like 'descrigion'
Verdict: GPT Image 1 followed the prompt's layout requirements much more effectively, producing a clean, legible, and professional grid design that feels like a real menu. While FLUX1.1 [pro] created a more artistic lifestyle shot, its text is completely unreadable and the food grid is less organized. GPT Image 1 is the clear winner for its functional design and adherence to the layout instructions.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photorealistic texture on the meat and bun
- + Clean typography and professional layout design
- + Realistic lighting integration between the burger and the background flames
- − Failed the 'exploded' instruction as the layers are stacked together
- − One instances of neon blue text contradicts the 'fiery glowing' prompt
- − Missing the requested starburst for the price
GPT Image 1
- + Followed the exploded view instruction perfectly with clear separation of ingredients
- + Strict adherence to the fiery glowing text aesthetic across all elements
- + Included the requested price starburst with the correct currency symbol
- − The price text is incorrect, showing €0.99 instead of €6.99
- − Overall image clarity and texture detail is slightly lower than the competitor
- − Composition is a bit crowded at the top and bottom borders
Verdict: GPT Image 1 followed the complex prompt instructions much more accurately, successfully creating an exploded view and consistent fiery text effects, even including the requested starburst. FLUX 1.1 [pro] produced a more professional-looking commercial image with superior textures, but it failed the primary 'exploded' composition and missed several specific text style requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic depth of field and background bokeh
- + Authentic cursive style for the title
- − Failed the prompt instructions by splitting one menu item into two lines with incorrect text and enormous prices ($228, $98)
- − Included spelling errors like 'Chipkies' and 'Herbss'
- − Text looks like a digital overlay rather than physical chalk grit
GPT Image 1
- + Perfect adherence to the menu list and prices requested
- + Superior chalk texture that looks like real dust on a board
- + Correct spelling for all items including the cut-off 'Cookies'
- − The handwriting is somewhat uniform and lacks the 'elegant cursive' requested for the title
- − The lighting is flat compared to the atmospheric lighting in Model A
Verdict: While FLUX1.1 [pro] creates a more visually pleasing and atmospheric cafe scene, it fails significantly on the prompt's specific text requirements, including hallucinations of prices and incorrect item names. GPT Image 1 followed every instruction perfectly, capturing the specific chalk texture and the exact menu items requested without error.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic lighting and dynamic composition.
- + Highly detailed textures on the spacesuit and horse's mane.
- + Vibrant and surreal atmosphere with beautiful celestial clouds.
- − Completely failed the negative constraint to have the horse on top.
- − Anatomical issues with the horse's front legs merging into its chest.
- − The horse appears to be Wearing an astronaut suit pieces in a nonsensical way.
GPT Image 1
- + Strong image clarity and clean rendering.
- + Good spatial composition with the curve of the planet.
- + Consistent horse anatomy and realistic fur texture.
- − Completely failed the negative constraint to have the horse on top.
- − Less creative interpretation of the 'surreal' aspect compared to Model A.
- − A more standard, less cinematic color palette.
Verdict: Both FLUX1.1 [pro] and GPT Image 1 completely failed the specific spatial instruction to place the horse on top of the astronaut, both defaulting to the common trope of an astronaut riding a horse. FLUX1.1 [pro] is the preferred choice as it leans much more effectively into the 'surreal' and 'cinematic' keywords with its dramatic lighting and complex background elements, whereas GPT Image 1 feels flat and conventional.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic lighting and depth of field
- + High level of texture detail in the capybara's fur
- + Good text rendering on the hat
- − Compositional error puts both characters in the front or middle seat area
- − Failed to place the capybara's paws on the steering wheel correctly
- − The capybara is looking at the passenger rather than driving professionally
GPT Image 1
- + Perfect adherence to the requested composition with the capybara driving and passenger in the back
- + Accurately depicts both front paws on the steering wheel
- + Captures the bored, 'normal' expression of the businesswoman well
- − Anatomical issues with the capybara's paws looking like monkey hands
- − The capybara's face is slightly less detailed than in the competitor
- − Strange placement of the taxi light on top of the car, visible through the windshield
Verdict: GPT Image 1 followed the complex spatial instructions much better than FLUX1.1 [pro], correctly placing the capybara at the wheel and the passenger in the back seat. FLUX1.1 [pro] produced a more visually striking and detailed image, but failed on basic prompt requirements like the steering wheel and seat positioning.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Features cinematic lighting with high contrast and a vibrant glowing jack-o-lantern.
- + Detailed thorn border and atmospheric elements like the trees and moon.
- + Good use of vertical space despite the square aspect ratio.
- − Significant text errors including redundant lines and typos like 'univted'.
- − Information at the bottom is cluttered and repetitive.
- − Missed the 'Invitation' part of the title on the banner.
GPT Image 1
- + Excellent text legibility and accuracy with no spelling errors.
- + Clearly captures the vintage parchment texture requested.
- + Follows the layout instructions perfectly, including the scroll banner and clear footer details.
- − The lighting is a bit flat compared to the other model.
- − The aesthetic leans more modern-digital-vintage than truly gothic cinematic.
Verdict: GPT Image 1 is the clear winner because it correctly renders all of the requested text without spelling errors or repetitions, whereas FLUX1.1 [pro] fails significantly on the typography. GPT Image 1 also captures the requested vintage parchment feel more effectively, providing a polished and functional invitation design.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent variety of sushi types within the scene
- + Accurate 45-degree isometric perspective
- + High clarity with soft, appealing lighting
- − Missed the small flag icon requested in the prompt
- − Text is somewhat stylized rather than simply 'bold' as requested
GPT Image 1
- + Includes all elements including the small flag icon
- + Very clean, minimal composition that aligns well with the 'miniature' request
- + Perfect execution of the text hierarchy and bold font
- − Texture on the rice looks a bit more like clay than 'refined texture'
- − Background color is slightly darker than the requested 'light blue'
Verdict: GPT Image 1 followed the instructions more comprehensively by including every requested element, including the flag icon and the specific text layout. While FLUX1.1 [pro] offered more variety in the sushi models, GPT Image 1 captured the clean, miniature diorama aesthetic more effectively with better adherence to the background and layout constraints.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI judge analysis unavailable for this challenge.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX1.1 [pro]
- + Sophisticated vector illustration style with high artistic detail.
- + Excellent use of warm brown and cream tones with subtle texturing.
- + Includes a complex, ornate banner that fits the vintage aesthetic.
- − Spelling error in the main text rendering it as 'Caffé Franilian'.
- − The dome structure looks more like a building's cupola than a restaurant cloche dome.
GPT Image 1
- + Perfect spelling of 'Caffè Florian' and 'Est. 1720'.
- + Accurately represents a cloche dome as requested in the prompt.
- + Strong minimalist vector execution with a clear, balanced layout.
- − Failed the requirement for a light background, providing a black one instead.
- − The texture is very subtle and less noticeable than in the competing model.
Verdict: FLUX1.1 [pro] produced a beautiful, high-quality illustration but failed on basic text accuracy and interpreted the cloche dome as architectural. GPT Image 1 adhered much better to the specific 'minimalist' and 'cloche dome' keywords with perfect spelling, although it ignored the request for a light background. GPT Image 1 is the winner for functional logo use due to correct spelling and prompt adherence.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX1.1 [pro]
- + Clean aesthetic with a professional layout and hierarchy.
- + Effective use of negative space and a wide color range from the NASA palette.
- + Captures the sense of a complete infographic poster.
- − Text is largely illegible gibberish after the main title.
- − The numbering and sequence of steps are logically broken and repetitive.
- − Iconography is inconsistent and does not clearly match the specific requested items.
GPT Image 1
- + Excellent adherence to specific icon requests like the lunar module and trajectory arc.
- + Highly legible typography with correct spelling of mission names (mostly).
- + Strict adherence to the flat-vector style and NASA-inspired palette.
- − Composition feels more like an icon set than a cohesive infographic poster.
- − One spelling error with 'EARLLUNAR'.
- − The layout lacks a clear flow or connective lines between the steps.
Verdict: GPT Image 1 is the superior choice for its high level of adherence to the specific iconography and labels requested in the prompt, despite a slightly disjointed layout. Conversely, FLUX1.1 [pro] creates a more visually appealing poster at a glance, but fails completely on informational accuracy, featuring repetitive numbers, broken logic, and nonsensical text.
Explore each model
OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs