Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
GPT Image 1
#32 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0.0%
win rate
Ties
0.0%
GPT Image 1
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light and caustic reflections on the wooden surface.
- + Hyper-realistic glass textures including subtle distortions.
- + The blue sphere has a unique, tactile texture that reacts to the lighting.
- − The plant is quite far in the background and less 'behind' the cube compared to the other model.
- − The text on the book spine is gibberish.
GPT Image 1
- + Strong composition with the plant clearly visible through the glass as requested.
- + Very clean, minimal aesthetic with high clarity.
- + Correct placement and scale of all requested objects.
- − The lighting is very flat and lacks the 'left window' directional quality of the other model.
- − The glass cube appears more like a wireframe with glass panels rather than a solid glass object.
Verdict: FLUX.1 Kontext [max] produces a much more realistic image with sophisticated lighting and caustic effects that ground the objects in the scene. While GPT Image 1 follows the spatial instructions perfectly (specifically looking through the glass at the plant), it suffers from flatter lighting and a less convincing glass material.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of rain with visible streaks and splashing on the ground
- + Highly realistic reflections and wet road texture
- + Anatomically correct hands and bicycle mechanics
- − The subject's ethnicity is somewhat ambiguous
- − Rain streaks are a bit uniform across the frame
GPT Image 1
- + Perfect adherence to the 'Japanese man' ethnic description
- + Cinematic color grading and authentic street atmosphere
- + Good application of shallow depth of field
- − Minimal visual evidence of rain despite the wet pavement
- − The mechanical details of the bicycle's rear hub are slightly jumbled
- − Lacks the requested motion blur from passing cars
Verdict: FLUX.1 Kontext [max] produces a technically superior image with highly realistic environmental effects like rain splashes and reflections, although the subject's ethnicity is less specific. GPT Image 1 captures the character and cinematic mood better but fails to include the requested rain streaks and motion blur. FLUX.1 Kontext [max] is the winner for its superior prompt adherence regarding the weather and physical environment.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Intricate engraving on the plate armor with excellent metallic reflections
- + High-detail skin texture and intense, glowing eyes
- + Beautiful bokeh sparks adding to the atmospheric depth
- − Missed the request for beads in the hair braids
- − Skin looks a bit overly saturated/reddish
GPT Image 1
- + Excellent adherence to the 'battle-worn' prompt with realistic dirt and faint scars
- + Successfully included small beads within the braided hair
- + Highly detailed engraving texture on the armor
- − Lighting is slightly flatter compared to Model A
- − The eyes are less striking than those in the first image
Verdict: Both models performed exceptionally well on this complex prompt. FLUX.1 Kontext [max] delivered a more cinematic image with superior lighting and ocular detail, while GPT Image 1 followed the specific structural details of the prompt better, particularly regarding the beads in the braids and the battle-worn facial textures. GPT Image 1 is slightly preferred for its comprehensive prompt adherence.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional flyer layout with a clear header and branding area.
- + Good grid organization and usage of white space.
- + Includes a wide variety of food images consistent with a full menu.
- − Text is mostly gibberish or illegible at smaller scales.
- − The font choice for headers is scripted, contrary to the 'sans-serif' request.
- − Almost all food items shown are pizzas, lacking diversity in dish types.
GPT Image 1
- + Perfectly legible English text for headers and prices.
- + High-quality, distinct food photos representing different categories like salad, pizza, and mains.
- + Adheres strictly to the sans-serif font and vibrant accent request.
- − The layout is zoomed in, showing only a portion of a menu rather than a full page.
- − The repetitive 'Apperoiation descrigion' text is a minor spelling artifact.
Verdict: GPT Image 1 is the superior choice because it delivers legible, professional typography and diverse food photography that matches the 'appetizers/pizza/mains' categories perfectly. While FLUX.1 Kontext [max] has a better overall page composition, its text is completely illegible and it failed to use the requested sans-serif font for headers.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and clean font choices
- + High-resolution textures on the meat and bun
- + Successfully integrated the euro symbol and price correctly
- − The burger is not truly 'exploded' as most components are still stacked in the center
- − Failed to include a starburst for the price
GPT Image 1
- + Perfect interpretation of the 'exploded' request with clear vertical separation of layers
- + Beautiful fiery glow effect on all text elements
- + Included the starburst element requested for the price
- − The price text is incorrect, displaying '.99' instead of '6.99'
- − The 'MAGIC BURGER' title is partially cropped at the top
Verdict: GPT Image 1 followed the composition instructions much better, delivering a true 'exploded' view and including the requested starburst. However, FLUX.1 Kontext [max] produced more accurate text and a more realistic, detailed patty, whereas GPT Image 1 failed on the specific price value and layout cropping.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text rendering with no spelling errors.
- + The chalk handwriting looks very natural with realistic smudges on the board.
- + Strong environmental context with a visible café background and board frame.
- − The title is in print blocks rather than the requested 'elegant cursive' style.
- − The handwriting looks slightly more like a whiteboard marker in some strokes than dry chalk.
GPT Image 1
- + Successfully captured the grainy, textured feel of real chalk strokes.
- + The text is very legible and well-centered.
- + Followed the layout requirements for the price placements and menu items.
- − Missed the request for 'elegant cursive' handwriting for the title.
- − The text style looks slightly more like a digital chalk font rather than organic hand-drawn letters.
- − Formatting error on the final price: 'Cookies 9' instead of '$9'.
Verdict: Both models struggle with the specific 'cursive' instruction for the title, both opting for print-style lettering instead. FLUX.1 Kontext [max] provides a much more realistic image overall, with natural handwriting variations and a convincing café environment. GPT Image 1 feels flatter and more like a digital graphic, and it fails to include the dollar sign on the final item price.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent lighting and cinematic color grading.
- + Crispy sharp details on the horse's mane and the space suit fabric.
- + Realistic interaction between the astronaut's hands and the reins.
- − The position of the astronaut on the horse is traditional, missing the 'surreal' potential of the prompt.
GPT Image 1
- + Strong atmospheric mood with a painterly, classical feel.
- + Good composition with the planet curvature in the bottom corner.
- + Accurate rendering of a professional space suit.
- − The horse's hind legs have a strange, elongated anatomical distortion.
- − The reins appear to pass through the horse's neck rather than around it.
Verdict: Both models struggled with the negative constraint 'horse on top, not vice versa,' opting instead for the standard astronaut-on-horse interpretation. FLUX.1 Kontext [max] is the winner due to significantly better limb anatomy and superior technical image quality whereas GPT Image 1 suffered from anatomical glitches and clipping issues.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the capybara fur
- + High clarity and vibrant colors in the city background
- + Accurate depiction of a New York taxi exterior and interior
- − The passenger is sitting in the front passenger seat, not the back seat as requested
- − Only one paw is visible on the steering wheel
GPT Image 1
- + Perfect adherence to the 'back seat' positioning for the passenger
- + Both front paws are clearly placed on the steering wheel
- + Captures the requested 'bored' expression on the businesswoman perfectly
- − Lighting is a bit moody and dark, obscuring some details
- − The capybara's head shape is slightly less natural than Model A
Verdict: While FLUX.1 Kontext [max] has superior lighting and fur texture, GPT Image 1 followed the complex spatial instructions much better. GPT Image 1 correctly placed the passenger in the back seat and showed both paws on the wheel, making it a more accurate realization of the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography style that matches the 'gothic' request perfectly.
- + Intricate border details with clearly defined webs and thorns.
- + High-contrast cinematic lighting on the jack-o-lantern.
- − Redundant text at the bottom repeating 'The Arches, NYC' twice.
- − The date uses commas '30,10,2026' instead of the requested dots.
GPT Image 1
- + Perfect text accuracy including the dots in the date.
- + Clean and legible layout that feels like a professional poster.
- + Includes the requested scroll banner more effectively than the competitor.
- − The font is a standard serif rather than 'elegant gothic' as requested.
- − The lighting is flatter and less 'cinematic' than Model A.
Verdict: FLUX.1 Kontext [max] captured the artistic gothic aesthetic much better with its choice of font and mood, although it suffered from slight text redundancy and a minor punctuation error. GPT Image 1 followed the technical text instructions perfectly but failed to deliver the specific 'gothic' elegant typography and moody atmosphere requested in the prompt.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent preservation of the original facial features and bone structure.
- + Realistic hair texture and color that matches the existing beard.
- + Maintains the original lighting and background with high fidelity.
- − The hairline on the forehead is slightly too straight and clean for a completely natural look.
GPT Image 1
- + Successfully adds a thick head of hair with significant volume.
- + Preserves the overall composition and background of the original image.
- − Significantly alters the person's face, making them look like a different individual.
- − The hair looks somewhat like a wig or an overlay rather than growing naturally from the scalp.
- − Loss of specific skin details and eye characteristics from the source.
Verdict: FLUX.1 Kontext [max] performed significantly better as it successfully added the hair while keeping the person's face identical to the source image. In contrast, GPT Image 1 failed to preserve the subject's identity, resulting in a face that only roughly resembles the original and a hair texture that feels less integrated with the existing beard.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text rendering with a stylized 3D bubble effect
- + High-quality material rendering on the fish textures
- + Clear adherence to the isometric perspective
- − Missed the small flag icon request
- − Chopsticks are slightly merged into the base
GPT Image 1
- + Successfully included the Japanese flag icon as requested
- + Clean and diverse variety of sushi items including nigiri and a maki roll
- + Subtle, soft textures that fit the 3D cartoon aesthetic well
- − Typography is a bit flatter compared to the 3D style of Model A
- − The base has a slightly grainy texture compared to the smooth background
Verdict: Both models followed the technical instructions for an isometric diorama very well. GPT Image 1 is slightly better for prompt adherence as it included the requested flag icon, while FLUX.1 Kontext [max] produced more impressive 3D typography and more realistic material shaders on the sushi itself.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent preservation of clothing details like the denim shirt and black undershirt
- + Clear representation of the news anchor setting with microphone and desk
- + Clean, modern cartoon aesthetic
- − The hockey element is minimal and poorly integrated (just a floating stick blade)
- − Added glasses which were not in the source image
- − Caricature style is a bit generic and less 'exaggerated' in the facial features
GPT Image 1
- + Strong caricature style with exaggerated facial features that still resemble the source woman
- + Very creative integration of all elements, including a dog playing hockey on the TV screen
- + Matches the traditional hand-drawn watercolor caricature style perfectly
- − The denim shirt loses some of the specific button and seam details from the source
- − The desk perspective is slightly warped
Verdict: GPT Image 1 is the clear winner for its creative and humorous integration of all prompt elements, specifically the clever way it combined the hockey and dog themes on the background monitor. While FLUX.1 Kontext [max] preserved the source clothing better, it failed to meaningfully incorporate the hockey theme and added unnecessary accessories like glasses.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Features a very vibrant, lush meadow with high floral density
- + Excellent rendering of translucent light in the bunny's ears
- + Soft, dreamlike aesthetic that matches the 'wholesome' part of the prompt
- − The characters are mostly sitting rather than 'playfully chasing' and 'tumbling'
- − The anatomy of the bunny's paws is slightly unclear
- − The 'god rays' appear as bright sparkles rather than defined beams
GPT Image 1
- + Perfectly captures the 'playfully chasing' and 'tumbling' motion requested
- + Features distinct, well-defined god rays from the sunrise
- + Excellent fur texture and realistic anatomy for all four baby animals
- − The meadow is less dense with wildflowers compared to Image A
- − The lighting on the kitten's face is a bit flat compared to the puppy and fox
Verdict: GPT Image 1 is the superior image because it accurately depicts the dynamic action of the animals 'chasing' and 'tumbling' as requested in the prompt, whereas FLUX.1 Kontext [max] shows them in a relatively static pose. GPT Image 1 also features more realistic anatomy and better execution of the 'god rays' lighting effect.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent preservation of the source image's composition and character details.
- + Successfully captures the specific Studio Ghibli cel-shaded character style.
- + The plaid pattern on the shirt and the background structures are maintained with high fidelity.
- − The eyes on the central male character feel a bit too 'modern anime' compared to Ghibli's specific style.
GPT Image 1
- + Beautiful use of warm, nostalgic, and dreamy lighting as requested.
- + Strong hand-painted texture that feels very artistic and organic.
- + The color palette perfectly matches the 'soft pastel' requirement.
- − Loses significant detail from the source, such as the plaid pattern on the shirt.
- − The character in the foreground has closed or semi-closed eyes, changing the expression from the source.
Verdict: FLUX.1 Kontext [max] is the winner for its superior preservation of the source image's identity, including the complex plaid pattern, while perfectly translating it into a high-quality Ghibli aesthetic. GPT Image 1 captures the 'dreamy' lighting and texture better, but it sacrifices too much detail and changes the facial expressions of the subjects significantly.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully added blowing hair and flying leaves.
- + Maintained high image clarity and sharpness.
- + Preserved the original woman's facial features and clothing well.
- − Anatomical error with the left hand being unusually long and reaching for the dog's head incorrectly.
- − The motion of the hair feels somewhat stiff compared to the instruction.
GPT Image 1
- + Excellent dynamic hair movement that feels very natural and wind-blown.
- + Abundance of flying leaves enhances the 'lively' feel requested.
- + Better preservation of the original pose without adding extra limbs or distorted hands.
- − Slight loss in facial sharpness compared to the original image.
- − A few leaves appear slightly like digital artifacts rather than integrated objects.
Verdict: GPT Image 1 followed the instructions more effectively by creating a more convincing sense of motion in the hair and adding more environmental elements like leaves. While FLUX.1 Kontext [max] kept the image sharper, it introduced a significant anatomical error by distorting the woman's left hand to touch the dog.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography including the correct grave accent on 'Caffè'
- + Perfect adherence to the light background and brown/cream color palette
- + High-quality vector-style texture and consistent line work
- − None notable
GPT Image 1
- + Accurate text rendering for both name and date
- + Clean silhouette of the cloche dome
- − Failed the requirement for a light background by using black
- − Lower level of detail and character in the typography compared to model_a
- − The 'Est. 1720' banner feels less integrated into the vintage aesthetic
Verdict: FLUX.1 Kontext followed all prompt instructions perfectly, including the specific color scheme and light background texture. GPT Image 1 generated a solid logo but failed to follow the background color requirement, resulting in a less authentic 'vintage' feel than FLUX.1 Kontext.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully captures a more complex, illustrative scene with depth.
- + Uses a NASA-inspired color palette effectively.
- + Includes detailed vector interpretations of the Saturn V and Lunar Module.
- − Text labeling is highly nonsensical and contains many spelling errors (e.g., 'LUNEAR MODULLE').
- − Mismatches labels with icons (e.g., pointing to Earth and labeling it 'MOON').
- − Includes irrelevant elements like a ringed planet (Saturn) which wasn't part of the Apollo 11 mission steps.
GPT Image 1
- + Excellent adherence to the 'infographic' layout with clear, logical steps.
- + Perfect spelling on major names like Armstrong, Aldrin, and Collins.
- + Consistent flat-vector iconography that aligns well with the prompt's requested style.
- − One minor spelling error in a sub-label ('EARLLUNAR').
- − The Saturn V rocket silhouette is simplified but slightly less detailed compared to Model A.
Verdict: GPT Image 1 followed the infographic constraints much better, providing a clear, step-by-step layout with accurate name spelling and consistent iconography. FLUX.1 Kontext [max] produced a visually interesting illustration but failed significantly on the 'infographic' aspect, with nonsensical labels and factual errors in the diagram's logic.
Explore each model
OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs