Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
GPT Image 2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light and caustic reflections on the wooden table.
- + Highly realistic glass transparency and thickness.
- + The plant is clearly visible through the glass as requested.
- − The sphere has a textured, glitter-like surface rather than a smooth finish.
- − The text on the book spine is nonsensical.
GPT Image 2
- + Perfect adherence to all spatial instructions and object descriptions.
- + The sphere has a clean, matte blue finish that feels very cohesive.
- + Excellent composition with the window clearly visible to establish the light source.
- − The glass cube edges appear slightly thinner and less physically substantial than Model A.
- − Slightly less complex light interaction on the table surface.
Verdict: Both models followed the complex spatial instructions perfectly. FLUX.1 Kontext [max] produced a more photographically impressive image with superior lighting and glass physics, while GPT Image 2 provided a cleaner, more literal interpretation of the objects with better background context. FLUX.1 Kontext [max] is the winner due to the tangible realism of the reflections and materials.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of rain with realistic streaks and wet surface reflections.
- + Dynamic composition with 'imperfect' cinematic framing requested in the prompt.
- + Superior textures on the man's hands and face, feeling very natural.
- − The man's ethnicity appears somewhat ambiguous rather than specifically Japanese.
- − The rain effects are a bit heavy, slightly obscuring the background motion blur.
GPT Image 2
- + Successfully captures the specific Japanese setting including legible kanji/signage.
- + Good motion blur on the passing white car in the background.
- + Includes additional realistic elements like a tool kit and a worn, weathered bicycle.
- − The 'light rain' is barely visible, lacking the atmospheric quality of Image A.
- − The pavement doesn't look particularly wet or reflective.
- − The lighting is a bit flat and less cinematic than requested.
Verdict: FLUX.1 Kontext [max] creates a much more atmospheric and cinematic image that captures the 'feeling' of rain and the specified photographic style perfectly. While GPT Image 2 better captures the specific cultural setting and minor details like the tool kit, it fails to convincingly render the rain and wet pavement requested in the prompt. FLUX.1 Kontext [max] is the winner for its superior texture and adherence to the requested mood.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of warm torchlight reflecting off the armor and skin.
- + Highly detailed engraving on the plate armor with convincing metal textures.
- + Strong adherence to the 'battle-worn' descriptor with visible sweat, grime, and intense eyes.
- − The beads in the hair are missing or very indistinct.
- − Skin texture on the forehead appears slightly oversaturated or unnaturally shiny.
GPT Image 2
- + Perfect adherence to the 'hair braided with small beads' requirement.
- + Exceptional facial detail with realistic freckles, faint scars, and dirt.
- + Clearly visible textured leather straps and cloth underlayers as requested.
- − The lighting is a bit more muted than the 'warm torchlight' requested.
- − The depth of field is less shallow compared to Model A.
Verdict: While FLUX.1 Kontext [max] captures the intense atmosphere and lighting of a battle-worn scene more effectively, GPT Image 2 adheres more closely to the specific technical details of the prompt, such as the beads in the hair and the distinct layering of leather and cloth. GPT Image 2's facial realism and subtle scar textures provide a more nuanced interpretation of the character.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Clean layout with effective use of white space.
- + Maintains a consistent aesthetic for a pizza-focused casual dining environment.
- − Text is largely illegible gibberish.
- − The grid is repetitive, showing almost exclusively pizza even for different sections.
- − The font choice for section headers is a script font rather than bold sans-serif.
GPT Image 2
- + Perfect adherence to section requirements (Appetizers, Pizza, Mains) with distinct food photos for each.
- + Exceptional text rendering with coherent item names, descriptions, and prices.
- + Follows all stylistic cues including bold sans-serif fonts and vibrant accents.
- − The layout is slightly more dense than Model A, though still professional.
Verdict: GPT Image 2 is the clear winner as it produced a fully functional, legible menu that strictly followed all prompt instructions including specific food sections. FLUX.1 Kontext [max] failed to provide readable text and lacked the diversity of food photos requested, defaulting mostly to pizza images.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and clean font rendering.
- + Good photorealistic texture on the central burger patty and bun.
- − Fails to create a truly 'exploded' view, with most ingredients still stacked.
- − Missing the starburst element for the price tag.
GPT Image 2
- + Perfect adherence to the 'exploded' view requirement with vertical separation.
- + Dynamic fiery text effects and starburst element exactly as requested.
- + Thorough integration of all components including sauces and glowing embers.
- − Slightly cluttered composition with the text overlapping some sparks.
- − The price text is a bit thick, making the decimal point small.
Verdict: GPT Image 2 followed the prompt much more accurately, successfully delivering the 'exploded' burger layout and the specific starburst price tag that FLUX.1 Kontext [max] omitted. While FLUX.1 Kontext [max] produced a clean image, it failed to interpret the dynamic motion of the ingredients, whereas GPT Image 2 captured the energy and fire theme throughout all elements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfectly legible and accurate text rendering for the entire prompt.
- + Good chalk smudging effects on the background for realism.
- + Clear, high-contrast composition.
- − The text looks a bit too clean, like a digital chalk-style font rather than organic handwriting.
- − Failed to follow the instruction for 'elegant cursive' for the title.
GPT Image 2
- + Excellent chalk texture and realistic varied handwriting style.
- + Successfully followed the request for elegant cursive in the header.
- + Warm, cozy café lighting enhances the atmospheric quality.
- − The text is slightly less legible due to the thinness of the chalk strokes.
- − One small typo in the word 'Special' (looks like 'Specials' with a slightly merged 's').
Verdict: GPT Image 2 is the superior choice because it captures the authentic 'handwritten' aesthetic and followed the specific stylistic instruction for cursive heading, whereas FLUX.1 Kontext [max] used text that felt more like a uniform digital font. GPT Image 2 also provided a much more realistic chalk texture and a warmer, more thematic background environment.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent anatomical rendering for both the horse and astronaut
- + High-quality cinematic lighting and background composition
- + Clear, high-resolution textures
- − Failed to follow the logical reversal requested in the prompt ("horse on top")
GPT Image 2
- + Perfectly adhered to the complex prompt instruction of having the horse on top of the astronaut
- + Highly surreal and creative interpretation of the concept
- + Detailed textures on the spacesuit and lunar surface
- − Mismatched scale between the horse's upper body and its tiny hooves
- − Awkward anatomical merger where the horse's legs meet the astronaut's shoulders
Verdict: While FLUX.1 Kontext produced a much more aesthetically pleasing and high-quality image, it completely ignored the specific instruction to place the horse on top of the astronaut. GPT Image 2 successfully followed the difficult prompt logic, creating a truly surreal image that matches the requested description, despite some anatomical inconsistencies in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + High resolution and realistic fur texture on the capybara.
- + The yellow taxi cab lighting on the vehicle frame is very convincing.
- − The woman in the back is holding a phone to her ear like a call, rather than looking at it as requested.
- − The capybara's paws are not clearly placed on the steering wheel as requested.
GPT Image 2
- + Perfect adherence to the pose, showing both paws on the steering wheel.
- + Excellent depiction of the passenger looking at her phone with a bored expression.
- + The interior composition feels more like a real taxi ride with a clearer view of the back seat.
- − The capybara's fur looks slightly less defined compared to Model A.
- − The proportions of the capybara's head to the steering wheel are a bit oversized.
Verdict: GPT Image 2 is the clear winner as it followed every detail of the prompt, including the specific boredom and action of the passenger and the placement of the capybara's paws. While FLUX.1 Kontext [max] has slightly better lighting and texture, it failed to depict the passenger correctly and missed the specific paw placement requirement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and size hierarchy
- + Very clean border design that incorporates webs and thorns effectively
- + Perfect adherence to the specified date and color contrast
- − The 'Arches' location is repeated twice in the text block
- − The background is slightly less detailed than the competitor
GPT Image 2
- + Beautifully detailed gothic illustration including a castle, bridge, and full moon
- + The scroll banner is more stylistic and aesthetically pleasing
- + Text rendering is highly accurate with no repetitions
- − The event details at the bottom are quite small and harder to read at a glance
- − The border, while intricate, feels a bit cluttered compared to the main title
Verdict: Both models followed the prompt exceptionally well, producing high-quality invitations with perfect text. FLUX.1 Kontext [max] produced a cleaner, more practical invitation layout, but GPT Image 2 is the winner due to its superior artistic detail, including the addition of a full moon and a thematic NYC-style bridge ('The Arches') within the background art.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'minimal garnish' request
- + Smooth, refined 3D textures that match the 'soft cartoon' aesthetic well
- + Clean text rendering for 'JAPAN' and 'SUSHI'
- − Failed to include the requested flag icon
- − The background is slightly more saturated than 'light blue'
GPT Image 2
- + Perfect adherence to all prompt elements, including the flag icon
- + Higher detail in materials and textures while maintaining the diorama look
- + Very clean, bold typography that pops against the background
- − Ignores the 'minimal garnish' instruction by adding many extra elements
- − The lighting is slightly more artificial and harsh compared to Model A
Verdict: GPT Image 2 is the more successful output as it included every specific prompt requirement, including the flag icon which FLUX.1 Kontext [max] omitted. While GPT Image 2 ignored the 'minimal' constraint in favor of a busier scene, its overall execution of the isometric diorama style and high-clarity materials is superior for this specific challenge.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully includes all four requested species with clear, distinct features.
- + Captured a soft, 'dreamy' lighting style that enhances the wholesome vibe.
- + Excellent fur textures and expressive eye glints.
- − The composition is quite static, with animals sitting rather than 'playfully chasing' or 'tumbling'.
- − The butterflies appear somewhat flat and glowy rather than realistic.
GPT Image 2
- + Excellent depiction of movement, with animals leaning forward and paws raised to simulate chasing.
- + The lighting and god rays are very dramatic and well-integrated into the scene.
- + High level of detail in the grass and wildflowers providing a sense of depth.
- − The fox kit has a slightly unusual anatomy on its front right paw.
- − The kitten's tail is positioned oddly, almost appearing to sprout from its upper back.
Verdict: GPT Image 2 captured the dynamic energy of the prompt much better, showing the animals actively chasing and moving through the meadow. While FLUX.1 Kontext produced a cleaner, more stable image, it felt more like a posed portrait rather than the playful scene requested. GPT Image 2 is the winner for better prompt adherence regarding action and lighting.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'minimalist' keyword with clean forms
- + Perfect text rendering of the brand name and establishment year
- + Subtle, effective texture that enhances the vintage feel without clutter
- − The steam element is a bit thick and stylized compared to the cloche
- − Slightly less 'grand' than the historic nature of the real restaurant
GPT Image 2
- + Beautiful, intricate linework on the cloche and frame
- + Sophisticated typography that feels very high-end and classical
- + Warm color palette and parchment-style background are executed perfectly
- − Fails the 'minimalist' requirement by including heavy framing and flourishes
- − The font weights between the two names are slightly unbalanced
Verdict: FLUX.1 Kontext [max] delivered a much more accurate response to the 'minimalist' and 'vector emblem' part of the prompt, creating a clean logo that would be practical for modern use. While GPT Image 2 produced a more visually stunning and decorative illustration, it ignored the minimalist constraint in favor of a complex, ornate design. Both models handled the text and specific required elements like the cloche and banner flawlessly.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Clean minimalist color palette and background.
- + Includes clear silhouette icons for the three astronauts.
- + The overall flat-vector style is consistent throughout the image.
- − Misses several steps requested in the prompt (only labels 4 areas).
- − Displays generic or nonsensical icons instead of the specific Saturn V and Lunar Module requested.
- − Layout is disorganized and text labels are poorly placed.
GPT Image 2
- + Exemplary adherence to the step-by-step instructions and iconography.
- + High-quality vector illustrations of the Saturn V and Lunar Module.
- + Superior layout organization with clear headings and consistent graphic design.
- − The eagle mascot in the top right corner is slightly busy compared to a strictly flat-vector style.
- − Minor lighting effects on the lunar surface deviate slightly from the 'subtle gradients only' constraint.
Verdict: GPT Image 2 is the clear winner as it perfectly follows the instructional sequence requested in the prompt, including all six specific steps with accurate icons for each. FLUX.1 Kontext [max] failed to provide the requested timeline, used unrecognizable icons for the space craft, and had a fragmented composition.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following