Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
GPT Image 1.5
#7 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
GPT Image 1.5
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light and shadows, creating a realistic depth.
- + Accurately places the plant behind the cube as requested.
- + Realistic textures on the wooden table and sparking sphere.
- − The sphere appears to have a textured or glittery surface rather than a smooth one.
- − The glass cube looks more like a shallow container with an open top where the book rests.
GPT Image 1.5
- + Clearer representation of a 'glass cube' structure.
- + Good adherence to all spatial instructions in the prompt.
- + The blue sphere is smooth and geometrically perfect.
- − The plant appears to be inside the cube's reflections rather than logically placed behind it.
- − Lighting is a bit flat compared to the other model.
Verdict: Both models followed the prompt instructions perfectly regarding objects and placement. FLUX.1 Kontext [max] produced a more photographically realistic image with sophisticated lighting, whereas GPT Image 1.5 followed the geometric literalism of a 'sphere' and 'cube' more strictly but with less realistic lighting and texture.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of rain streaks that feel integrated into the 3D space.
- + Strong cinematic lighting with realistic reflections on the pavement.
- + Captures a very natural, aged skin texture on the subject's hands and face.
- − The anatomy of the bicycle is highly distorted, specifically the chain and crankset area.
- − The subject does not look distinctly Japanese, appearing more generically elderly.
GPT Image 1.5
- + Much better bicycle anatomy including a realistic chain, gear set, and accessories like a basket.
- + The subject perfectly matches the demographic requested in the prompt.
- + The 'imperfect framing' and composition feel more like a genuine candid street photo.
- − The rain effect is very subtle, almost difficult to see compared to the puddles.
- − Slightly less 'cinematic' lighting compared to the first model.
Verdict: GPT Image 1.5 is the clear winner for its superior attention to the subject's ethnicity and the mechanical accuracy of the bicycle. While FLUX.1 Kontext [max] creates a more dramatic atmosphere with rain streaks and lighting, its failure to generate a functional-looking bicycle and the more generic appearance of the man make it less effective at fulfilling the specific prompt details.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent metal texture with realistic engraving and dirt patterns.
- + Strong dramatic lighting that creates a moody atmosphere.
- − The hair braids are thick and lack the requested beads.
- − Skin texture looks slightly overly smooth/metallic compared to the armor.
GPT Image 1.5
- + Perfect adherence to the 'beads in braids' and 'faint scars' prompt elements.
- + Highly detailed texture on the leather straps and multiple layers of clothing.
- + Excellent facial detail and lifelike expression with realistic dirt smudges.
- − The depth of field is slightly less pronounced than Model A's composition.
- − The warm torchlight highlights are a bit harsh in some areas of the plate.
Verdict: GPT Image 1.5 captured every specific detail of the prompt, including the small beads and the layered clothing, which FLUX.1 Kontext [max] missed. While FLUX.1 Kontext [max] has a very compelling cinematic atmosphere, GPT Image 1.5 is the superior image for its intricate detail and better fulfillment of the prompt's character description.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Strong minimalist aesthetic with a clean white professional layout
- + Food photography is well-integrated and looks high-quality
- + Excellent use of negative space for a modern brand feel
- − Text is mostly gibberish or contains severe misspellings
- − Does not follow the grid structure for food photos as strictly
- − Contrast is a bit low on some font elements
GPT Image 1.5
- + Perfect text rendering with realistic restaurant item names and prices
- + Matches the specific grid layout requested in the prompt
- + Clearly defines sections for Appetizers, Pizza, and Mains as requested
- − Layout is slightly more generic and less stylish than Image A
- − The horizontal lines between items feel a bit heavy for a 'minimalist' design
Verdict: GPT Image 1.5 is the clear winner for its superior text rendering and strict adherence to the requested content sections (Appetizers/Pizza/Mains). While FLUX.1 Kontext [max] captures a more high-end minimalist brand aesthetic, it fails to produce legible text and misses the specific organizational categories requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and clean graphic design.
- + High photorealistic detail on the burger patty and lettuce texture.
- + Sophisticated lighting on the environment and food.
- − The main burger is mostly assembled rather than properly 'exploded'.
- − Missing the starburst element for the price.
- − Strange floating orange bun-like shapes on the sides.
GPT Image 1.5
- + Perfectly follows the 'exploded' mechanical instruction with distinct layers.
- + Includes all requested text elements, including the starburst and fiery effect.
- + Strong sense of motion and energy with food and sparks.
- − The 'MAGIC BURGER' text is slightly cut off at the top.
- − Over-sharpened visual style that looks more like a composite than a single photo.
Verdict: GPT Image 1.5 adhered much more closely to the specific prompt requirements, successfully rendering the 'exploded' burger layout and the price within a starburst. While FLUX.1 Kontext [max] has cleaner typography and higher literal photorealism, it failed to separate the burger components as requested and included nonsensical floating shapes.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text legibility and spelling accuracy
- + Realistic chalkboard smudges and background details of the café Environment
- + Includes additional appropriate context-related text at the bottom
- − The handwriting style is a bit too clean and uniform, bordering on look-alike font
- − Missing the requested 'elegant cursive' for the title line
GPT Image 1.5
- + Successfully captured the 'elegant cursive' style for the title as requested
- + Highly realistic chalk texture with authentic varying opacity and dustiness
- + Handwriting variations feel more organic and less like a digital font
- − Text is slightly less legible due to the heavy chalk grain
- − The background is just a flat board with less environmental context than Model A
Verdict: GPT Image 1.5 followed the stylistic prompts more accurately, particularly by rendering the title in cursive as requested and creating a much more convincing dry-chalk texture. While FLUX.1 Kontext [max] produced a cleaner and more legible image, its handwriting looked slightly more like a digital font preset than a hand-drawn chalkboard.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent anatomical rendering of the horse and astronaut.
- + Clean, cinematic lighting with a minimalist and surreal background.
- + High-fidelity textures on the spacesuit and horse's coat.
- − Fails to follow the tricky 'horse on top' spatial instruction.
- − Background is slightly generic compared to the detailed foreground.
GPT Image 1.5
- + Rich, busy composition with multiple celestial elements like Saturn and a lunar lander.
- + Dynamic sense of motion with the dust kicking up on the lunar surface.
- + Strong cinematic atmosphere with complex lighting.
- − Fails to follow the 'horse on top' spatial instruction.
- − Anatomical issues with the horse's front legs and hooves.
Verdict: Both models failed the specific 'horse on top' logic test, which was a trick prompt to see if they could invert the usual relationship; both defaulted to an astronaut riding a horse. FLUX.1 Kontext [max] is the superior image due to its significantly higher anatomical accuracy and cleaner, more professional lighting, whereas GPT Image 1.5 suffers from cluttered elements and distorted animal anatomy.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture and lighting
- + Clean, high-resolution details on the capybara's fur and the taxi cap
- + The composition creates a cinematic depth of field
- − The passenger is holding a phone to her ear like a call instead of looking at it as requested
- − Only one paw is clearly on the steering wheel
GPT Image 1.5
- + Perfect adherence to the pose prompt with both paws on the steering wheel
- + Captures the bored expression of the passenger looking at her phone accurately
- + Authentic taxi interior details like the meter and dashboard
- − Image quality is slightly grainier compared to the alternative
- − Clothing textures on the capybara are a bit muddy
Verdict: GPT Image 1.5 followed the specific prompt instructions better, accurately depicting both paws on the wheel and the passenger looking at her phone with a bored expression. FLUX.1 Kontext [max] produced a more visually stunning and 'clean' image with superior realism, but it missed several key behavioral details requested in the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent font choice that feels authentically gothic
- + Highly legible text with strong contrast
- + Clever integration of cobwebs into the corners
- − Redundant text at the bottom repeats the location and date differently
- − The 'parchment' texture is very subtle and looks more like a modern poster
- − Jack-o-lantern rendering is a bit flat compared to the background
GPT Image 1.5
- + Beautiful sepia-toned parchment aesthetic with rich textures
- + Excellent composition with a detailed graveyard background and moon
- + More intricate border work featuring thorns and webs
- − Minor text artifacts in the cursive banner
- − The main title font is slightly less 'gothic' than Model A
Verdict: GPT Image 1.5 is the clear winner as it more successfully captures the requested 'vintage' and 'parchment' aesthetic with superior textures and a cohesive atmospheric background. FLUX.1 Kontext [max] has very clean text but suffers from redundant information at the bottom and a less interesting background overall.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Satisfies the prompt for a full, thick head of hair
- + Maintains original lighting and jacket details
- + Natural hair color matching the beard
- − Noticeably alters the man's facial features and eye shape
- − Hairline and sideburn integration looks slightly unnatural or layered on top
- − Changed the frame color and thickness of the glasses
GPT Image 1.5
- + Excellent preservation of original facial features and expression
- + Highly realistic hair texture and natural integration with the existing sideburns
- + Perfectly maintains the background and accessories like the glasses
- − The hair density is slightly less 'full' compared to Model A, though still realistic
Verdict: GPT Image 1.5 is the clear winner as it successfully adds realistic hair while perfectly preserving the identity of the person in the source image. FLUX.1 Kontext [max] provides the requested hair but significantly alters the person's face and eyes, failing the core requirement of source preservation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent soft 3D cartoon style consistent with the 'miniature' prompt.
- + Clean, minimalist composition that adheres to the specific texture requests.
- + Accurate text rendering for both requested words.
- − Missed the request for a small flag icon.
- − The isometric angle is slightly shallow compared to a standard 45° view.
GPT Image 1.5
- + Successfully included all elements, including the flag icon.
- + Outstanding PBR material rendering, especially on the wood and ceramic surfaces.
- + Perfect 45° isometric perspective.
- − More cluttered than the requested 'minimal garnish' scene.
- − The text style is a bit flatter compared to the 3D aesthetic of the objects.
Verdict: While FLUX.1 Kontext [max] captured the soft cartoon aesthetic more effectively, GPT Image 1.5 is the winner for following every part of the prompt, including the flag icon and the specific 45-degree angle. GPT Image 1.5 also exhibited superior material work, particularly the wood grain and soy sauce bottle, which better fit the 'realistic PBR materials' requirement.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Maintains the subject's denim clothing from the source image
- + Clear caricature art style with bold lines
- + Good integration of the hockey stick in the background
- − Added glasses which were not in the source image
- − The caricature features look generic and less like the specific person
- − The dog is small and somewhat obscured in the background
GPT Image 1.5
- + Exaggerated features strongly resemble the subject's face
- + Highly detailed news studio environment with a professional broadcast look
- + Creative and humorous integration of hockey (the screen and the small dog in a helmet)
- − Swapped the denim shirt for a red dress
- − The large dog's left paw/leg is anatomically messy
- − The text on the lower third is slightly stylized but readable
Verdict: GPT Image 1.5 provides a much more effective caricature by capturing the specific facial features of the woman in the source image, whereas FLUX.1 Kontext [max] creates a more generic cartoon character. GPT Image 1.5 also excelled at the humorous prompt requirements, including a dog in a hockey helmet and a background hockey game, making it the more creative and successful interpretation.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'golden sunrise' light with beautiful god rays
- + Extremely clean and high-definition fur textures on all animals
- + Well-composed and organized group portrait that includes all requested animals clearly
- − The animals are largely static and sitting rather than 'tumbling and chasing' as requested
- − The lighting on the butterfly wings feels a bit flat compared to the backgrounds
GPT Image 1.5
- + Perfectly captures the 'tumbling' and 'playfully chasing' action described in the prompt
- + Captures the 'dew sparkles' effect very effectively across the meadow
- + Dynamic poses create a more wholesome and joyful vibe
- − The fox kit has structural issues with its anatomy, specifically the placement of its paws and neck
- − The kitten's pose is a bit chaotic and slightly distorts its facial features
Verdict: While FLUX.1 Kontext [max] produces a cleaner, more technically perfect masterpiece with superior fur rendering, GPT Image 1.5 much better captures the spirit of the prompt's action requirements. FLUX.1 Kontext [max] is the winner because its anatomical consistency and lighting quality are significantly more polished and photorealistic.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent preservation of the original composition and poses.
- + Captures the iconic hand-painted watercolor texture of Studio Ghibli backgrounds.
- + Maintains clear facial expressions that match the source image context.
- − The character design looks more like modern TV anime rather than the specific Soft Ghibli style.
GPT Image 1.5
- + Successfully captures the 'warm, nostalgic mood' with a beautiful golden-hour lighting effect.
- + Effective use of soft focus and pastel tones to create a dreamy atmosphere.
- + Preserves the original arrangement of characters perfectly.
- − The woman in the foreground is excessively blurred, losing too much detail.
- − The lighting effect, while pretty, borders on a generic 'shoujo' filters rather than the clean look of Ghibli.
Verdict: Both models did an excellent job of maintaining the layout and poses of the famous 'distracted boyfriend' meme while translating it into an illustration. FLUX.1 Kontext [max] is the winner because its textures feel more like an authentic hand-painted background and its foreground characters remain sharp and expressive, whereas GPT Image 1.5 relies too heavily on a blur filter that washes out the subject.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully added wind-blown hair effect
- + Includes leaves flying in the air as requested
- + Adds a dynamic gait to the walking pose
- − Anatomy of the left hand is significantly distorted
- − Changes the background and overall composition more than necessary
GPT Image 1.5
- + Excellent depiction of wind-blown hair that looks natural
- + Abundant flying leaves create a high sense of energy
- + Preserves the original appearance of the person and dog very accurately
- − Some leaves appear to be floating or lack motion blur, reducing the realism of the wind
Verdict: GPT Image 1.5 is the clear winner as it successfully follows all edit instructions while maintaining the identity and anatomy of the original subjects perfectly. FLUX.1 Kontext [max] introduces significant anatomical errors to the hands and changes the background elements, making it a poor choice for a preservation-focused edit.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfectly captures the 'light background' and 'subtle texture' requested in the prompt.
- + Excellent typography with accurate spelling of 'Caffè' including the accent mark.
- + Clean, minimalist vector design that functions well as a logo.
- − The 'Est. 1720' text is slightly off-center within the banner.
GPT Image 1.5
- + More detailed illustration of the cloche with effective use of highlights and shading.
- + Good use of classic, elegant typography for the restaurant name.
- + Accurate banner design with correctly rendered 'Est. 1720' text.
- − Failed to follow the 'light background' instruction, providing a black background instead.
- − The typography style for 'Caffè' is slightly disconnected from the 'Florian' font choice.
- − The overall design feels more like a detailed illustration than a minimalist vector logo.
Verdict: FLUX.1 Kontext [max] followed the prompt more precisely, particularly the light background and minimalist vector style requirements. While GPT Image 1.5 produced a visually appealing illustration, it ignored the color scheme and background constraints, and FLUX.1 Kontext [max] handled the specific Italian accent mark in 'Caffè' more naturally.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent font usage for 'APOLLO 11' title.
- + Features high-quality silhouettes of the crew members.
- + Strong flat-vector aesthetics that feel cohesive.
- − Confusing informational flow with arrows pointing in illogical directions.
- − Labels do not match the icons shown (e.g., 'Translunar' is under a moon icon).
- − Contains significant gibberish text in the center.
GPT Image 1.5
- + Perfect adherence to the 6-step logical sequence requested.
- + Very clean, legible text and clearly labeled sections.
- + Consistency in iconography across all six stages of the mission.
- − Spacecraft designs are slightly more generic and less detailed than Image A.
- − The 'Saturn V' rocket in the launch phase is a bit simplified.
Verdict: GPT Image 1.5 is the clear winner as it successfully follows the requested 6-step infographic structure with accurate labeling and a logical flow. FLUX.1 Kontext [max] has a more sophisticated individual art style but fails significantly on the infographic functionality, providing mismatched labels and nonsensical flow arrows.
Explore each model
OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts