Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
GPT Image 1.5
#7 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0.0%
win rate
Ties
0.0%
GPT Image 1.5
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent photographic realism with a high-end lens feel.
- + Perfect handling of refractive glass and reflections.
- + Strong adherence to the 'light from the left' instruction with realistic shadows.
- − The glass cube is missing a front face, appearing more like a glass stand or open box.
GPT Image 1.5
- + Successfully renders all six sides of the glass cube.
- + Follows all spatial instructions including the sphere inside and plant behind.
- + Good material texture on the wooden table.
- − The plant visibility through the glass is slightly murky.
- − The lighting lacks the depth and contrast found in the competing image.
Verdict: Both models followed the prompt's spatial instructions perfectly. FLUX.1 Kontext [dev] produced a more visually stunning, photorealistic image with superior lighting, though the glass cube lacks a front pane; GPT Image 1.5 successfully rendered a fully enclosed cube but with a flatter, more synthetic overall appearance.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent color vibrance and clear rainy day atmosphere.
- + Realistic skin texture and facial features during a candid-style shot.
- − The subject is posing with or sitting on the bike rather than repairing it.
- − The bicycle is missing its chain, which is a significant structural error.
- − Lacks the requested motion blur from passing cars; cars appear static.
GPT Image 1.5
- + Strong adherence to the 'repairing' action with a squatting pose and toolkit.
- + Excellent use of shallow depth of field and motion blur on the background car.
- + Perfectly captures the 'imperfect framing' and 'candid' street photography aesthetic.
- − The bike's rear structure and kickstand area are a bit cluttered/messy in detail.
Verdict: GPT Image 1.5 is the clear winner as it followed every technical instruction in the prompt, including the specific action of 'repairing' and the difficult 'motion blur' request. FLUX.1 Kontext [dev] produced a high-quality portrait, but failed on the core activity and contained a glaring anatomical error with the missing bicycle chain.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent high-relief engraving on the chestplate
- + Strong cinematic lighting with dramatic shadows
- + Sharp texture on the fabric underlayer
- − Missed the request for braided hair with beads
- − The face appears too clean and lacks the requested dirt and grit
- − Armor looks more like cast bronze than battle-worn plate
GPT Image 1.5
- + Perfect adherence to all prompt details including braids with beads, dirt, and scars
- + Superior texture on leather straps and cloth elements
- + Extremely lifelike eyes and realistic skin micro-details
- − Slightly cluttered composition due to the amount of hair and detail
- − Reflection on the metal is quite bright, bordering on overexposed in small spots
Verdict: GPT Image 1.5 is the clear winner as it followed every specific detail of the prompt, including the braided hair with beads and the battle-worn grit on the skin, which FLUX.1 Kontext [dev] ignored. While FLUX.1 Kontext [dev] produced a clean, cinematic portrait, GPT Image 1.5 delivered a much more realistic and texture-rich interpretation of a battle-hardened character.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong minimalist aesthetic with bold typography
- + High-quality food photography that feels cohesive
- + Unique grid layout that creates a high-end editorial feel
- − Text is nonsensical and includes severe artifacts or spelling errors
- − Doesn't clearly differentiate several of the specific food sections requested
- − Some graphic elements overlap awkwardly with text
GPT Image 1.5
- + Excellent text readability with no spelling errors
- + Accurately includes all requested sections (Appetizers, Pizza, Mains)
- + Logical layout with Pricing and descriptions following real-world menu conventions
- − Composition feels slightly crowded compared to a truly 'minimalist' design
- − Grid of pictures is less integrated into the white space than Model A
Verdict: While FLUX.1 Kontext [dev] captures a more modern and artistic minimalist aesthetic, it fails to produce legible or meaningful text. GPT Image 1.5 is the clear winner for this task as it provides a fully functional menu design with perfect adherence to the required sections, pricing, and high-quality food photography that matches the descriptions.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering for 'MAGIC BURGER'.
- + Good use of the starburst element for the price.
- + High resolution and clean composition.
- − Failed the main request for an 'exploded' burger, showing a fully assembled one instead.
- − Spelling error in 'LIMITED TIME ONLY' (rendered as 'LNHLY').
- − The background coals look a bit generic and static compared to the dynamic request.
GPT Image 1.5
- + Perfect adherence to the 'exploded' burger request with dynamic suspended components.
- + High level of texture detail on the patty, bun, and vegetables.
- + Successfully applied the fiery, glowing effect to both the background and the text.
- − The pricing starburst is a bit jagged and overlaps the burger.
- − Some visual clutter from the excessive ember sparks near the top text.
Verdict: GPT Image 1.5 is the clear winner as it followed the complex 'exploded burger' instruction and maintained high realism across all food components. FLUX.1 Kontext [dev] failed the primary layout request by showing a standard assembled burger and included a spelling typo in the secondary text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Clean wood frame composition.
- + Consistent handwriting style across different sections.
- − Significant spelling errors throughout including 'Mashroom', 'Risoktso', and 'Octpus'.
- − Date text at the top is largely illegible and jumbled.
GPT Image 1.5
- + Perfect text rendering with zero spelling errors.
- + Excellent chalk texture and smudging details for high realism.
- + Strict adherence to the cursive style requested for the title.
- − The bottom line of text uses a slightly different, more uniform weight that looks a bit more digital than the main items.
Verdict: GPT Image 1.5 is the clear winner as it followed every complex prompt instruction perfectly, including the specific date and difficult food items without a single typo. FLUX.1 Kontext [dev] struggled significantly with the text rendering, resulting in multiple misspellings and an illegible date.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the 'horse on top' spatial instruction
- + Clear and clean visual style with high resolution
- + Captures the surreal nature of the prompt well
- − The composition is a bit static compared to a cinematic style
- − The astronaut's suit looks more like a modern jump suit than a technical space suit
GPT Image 1.5
- + Dynamic and cinematic lighting with rich textures
- + High level of detail on the space suit and celestial bodies
- − Complete failure to follow the core instruction of 'horse on top'
- − Anatomical issues with the horse's front legs/hooves
- − Overly busy background creates visual clutter
Verdict: FLUX.1 Kontext [dev] is the clear winner because it successfully interpreted the challenging spatial instruction to put the horse on top of the astronaut. GPT Image 1.5 provided a high-quality but generic image of an astronaut riding a horse, failing the specific 'not vice versa' requirement of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent fur texture and sharp lighting
- + High resolution with clean, professional composition
- + Human facial expression is very grounded and realistic
- − Failed the instruction to have both paws on the steering wheel
- − The cap is a modern baseball style rather than a traditional taxi driver cap
- − The capybara's face looks slightly anthropomorphized compared to a real animal
GPT Image 1.5
- + Perfectly followed the instruction for both paws on the steering wheel
- + Captures a very authentic 'New York taxi driver' cap and vintage interior aesthetic
- + Features a more realistic animal face that contrasts well with the absurd scenario
- − The passenger in the back is slightly out of focus and has some minor artifacts on her hands/phone
- − The 'TAXI' text on the cap is a bit generic
Verdict: GPT Image 1.5 is the winner as it followed all specific prompt instructions, including the difficult 'both paws on steering wheel' requirement and the correct style of taxi cap. FLUX.1 Kontext [dev] produced a higher resolution image with better lighting, but failed several key details of the prompt like paw placement and the boredom of the passenger.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent large title text readability
- + Clean and simple graphic composition
- + Follows the square format request
- − Significant spelling errors in the scroll banner and location text
- − Lacks the 'vintage parchment' texture requested, appearing more like a flat digital poster
- − Bats look like simple cartoon icons rather than cinematic elements
GPT Image 1.5
- + Perfect text rendering for all requested details, including the scroll and event info
- + Strong adherence to 'vintage gothic parchment' and 'moody night sky' aesthetic
- + Highly detailed thorns and webs border that adds to the cinematic quality
- − The 'Arches' location text is slightly less prominent than other details
- − Composition is a bit crowded with many competing atmospheric elements
Verdict: GPT Image 1.5 is the clear winner as it perfectly executed the 'vintage parchment' and 'cinematic' requirements while maintaining flawless spelling for all requested text. In contrast, FLUX.1 Kontext [dev] failed significantly on text accuracy for the small banner and location details, and lacked the textural depth requested in the prompt.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent addition of thick, dark hair with defined texture.
- + Preserves the overall composition and lighting of the background.
- − Significantly alters facial features, making the person look like a different individual.
- − The hair texture looks slightly airbrushed/artificial compared to the original beard.
GPT Image 1.5
- + Successfully preserves the original facial structure, glasses, and skin details.
- + Matches the hair texture and color perfectly to the existing beard for a cohesive look.
- + Maintains the grainy, film-like quality of the source image.
- − The hairline transition is slightly visible upon close inspection, but still realistic.
Verdict: GPT Image 1.5 is the clear winner for this task as it successfully fulfills the edit request while perfectly preserving the subject's identity and the image's overall aesthetic. FLUX.1 Kontext [dev] fails as an editing tool in this instance because it completely redesigns the man's face, resulting in a different person entirely.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent soft 3D cartoon aesthetic
- + Very clean and minimalist composition
- + Large, legible text
- − The flag icon is a strange geometric shape instead of the Japanese flag
- − The sushi anatomy is slightly confusing with the black band placement
GPT Image 1.5
- + Detailed and accurate Japanese flag icon
- + Beautiful PBR materials with realistic textures
- + Excellent interpretation of a 'miniature diorama' on a base
- − Ignored the 'minimal garnish' instruction by adding a teapot, soy sauce, and multiple items
- − Text is smaller and less bold compared to Model A
Verdict: GPT Image 1.5 produced a much more sophisticated diorama with accurate cultural symbols like the flag, though it ignored the instruction for minimal garnish. FLUX.1 Kontext [dev] captured the 'cartoon' and 'minimal' aesthetic better but failed significantly on the flag icon. GPT Image 1.5 is the winner for its superior visual quality and material rendering.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Provides a comic-book style caricature that translates physical features well.
- + Includes a small dog and a television set to represent the job.
- − Completely misses the hockey requirement.
- − The text 'JOB' and other elements are unnecessary and low-quality.
- − Poorly rendered hand/paw at the bottom right.
GPT Image 1.5
- + Successfully incorporates all elements: TV anchor, multiple dogs, and hockey (background and helmet).
- + High visual quality with a professional digital caricature aesthetic.
- + Excellent text rendering of 'BREAKING NEWS' and clear facial expressions.
- − Changes the subject's clothing from the original denim to a red dress.
- − The hockey stick held by the small dog is floating/awkwardly placed.
Verdict: GPT Image 1.5 is the clear winner as it followed every part of the prompt, including the hockey requirement which FLUX.1 Kontext [dev] ignored entirely. While both models captured the subject's likeness well, GPT Image 1.5 created a much more cohesive and professional-looking caricature with high-quality details.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent sense of motion and action/chasing.
- + Clean, high-quality digital aesthetic with well-defined butterflies.
- + Strong lighting and color saturation that fits the 'joyful' theme.
- − Failed to include all four animals (missing the bunny and fox).
- − The animals look more like generic 3D assets than photorealistic fur.
- − Incorrectly generated two cats instead of a kitten and fox.
GPT Image 1.5
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox.
- + Beautifully rendered 'god rays' and dew sparkles as requested in the prompt.
- + Highly detailed fur texture and complex wildflower interaction.
- − The composition is a bit crowded with the animals overlapping awkwardly.
- − The fox's paw in the bottom right corner has some anatomical blurring.
- − The lighting on the puppy's face is slightly flat compared to the dramatic background rays.
Verdict: While FLUX.1 Kontext [dev] has a more dynamic composition with clear movement, it failed significantly on the prompt adherence by missing two of the four required animals. GPT Image 1.5 correctly included the golden retriever, tabby kitten, bunny, and fox kit while capturing the specific atmospheric details like god rays and dew sparkles, making it the superior choice for this specific challenge.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent preservation of the original image's composition and poses
- + Very clean line art and clear character expressions
- + Matches the modern anime aesthetic common in 2D digital art
- − Does not capture the painterly Ghibli aesthetic well
- − Colors are too saturated and lack the requested soft pastel quality
GPT Image 1.5
- + Perfectly captures the Studio Ghibli 'hand-painted' texture and soft lighting
- + Uses the requested pastel color palette and warm, nostalgic mood
- + Strong adherence to the 'dreamy background' requirement
- − Slight loss of sharpness in facial features compared to the source
- − The girl in the foreground loses some of her distinctive facial proportions
Verdict: While FLUX.1 Kontext [dev] preserved the structure of the original meme more accurately, it failed to deliver on the specific artistic style requested, looking more like standard digital anime. GPT Image 1.5 successfully transformed the image into a Ghibli-inspired watercolor painting with the correct textures, lighting, and mood, making it the superior choice for this specific creative edit.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent source preservation, maintaining the woman's face and the dog's features accurately.
- + Subtle and realistic hair movement that looks natural for a breeze.
- − The 'flying leaves' are barely noticeable and look more like green specks than leaves.
- − Overall energy shift is very minimal compared to the source.
GPT Image 1.5
- + Strong adherence to the request for 'energetic and lively' through dramatic wind and leaf effects.
- + Successfully captures many flying leaves with motion blur, creating a sense of depth.
- − Significant loss of detail in the woman's face and hair texture compared to the source.
- − Slight alteration of the dog's facial features and eyes.
Verdict: FLUX.1 Kontext [dev] focuses on high-fidelity source preservation but fails to meaningfully execute the request for flying leaves. GPT Image 1.5 much more effectively captures the 'dynamic motion' requested with dramatic wind effects, though it makes noticeable changes to the woman's face in the process.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Clean vector aesthetic suitable for a minimalist brand
- + Excellent text legibility and font choice
- + Perfectly centered and balanced composition
- − Missed the 'banner' requirement for the established date
- − The cloche icon looks more like a building dome or a cupcake
GPT Image 1.5
- + Successfully included the 'banner' for the 'Est. 1720' text
- + Very clear and well-illustrated cloche dome
- + Excellent use of the 'subtle texture' and 'vintage' style requested
- − Failed the 'light background' requirement, opting for black
- − The word 'Caffè' is slightly less legible due to the scripted font style
Verdict: GPT Image 1.5 adhered better to the specific iconography requirements like the banner and the actual appearance of a cloche, though it failed to provide the light background requested. FLUX.1 Kontext [dev] produced a much cleaner minimalist vector, but the cloche was less recognizable and it missed the banner element.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Follows the color palette closely with a navy background.
- + Captures a stylized, abstract vector feel.
- − Text is heavily garbled and contains numerous spelling errors like 'APOLO'.
- − The icons are messy, non-representative of the prompt steps, and difficult to interpret.
- − Failed to follow the logical sequence of the mission steps.
GPT Image 1.5
- + Excellent text rendering with correct spelling for all mission steps and astronaut names.
- + Perfect adherence to the 6-step logical sequence with clear, relevant icons for each.
- + Clean, professional composition that looks like a real infographic poster.
- − Added green to the Earth icon which was not part of the requested NASA-inspired palette.
- − Slightly more complex than 'flat-vector' due to some layered shapes, though still highly effective.
Verdict: GPT Image 1.5 is the clear winner as it successfully followed the complex multi-step prompt and rendered all text accurately. FLUX.1 Kontext [dev] produced garbled text and abstract icons that did not clearly represent the 6 mission stages requested.
Explore each model
OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts