Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0.0%
win rate
Ties
0.0%
GPT Image 2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to lighting instructions with a clear soft window light from the left.
- + Highly realistic textures on the wood grain and glass reflections.
- + Vibrant and saturated colors.
- − The sphere appears more like a large marble than a 'small sphere'.
- − The glass cube has an inconsistent open-bottomed structure that doesn't fully enclose the sphere.
GPT Image 2
- + Perfectly captures the scale of a 'small' blue sphere relative to the cube.
- + Strong composition with clear layering of the foreground objects and the plant behind.
- + The glass cube has a more realistic and defined structural geometry.
- − The blue sphere has a slightly flat, matte texture compared to the highly detailed surrounding materials.
- − The 'soft window light' is less pronounced than in the competitor's image.
Verdict: Both models followed the spatial and color instructions perfectly. FLUX.1 Kontext [dev] produced a more photographic and aesthetically pleasing image with superior lighting, while GPT Image 2 better managed the relative scale of the objects (making the sphere small rather than large). FLUX.1 Kontext [dev] is the winner due to the higher quality of light and material rendering.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent depiction of raindrops and realistic wet pavement reflections
- + High quality rendering of the subject's face and clothing
- + Cinematic color grading and lighting
- − Failed the core action: the man is sitting on/straddling the bike rather than repairing it
- − Composition is very centered and posed, missing the requested 'imperfect framing' and 'candid' feel
- − Background cars are static and sharp rather than having motion blur
GPT Image 2
- + Perfect adherence to the 'repairing' action and 'candid' street photography style
- + Accurately captured motion blur on passing vehicles and imperfect framing
- + Highly realistic skin textures and details, including a toolbox and realistic bicycle wear
- − The light rain is very subtle and less visible than in Model A
- − A slight blurring artifact on the subject's right hand involved in the repair
Verdict: While FLUX.1 Kontext [dev] produced a more traditionally beautiful and cinematic image, it failed to follow several technical instructions, specifically the action of repairing the bike and the motion blur. GPT Image 2 followed the prompt almost perfectly, capturing the candid nature, the repair task, the motion blur, and the 'imperfect' framing of a real street photo.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Ornate engraved plate armor is highly defined and clear
- + Dramatic warm torchlight lighting with effective bokeh sparks
- + High contrast and sharp facial features
- − Missed the request for braided hair with beads
- − Skin appears relatively clean with very minimal 'battle-worn' texture compared to the prompt
- − Leather and cloth textures are somewhat obscured by shadows
GPT Image 2
- + Excellent adherence to the 'braided hair with small beads' requirement
- + Superb skin texture showing dirt, freckles, and a realistic battle-worn look
- + Highly detailed rendering of worn metal, leather straps, and cloth underlayers
- − Lighting feels more like natural daylight with a small backlight rather than warm torchlight
- − The bokeh sparks are very subtle compared to Model A
Verdict: GPT Image 2 is the superior response because it captured almost every specific detail of the prompt, including the complex hair braids and the gritty skin textures. While FLUX.1 Kontext [dev] produced a more dramatic lighting effect, it failed to include the requested beads and braids, and the character looks too pristine for the 'battle-worn' description.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent high-resolution food photography that feels authentic
- + Strong modern grid layout that maximizes visual impact of the food
- − Text consists of gibberish words and incoherent characters
- − Layout lacks the functional structure of a usable menu
GPT Image 2
- + Perfectly legible English text with logical descriptions and pricing
- + Highly organized sections for Appetizers, Pizza, and Mains as requested
- + Professional use of iconography, branding, and social media footers
- − The food photos look slightly more like stock 3D renders compared to Model A
- − Layout is quite dense, bordering on a flyer rather than a minimalist menu
Verdict: GPT Image 2 is the superior choice because it functions as an actual menu with legible, accurate text and clearly defined sections that match the prompt exactly. While FLUX.1 Kontext [dev] produces more realistic food photography, the text is unreadable nonsense and the layout fails to convey the necessary information for a restaurant setting.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Clean and legible text rendering
- + Vibrant colors and high-contrast lighting
- − Failed to create an 'exploded' burger, showing multiple whole burgers instead
- − Used the wrong currency symbol (£ instead of €)
- − The composition feels static rather than dynamic
GPT Image 2
- + Perfect adherence to the 'exploded burger' layout
- + Energetic composition with great sense of motion and detail
- + Beautifully integrated fiery text effects and the correct currency symbol
- − The 'LIMITED TIME ONLY' text is slightly less crisp than the other elements
- − The top bun is slightly distorted in perspective
Verdict: GPT Image 2 followed the complex prompt requirements much better than FLUX.1 Kontext [dev], accurately depicting the exploded burger components and the specific currency symbol requested. FLUX.1 Kontext [dev] produced a cleaner aesthetic but failed the core creative concept of an exploded view, instead duplicating the product.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Features a clear, well-framed chalkboard with high contrast.
- + Good imitation of chalk texture on the 'Today Specials' header.
- − Numerous spelling errors including 'Mashroom', 'Risoktso', 'Octpus', and garbled date text.
- − Multiple repetitions of words like 'with with' and duplicated price symbols.
- − The text style looks like a digital marker font rather than natural chalk.
GPT Image 2
- + Excellent adherence to the requested cursive elegant header style.
- + Perfect spelling and grammar for all menu items and the date.
- + Highly realistic chalk texture with natural-looking dust and imperfections on the board.
- − The ambient lighting is slightly dim, making the bottom text a bit harder to read compared to the top.
Verdict: GPT Image 2 is significantly superior, executing the complex text prompt with perfect spelling and a highly realistic handwritten chalk aesthetic. In contrast, FLUX.1 Kontext [dev] contains multiple glaring typos, word repetitions, and failed to render the date correctly.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent anatomical rendering of both the human face and the horse.
- + Clean, cinematic lighting and high resolution.
- − Failed the primary instruction to have the horse riding the astronaut; the horse is simply floating behind him.
- − The composition feels static and less surreal than requested.
GPT Image 2
- + Perfect adherence to the complex prompt instruction of 'horse on top'.
- + High level of surrealism with a saddle placed on the astronaut's back.
- + Strong detailed texture on the moon's surface and the spacesuit.
- − The astronaut's hands have too many fingers (6-7 visible on each hand).
- − The horse's front legs and harness attachment are anatomically confusing.
Verdict: While FLUX.1 Kontext [dev] produced a much cleaner and more realistic image, it completely failed the logic of the prompt. GPT Image 2 successfully captured the surreal concept of a horse riding an astronaut, making it the winner for prompt adherence despite the anatomical errors in the hands.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent fur texture rendering
- + Clear, high-quality lighting inside the cabin
- + Realistic human features
- − The capybara only has one paw on the wheel while the other is on its lap
- − The capybara's face looks slightly more like a general rodent/groundhog than a distinct capybara
GPT Image 2
- + Perfect adherence to the pose prompt with both paws on the steering wheel
- + Very accurate capybara facial anatomy
- + Authentic New York taxi driver hat design
- − Slightly lower background detail clarity
- − The lighting is a bit muddy compared to Model A
Verdict: Model B (GPT Image 2) followed the prompt instructions more closely, correctly placing both paws on the steering wheel and capturing the specific look of a capybara and a NYC taxi driver cap perfectly. While Model A (FLUX.1 Kontext [dev]) has higher overall image sharpness and better lighting, it failed the specific pose instruction and the hat looks more like a standard baseball cap.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong, glowing central jack-o-lantern.
- + Border of thorns is clearly defined.
- − Several spelling errors in the scroll banner and location text.
- − Lacks the 'vintage parchment' texture requested, appearing more like a modern digital graphic.
- − The background is mostly flat black rather than a detailed night sky.
GPT Image 2
- + Perfect text rendering for all requested titles, banners, and event details.
- + Excellent gothic atmosphere with rich textures, parchment effects, and cinematic lighting.
- + High level of detail in the border, moon, and background scenery.
- − The pumpkin's internal glow is slightly less intense than Model A's.
Verdict: GPT Image 2 is the clear winner as it followed every instruction perfectly, including complex text rendering with zero spelling errors. While FLUX.1 Kontext [dev] captured the individual elements, it failed to produce the vintage parchment aesthetic and struggled significantly with the text on the scroll and location, whereas GPT Image 2 delivered a professional, high-quality invitation layout.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Clean, minimalist 2D/3D hybrid aesthetic
- + Accurate text rendering for 'JAPAN' and 'SUSHI'
- + Strict adherence to the 'minimal garnish' instruction
- − The flag icon is unrecognizable and abstract
- − The sushi model is overly simplified, lacking realistic PBR textures requested
GPT Image 2
- + Excellent miniature 3D diorama execution with high detail
- + Correct Japanese flag icon and well-styled 3D typography
- + Superior material rendering with texture on the fish and wood
- − Ignored the 'minimal garnish' instruction, including a complex scene
- − The text layout is slightly crowded at the top
Verdict: While FLUX.1 Kontext [dev] followed the 'minimalist' instruction better, GPT Image 2 produced a much higher quality 3D miniature diorama that aligned with the 'realistic PBR materials' and style requirements. GPT Image 2 also successfully rendered a recognizable Japanese flag, whereas FLUX.1 Kontext [dev] failed that specific detail.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong bokeh effect and warm lighting.
- + Action-oriented poses showing joyful movement.
- − Failed to include the bunny and red fox kit, showing only three animals (dog and two kittens/cats).
- − The 'tabby' kitten is depicted as white and orange rather than distinct tabby markings.
- − Anatomical issues with the animals' paws and the floating butterfly wing in the background.
GPT Image 2
- + Excellent prompt adherence, including all four specific animals: puppy, tabby kitten, bunny, and fox kit.
- + Realistic fur textures and clear, expressive eyes in alignment with the prompt.
- + Superior lighting effects featuring clear god rays and dew-like sparkles.
- − The fox kit’s leg and paw anatomy is slightly distorted during the run.
Verdict: GPT Image 2 is the clear winner as it successfully follows the complex prompt requiring four specific animal types, whereas FLUX.1 Kontext [dev] omitted the bunny and fox entirely. GPT Image 2 also delivers much better technical detail in the fur textures and environmental lighting (god rays).
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Perfectly follows the minimalist and vector emblem style instruction
- + Clean, high-contrast typography
- + Accurate spelling and accented character rendering
- − Missed the 'banner' requirement for the 'Est. 1720' text
- − The steam effect is very abstract and less recognizable
GPT Image 2
- + Excellent adherence to the 'vintage' and 'banner' requirements
- + Beautifully textured background and sophisticated engraving style
- + Accurate spelling with elegant serif typography
- − Less 'minimalist' than requested, leaning more into high-detail ornamental design
Verdict: While FLUX.1 Kontext [dev] followed the 'minimalist' keyword more strictly, GPT Image 2 captured the aesthetic essence of a vintage logo much more effectively, including the specific banner element which FLUX missed. GPT Image 2's use of texture and classical typography feels more authentic to a historic establishment like Caffè Florian.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully captures a flat, minimalist vector style
- + Adheres to the muted red and navy color palette
- + Uses interesting abstract iconography
- − Significant text corruption and spelling errors (e.g., 'APOLO')
- − Icons are too abstract and do not clearly represent the requested mission steps
- − The composition feels cluttered and lacks a logical flow
GPT Image 2
- + Excellent layout with a clear chronological timeline
- + High text legibility and accurate spelling of names and stages
- + Includes detailed, recognizable icons for the Saturn V and Lunar Module
- − Stylistically more detailed/illustrative than the requested 'flat-vector' style
- − The 'Earth Orbit' illustration shows the command module relative to the planet in an oversized scale
Verdict: GPT Image 2 is the superior infographic because it follows the logical structure of the prompt, providing all six requested steps in a clear, readable timeline. While FLUX.1 Kontext [dev] captures the 'flat' art style more closely, its total failure in text rendering and icon clarity makes it useless as an educational poster.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following