Black Forest Labs' 12-billion parameter multimodal flow transformer for in-context image generation and editing with character consistency, typography handling, and commercial-ready quality
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [pro]
#41 of 62 in Text-to-Image
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [pro]
0%
win rate
Ties
0%
GPT Image 2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent photographic realism with natural lens blur (bokeh)
- + Clean, precise geometry for the glass cube
- + Accurate representation of the requested lighting from the left
- − The blue sphere has a fuzzy, felt-like texture that maybe wasn't intended
- − The sphere appears to be floating slightly above the bottom surface of the cube
GPT Image 2
- + Strong prompt adherence for all spatial relationships
- + Highly realistic glass refraction and thickness
- + Vibrant colors and sharp textures on the book and sphere
- − The red book is slightly wider than the cube, creating a minor balance oddity
- − The green plant is a bit more crowded in the frame compared to the first model
Verdict: Both models followed the prompt perfectly, placing all objects in their correct relative positions. GPT Image 2 is the slight winner due to superior material rendering, specifically the convincing thickness of the glass and the solid, grounded appearance of the blue sphere, whereas FLUX.1 Kontext [pro] shows the sphere slightly hovering.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent depiction of raining atmosphere and wet textures.
- + Highly realistic skin details and naturalistic lighting.
- + Strong bokeh effect that isolates the subject well.
- − The man appears to be riding or leaning on the bike rather than 'repairing' it.
- − Rain streaks look a bit like a static filter overlay in some areas.
GPT Image 2
- + Accurately depicts the 'repairing' action with tools visible.
- + Captures the motion blur of passing cars effectively.
- + The framing feels more like a spontaneous 'candid' shot as requested.
- − Significant anatomical errors with the hands (extra fingers/merged digits).
- − Wet pavement reflections are present but less cinematic than Model A.
- − The bike's handlebars and seat geometry are slightly warped.
Verdict: FLUX.1 Kontext [pro] creates a much more visually stunning and photorealistic image with superior skin textures and lighting, though it fails to show the 'repairing' action clearly. GPT Image 2 follows the specific narrative details of the prompt better—including the tools and active repair—but suffers from poor hand rendering and less convincing realism. FLUX.1 Kontext [pro] is the preferred choice for its professional photographic quality and adherence to the 'cinematic but realistic' requirement.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent skin texture and realistic facial features
- + Ornate engraving on the armor is clear and prominent
- + Strong lighting coherence with the torchlight effect
- − Missed the request for beads in the hair braids
- − The dirt and scars are very faint, making the character look perhaps too clean for 'battle-worn'
GPT Image 2
- + Strict adherence to all prompt elements including hair beads and dirt/grime
- + Beautifully detailed engraving and weathered texture on the armor
- + Dynamic lighting and effective use of bokeh
- − Slightly less lifelike eye rendering compared to model A
- − Composition feels slightly tighter on the shoulder than the face
Verdict: While FLUX.1 Kontext [pro] produces a very clean and realistic portrait, GPT Image 2 follows the specific prompt details much more closely, including the beads in the braids and a more convincing 'battle-worn' appearance. GPT Image 2 captures the gritty atmosphere of the prompt while maintaining high levels of detail in the engraving and lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Features very clean, minimalist layout consistent with 'modern' design trends.
- + Accurately represents the requested sections (Appetizers, Pizza, Mains).
- + High photographic quality for the food items shown.
- − Text consists of nonsensical, garbled characters.
- − Food images do not always align with the section titles (e.g., a salad-like dish under Pizza and a pizza under Mains).
GPT Image 2
- + Exceptional text rendering with coherent, legible dish names and descriptions.
- + Perfectly executes the grid layout with a high variety of colorful, professional food photos.
- + Strong branding elements and icons that enhance the professional casual dining theme.
- − Slightly more cluttered than typical 'minimalism' due to the high density of photos.
Verdict: GPT Image 2 is significantly superior as it provides a fully functional, legible, and professional menu design with high-quality photography that matches the text. While FLUX.1 Kontext [pro] captures a more 'minimalist' aesthetic, its garbled text and mismatched food-to-category placement make it far less useful for the specific prompt.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent photorealistic texture on the bun and melting cheese.
- + Clean and legible typography for the main title.
- + High aesthetic quality in the 'magical' ember lighting.
- − Failed to include the starburst element for the price.
- − The price text is redundant, appearing twice.
- − The burger explosion is very static and lacks the requested sense of motion.
GPT Image 2
- + Successfully followed all prompt instructions, including the starburst and specific text placements.
- + Dynamic 'exploded' composition with ingredients tilted and sauce splashing for a high sense of motion.
- + Impressive fiery texture applied directly to the font as requested.
- − The 'LIMITED TIME ONLY' text is slightly less polished than the other text elements.
- − Some minor lighting inconsistencies between the burger and the background flames.
Verdict: GPT Image 2 is the clear winner as it followed every detail of the complex prompt, including the specific 'starburst' for the price and the 'fiery' rendering of the text. While FLUX.1 Kontext [pro] produced a very clean, high-quality image, it failed to include the starburst and lacked the dynamic motion requested in the exploded view.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent chalk texture and realistic graining on the letters
- + Perfectly legible text including the small footer note
- − Failed to render the title in the requested 'elegant cursive' style
- − The handwriting looks somewhat digital and uniform compared to natural chalk
GPT Image 2
- + Successfully rendered the title in elegant cursive as requested
- + High level of prompt adherence for all menu items and prices
- + Enhanced cozy café atmosphere with better environmental detail and lighting
- − None notable
Verdict: GPT Image 2 is the clear winner as it followed all specific formatting instructions, including the cursive title requirement which FLUX.1 Kontext [pro] ignored. GPT Image 2 also provided a more immersive 'cozy café' setting while maintaining high text accuracy and a natural chalk aesthetic.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent cinematic lighting and composition
- + High level of detail in the astronaut suit and horse anatomy
- + Creative surrealism by including a second small astronaut riding the horse
- − The astronaut has horse hooves instead of boots
- − The physics of the reins and the way the horse is balanced on the astronaut are a bit incoherent
GPT Image 2
- + Perfectly literal adherence to the 'horse on top' spatial instruction
- + High texture quality on the lunar surface and spacesuit
- + Clear and legible NASA-style branding
- − The composition is very centered and less cinematic than the competitor
- − The astronaut's hands have too many fingers (6-7 per hand)
Verdict: FLUX.1 Kontext [pro] provides a much more cinematic and visually interesting image, though it suffers from anatomical artifacts like hooves on the astronaut. GPT Image 2 is more literal in its interpretation of the 'horse on top' request and captures the textures well, but is marred by significant lobster-like hand artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent fur texture and lighting on the capybara.
- + The capybara's paw placement on the steering wheel is very natural.
- − The passenger is on a phone call rather than 'looking at her phone' as requested.
- − The 'TAXI' text on the hat is partially obscured or oddly cropped.
GPT Image 2
- + Perfect adherence to the passenger's action of looking at her phone.
- + The wide-angle composition better captures the 'New York at night' atmosphere through multiple windows.
- + The cap and jacket look more like authentic driver uniforms.
- − The capybara's paws look slightly more like human fingers in gloves than animal paws.
- − Minor distortion on the car door frame in the foreground.
Verdict: GPT Image 2 is the superior choice because it captures all elements of the prompt, specifically the businesswoman's bored expression while looking at her phone. While FLUX.1 Kontext [pro] has slightly better fur rendering, it fails the specific passenger interaction and uses a tighter crop that shows less of the New York environment.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Perfect text rendering for the main titles and banner
- + Clean, high-contrast composition that is easy to read
- + Strong cinematic lighting on the jack-o-lantern
- − Included a line of hallucinated nonsensical text ('Your: Vorkleat...')
- − Background is a bit sparse compared to the 'parchment' request
GPT Image 2
- + Exceptional vintage parchment aesthetic with intricate thorn and web borders
- + Accurate rendering of all requested text elements without hallucinations
- + Richly detailed background featuring the skyline, arches, and gothic architecture
- − The scrolling on the 'Invitation' text is slightly less clean than model A
- − The jack-o-lantern is a bit darker, making it pop slightly less against the background
Verdict: GPT Image 2 is the superior choice because it followed all text instructions perfectly without adding hallucinated words, whereas FLUX.1 Kontext [pro] added a line of gibberish. GPT Image 2 also captured the 'vintage gothic parchment' aesthetic much more effectively with its highly detailed borders and thematic background elements like the arches and the NYC skyline.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography style that matches the cartoon aesthetic
- + Clean, minimalist composition that adheres strictly to the simple diorama request
- + Soft, clay-like textures create a cohesive miniature look
- − The flag icon is stylized to the point of being a slightly odd shape
- − The rice grains look more like foam balls than realistic rice
GPT Image 2
- + High visual complexity with detailed PBR materials on the sushi and stone
- + Precise rendering of the Japanese flag icon
- + Stronger sense of a 'miniature 3D scene' with the garden and lantern details
- − The text layout is slightly disjointed with large gaps between elements
- − Ignored the request for 'minimal garnish' by including a large garden and multiple leaves
Verdict: FLUX.1 Kontext [pro] followed the minimalist aesthetic and text layout instructions more accurately, resulting in a cleaner graphic design. However, GPT Image 2 produced a much more impressive 3D miniature scene with superior material textures and variety, despite failing the 'minimal garnish' constraint.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent soft lighting and dreamy golden hour atmosphere.
- + High level of fur detail and consistent 'cute' aesthetic across all animals.
- − Failed to include a distinct bunny, instead creating two cat-like hybrids in the middle.
- − The animals are sitting statically rather than 'playfully chasing' and 'tumbling' as requested.
GPT Image 2
- + Accurately included all four distinct species: puppy, kitten, bunny, and fox.
- + Captured the requested action of chasing and tumbling, providing a much more dynamic composition.
- + Better implementation of 'god rays' and backlighting on the animals' fur.
- − The fox's front paw has some anatomical warping.
- − The kitten's tail is positioned somewhat awkwardly in the air.
Verdict: While FLUX.1 Kontext [pro] creates a beautiful, soft image, it fails the prompt by merging the species into indistinguishable cat-like creatures and ignoring the action requirements. GPT Image 2 successfully delivers on every part of the prompt, showing distinct recognizable animals in motion with excellent lighting and a true playful vibe.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent minimalist execution that feels modern yet vintage
- + Bold and readable classic typography
- + Subtle but effective paper texture on the background
- − Small spelling error in 'EEST' instead of 'EST'
- − The steam element is a bit overly simplified
GPT Image 2
- + Highly detailed engraving style with excellent shading on the cloche
- + Correct spelling of 'EST. 1720'
- + Beautifully nested composition within a decorative frame
- − May be considered too complex for a 'minimalist' request
- − The 'F' in Florian is slightly disconnected in its script style
Verdict: While FLUX.1 Kontext [pro] captures the 'minimalist' aspect of the prompt more effectively with its clean vector style, it fails on a basic text requirement by misspelling 'EST' as 'EEST'. GPT Image 2 provides a much richer and more professional-looking emblem with perfect spelling and sophisticated textures, making it the better choice for a boutique brand logo despite being less minimalist.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Features a more whimsical, artistic vector style.
- + Correctly captures the requested navy and muted red palette.
- − Fails to follow the numbered sequence and specific steps requested.
- − Produces nonsensical labels and inaccuracies, such as placing Saturn near the rocket and mislabeling the Moon as Earth.
- − Text rendering is poor and illegible in several places.
GPT Image 2
- + Strictly follows all 6 requested steps with perfect text rendering.
- + Expertly executes the clean, modern infographic layout with clear iconography.
- + Adds creative details like the crew names and landing site which enhance the poster theme.
- − The illustrations are slightly more Detailed/3D than a 'flat-vector' style would typically dictate.
- − The Saturn V rocket silhouette in step 1 is a bit squat compared to the real vehicle.
Verdict: GPT Image 2 is the clear winner as it followed every instruction in the prompt, including the specific sequence of six steps and the inclusion of specific icons. FLUX.1 Kontext [pro] failed significantly on coherence, mislabeling celestial bodies and ignoring the required logical flow of the infographic.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following