Head to head
Esc

Models · slot A

to navigate to pick

FLUX.2 [flex] Black Forest Labs GPT Image 2 OpenAI

Settled by community votes across 14 shared challenges, with an AI judge weighing in on each.

FLUX.2 [flex]

24.8 arena score

#14 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.2 [flex]

0%

win rate

Ties

0%

GPT Image 2

0%

win rate

Shared challenges 14

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent adherence to lighting instructions with a soft, natural window glow.
  • + Precise object alignment and modern, clean aesthetic.
  • + Natural-looking reflections on the tabletop.
  • The sphere is quite large relative to the 'small blue sphere' request.
  • The cube lacks a top glass panel, acting more like an inverted container.

GPT Image 2

  • + Follows the 'small' scale of the sphere more accurately.
  • + Highly realistic textures on the book cover and the wooden table surface.
  • + Perfectly captures the refraction and light green tint typical of thick glass edges.
  • The lighting feels slightly flatter compared to Model A.
  • The plant visibility through the glass is less pronounced than in the other version.

Verdict: GPT Image 2 is the winner as it more accurately captures the scale of the objects, specifically the 'small' blue sphere, and provides superior texture work on the wood and book. FLUX.2 [flex] produced a beautiful image with better lighting, but its sphere was oversized and the glass cube appeared to have no top surface.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent atmospheric lighting with cinematic bokeh
  • + Realistic skin textures and weather effects
  • + Superior adherence to the 50mm shallow depth of field request
  • The structural anatomy of the bicycle is slightly nonsensical and simplified
  • The background motion blur feels a bit generic compared to the foreground detail

GPT Image 2

  • + High level of technical detail on the bicycle and tool kit
  • + The 'imperfect framing' is captured well with objects cutting into the edges
  • + Accurate rendering of Japanese text on the sign
  • Lighting is flat and lacks the cinematic, atmospheric quality requested
  • Does not properly execute the 'motion blur from passing cars' instruction
  • Less 'shallow' depth of field than requested

Verdict: FLUX.2 [flex] wins on atmospheric quality and prompt adherence regarding the visual 'vibe', delivering a beautifully cinematic shot with realistic skin and rain. GPT Image 2 is more technically literal with the objects (bicycle parts and tools) and text, but it fails to capture the requested motion blur and cinematic lighting, resulting in a flatter, less dynamic image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent execution of ornate plate armor engravings
  • + Superior detail on leather straps and clothing textures
  • + Strong bokeh effect with distinct floating sparks
  • The scars look a bit painted on rather than integrated into the skin
  • The lighting on the face is slightly flat compared to the dramatic armor highlights

GPT Image 2

  • + Natural, lifelike skin texture with integrated dirt and scars
  • + Highly realistic lighting and cinematic color grading
  • + Intricate hair braiding with small beads perfectly integrated
  • The armor detail, while good, is slightly less crisp than Model A
  • The 'bokeh sparks' are less prominent than requested

Verdict: Both models followed the prompt exceptionally well, but GPT Image 2 creates a more believable, lifelike character with superior skin and hair integration. FLUX.2 [flex] excels in the technical rendering of the armor and specific prompt details like the leather straps, but it feels slightly more like a high-end CGI render than a photograph.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent photographic quality for the food items.
  • + Clean, modern use of white space and color blocks.
  • Nonsensical placeholder text under headers.
  • Logic error where the main title 'Appetizers' is used for the entire page.

GPT Image 2

  • + Perfectly legible and accurate English text throughout.
  • + Highly organized horizontal grid layout that matches the prompt perfectly.
  • Visuals are slightly more 'illustrative' than realistic photos.
  • Layout feels a bit crowded at the bottom.

Verdict: GPT Image 2 is significantly more functional as a menu design because it features actual legible text and properly categorized sections, whereas FLUX.2 focuses on photographic quality but fails with garbled text and a confusing titling structure. GPT Image 2 better adheres to the specific request for distinct sections for appetizers, pizza, and mains.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent typography with a clean, professional fiery glow
  • + High degree of photorealism in the textures of the patty and bun
  • + The starburst design is integrated naturally into the ad layout
  • The 'exploded' effect is less dynamic than Model B, with components still feeling somewhat stacked

GPT Image 2

  • + Superior 'exploded' effect with dynamic tilt and suspended ingredients
  • + Includes more ingredients like onions to add to the visual complexity
  • + The fiery texture on the text is very intense and matches the background perfectly
  • The layout is a bit cluttered, with text overlapping the image elements on the left
  • The bottom bun's sauce splash looks slightly more digital and less photorealistic than Image A

Verdict: Both models followed the prompt exceptionally well, but GPT Image 2 captured the 'dynamic, exploded' request more effectively through its tilted composition and scattered ingredients. FLUX.2 (flex) produced a much cleaner, more professional ad layout and arguably better food textures, but it feels slightly more static by comparison.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent layout with centered text and realistic chalk smudges
  • + High contrast and very legible text
  • + Includes chalk pieces at the bottom of the tray for realism
  • The handwriting looks slightly too uniform, more like a font than a hand-drawn sign
  • Missing the requested underlining flourish on the header

GPT Image 2

  • + Perfectly captures the authentic grainy chalk texture and varying pressure
  • + Handwriting looks truly organic with natural slanting and letter variations
  • + Includes the elegant cursive flourishes and decorative underline requested
  • The date is slightly crowded against the right edge of the board
  • Overall lighting is a bit dimmer compared to Model A

Verdict: Both models followed the complex text instructions perfectly, but GPT Image 2 captured the 'handwritten' aesthetic far more convincingly with realistic chalk grain and organic letter variations. FLUX.2 [flex] produced a very clean and attractive image, but the text appears slightly too much like a digital font despite the chalk-style rendering.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent character preservation including face, sunglasses, and the specific scarf from Image 2.
  • + Matches the yellow background and red ottoman studio environment perfectly.
  • Fails to match the exact pose, providing a crouching stance rather than the cross-legged lean from Image 1.
  • The right foot has anatomical issues with six toes and poor blending.

GPT Image 2

  • + Captures the precise, difficult body position and cross-legged pose from Image 1 very accurately.
  • + Successfully integrates the clothing and scarf from Image 2 onto the new pose.
  • + Preserves the environment and lighting of the source pose image perfectly.
  • The right hand has red-lacquered fingernails, which is a carry-over artifact from the female hand in Image 1.
  • The face and hair are slightly less sharp in resolution compared to the body.

Verdict: GPT Image 2 is the superior choice because it successfully followed the core instruction to replicate the 'exact pose' from Image 1, which FLUX.2 [flex] failed to do. While GPT Image 2 has a minor artifact (red fingernails from the original image), it managed the complex skeletal structure and clothing physics of the request whereas FLUX.2 [flex] simplified the pose into a generic crouch.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent cinematic lighting and color palette
  • + Highly detailed anatomy and textures
  • + Creative interpretation where the astronaut is carrying the horse.
  • Does not strictly show the horse 'riding' in a traditional pose; it looks more like the astronaut is a pack animal or carrying it.
  • The horse's legs are floating oddly around the astronaut's shoulders.

GPT Image 2

  • + Perfect adherence to the prompt with the horse literally riding the astronaut with a saddle.
  • + Good clarity on the textures of the spacesuit and horse fur.
  • + Consistent lighting between the foreground and background.
  • The astronaut's hands/gloves have an incorrect number of fingers and look morphologically strange.
  • The composition is a bit stiff and centered compared to the more dynamic Model A.

Verdict: GPT Image 2 followed the specific instruction of 'horse on top' much more literally and successfully by including a saddle and a riding posture. While FLUX.2 [flex] produced a more visually stunning and high-quality artistic piece, it interpreted the 'horse on top' as the astronaut carrying the horse rather than being ridden by it.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent anatomical interpretation of paws gripping a steering wheel
  • + Clear, vibrant composition with distinct Manhattan street lighting
  • + High levels of detail in the capybara's fur and the taxi interior hardware
  • The passenger is slightly less integrated into the background lighting

GPT Image 2

  • + Natural, moody cinematic lighting that blends the characters with the environment
  • + Captures the bored, detached expression of the passenger perfectly
  • + The capybara's professional expression is very convincing
  • Paws are rendered as amorphous furry bumps rather than having distinct toes/fingers
  • The taxi interior is quite dark and lacks detail in the dashboard area

Verdict: Both models followed the complex prompt very well, but FLUX.2 [flex] edges ahead due to the superior rendering of the capybara's paws on the steering wheel, which was a specific requirement. GPT Image 2 has a more cinematic and atmospheric feel with better passenger expressions, but it fails on the finer details of the capybara's anatomy.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent typography legibility and clean layout.
  • + Accurately follows the parchment paper aesthetic with torn edges.
  • + Clean, predictable composition that works well for a functional invitation.
  • The background is a bit sparse and lacks the 'cinematic' depth requested.
  • The 'border' is quite thin and simple compared to the gothic request.

GPT Image 2

  • + Strong gothic atmosphere with intricate details, webs, and thorns.
  • + Creative inclusion of 'The Arches' and a NYC skyline in the background to match the prompt details.
  • + Highly artistic and cinematic lighting that feels more like a vintage poster.
  • The typography is slightly crowded with the scroll overlapping the jack-o-lantern stem.
  • Darker composition makes some of the text slightly harder to read at a glance.

Verdict: While FLUX.2 provides a very clean and functional layout with perfect text, GPT Image 2 is the superior choice for its artistic depth and creative interpretation of the location details. GPT Image 2 successfully incorporates a visual representation of 'The Arches' and the NYC skyline into a dark, gothic aesthetic that feels much more 'cinematic'.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent adherence to the 'minimal' request with a clean, simple layout
  • + Accurate 45-degree isometric perspective and soft lighting
  • + Perfectly centered and clean text placement
  • The textures look more like clay than realistic PBR materials
  • The flag icon is quite large and simple compared to the 3D scene

GPT Image 2

  • + Impressive realistic PBR textures on the fish and materials
  • + Highly detailed modeling including wood grain and stone textures
  • + Creative interpretation of the diorama base with garden elements
  • Ignored the 'minimal' garnish prompt, resulting in a cluttered scene
  • The text styling is a bit bulky and has shadow artifacts

Verdict: FLUX.2 [flex] adhered much better to the specific 'minimal' and 'soft texture' constraints of the prompt, creating a clean graphic design. However, GPT Image 2 produced significantly better material rendering (PBR) and visual complexity, though it largely ignored the request for minimalism. FLUX.2 is the better choice for a logo or clean icon style, while GPT Image 2 is a better 3D render.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent handling of god rays and soft morning mist
  • + Realistic level of detail in the fur texture
  • + Dynamic pose that feels like the animals are in motion
  • The fox's front right leg has an anatomically awkward black marking extension
  • The kitten is slightly smaller than naturally expected relative to the puppy

GPT Image 2

  • + Beautiful backlighting and vibrant color palette
  • + Good capture of the 'tumbling' interaction between animals
  • + Detailed rendering of wildflowers and foreground grass
  • The fox has an extra paw or strangely detached appendage visible in the lower center
  • The kitten's facial structure is slightly distorted and less realistic than its counterparts

Verdict: Both models captured the heartwarming, high-detail requirements of the prompt well. FLUX.2 [flex] is the likely winner because it avoids the anatomical glitches found in GPT Image 2, particularly the extra/misplaced limb on the fox, and provides a more cohesive sense of atmosphere with the morning mist.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Excellent adherence to the minimalist aspect of the prompt
  • + Clean vector lines suitable for an actual logo
  • + Perfect spelling and typography.
  • The 'subtle texture' requested is nearly invisible
  • The design is perhaps too simplified, bordering on generic clip-art style.

GPT Image 2

  • + Beautiful vintage aesthetic with rich engraving-style details
  • + Excellent use of the background texture and warm color palette
  • + Professional composition that feels like a high-end brand identity.
  • Fails the 'minimalist' requirement of the prompt
  • Slightly less 'vector' in appearance compared to Model A.

Verdict: While FLUX.2 [flex] followed the 'minimalist' constraint more strictly, the result is somewhat basic. GPT Image 2 produced a much more sophisticated and visually appealing vintage emblem that captures the 'Est. 1720' historical feel of Caffè Florian perfectly, despite ignoring the minimalist instruction.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.2 [flex]
GPT Image 2

AI Judge Analysis

FLUX.2 [flex]

  • + Clean minimalist design with a true vector aesthetic.
  • + Accurate NASA-inspired muted color palette.
  • + Legible typography with no spelling errors.
  • Missing step 6 (Landing) entirely.
  • The layout is a bit sparse with significant empty space.

GPT Image 2

  • + Successfully includes all 6 requested steps including the final Landing.
  • + Excellent adherence to the 'NASA-inspired' and informational request including crew and landing site.
  • + Strong visual hierarchy and composition for an infographic.
  • Icons are slightly more complex than the requested 'flat-vector' style.
  • Small artifacts present in the crew silhouettes and some text elements.

Verdict: GPT Image 2 is the superior choice because it fully adhered to the sequential instructions, including all six specific mission phases whereas FLUX.2 [flex] stopped at step five. While FLUX.2 [flex] nailed the minimalist flat-vector aesthetic, GPT Image 2 provided a much more comprehensive and educational infographic that feels like a completed project.

Next steps

Explore each model