Head to head
Esc

Models · slot A

to navigate to pick

FLUX.2 [dev] Black Forest Labs GPT Image 2 OpenAI

Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.

FLUX.2 [dev]

24.5 arena score

#18 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 2

28.1 arena score

#3 of 62 in Text-to-Image

Top 3 in Text-to-Image
Vote tally

Where the votes landed

FLUX.2 [dev]

0%

win rate

Ties

0%

GPT Image 2

0%

win rate

Shared challenges 15

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent adherence to lighting requirements with a warm sunlit effect
  • + Superior glass physics and reflections on the sphere and table
  • + Realistic wood texture with visible grain and depth
  • The glass cube has double lines on the top edge that look slightly unrealistic

GPT Image 2

  • + Perfectly centered and clean composition
  • + Crisp focus on the plant leaves behind the cube
  • + Highly detailed book texture and paper edges
  • The blue sphere lacks realistic reflections and highlights for a glass environment
  • The lighting is more diffuse and lacks the specific 'left window light' directional quality

Verdict: Both models followed the spatial requirements of the prompt perfectly, placing all objects in their correct relative positions. FLUX.2 [dev] is the winner because its rendering of light, glass refraction, and surface reflections is significantly more realistic than GPT Image 2, which feels slightly flatter in its lighting and materials.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent adherence to the 'motion blur from passing cars' prompt requirement.
  • + Highly realistic skin textures and wet fabric details.
  • + Strong cinematic atmosphere with realistic light reflections on the pavement.
  • The mechanical structure of the bike near the handlebars is distorted and illogical.
  • The hands are slightly cluttered with anatomical artifacts.

GPT Image 2

  • + Good inclusion of props like the toolbox to emphasize the 'repairing' action.
  • + The Japanese text on the sign adds to the locational authenticity.
  • + Clearer overall composition that shows more of the setting.
  • Failed to include the requested motion blur for passing cars.
  • Lower photographic realism with slightly smoother, flatter skin and hair textures.
  • Lighting feels a bit too bright and even, missing the 'cinematic' mood requested.

Verdict: FLUX.2 [dev] followed the technical requirements of the prompt much better, specifically capturing the motion blur and the cinematic, moody lighting of a rainy day. While GPT Image 2 provided a clean composition with good environmental details, it ignored the motion blur instruction and lacked the realistic photographic texture present in the FLUX.2 [dev] output.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent adherence to the 'beads' in hair requirement
  • + Clear depiction of ornate engraving and battle-worn features like scars
  • + Strong lighting effects from the torch within the composition
  • The leather straps look slightly flatter than Model B
  • The torch itself is a bit distracting at the edge of the frame

GPT Image 2

  • + Exceptional skin texture and realistic, lifelike eyes
  • + High-quality rendering of tarnished metal and weathered leather
  • + Sophisticated use of shallow depth of field and soft lighting
  • The beads in the hair are very subtle/metallic and less distinct than requested
  • Missing the explicit 'scars' mentioned in the prompt, focusing more on dirt

Verdict: FLUX.2 [dev] provides better adherence to specific prompt details like the scars and the colorful beads, while GPT Image 2 offers a more photorealistic and aesthetically pleasing composition with superior skin and armor textures. However, FLUX.2 [dev] captures the 'battle-worn' essence more accurately through the visible wounding and clearer torch interaction.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent photographic quality and lighting in the food images
  • + Clean dual-page spread layout
  • + Consistent colorful accent bars that define the style
  • Nonsensical text and incorrect section headers like 'PIZZAU'
  • Poor grid alignment between photos and corresponding text list items

GPT Image 2

  • + Exceptional text legibility and semantic relevance
  • + Perfect adherence to all requested sections (Appetizers, Pizza, Mains)
  • + Superior layout design with clear price points and descriptors
  • Slightly lower fidelity in food photography compared to FLUX.2
  • Small social media icons at the bottom are slightly distorted

Verdict: GPT Image 2 is significantly better for this task as it produces a functional, legible menu with accurate text and clear categorization. While FLUX.2 [dev] produces higher quality individual food images, its inability to render sensible text or align menu items with their images makes it less useful for a design prompt.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent text legibility and clean graphic design
  • + Highly realistic food textures, especially the lettuce and bun
  • + Very accurate adherence to all specific text constraints
  • The 'exploded' effect is a bit static compared to the dynamic angles of the competitor
  • Composition feels a bit safe and centered

GPT Image 2

  • + Superior dynamic composition with tilted angles creating a sense of energy
  • + Excellent fiery glowing texture on the text elements
  • + Higher level of detail in the food components like onions and sauce splashes
  • The '6' in the price has some internal graphical artifacts or speckling
  • The top bun's perspective is slightly warped due to the extreme angle

Verdict: Both models followed the prompt exceptionally well, but GPT Image 2 wins on artistic energy and dynamic composition, making the burger feel truly 'exploded' and exciting. FLUX.2 [dev] produced a more professional, clean advertisement look with better text clarity, but it lacked the intensity requested in the motion and background descriptions.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent chalk texture and realistic smudge marks on the board.
  • + Strong visual clarity while maintaining a authentic handwriting feel.
  • Includes some ghosting/smudging artifacts in the middle of the 'Grilled Octopus' line.
  • The handwriting style is slightly thicker and less 'elegant' than the style requested.

GPT Image 2

  • + Highly consistent and elegant handwriting style that perfectly matches the 'cursive' prompt.
  • + Excellent layout within the wooden frame that feels more like a completed café scene.
  • + Perfect text accuracy and spacing for all listed menu items.
  • The chalk texture on the letters is slightly more uniform than the varied texture found in Model A.

Verdict: Both models followed the complex text instructions perfectly, but GPT Image 2 produced a superior result due to its better layout and more elegant handwriting style. FLUX.2 [dev] had some minor rendering artifacts (smudges) that obscured a small portion of the text, whereas GPT Image 2 was clean and professional throughout.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Succesfully incorporates specific clothing items like the scarf and black sweater from the reference.
  • Catastrophic failure in anatomy, merging the two characters into a multi-headed or layered monstrosity.
  • Poor image coherence and visible artifacts on the left side of the character.

GPT Image 2

  • + Perfectly replicates the complex pose from Image 1.
  • + Excellent character consistency, including the face, hair, and accessories from Image 2.
  • + Seamless lighting and high visual quality.
  • The scarf's pattern is slightly simplified compared to the source image.

Verdict: FLUX.2 [dev] failed significantly by merging the two source images into a single distorted figure with multiple heads and limbs. In contrast, GPT Image 2 perfectly executed the request, placing the character from Image 2 into the exact dynamic pose of Image 1 while maintaining natural anatomy and high visual fidelity.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + High visual quality and resolution.
  • + Cinematic lighting and composition.
  • Completely failed the prompt's core constraint of having the horse on top of the astronaut.

GPT Image 2

  • + Excellent adherence to the tricky prompt logic, correctly placing the horse on top of the astronaut.
  • + Highly detailed textures on the spacesuit and the lunar surface.
  • + Great attention to detail with the reflection in the astronaut's visor.
  • The horse's front legs look slightly awkward resting on the astronaut.

Verdict: The prompt included a specific constraint: 'horse on top, not vice versa'. FLUX.2 [dev] failed this constraint completely, generating a standard astronaut riding a horse, whereas GPT Image 2 understood the surreal request perfectly, depicting a horse riding an astronaut with great detail and textures.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Successfully replicates the scarf pattern and coat style
  • + Includes extra jewelry details not in model B
  • Completely fails to preserve the face and skin of the person from Image 1, merging facets of the person in Image 2 instead
  • Changes the hair style and facial features significantly

GPT Image 2

  • + Excellent preservation of the person's face, skin markings, and hair from Image 1
  • + Accurately places the clothing layers from Image 2 onto the target body shape while maintaining the background elements
  • The scarf pattern is slightly less precise compared to Image 2 than what Model A achieved

Verdict: GPT Image 2 is the clear winner because it successfully followed the core editing instruction to keep the person's face, hair, and skin markings from the original image completely unchanged while applying the new clothes. FLUX.2 failed this task fundamentally by generating a hybrid person that looks more like the man in the second reference image, losing the unique vitiligo patterns and facial structure of the subject in Image 1.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent photorealistic lighting and skin/fur textures
  • + Perfect execution of the 'bored' expression on the passenger
  • + High-quality rendering of the capybara's paws on the steering wheel
  • The capybara's jacket looks slightly like it's merging with the seat/door

GPT Image 2

  • + Strong cinematic composition through the open door/window area
  • + Very detailed taxi driver hat with NYC branding
  • + Accurate depiction of a rainy/glowing Manhattan night background
  • The passenger's expression is a bit too drowsy rather than 'bored'
  • The scale of the capybara relative to the interior feels slightly too large

Verdict: Both models followed the complex prompt with high accuracy, but FLUX.2 [dev] stands out for its superior photorealism and better handle on the specific facial expressions requested. While GPT Image 2 has a great cinematic atmosphere, FLUX.2 [dev]'s rendering of the capybara's paws and the woman's casual phone usage feels more grounded and cohesive.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent legibility and clean graphic design
  • + Perfect adherence to the specified date, time, and location text
  • + Well-defined thorns and spiderwebs in the border
  • The parchment effect is a bit too clean and modern for a vintage gothic look
  • The composition feels a bit more like a digital graphic than a cinematic poster

GPT Image 2

  • + Rich, atmospheric vintage gothic aesthetic with high level of detail
  • + Creative interpretation of 'The Arches' showing a bridge and cityscape in the background
  • + Superior cinematic lighting and texture on the pumpkin and parchment
  • The text 'The Arches' is slightly cramped near the bottom border
  • More visual clutter makes the small banner text harder to read than in the other version

Verdict: FLUX.2 [dev] provides a very clean, functional invitation with perfect typography and clear layout. However, GPT Image 2 offers a much more immersive 'vintage gothic' atmosphere with superior textures, better lighting, and a clever visual nod to the NYC location. While both models followed the text instructions perfectly, GPT Image 2 is the winner for its artistic depth and cinematic quality.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent adherence to 'minimal garnish' instruction.
  • + Clean and soft lighting that matches the 'cartoon 3D' aesthetic perfectly.
  • + Perfect text layout and flag icon rendering.
  • The 'diorama base' is extremely simple compared to a miniature scene.

GPT Image 2

  • + Incredibly detailed PBR materials and high-quality textures.
  • + Creative diorama base including a stone lantern and garden details.
  • + Very high visual clarity and vibrant colors.
  • Failed the 'minimal garnish' instruction with a cluttered scene.
  • The text and flag have heavy drop shadows that diverge slightly from the clean 2D clean-text style.

Verdict: FLUX.2 [dev] followed the prompt's stylistic cues for minimalism and soft cartoon textures more accurately, whereas GPT Image 2 ignored the 'minimal' constraint in favor of a complex, busy scene. However, GPT Image 2 showcased much higher material realism and a more interesting diorama interpretation, making it more visually impressive as a 3D miniature.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent anatomical accuracy on the fur textures and animal features.
  • + Beautiful backlight and rim lighting that integrates the animals into the scene.
  • + Clean, high-resolution rendering with a soft, cinematic bokeh.
  • Static composition where animals are sitting rather than 'playfully chasing'.
  • Includes an extra animal (two rabbits) not requested in the prompt.

GPT Image 2

  • + Captures the action of 'playfully chasing' and 'tumbling' much better than Image A.
  • + Dynamic composition with a more varied and colorful wildflower meadow.
  • + Follows the count and type of animals exactly as requested.
  • Visible anatomical artifacts, particularly on the fox's front right leg which appears disjointed.
  • Slightly more 'AI-plastic' look to the fur compared to the realism of Flux.

Verdict: FLUX.2 [dev] produces a much higher quality, photorealistic image with superior lighting and texture, but it fails to capture the 'playful chasing' movement and incorrectly adds a second rabbit. GPT Image 2 captures the energy and specific animal count of the prompt perfectly, but the technical execution contains notable anatomical glitches and less realistic fur. FLUX.2 [dev] is the likely winner for its overall aesthetic professionalism and mastery of the '8K masterpiece' requirement.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent minimalist vector style suitable for a modern logo.
  • + Perfect typography and spelling of all requested text.
  • + Clean, bold shapes that are highly scalable.
  • Minimalist approach results in very simple steam rendering.
  • The banner design is a bit clunky compared to the cloche.

GPT Image 2

  • + Beautiful sophisticated typography with elegant serifs.
  • + Excellent use of hatching and subtle texture for a premium vintage feel.
  • + Well-balanced composition with an ornate frame.
  • Less 'minimalist' than requested, leaning more towards detailed illustration.
  • The banner is at the bottom, separating it from the main brand name.

Verdict: Both models followed the prompt exceptionally well, with perfect text rendering. FLUX.2 [dev] delivered a truly minimalist vector emblem that is ready for logo use, while GPT Image 2 provided a much more ornate and detailed vintage illustration. FLUX.2 [dev] is the winner for better adhering to the 'minimalist' keyword while maintaining a professional logo aesthetic.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.2 [dev]
GPT Image 2

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent vector art style that feels clean and modern
  • + Creative inclusion of the crew names in the header
  • + Rich, saturated color palette that follows the prompt instructions
  • Confused layout with haphazard icon placement
  • Major text legibility issues and spelling errors (e.g., 'Trarquility')
  • Failed to follow the requested number and order of steps

GPT Image 2

  • + Perfect infographic layout with a logical progression of steps
  • + Accurate and legible text rendering throughout the image
  • + Authentic NASA/Apollo 11 mission patch and branding details
  • + Correct icons for every specific mission phase requested
  • Slightly more complex illustrative style than a strict 'flat-vector' request
  • The crew silhouettes are a bit generic

Verdict: GPT Image 2 is significantly better as a functional infographic, providing a clear, chronological sequence of the mission steps with perfect legibility and adherence to the prompt's structural requirements. FLUX.2 [dev] produces attractive individual vector assets but fails to organize them into a coherent or readable layout, and it struggles with basic text rendering.

Next steps

Explore each model