Head to head
Esc

Models · slot A

to navigate to pick

FLUX.2 [dev] Black Forest Labs GPT Image 1.5 OpenAI

Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.

FLUX.2 [dev]

24.5 arena score

#18 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Vote tally

Where the votes landed

FLUX.2 [dev]

0.0%

win rate

Ties

16.7%

GPT Image 1.5

83.3%

win rate

0.0% 16.7% ties 83.3%
Shared challenges 20

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.2 [dev]
GPT Image 1.5
0% wins 25% ties 75% wins

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent photographic quality with realistic soft lighting and depth of field.
  • + Accurate spatial arrangement of objects.
  • + High rendering quality of the glass textures and reflections.
  • The blue sphere appears to be floating inside the cube rather than resting on the bottom.

GPT Image 1.5

  • + Strong adherence to all prompt elements including the green plant visible through the glass.
  • + The glass cube has realistic beveling and an interesting mirrored base.
  • The lighting is a bit flat compared to Model A.
  • The blue sphere is quite large, pushing the definition of 'small sphere'.

Verdict: Both models followed the complex spatial instructions perfectly. FLUX.2 [dev] produced a more aesthetically pleasing, professional-grade photograph with superior lighting, though GPT Image 1.5 offered better clarity for the plant behind the glass. FLUX.2 [dev] is the winner due to its better overall visual composition and realism.

Man and Car in California

Editing
Edit instruction

“Make a photo of the man driving the car down the California coastline”

Source
FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent preservation of the subject's clothing and hairstyle from the source image.
  • + Accurate rendering of the Rolls-Royce interior details, including the wood grain dashboard.
  • + Realistic lighting and appropriate motion blur on the road.
  • The man's profile facial features are slightly altered compared to the original photo.

GPT Image 1.5

  • + Successfully captures the man's facial expression and likeness from the source image.
  • + Clear and iconic representation of the California coastline with palm trees and winding roads.
  • The car is rendered as a right-hand drive vehicle, which is incorrect for a California coastline setting.
  • Significant anatomical error with the hand placement on the steering wheel appearing disconnected from the body.
  • Loss of detail on the man's scarf and plaid coat compared to the source.

Verdict: FLUX.2 [dev] is the winner because it maintains high consistency with the source materials, perfectly preserving the man's specific outfit and the luxury car's interior. GPT Image 1.5 has better facial likeness but fails on logic by placing the driver on the right side of the car for an American road and exhibits distracting anatomical glitches with the hands.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent execution of motion blur from passing cars
  • + High realism in facial textures and clothing folds
  • + Successfully captures several technical prompts like 50mm feel and reflections
  • The structural anatomy of the bicycle is incoherent near the handlebars and seat
  • The man appears to be sitting on air or a floating seat

GPT Image 1.5

  • + Stronger composition with the subject grounded in a crouching position
  • + Great attention to detail with the toolkit and puddle reflections
  • + Bicycle geometry is much more realistic and logical
  • Lacks the requested motion blur from passing cars
  • The car in the background looks slightly static despite the rain

Verdict: FLUX.2 [dev] followed the technical camera prompts more closely, especially the motion blur of passing cars and the candid framing, but failed significantly on the physical logic of the bicycle. GPT Image 1.5 produced a much more coherent and grounded scene with better object details, though it missed the specific kinetic energy of the motion blur request. GPT Image 1.5 is preferred for its overall visual consistency and anatomical accuracy.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.2 [dev]
GPT Image 1.5
0% wins 0% ties 100% wins

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent depiction of ornate engraved plate armor with high-quality metal textures.
  • + Includes clearly visible beads in the braided hair as requested.
  • + Realistic lighting from the torch source reflecting on the metallic surfaces.
  • The facial skin texture is slightly smoother and less 'battle-worn' than Model B.
  • The beads on the braids look a bit like modern jewelry compared to the medieval setting.

GPT Image 1.5

  • + Incredible skin texture with realistic pores, grime, and gritty battle-worn detail.
  • + Stronger 'paladin' aesthetic with the cross insignia and rugged cloth underlayers.
  • + Exceptional eye detail and lifelike expression.
  • The beads in the hair are less prominent and look more like metallic rings/bands than beads.
  • A bit more lens flare/haze which slightly obscuring the fine engravings on the central armor piece.

Verdict: Both models followed the prompt exceptionally well, but GPT Image 1.5 wins due to the superior 'battle-worn' skin texture and the more thematic paladin aesthetics. While FLUX.2 [dev] did a better job with the specific request for 'beads' in the hair, GPT Image 1.5's overall realism, cinematic lighting, and detailed leather and cloth textures make for a more compelling portrait.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Strong aesthetic alignment with modern neon/vibrant design trends
  • + Excellent grid layout and spacing
  • + Consistent high-quality photography
  • Text is mostly gibberish despite having a professional font appearance
  • Randomly repeats 'Pizza' sections twice
  • Pricing values are unrealistic for casual dining

GPT Image 1.5

  • + Perfectly legible, relevant text for all items
  • + Accurate representation of food items mentioned in text
  • + Clean, functional minimalist layout
  • Visuals are a bit generic compared to the 'vibrant' requirement
  • Slightly less 'bold' in its design execution

Verdict: GPT Image 1.5 is the clear winner for a design task because it produces fully legible, professional text and coherent menu items that match the photos. While FLUX.2 [dev] has a more stylish, vibrant aesthetic, the text is nonsensical, making it unusable as an actual menu design.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent typography with a clear fire/glow effect and perfect spelling.
  • + Clean and professional composition with distinct, un-cluttered layers.
  • + Accurate rendering of the starburst element as requested.
  • The lighting on the burger feels a bit flat compared to the fiery background.
  • The sauce droplets look slightly static rather than dynamic.

GPT Image 1.5

  • + Outstanding photorealistic textures on the charred bun and juicy patty.
  • + High sense of motion with flying embers and dripping sauces.
  • + Dynamic lighting that realistically reflects the fire onto the food surfaces.
  • The 'LIMITED TIME ONLY' text is small and positioned at the very bottom, losing prominence.
  • The starburst design is less integrated and looks slightly more organic/jagged than a clean graphic element.

Verdict: Both models followed the prompt exceptionally well, but FLUX.2 [dev] produced a more balanced advertisement layout with superior typographic clarity. While GPT Image 1.5 offered more grit and better food textures, FLUX.2 [dev] felt more like a professional marketing piece with its clean execution of the glowing text and starburst.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent chalk texture including realistic smudges and strokes.
  • + Accurate text rendering for all requested items.
  • + Dynamic composition with varying font weights that look hand-drawn.
  • The word 'with' in the octopus item has an odd artifact/overlapping ghosting behind it.
  • The cursive for the title is more of a print-script hybrid rather than elegant cursive.

GPT Image 1.5

  • + Perfect adherence to all requested text including the truncated 'Brown But...' prompt.
  • + The 'TODAY'S SPECIALS' title features a more elegant, fluid cursive style as requested.
  • + Very clean and legible handwriting that maintains a consistent slant.
  • The chalk texture is a bit too uniform, lacking the heavy-stroke realism seen in the other model.
  • The composition is a bit static with very even spacing between all lines.

Verdict: Both models followed the complex text instructions perfectly. FLUX.2 [dev] excels in capturing the physical medium of chalk, providing a more authentic 'messy' café feel, though it suffered from a minor visual artifact on the second menu item. GPT Image 1.5 followed the stylistic request for 'elegant cursive' much better and resulted in a cleaner, more readable image, making it the slight favorite for overall aesthetic polish.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Successfully applied the character's clothing and accessories.
  • + Matches the specific yellow studio background and red ottoman.
  • Horrific anatomical failure with a second head and arm growing from the chest.
  • The character is floating above the ottoman instead of standing on it.

GPT Image 1.5

  • + Excellent anatomical coherence and character consistency.
  • + Flawless integration of the scarf and sunglasses into the new pose.
  • + Correctly places the character's feet on the ottoman with natural weight distribution.
  • Misses the extreme lean/tilt of the torso present in the pose reference.
  • The hair is slightly simplified compared to the source material.

Verdict: FLUX.2 [dev] suffered a catastrophic failure in image composition, generating a multi-headed anatomical nightmare. GPT Image 1.5 successfully transferred the character into the new pose with high visual quality and perfect logic, despite slightly softening the extreme angle of the reference pose.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent visual clarity and clean composition
  • + Cinematic lighting with a clear horizon line
  • + Realistic rendering of the space suit and horse textures
  • Failed the negative constraint; the astronaut is riding the horse, not vice-versa

GPT Image 1.5

  • + Dynamic action with dust and high detail in the lunar landscape
  • + Strong aesthetic with many cosmic elements
  • + Distinctive cinematic style
  • Failed the negative constraint; the astronaut is riding the horse, not vice-versa
  • Anatomy of the horse's front legs is slightly distorted

Verdict: Both FLUX.2 [dev] and GPT Image 1.5 failed the specific logic test in the prompt, which requested the horse to be on top of the astronaut ('horse on top, not vice versa'). Both models defaulted to the standard trope of an astronaut riding a horse. FLUX.2 [dev] is slightly preferred for its superior image quality and cleaner lighting, though both failed the core creative challenge.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Successfully replicates the specific scarf pattern, coat style, and sunglasses.
  • + Integrates the lighting and shadows appropriately for the outdoor beach setting.
  • + Applies the requested clothing layers and accessories such as the belt and jewelry.
  • Fails to preserve the identity of the person in Image 1, instead merging their features with the person from Image 2.
  • Does not maintain the exact face and hair of the original subject as instructed.

GPT Image 1.5

  • + Successfully maintains the clothing style, including the coat and plaid scarf.
  • + Correctly identifies and keeps the lower half of the person's face/skin tone closer to Image 1.
  • The image is awkwardly cropped, cutting off the person's head and failing to show the face/hair preservation.
  • The hands and pockets exhibit significant anatomical artifacts and lack of detail.
  • Misses key accessories like the sunglasses requested from the reference.

Verdict: Both models struggled with the complex instruction of preserving identity while swapping outfits. FLUX.2 [dev] produced a much higher quality image with better clothing detail, but it failed the core instruction of keeping the person's face identical to Image 1, instead creating a hybrid person. GPT Image 1.5 failed the formatting by cropping out the head entirely and produced poor details on the hands, though it captured the lower face tone better.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent fur texture and lighting on the capybara's face.
  • + The businesswoman in the background has a perfect 'bored' expression as requested.
  • + Very clean, high-resolution rendering with cinematic bokeh.
  • The capybara's paws look somewhat human-like and uncanny.
  • The camera angle is slightly outside the car, whereas the prompt suggested a scene 'inside'.

GPT Image 1.5

  • + Correctly captures a more natural paws-on-wheel pose for a capybara.
  • + The camera placement feels more like a passenger's perspective or a dashboard cam inside the car.
  • + The hat design with 'TAXI' and checkerboard pattern is very thematic and accurate.
  • Slightly noisier image quality compared to Model A.
  • The businesswoman's face is quite blurry and lacks the distinct bored detail of the other model.

Verdict: Both models followed the prompt exceptionally well, capturing the surreal humor of a capybara taxi driver. FLUX.2 [dev] provides a sharper, more polished image with a better-realized passenger, while GPT Image 1.5 offers a more immersive internal perspective and better paw anatomy. FLUX.2 [dev] is the winner due to the higher visual fidelity and the passenger's expression, which sells the 'normalcy' of the prompt more effectively.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent text legibility and clean font choices.
  • + Well-defined thorns and spiderweb border elements.
  • + Balanced composition with high contrast.
  • The parchment effect is a bit simple.
  • The lighting on the pumpkin feels a bit flat compared to the background.

GPT Image 1.5

  • + Beautifully textured vintage parchment aesthetic.
  • + Atmospheric cinematic lighting with a glowing moon and embers.
  • + Highly detailed and intricate border and background elements.
  • Text layout at the bottom is slightly less organized than Model A.
  • The 'You are invited' scroll has some minor graphical glitches on the right edge.

Verdict: Both models followed the prompt exceptionally well, producing high-quality invitations with accurate text. FLUX.2 [dev] offers a cleaner, more readable design, while GPT Image 1.5 excels at creating a moody, atmospheric, and highly textured vintage aesthetic that feels more cinematic.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
FLUX.2 [dev]
Before After
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent preservation of the original face, clothing, and background.
  • + Highly realistic hair texture and detailed individual strands.
  • + Seamless blending with the original lighting and shadow on the face.
  • The hair volume is quite extreme (large afro), which may be more than requested for a 'natural' look.
  • The color of the hair is significantly darker than the beard, creating a slight mismatch.

GPT Image 1.5

  • + Natural and realistic volume and hairstyle for the character.
  • + Good hair color matching with the existing beard.
  • + Excellent preservation of the source image identity and environment.
  • Slight blurring/smudging on the forehead where the new hairline meets the skin.
  • Some subtle changes to the shape of the glasses frame compared to the original.

Verdict: Both models performed exceptionally well in preserving the source image. FLUX.2 [dev] provides better fine-detail texture and lighting integration but opted for a very large afro style, whereas GPT Image 1.5 provided a more conventional hairstyle that matches the beard color perfectly, though it has minor artifacts along the hairline.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent adherence to the 'minimal' requirement
  • + Clean, soft textures that match the 'cartoon scene' style perfectly
  • + Flawless text rendering and alignment
  • Lighting is a bit flat compared to the other model

GPT Image 1.5

  • + Rich, realistic material textures for the wood and ceramics
  • + Dynamic composition with extra details like the teapot and soy sauce
  • + Good use of shadows and highlights to create depth
  • Failed the 'minimal garnish and plate' instruction by adding many extra objects
  • Small artifacts in the text rendering, specifically the 'S' in SUSHI
  • The diorama base is more complex than requested

Verdict: FLUX.2 [dev] followed the prompt more accurately, delivering a clean, minimal, and perfectly centered composition that matches the requested 3D cartoon aesthetic. GPT Image 1.5 produced a higher level of detail and texture, but it ignored the instruction for a minimal scene by adding a teapot, soy sauce, and multiple tiers to the base.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent caricature style with exaggerated proportions.
  • + Effectively combines all three themes (anchor, dogs, hockey) into a single cohesive scene.
  • + Maintains recognizable facial features from the source image.
  • Garbled and unreadable text on the anchor desk.
  • The background characters are a bit cluttered and distracting.

GPT Image 1.5

  • + High-quality digital illustration style with vibrant colors.
  • + Cleverly includes a dog in a hockey helmet and holding a stick.
  • + Clean and readable text on the news ticker.
  • The facial likeness is significantly less accurate to the source than Model A.
  • The caricature feels slightly more generic and less 'exaggerated' in the classical sense.

Verdict: Both models followed the instructions well, but FLUX.2 [dev] managed to maintain a much stronger facial likeness to the source woman while embracing a more traditional caricature art style. GPT Image 1.5 produced a polished image with fun details like the dog in a helmet, but the subject's face looks like a generic character rather than the person in the photo.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent realism in fur textures and backlighting.
  • + Correctly includes all requested animals (including a second bunny).
  • + Balanced composition with clear, sharp details on the butterflies and flowers.
  • The animals are mostly sitting rather than 'playfully chasing' or 'tumbling'.
  • The kitten's eye anatomy is slightly off (one pupil is larger/distorted).

GPT Image 1.5

  • + Better captures the 'tumbling' and 'playful chasing' aspect of the prompt.
  • + Highly expressive facial expressions that fit the 'joyful vibe'.
  • + Strong use of 'god rays' and bokeh effects to create a magical atmosphere.
  • Anatomical errors such as the kitten having five visible paws/limbs/appendages.
  • The fox's front right paw is an indistinct dark mass that looks disconnected.
  • The dog's left ear merges awkwardly into its body.

Verdict: While GPT Image 1.5 does a much better job of capturing the active 'tumbling' motion and joyful energy requested in the prompt, it suffers from significant anatomical errors, particularly with the kitten's limbs. FLUX.2 [dev] produces a much cleaner, more physically coherent image with superior texture rendering, even if the composition is more static than requested.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent preservation of the character poses and clothing patterns.
  • + Captures the iconic Studio Ghibli fine-line art style very effectively.
  • + Cleverly transforms the urban background into a beautiful floral meadow consistent with the Ghibli aesthetic.
  • The color of the red dress is slightly washed out compared to the original.
  • The man's facial expression is a bit more neutral than the original 'distracted' look.

GPT Image 1.5

  • + Successfully captures the warm, nostalgic, and glowy lighting requested.
  • + Preserves the urban street setting while giving it a dreamy, painterly texture.
  • + Maintains the emotional intensity of the facial expressions from the source meme.
  • The woman in the foreground is excessiveley blurry compared to the source.
  • The art style leans more toward modern digital anime than the specific hand-painted Ghibli look.

Verdict: Both models do an excellent job of translating the 'Distracted Boyfriend' meme into an illustrative style. FLUX.2 [dev] produces a much more accurate Studio Ghibli aesthetic with its clean line work and watercolor-esque backgrounds, whereas GPT Image 1.5 creates a more generic, glowy anime style. FLUX.2 [dev] is the winner for better adhering to the specific stylistic request while maintaining the structural integrity of the original image.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
FLUX.2 [dev]
Before After
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent preservation of the subject's face and clothing from the original image.
  • + The blowing hair effect looks more realistic and follows the direction of the wind.
  • + Maintains the background scenery and characters perfectly.

GPT Image 1.5

  • + Successfully adds the requested blowing hair and flying leaves.
  • + The leaves have a pleasant golden autumn color palette.
  • + Preserves the overall composition of the original photo well.
  • Noticeable distortion of the dog's left leg and paw compared to the source.
  • Hair segments appears a bit jagged and less naturally integrated than Image A.
  • The leash handle in the woman's hand is slightly altered/muddled.

Verdict: Both models followed the instructions well, but FLUX.2 [dev] is the winner due to its superior preservation of the source image's details. While GPT Image 1.5 added nice colors to the leaves, it introduced anatomical errors in the dog's legs and slightly altered the woman's hand, whereas FLUX.2 [dev] kept everything intact while effectively adding the motion effects.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Excellent vector emblem style with clean lines.
  • + Accurate text rendering and banner layout.
  • + Perfectly adheres to the light background with subtle texture requirement.

GPT Image 1.5

  • + Sophisticated vintage typography with custom flourishes.
  • + Detailed shading and highlighting on the cloche dome.
  • + High contrast aesthetic.
  • Failed the light background requirement by using a black background.
  • The steam looks a bit thick and less 'minimalist' than requested.
  • Slightly less like a vector logo and more like a detailed digital illustration.

Verdict: FLUX.2 [dev] followed the prompt more accurately, particularly regarding the light background and minimalist vector style. While GPT Image 1.5 produced very elegant typography, it completely ignored the instruction for a light background with subtle texture, making FLUX.2 [dev] the more successful output for the specific request.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.2 [dev]
GPT Image 1.5

AI Judge Analysis

FLUX.2 [dev]

  • + Includes all nine requested icons including the crew names.
  • + Displays a very high level of illustrative detail, especially on the Lunar Module.
  • The sequence of steps is completely scrambled and nonlinear.
  • Significant text rendering errors and gibberish labels for several icons.

GPT Image 1.5

  • + Strictly follows the requested 6-step sequence in the correct chronological order.
  • + Perfect text rendering for all mission phases and crew names.
  • + Excellent adherence to the 'flat-vector' style and NASA color palette.
  • The Saturn V rocket illustration is slightly simplified compared to the prompt's potential for detail.
  • The lunar module design is consistent but less complex than Model A's version.

Verdict: GPT Image 1.5 is the clear winner because it correctly follows the logical sequence of the infographic steps requested in the prompt, whereas FLUX.2 [dev] places them in a random, non-chronological order. GPT Image 1.5 also produces clean, legible text for all labels, while FLUX.2 [dev] suffers from significant spelling artifacts and garbled text.

Next steps

Explore each model