Black Forest Labs' distilled 9 billion parameter image generation model with sub-second inference and multi-reference support
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
FLUX.2 [klein] 9B
#13 of 32 in Image Editing
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [klein] 9B
40.0%
win rate
Ties
0.0%
GPT Image 2
60.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent handling of glass refraction and reflections
- + Realistic wooden table texture
- + Accurate soft lighting from the window side
- − The sphere appears to be floating unnaturally instead of resting on the bottom
- − The glass object looks more like a vase or jar than a perfect cube
GPT Image 2
- + Higher precision in geometric shapes, depicting a clear glass cube
- + Superior texture on the red book cover
- + Better composition with the plant clearly visible through the glass
- − The sphere's reflection on the bottom glass pane is a bit simplistic
Verdict: GPT Image 2 followed the prompt's geometric requirements more accurately, providing a distinct cube shape and a more detailed red book. FLUX.2 [klein] 9B excelled at realistic light refraction and material physics, but the main object felt more like a heavy glass container than a cube.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent shallow depth of field and realistic bokeh in the background.
- + Highly detailed, natural skin textures on the man's hands and face.
- + Strong rainy atmosphere with visible streaks and wet pavement reflections.
- − The bike anatomy is slightly off where the frame meets the handlebars.
- − The hand position is somewhat ambiguous regarding which part of the bike is being fixed.
GPT Image 2
- + Effective use of motion blur on the passing vehicle to show street activity.
- + Authentic Japanese street setting with appropriate signage and environmental details.
- + Good 'imperfect' framing that feels like a genuine candid street photo.
- − The image lacks the requested shallow depth of field, with the foreground and midground being equally sharp.
- − The skin texture and lighting look flatter and less photographic than Model A.
Verdict: FLUX.2 [klein] 9B captures the cinematic and photographic instructions much better, delivering a high-quality 50mm look with rich textures and atmosphere. GPT Image 2 succeeds in the 'imperfect framing' and 'motion blur' aspects of the prompt but fails to deliver the requested shallow depth of field and professional optical quality.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Exquisite engraving detail on the plate armor
- + Highly accurate rendering of leather and chainmail textures
- + Perfect execution of the braided hair with beads requirement
- − The scars look somewhat superficial and painted on
- − Slightly more 'digital' and less cinematic look than the competitor
GPT Image 2
- + Exceptional lifelike skin texture and realistic scarring/dirt
- + Superior cinematic lighting and atmosphere reflecting the torchlight
- + Very natural integrated hair braiding and beadwork
- − Armour engravings are slightly less prominent than Model A
- − The torch in the background is less distinct than Model A
Verdict: While both models followed the prompt excellently, GPT Image 2 creates a more believable, cinematic portrait with superior skin textures and lighting. FLUX.2 [klein] 9B excels in the mechanical and geometric details of the armor and braids, but the overall composition feels slightly more like a stiff character render compared to the lifelike depth of GPT Image 2.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Strong bold typography for headers
- + Clear minimalist white space
- + Vibrant colors in the food photography
- − Text is mostly gibberish or contains spelling errors like 'Resteuranbteteina'
- − Poor logical layout, showing pizza photos under 'Appetizers' and 'Mains' headers
- − Overlapping text elements and inconsistent pricing symbols
GPT Image 2
- + Perfect English text and high-quality legible typography
- + Excellent logical organization with photos corresponding to their category
- + High-quality professional graphic design with cohesive icons and branding
- − Slightly more cluttered than the minimal request due to footer icons
- − Small text for descriptions might be hard to read at lower resolutions
Verdict: GPT Image 2 is the clear winner as it produces a fully functional, professional-grade menu with legible descriptions and accurate food-to-category matching. FLUX.2 [klein] 9B fails significantly on text accuracy and displays pizza images under labels for appetizers and mains, making the design nonsensical.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent text legibility and clean graphic design layout.
- + Highly photorealistic textures on the bun and patty.
- + Well-integrated starburst graphic that matches the professional ad aesthetic.
- − The 'exploded' effect is safe and less dynamic than the competitor.
- − The secondary text lacks the requested fiery/glowing effect, appearing as flat white.
GPT Image 2
- + Highly dynamic 'exploded' composition with sauce droplets and flying ingredients.
- + Strong adherence to the fiery, glowing text effect throughout the image.
- + Complex layering of ingredients including onions and sauce that adds to the visual interest.
- − The texture of the meat patty looks slightly over-processed or crunchy rather than juicy.
- − The composition feels a bit crowded and chaotic compared to a professional advertisement.
Verdict: While FLUX.2 [klein] 9B produces a very clean and professional-looking advertisement with superior food textures, GPT Image 2 better followed the prompt's request for a dynamic 'exploded' look and applied the fiery effect to all text elements. GPT Image 2 is the winner because it captured the sense of motion and the specific stylistic requirements for the text more accurately.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent chalk texture and realistic smudge marks on the board.
- + Strong adherence to the 'cozy café' background setting.
- + Accurately rendered the specific date requested.
- − Numerous spelling errors including 'Riott', 'Risoto', 'Octoopus', and 'fress'.
- − Text layout is crowded with redundant price labels for the second item.
GPT Image 2
- + Perfect spelling for every item on the menu.
- + The handwriting style is much more consistent and realistic for a chalkboard.
- + Superior composition with balanced spacing and clean letterforms.
- − The lighting on the board is slightly uneven, though realistic for the setting.
- − The cursive title is a bit simpler than the 'elegant' descriptor might suggest.
Verdict: GPT Image 2 is the clear winner as it successfully rendered all text with perfect spelling and a very realistic chalk handwriting style. FLUX.2 [klein] 9B struggled significantly with spelling, introducing multiple typos and a cluttered layout despite having a good chalk texture.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Successfully replicates the exact pose and environment from Image 1.
- + Captures the character's facial features and clothing accurately.
- + Maintains the distinct scarf pattern and accessory details from Image 2.
- − The character's skin tone is significantly darker than in the source image.
- − The left hand is clenched in a fist rather than the open-palm pose from Image 1.
GPT Image 2
- + Excellent adherence to the pose and composition of Image 1.
- + High accuracy in recreating the facial features and skin tone of the character in Image 2.
- + Perfectly captures the clothing, sunglasses, and the specific hanging style of the scarf.
- − The fingers on the raised right hand are slightly poorly rendered.
Verdict: Both models followed the complex instructions exceptionally well by mapping the character from Image 2 onto the unique pose and environment of Image 1. FLUX.2 [klein] 9B performed well but altered the character's skin tone and missed the nuance of the hand position, whereas GPT Image 2 maintained much better consistency with the character's appearance and the specific hand modeling of the original pose.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent cinematic lighting and background detail
- + High-quality rendering of the cosmic environment
- + Clear, professional-looking composition
- − Failed the negative constraint: the astronaut is riding the horse instead of the horse on top
- − Included nonsensical AI-generated text at the bottom
GPT Image 2
- + Perfectly followed the unconventional request of having the horse on top of the astronaut
- + Highly detailed textures on the spacesuit and lunar surface
- + Creative interpretation of the harness and saddle setup
- − The astronaut's hands/gloves have an anatomical error with too many fingers
- − Less visually vibrant background compared to the competitor
Verdict: While FLUX.2 [klein] 9B produced a much more visually stunning and cinematic image, it completely failed to follow the specific 'horse on top' instruction. GPT Image 2 followed the complex prompt perfectly, creating a surreal and accurate representation of the requested scene despite some anatomical issues on the astronaut's hands.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent adherence to the requested accessories, including the gold necklaces and bracelets from Image 2.
- + Very high level of source preservation for the person's face, hair, and specific skin details.
- + Strong lighting matching with realistic highlights on the clothing that fit the beach environment.
- − Slightly altered the person's eye shape/expression compared to the original Image 1.
- − The added necklaces weren't explicitly central in Image 2 but were a good creative addition based on the prompt's request for 'accessories/jewelry'.
GPT Image 2
- + Perfect preservation of the person's face, eyes, and skin textures from Image 1.
- + Highly accurate recreation of the specific plaid pattern and texture of the scarf from Image 2.
- + Excellent integration of the coat's fit onto the subject's pose.
- − Missed the jewelry/accessories from Image 2 (watch, rings) mentioned in the prompt instructions.
- − The transition between the neck and the shirt collar is slightly less defined than in Image A.
Verdict: Both models performed exceptionally well at this complex image editing task. FLUX.2 [klein] 9B followed the instruction for 'all accessories and jewelry' more thoroughly by adding visible necklaces and bracelets, whereas GPT Image 2 captured the precise identity and facial expression of the subject in Image 1 with slightly more fidelity. FLUX.2 [klein] 9B is the winner for its more comprehensive adherence to the specific request for layers and jewelry while still maintaining high source preservation.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent photorealism in the capybara's fur and the woman's face.
- + Clear, sharp details in the car dashboard and city background.
- + Accurate representation of the requested boredom on the passenger's face.
- − Failed the spatial prompt as the businesswoman appears to be in the front passenger seat.
- − The text on the hat is gibberish.
- − The steering wheel grip looks slightly unnatural for the animal's paws.
GPT Image 2
- + Correct spatial placement of the passenger in the back seat.
- + Dynamic and cinematic lighting that feels more like a night scene in Manhattan.
- + Stronger adherence to composition by showing the distinction between front and back cabin.
- − The passenger's face is slightly blurry and lacks the detail seen in the foreground.
- − The perspective from outside the door makes the interior feel slightly cramped.
- − The capybara's paw looks more like a human hand in form than a paw.
Verdict: GPT Image 2 is the overall winner because it correctly placed the businesswoman in the back seat as requested, whereas FLUX.2 [klein] 9B placed her in the front passenger seat. While FLUX.2 had higher sharpness and better facial rendering for the woman, GPT Image 2 captured the requested scene structure and professional atmosphere much more effectively.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent typography legibility and spacing
- + Accurate execution of the thorn and web border
- + Clean layout that balances the central pumpkin with text elements
- − Visuals are a bit generic and feel like stock clip art
- − The scroll banner is slightly simple compared to the 'vintage' request
GPT Image 2
- + Strong 'vintage gothic' aesthetic with highly detailed textures
- + Creative inclusion of a spooky bridge and city skyline reflecting the NYC location
- + Atmospheric lighting and intricate border details
- − Text is slightly more cramped with some overlapping elements
- − The parchment looks heavily distressed, which might affect readability of smaller text
Verdict: Both models followed the prompt instructions perfectly, including all requested text and visual elements. FLUX.2 [klein] 9B produces a cleaner, more practical invitation layout, while GPT Image 2 creates a much more atmospheric and artistically detailed piece that better captures the 'vintage gothic' mood. GPT Image 2 is the winner for its superior creativity and thematic depth, especially the subtle inclusion of a gothic NYC skyline.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent adherence to the 'minimal' request with a clean, simple layout
- + Perfect rendering of the requested top-down 45-degree isometric angle
- + Soft, refined textures that match the '3D cartoon scene' description
- − The flag icon is incorrect, resembling the flag of Yemen rather than Japan
- − The salmon nigiri is placed awkwardly on top of a maki roll
GPT Image 2
- + Features a correct and well-designed Japanese flag icon
- + Higher detail in textures and PBR materials, especially for the fish and wooden board
- + Vibrant colors and a high-quality 3D rendered aesthetic
- − Failed the 'minimal' garnish prompt by including a stone lantern, foliage, and stones
- − The text 'JAPAN' is stylistic but less 'bold' than Model A's clear typography
Verdict: While GPT Image 2 has much better detail and a correct flag, it ignored the 'minimal' constraint by adding several unnecessary environment elements. FLUX.2 [klein] 9B followed the minimalist composition and 'small diorama' prompt perfectly, though it failed to generate the correct national flag. FLUX.2 is the winner for better prompt adherence regarding layout and simplicity.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Excellent fur texture and lighting on the fox and puppy.
- + Dynamic butterfly colors and sharp foreground flowers.
- + Consistent 'god rays' effects as requested.
- − Failed to include the baby bunny entirely.
- − Anatomy error where the puppy's paw appears to be growing out of or merged with the kitten's side.
- − Kitten has an artificial, wide-eyed look that borders on uncanny.
GPT Image 2
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox.
- + Natural poses that better reflect 'tumbling' and 'chasing' through the grass.
- + Superior anatomy with logical limb placement for all animals.
- − The fox's eyes appear slightly dark and less expressive than the other animals.
- − The kitten's tail is somewhat blurry and oddly shaped.
Verdict: GPT Image 2 is the clear winner because it followed the prompt instructions to include a baby bunny, which FLUX.2 [klein] 9B omitted. Additionally, GPT Image 2 avoids the significant anatomical merging error seen in FLUX.2's puppy and kitten, resulting in a much more coherent and realistic composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Clean vector-style execution
- + Excellent adherence to the minimalist aspect of the prompt
- + Perfectly legible text and accurate accent mark on 'Caffè'
- − The steam icon is a bit simplistic and less integrated into the art style
- − Composition feels slightly basic compared to professional branding
GPT Image 2
- + Beautiful vintage engraving style with rich crosshatching
- + Very sophisticated typography and ornate framing
- + Accurate text rendering for all requested elements
- − Less 'minimalist' than requested due to high level of detail
- − Slightly busier composition
Verdict: Both models followed the prompt instructions very well. FLUX.2 [klein] 9B followed the 'minimalist' instruction more closely with a modern vector aesthetic, but GPT Image 2 produced a much higher quality 'vintage' result that looks like an authentic historical emblem. GPT Image 2 is the winner for its superior artistic execution and sophisticated use of texture and line work.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [klein] 9B
- + Clean vector illustration style
- + Good use of the requested color palette
- − Numerous spelling errors including the main title
- − Nonsensical layout of steps that doesn't follow a logical order
- − Visual artifacts like the floating moon overlapping the lunar module
GPT Image 2
- + Perfect adherence to the 6-step logical flow requested
- + Excellent typography and text accuracy
- + Professional infographic layout with clear iconography
- − The rocket and lander icons are slightly more detailed than 'flat-vector' style
- − The Apollo 11 seal includes a generic eagle instead of the mission-specific design
Verdict: GPT Image 2 is significantly better, following the logical 6-step structure of the mission perfectly while maintaining high text legibility and accuracy. FLUX.2 [klein] 9B fails on both spelling and the sequential logic of the infographic, resulting in a confusing jumble of steps.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following