Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
LongCat-Image
#62 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
100.0%
win rate
Ties
0.0%
LongCat-Image
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photographic lighting with realistic window shadows on the table
- + High details on the blue sphere's texture
- + Perfect adherence to spatial instructions including the plant visibility
- − Nonsense text on the book spine
- − The plant pot is merged oddly with the reflection inside the cube
LongCat-Image
- + Natural and clean photographic style
- + Great rendering of the blue glass sphere's transparency
- + Accurate representation of all requested elements
- − Lighting is somewhat flat compared to the dramatic shadows in the other image
- − The cube edges have a Slight cyan tint that feels slightly digital
Verdict: Both models followed the complex spatial instructions perfectly. FLUX.1 Kontext [max] produced a more atmospheric image with impressive shadow work, while LongCat-Image produced a cleaner, more realistic blue sphere. FLUX.1 Kontext [max] is the slightly stronger choice for its superior lighting and texture detail.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent skin and fabric texture that looks very realistic.
- + Captures the interaction with the bicycle components in a much more believable and detailed way.
- + Effective use of rain droplets and reflections to create atmosphere.
- − The man's heritage appears somewhat ambiguous compared to the specific request for a Japanese man.
- − The rain effect manifests as vertical streaks that look slightly digital/uniform.
LongCat-Image
- + Stronger adherence to the 'Japanese man' prompt requirement.
- + Good inclusion of motion blur from passing cars in the background as requested.
- + Wide framing creates a nice sense of place.
- − The red bicycle is physically impossible, with a third wheel and nonsensical frame geometry.
- − Lower overall realism in skin textures and lighting compared to the competitor.
- − The man's connection to the ground and the bike feels slightly 'pasted in'.
Verdict: FLUX.1 Kontext [max] produces a much higher quality, more realistic image with superior textures and coherent object physics, despite the subject's features being less distinctly Japanese. LongCat-Image follows the subject description and background action more closely but fails significantly on the structural integrity of the bicycle and fine details like hand-object interaction.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Exceptional skin texture and hyper-realistic facial details
- + Accurately captures the 'close portrait' framing
- + Masterful use of lighting and bokeh sparks for atmosphere
- − Missed the request for small beads in the braided hair
- − Scars are very faint, making the 'battle-worn' look less obvious than in Model B
LongCat-Image
- + Perfect adherence to the 'beads in hair' and 'battle-worn' scar details
- + Excellent rendering of ornate engraved plate armor and leather straps
- + Strong composition showing the full gear setup
- − Composition is a medium shot rather than the requested 'close portrait'
- − Facial features look slightly more airbrushed/artificial compared to Model A
Verdict: FLUX.1 Kontext [max] provides a much more intimate and technically superior close-up portrait with stunning realism in the eyes and skin, though it missed the bead detail. LongCat-Image followed the prompt instructions more literally regarding specific accessories and scars, but the framing was further back than requested and the textures lack the same lifelike grit.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional layout that looks like a real menu.
- + High-quality, consistent food photography that fits a single restaurant theme.
- + Clean use of white space and readable (though gibberish) typography.
- − The font choice for section headers is script rather than the requested bold sans-serif.
- − The image grid focuses almost entirely on pizza, lacking variety for appetizers and mains.
LongCat-Image
- + Successfully used bold sans-serif fonts in vibrant colors as requested.
- + Good variety of food types represented in the photos including salads, burgers, and pizza.
- − The graphic design is cluttered and appears amateurish compared to Model A.
- − Technical artifacts in the text rendering make it look distorted and messy.
- − The 'vibrant' accents are a bit overwhelming and clash with the minimalist prompt.
Verdict: FLUX.1 Kontext [max] produces a much more believable and professional menu design with high-quality photography, whereas LongCat-Image's output feels like a rough, distorted draft with poor alignment. While LongCat-Image adhered better to the 'bold sans-serif' and 'vibrant' keywords, the overall aesthetic quality and realistic composition of FLUX.1 Kontext [max] make it the superior choice.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography style that feels integrated into the scene's lighting
- + Clean, high-resolution rendering of food textures
- + Dynamic debris and embers create a strong sense of motion
- − The burger is not truly 'exploded' as requested, remaining largely assembled
- − Missed the 'starburst' container for the price
- − Strange cross-section of buns floating on the sides
LongCat-Image
- + Accurately included the starburst element for the price and text
- + Higher energy 'fiery' effect on the typography
- + Good use of embers and charcoal in the foreground to establish the theme
- − The burger is fully assembled rather than a deconstructed/exploded view
- − Shadowing on the lettuce and tomatoes looks slightly flat compared to the lighting
- − The 'starburst' shape is a bit generic and clip-art-like
Verdict: Both models failed to deliver a truly 'exploded' burger, but FLUX.1 Kontext [max] produced a more professional and photorealistic ad layout with sophisticated lighting. LongCat-Image adhered better to the specific layout instructions like the starburst, but the overall composition feels more like a digital collage than a cohesive photograph.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text accuracy with perfect spelling for all requested menu items.
- + Realistic chalk texture with convincing dust smudges on the board.
- + Strong adherence to the requested tone and date.
- − The 'elegant cursive' requested for the title is rendered as simple print caps instead.
LongCat-Image
- + Great environmental background showing a cozy café interior.
- + Chalk texture on the board and frame is very realistic.
- − Severely garbled text and numerous spelling errors throughout the board.
- − Fails to render the specific requested menu items correctly.
- − Layout of prices and items is disorganized and messy.
Verdict: FLUX.1 Kontext [max] is the clear winner as it successfully rendered every word of the complex prompt with perfect spelling and a convincing chalk aesthetic. LongCat-Image failed significantly on prompt adherence, producing illegible or misspelled text that barely resembled the requested menu items.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent anatomical rendering of the horse and astronaut suit.
- + High-quality cinematic lighting and depth of field.
- + Very realistic textures on the horse's fur and the suit materials.
- − Failed the specific spatial instruction for the horse to be on top of the astronaut.
LongCat-Image
- + Includes interesting background elements like space stations and extra planets.
- + Decent composition with a clear horizon line.
- + Good representation of a surreal space environment.
- − Failed the specific spatial instruction for the horse to be on top of the astronaut.
- − Noticeable anatomy issues with the horse's front legs.
- − Poor blending and artifacting, especially around the aircraft in the upper left.
Verdict: Both FLUX.1 Kontext [max] and LongCat-Image failed the negative constraint/spatial instruction to place the horse on top of the astronaut. However, FLUX.1 Kontext [max] is the superior image due to its professional cinematic quality, realistic lighting, and correct anatomy, whereas LongCat-Image contains several visual glitches and awkward leg positioning.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the capybara fur
- + High quality cinematic depth of field
- + Natural lighting and shadows within the car interior
- − The passenger is on a phone call instead of looking at her phone as requested
- − Only one paw is clearly visible on the steering wheel
LongCat-Image
- + Follows the instruction for 'both front paws on the steering wheel' more accurately
- + Shows the passenger looking at her phone as requested
- + Wider composition allows more context of the Manhattan street
- − The passenger is duplicated/cloned in the back seat
- − The capybara's paw/hand has an uncanny, almost human-like structure with too many claws
- − The taxi sign on the roof contains garbled, nonsensical text
Verdict: FLUX.1 Kontext [max] produces a significantly more realistic image with superior lighting and texture, although it misses slight details like the passenger's specific action. LongCat-Image follows more of the prompt's structural requirements but suffers from a major logical error by duplicating the passenger and creating a distorted, eerie hand for the capybara.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent font choice following the gothic theme.
- + Near-perfect text rendering and spelling for both primary and secondary text.
- + Very strong atmospheric lighting and cohesive dark parchment aesthetic.
- − Redundant text at the bottom repeating 'The Arches, NYC'.
- − Uses commas instead of periods for the date format.
LongCat-Image
- + Successfully includes the requested thorns in the border.
- + Good use of the scroll banner element as requested.
- + Vibrant, clean illustration style.
- − Multiple spelling errors including '7mm', 'Armiees', and '7um'.
- − Text layout at the bottom is cluttered and contains gibberish words.
- − Border thorns look like a separate overlay rather than integrated art.
Verdict: FLUX.1 Kontext [max] produced a much more professional and thematic invitation with superior text legibility and atmosphere. While LongCat-Image captured more specific prompt elements like the thorns and scroll shape, its failure to render the location and time correctly makes it unusable as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent text rendering with a stylized 3D bubbly font that matches the scene.
- + Clean isometric composition with smooth, refined textures.
- + Included extra details like chopsticks and a soy sauce bowl that enhance the scene.
- − Failed to include the requested small flag icon.
- − The 'JAPAN' text is slightly off-center compared to the 'SUSHI' text.
LongCat-Image
- + Successfully included all elements including the small flag icon.
- + Higher contrast in lighting creates a more dynamic 3D feel.
- + Better vertical alignment of the text elements.
- − The textures on the fish look a bit more like plastic/clay than 'realistic PBR' materials.
- − Visible artifacts/blurriness around the flagpole and base of the sushi.
Verdict: FLUX.1 Kontext [max] produced a cleaner, more professional-looking image with superior 3D typography that fits the miniature aesthetic, despite missing the flag icon. LongCat-Image followed every prompt instruction including the flag, but the visual quality was slightly lower due to softer edges and less refined textures on the food.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Successfully included all four distinct animals (puppy, kitten, bunny, and fox kit).
- + Excellent soft lighting and atmospheric god rays that feel integrated into the scene.
- + High level of detail on the fur textures and consistent anatomy for all creatures.
- − The animals are sitting relatively still rather than actively 'tumbling' as requested.
LongCat-Image
- + Captures a more active 'chasing' and 'tumbling' motion in the puppy’s pose.
- + Vibrant butterfly details with clear patterns.
- − Failed to include the bunny as a separate animal, instead merging it with the kitten to create a 'bunny-cat' hybrid.
- − The lighting on the animals feels slightly artificial and overlayed compared to the background blossom.
- − Anatomical issues where the cat has rabbit ears.
Verdict: FLUX.1 Kontext [max] is the clear winner as it correctly rendered all four requested animals with high anatomical accuracy and beautiful atmospheric lighting. LongCat-Image failed a key part of the prompt by fusing the kitten and the bunny into a single creature with cat eyes and rabbit ears, which detracts from the realism and adherence.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with correct accents and spacing
- + Clean vector-style emblem execution
- + Accurate representation of a cloche dome
- − The steam element is very simple and minimal compared to the rest of the logo
LongCat-Image
- + Dynamic steam effects and sunburst design
- + Authentic vintage paper texture in the background
- − Confusing text repetition with 'Caffè' appearing twice
- − Text layout feels cluttered and overlaps the graphic elements
- − The cloche dome is poorly defined and looks morphed with the text
Verdict: FLUX.1 Kontext [max] produced a high-quality, professional logo that perfectly followed the prompt's requirements for a minimalist vector emblem. In contrast, LongCat-Image struggled with composition, resulting in repetitive text and a messy layout that lacked the clean aesthetic of a real logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Captures the requested flat-vector aesthetic with a cohesive color palette.
- + Includes a complex trajectory diagram that visually represents the mission steps.
- + Displays high-quality illustrations of the Saturn V and the Lunar Module.
- − Several labels are severely misspelled or nonsensical (e.g., 'Lunear Modulle', 'Ununy Modulle').
- − The Earth is mislabeled as 'Moon'.
- − Includes a random planet with rings (Saturn/Jupiter) which does not belong in an Apollo 11 Earth-Moon diagram.
LongCat-Image
- + Features a very clean, structured infographic layout with clear sections.
- + Strong adherence to the 'NASA-inspired' and modern vector style.
- + Iconography is sharp and consistent across the vertical layout.
- − Text is largely illegible or gibberish, even for the main title.
- − Does not follow the 6-step chronological sequence requested in the prompt.
- − The 'Saturn V' icon looks more like a generic shuttle or multiple rockets.
Verdict: FLUX.1 Kontext [max] does a better job of mapping out the mission steps in a single cohesive diagram, although its text and geographic labels are riddled with errors. LongCat-Image provides a cleaner, more professional infographic layout, but fails to follow the specific 6-step instructions and has worse text rendering. FLUX.1 Kontext [max] is the preferred choice for its more accurate depiction of the mission steps and craft, despite the labeling issues.
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images