Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
LongCat-Image
#62 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0%
win rate
Ties
0%
LongCat-Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to lighting instructions with a clear soft glow from the left.
- + The plant is highly visible through the glass as requested.
- + High visual fidelity on book textures and wood grain.
- − The glass cube lacks a physical bottom edge, making it appear more like a glass hood or open-bottom box.
- − The reflection on the bottom surface is slightly too mirror-like for a wooden table.
LongCat-Image
- + Superior glass rendering with realistic thickness and corner bevels.
- + The sphere has a nice semi-transparent glass quality that complements the cube.
- + More natural integration of the objects onto the wooden surface.
- − The plant is less visible through the glass compared to Model A, appearing more behind the cube than through it.
- − The lighting is a bit flatter and less directional than requested.
Verdict: Both models followed the complex spatial instructions perfectly. FLUX.1 Kontext [dev] captured the lighting and transparency of the plant through the glass better, while LongCat-Image produced a more convincing physical glass cube with distinct thickness and weight. FLUX.1 Kontext [dev] is the winner for its superior adherence to the 'soft light from the left' and the clarity of the plant through the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Natural skin texture and realistic color palette
- + Clean rendering of the bicycle and man's clothing
- − Fails the prompt as the man is standing with the bike rather than repairing it
- − Background cars are static rather than having motion blur
- − Composition is too centered and lacks the 'imperfect framing' requested
LongCat-Image
- + Strong adherence to the 'repairing' action and 'candid' street photography style
- + Effective use of motion blur on the passing vehicle
- + Excellent interaction with the wet environment and reflections
- − Anatomical and structural issues with the bike, including a phantom third wheel
- − Low-resolution artifacts and noise on the background car
Verdict: LongCat-Image captured the prompt's requested action and atmosphere much better, depicting a man actually squatting to repair the bike with realistic motion blur and candid framing. However, FLUX.1 Kontext [dev] produced a much cleaner and higher quality image, even though it ignored the specific 'repairing' instruction and opted for a generic pose.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong cinematic lighting with warm torchlight highlights
- + Highly detailed engraving on the plate armor
- + Very lifelike eyes and skin texture
- − Failed to include the hair braids and beads
- − The 'battle-worn' elements like scars are very subtle and look more like paint
LongCat-Image
- + Excellent adherence to the 'braids with beads' requirement
- + Strong texture on the leather straps and mail underlayer
- + Clear depiction of scars and battle-worn skin
- − The sparks look a bit like digital streaks rather than natural bokeh
- − The armor engraving is slightly less defined compared to the other model
Verdict: While FLUX.1 Kontext [dev] produced a more cinematically lit and intense portrait, it completely missed the requirement for braided hair and beads. LongCat-Image followed every aspect of the prompt, including the complex hair instructions and detailed layers of leather and mail armor, while still maintaining high visual quality.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Features a highly organized grid layout that aligns with modern minimalist aesthetics.
- + The imagery is sharp and high-resolution with realistic food textures.
- + Excellent use of white space and bold typography for readability.
- − Text consists of gibberish characters and some graphical artifacts in the font.
- − The food items are somewhat abstract and hard to identify as specific menu items.
LongCat-Image
- + Incorporates vibrant color accents as requested in the prompt.
- + Recognizable food icons like pizza are included to match the prompt context.
- + Layout is busy but informative with various pricing and description placeholders.
- − The font is stylized but lacks the requested bold sans-serif clean look.
- − Image quality is lower with softer edges and visible AI distortions around the text.
Verdict: FLUX.1 Kontext [dev] delivers a superior professional layout that truly feels like a modern minimalist menu, despite the nonsensical text. LongCat-Image attempts more of the prompt's specific culinary categories but suffers from poor resolution and a cluttered aesthetic that misses the 'clean' requirement.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering with clean, readable fonts
- + Warm, appealing color palette that fits the fire theme
- + Perfectly includes all required text elements
- − Missed the 'exploded' instructions; the burger is fully assembled
- − Spelling error in 'ONLY' (rendered as LNHLY)
LongCat-Image
- + High-quality photorealistic textures on the patty and bun
- + Excellent fiery/glowing effect on the letters
- + Perfect spelling for all requested text components
- − Missed the 'exploded' or 'suspended components' instruction as the burger is mostly assembled
- − The starburst graphic is a bit cluttered with too much text
Verdict: Both models failed to create the 'exploded' view where components are separated in mid-air, but both successfully captured the fiery aesthetic. LongCat-Image is the winner because it successfully spelled all text correctly and provided superior photorealistic textures, whereas FLUX.1 Kontext [dev] had a significant spelling error in the primary advertising text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent chalk texture and realistic handwriting style.
- + High degree of text legibility and accuracy compared to the prompt.
- + Superior composition with a clean, centered layout that mimics a real chalkboard.
- − Minor spelling errors and repetitions like 'Mushroom Mashroom' and 'with with'.
- − The date 'APRIL' is poorly rendered as 'HIVRIII'.
LongCat-Image
- + Natural background environment of a cozy café adds context.
- + Effective use of chalk dust and texture on the board edges.
- − Significant spelling failures and garbled text across all menu items.
- − Poor layout and spacing, making the board look cluttered and unorganized.
- − Failed to render the full title correctly.
Verdict: FLUX.1 Kontext [dev] is the clear winner as it successfully rendered most of the requested text with a highly realistic chalk texture, despite a few minor spelling glitches. LongCat-Image captured the café environment well but failed significantly on text legibility, producing mostly gibberish that did not follow the prompt's specific menu requirements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent adherence to the 'horse on top' spatial instruction
- + High quality rendering of the astronaut's face and suit textures
- + Clean, cinematic lighting that creates a focused surreal mood
- − The position of the horse's legs behind the astronaut is a bit anatomically ambiguous
- − The 'space' background is relatively sparse with simple spheres
LongCat-Image
- + Dynamic composition with a lot of environmental detail
- + Good rendering of the horse's mane and anatomy
- − Completely failed the negative constraint: horse is being ridden by the astronaut, not vice versa
- − Presence of strange artifacts like the floating plane and distorted satellites
- − Cluttered background distracts from the main subject
Verdict: FLUX.1 Kontext [dev] successfully interpreted the difficult spatial prompt, placing the horse in a position that suggests it is 'riding' the astronaut. LongCat-Image completely ignored the specific inversion requested in the prompt, providing a standard astronaut-on-horse image with several nonsensical background artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent close-up composition with high detail on fur and the driver's cap.
- + Very realistic lighting and dark cinematic tone appropriate for a New York taxi at night.
- + The passenger's bored expression perfectly captures the 'normalcy' requested in the prompt.
- − The capybara only has one hand on the wheel instead of both as requested.
- − The animal's face looks slightly more like a groundhog or specialized rodent rather than a classic capybara.
LongCat-Image
- + Successfully places the capybara's hands (paws) on the steering wheel as requested.
- + Broad street view provides good context with recognizable Manhattan-style lighting and blurred traffic.
- + The capybara's physical proportions and fur texture are highly accurate.
- − The passenger is duplicated/hallucinated, with two similar women appearing in the backseat.
- − The external angle through the door window makes it less of a 'scene inside' the taxi compared to Model A.
Verdict: FLUX.1 Kontext [dev] creates a much more atmospheric and high-quality image with a better focus on the passenger's expression, though it missed the specific detail about both paws on the wheel. LongCat-Image captured the 'both paws' instruction but suffered from significant prompt adherence issues, such as duplicating the passenger and ignoring the specific requested interior perspective. FLUX's superior lighting and photorealistic rendering make it the stronger overall choice.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Strong composition with a central focus
- + Includes the thorn border and twisted trees as requested
- + Clean rendering of the main headline text
- − Secondary and detail text is garbled and misspelt
- − Fails to include the 'parchment' texture, opting for a black background
LongCat-Image
- + Excellent adherence to the 'parchment' and 'gothic' aesthetic
- + Highly accurate text rendering for both the banner and the headline
- + Atmospheric background with moody sky and cinematic lighting
- − Bottom event details contain minor spelling errors and repetitions
- − The layout of the event details is slightly cluttered
Verdict: LongCat-Image is the clear winner as it successfully incorporates nearly all elements of the prompt, including the specific parchment texture and the moody background sky which FLUX.1 Kontext [dev] ignored. Furthermore, LongCat-Image's text legibility for the middle banner is significantly superior to the distorted text in FLUX.1 Kontext [dev].
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography with very clean, bold text
- + Smooth, minimalist aesthetic that fits the 'cartoon' request
- + Perfectly centered and clean composition
- − The flag icon is unrecognizable and lacks the Japanese motif
- − The sushi anatomy is slightly confusing with the black band wrapping high around the fish
LongCat-Image
- + Extremely accurate Japanese flag icon
- + Higher detail in the 3D model textures like the wood grain and rice grains
- + Better representation of sushi with more realistic variety (salmon and tuna)
- − Text alignment is slightly off-center compared to the central image
- − Small artifacts in the 'JAPAN' text edges
Verdict: While FLUX.1 Kontext [dev] produced cleaner typography and a more polished minimalist look, LongCat-Image followed the prompt's detail requirements more closely, providing a recognizable Japanese flag and significantly better sushi textures. LongCat-Image is the winner for its superior 3D material rendering and adherence to the specific iconography requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Dynamic sense of motion with the animals pouncing and running
- + Consistent lighting and soft bokeh effect
- + Anatomically coherent animals for the most part
- − Failed to include the fox and the bunny
- − Includes two kittens and a puppy instead of the four distinct species requested
- − Fur detail is slightly smoothed over compared to the 'ultra-detailed' prompt requirement
LongCat-Image
- + Successfully included all four animals: golden retriever, tabby kitten, bunny, and fox
- + Excellent 'god rays' lighting and dew sparkle effects
- + High texture detail in the fur and butterfly wings
- − Odd anatomical merging where the kitten has bunny ears
- − The composition feels a bit cramped and static compared to the 'tumbling' prompt
- − Floating water droplets/bokeh look slightly artificial
Verdict: LongCat-Image adheres much better to the prompt by attempting to include all four requested species, whereas FLUX.1 Kontext [dev] only generated three animals and missed two of the requested types. While LongCat-Image has a strange anatomical error where it combined the cat and bunny features, its superior lighting, detail, and prompt adherence make it the more successful image for this specific task.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography and spelling accuracy.
- + Very clean minimalist aesthetic suitable for a modern logo.
- + Perfect adherence to the vector emblem style requested.
- − The 'cloche' looks more like a dome or a cupcake wrapper than a traditional food cloche.
- − The 'Est. 1720' is plain text rather than the requested banner.
LongCat-Image
- + Excellent vintage texture and detailed line work.
- + Correctly implemented the banner for 'Est. 1720'.
- + Creative and dynamic composition with radiant lines.
- − Text is repetitive and messy, appearing to say 'Caffè Caffè Florian'.
- − Typography is inconsistent with different weights and scales layered over each other.
- − Lacks the 'minimalist' quality requested in the prompt.
Verdict: FLUX.1 Kontext [dev] produced a much more professional and usable logo with clean lines and perfect spelling, though it was perhaps too simple for some of the prompt's decorative requests. LongCat-Image captured the 'vintage' and 'banner' descriptors much better, but failed significantly on text legibility and the 'minimalist' constraint. FLUX.1 Kontext [dev] is the winner for providing a coherent, high-quality vector-style emblem.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Closer adherence to the requested color palette
- + Included more sequence steps as requested
- + Clean layout that feels more like an infographic
- − Poor spelling in the title (APOLO 11) and messy, illegible supporting text
- − Icons are abstract to the point of being unrecognizable for several steps
LongCat-Image
- + High-quality vector styling with clean line work
- + Excellent illustrative detail for the lunar module and Earth
- + Better overall visual appeal and balance for a poster
- − Failed to provide the specific 6-step sequence requested
- − Title text is nonsensical (Aaa o 11)
- − Includes space shuttle-like icons which are historically inaccurate for Apollo missions
Verdict: While FLUX.1 Kontext [dev] followed the instruction for a multi-step sequence more accurately, the individual icons are confusing and the text is extremely garbled. LongCat-Image produced a much more professional and aesthetically pleasing vector illustration, but it failed the core requirement of showing the specific mission progression steps and included anachronistic spacecraft. LongCat-Image is the slight favorite purely for its superior visual execution of the 'vector infographic' style.
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images