Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
LongCat-Image
#62 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
100.0%
win rate
Ties
0.0%
LongCat-Image
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent surface textures and realistic reflections on the table.
- + High image clarity and vibrant colors.
- + Accurately places the plant behind the objects as requested.
- − Failed the spatial logic by adding an extra blue sphere on top of the book which was not requested.
- − The sphere inside the cube appears to be floating without a contact point.
LongCat-Image
- + Perfect adherence to all spatial instructions, including placing the book on top of the cube.
- + Realistic light diffusion from the window on the left.
- + Naturalistic glass thickness and refractive qualities.
- − The plant in the background is slightly blurry, though this fits the depth of field.
- − Slightly less punchy colors compared to the other model.
Verdict: LongCat-Image is the clear winner as it followed every instruction perfectly, whereas FLUX.1 [schnell] hallucinated an additional blue sphere on top of the book. LongCat-Image also demonstrated better physical logic, showing the sphere resting on the bottom of the cube rather than hovering in the center.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent skin texture and elderly character detail.
- + Strong application of shallow depth of field.
- − Failed to include visible rain and motion blur for cars.
- − Bicycle structure has slight clipping on the rear gears.
LongCat-Image
- + Successfully captured visible rain and wet pavement reflections.
- + Accurate crouching pose makes the repair action feels authentic.
- − The bicycle frame has severe anatomical errors and extra wheels.
- − The motion blur on the cars looks like a static glow.
Verdict: FLUX.1 [schnell] creates a more realistic portrait with superior skin detail, though it misses the specific atmospheric cues like visible rain. LongCat-Image captures the mood and setting better but suffers from major structural failures in the bicycle geometry. FLUX.1 [schnell] is the better image due to higher coherence and fewer artifacts.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely high skin texture detail and realistic eyes
- + Excellent shallow depth of field with cinematic lighting
- + Intense, expressive facial performance
- − Missed the 'small beads' request in the hair
- − Armor is mostly out of frame, losing the 'ornate engraved' detail
- − Fails to include the 'bokeh sparks'
LongCat-Image
- + Strong adherence to all prompt elements including beads, sparks, and chainmail layers
- + Excellent rendering of ornate engraved plate armor
- + Good composition showing the leather straps and garment textures
- − Face looks slightly clean/plastic compared to the gritty prompt
- − The battle-worn scars look a bit like digital paint stamps
- − Skin texture lacks the realism of the competitor
Verdict: While FLUX.1 [schnell] produces a more technically impressive and lifelike portrait with superior skin textures, LongCat-Image followed the prompt much more accurately. LongCat-Image captured the specific details of the beads, the ornate armor engravings, and the bokeh sparks, whereas FLUX.1 [schnell] focused almost entirely on the face and ignored several key descriptors.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent typography rendering with almost perfectly legible text
- + Refined minimalist aesthetic with great negative space
- + Accurate sections for 'Pizza', 'Appetizers', and 'Mains'
- − The food items in the grid look a bit repetitive (multiple pizzas)
- − Lower resolution overall compared to Model B
LongCat-Image
- + High resolution and vibrant color palette
- + Dynamic layout with more varied food imagery
- + Stronger 'casual dining' vibe with bold graphic elements
- − Text is largely illegible gibberish
- − The 'grid' is messy and overlapping, lacking professional alignment
- − Fails the 'minimalist' and 'white background' requirement due to heavy blue/yellow blocks
Verdict: FLUX.1 [schnell] is the winner because it successfully followed the 'minimalist' and 'white background' instructions while producing remarkably coherent and legible text. While LongCat-Image has high-quality image assets, it failed the specific layout constraints and produced nonsensical typography that would require significant manual editing for actual use.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features a high level of detail in the burger texture.
- + Good use of embers and floating food particles to create motion.
- − Significant text errors including missing letters ('AGIC BURGER') and repeated prices.
- − Failed to create an 'exploded' burger, showing a mostly assembled one instead.
- − The starburst design is messy and lacks a glowing effect.
LongCat-Image
- + Excellent typography with perfect spelling and the requested fiery, glowing effect.
- + Followed all text integration instructions, including the starburst for the price.
- + High visual quality with a clear, professional ad composition.
- − Failed to provide the 'exploded' view where components are suspended separately.
- − The burger looks a bit too 'perfect' or plastic compared to the more raw textures of Model A.
Verdict: LongCat-Image is the clear winner due to its superior handling of text and graphic design elements, perfectly executing the complex typography requirements despite missing the 'exploded' layout. FLUX.1 [schnell] suffered from significant spelling errors and a cluttered footer that detracted from the professional quality of the advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + The text is relatively legible and follows the requested menu layout.
- + Good handheld marker/chalk aesthetic for the main body text.
- − Missed the 'elegant cursive' requirement for the title.
- − Significant spelling errors such as 'Pril', 'Taffle', and 'Mushmnctiom'.
- − The chalk texture is very clean, bordering on looking like a digital font.
LongCat-Image
- + Strong chalk texture with realistic dust and smudging on the board.
- + Better attempt at a stylized, cursive-like decorative title.
- + Excellent cozy café background composition with warm lighting and depth of field.
- − Severe spelling hallucinations and garbled characters in the text.
- − Failed to include the specific dish names requested in the prompt.
- − The layout is cluttered with strange character stacking in the title.
Verdict: FLUX.1 [schnell] follows the prompt instructions more closely regarding the specific text to be written, although it still suffers from spelling errors. LongCat-Image creates a much more atmospheric and visually convincing 'chalk' texture and café environment, but the text is largely illegible and ignores the specific menu items requested. FLUX.1 [schnell] is the winner for better prompt adherence despite its plainer aesthetic.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully placed the horse on top of the astronaut
- + Strong cinematic lighting and high-quality textures
- − Major anatomical errors with the horse having two heads
LongCat-Image
- + Clear composition and coherent rendering
- + Realistic space suit lighting and environment
- − Failed the negative constraint to put the horse on top of the person
Verdict: FLUX.1 [schnell] adhered to the difficult surreal constraint of a horse riding an astronaut, though it suffered from severe anatomical artifacts. LongCat-Image ignored the specific positioning instruction entirely and provided a standard astronaut-on-horse image, making FLUX.1 [schnell] the winner for prompt adherence.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Natural composition and lighting
- + Highly accurate capybara facial anatomy
- − Capybara only has one paw on the wheel
LongCat-Image
- + Excellent depiction of the taxi exterior and lights
- + Accurate placement of paws on the steering wheel
- − Duplicate passenger in the back seat
- − Capybara neck anatomy looks unnatural
Verdict: FLUX.1 [schnell] creates a more believable, cinematic scene with superior lighting and realistic animal fur textures. While LongCat-Image manages better adherence to the steering wheel prompt, the presence of an extra passenger and anatomical awkwardness makes it less refined.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Stronger atmospheric lighting and silhouettes.
- + Excellent central jack-o-lantern illustration.
- − Significant text repetition and spelling errors in the footer.
- − Lacks the parchment texture requested in the prompt.
LongCat-Image
- + Accurately depicts the requested parchment and thorn border.
- + Overall layout feels more like a physical invitation.
- − Main title text has some distorted character scaling.
- − The detail text for location and date is slightly garbled.
Verdict: FLUX.1 [schnell] creates a more cinematic and spooky atmosphere but fails significantly on text legibility and following the parchment material requirement. LongCat-Image adheres much better to the structural requirements of the prompt, including the thorn border and parchment style, despite slightly lower aesthetic cohesion.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean isometric perspective with a stylized diorama base.
- + Refined textures that give a high-clarity 3D render feel.
- − Missed the 'SUSHI' text requirement entirely.
- − White text on a very light background has poor contrast.
LongCat-Image
- + Included all requested text elements and flag icon correctly.
- + Excellent tactile, clay-like textures with vibrant colors.
- − Isometric angle is slightly flatter than the requested 45 degrees.
- − Garnish is more cluttered than the 'minimal' request specified.
Verdict: LongCat-Image is the superior choice because it followed all text-based instructions, whereas FLUX.1 [schnell] omitted the word 'SUSHI'. While FLUX.1 [schnell] had a cleaner 3D aesthetic, LongCat-Image captured the miniature cartoon style with much better clarity and completeness.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent soft lighting and atmosphere
- + Naturally blended group composition
- + Beautifully detailed butterflies and meadow
- − Failed to generate a distinct bunny, creating a cat-hybrid instead
- − Fox kit looks slightly like a plush toy
LongCat-Image
- + Strong 'god rays' lighting effect as requested
- + Includes all requested animal types clearly
- + Good focus and sharpness on the main subjects
- − The bunny is a bizarre cat-hybrid with bunny ears
- − Anatomical issues with the fox's front paw being a human-like hand
- − Floating dew drops look like artifacts
Verdict: Both models struggled with the complex prompt, specifically failing to render a distinct bunny and instead creating 'cat-bunnies'. FLUX.1 [schnell] produced a more aesthetically pleasing and coherent scene with superior fur texture, while LongCat-Image had more literal lighting but suffered from significant anatomical errors like the fox's human-like hand.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector aesthetic
- + Professional balanced composition
- + Clear and legible typography
- − Major spelling errors in the brand name ('Framilan') and date ('7720')
- − Missing the requested steam element
LongCat-Image
- + Perfect text accuracy including the name and establishment date
- + Includes the requested steam and banner elements
- + Excellent hand-drawn vintage texture
- − Slightly cluttered composition with redundant text ('Caffé Caffé')
- − Less 'minimal' than the vector style requested
Verdict: LongCat-Image followed the prompt much more successfully, correctly spelling 'Caffé Florian' and 'Est. 1720', while FLUX.1 [schnell] failed significantly on text accuracy. While FLUX.1 [schnell] produced a cleaner vector logo, the factual errors in the brand name make it unusable compared to the detailed and accurate illustration from LongCat-Image.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the clean, flat-vector style with professional composition.
- + Follows the NASA-inspired color palette perfectly.
- + Structure suggests a clear, logical sequence of steps.
- − Text consists of illegible gibberish, failing to provide actual information.
- − Icons for specific steps like 'Saturn V' are abstract and don't quite resemble the real craft.
LongCat-Image
- + Includes a recognizable lunar module on the surface for the landing step.
- + Captures the NASA aesthetic through specific iconography like the American flag and bold typography.
- − Fails to include the requested 6-step sequence, showing only a few disjointed icons.
- − The line work is less 'clean' than Model A, with some fuzzy borders and inconsistent scaling.
- − Includes text that is nearly legible but still garbled (e.g., 'Tranquilty' vs 'Tranquility').
Verdict: FLUX.1 [schnell] creates a much more aesthetically pleasing and professionally balanced infographic that adheres to the requested vector style and color palette. While LongCat-Image attempts to include more literal mission elements like flags and a lunar module, it fails the instruction to provide 6 specific steps and lacks the layout sophistication of FLUX.1 [schnell].
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images