6B parameter image generation model excelling at rendering multilingual text directly in generated images
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
LongCat-Image
#62 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
LongCat-Image
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
LongCat-Image
- + Excellent clean glass rendering with realistic reflections and refractions
- + Accurate spatial arrangement as per the prompt
- + Clean, modern aesthetic with sharp focus on the central objects
- − The glass cube has no visible top face although the book is resting on it
- − The lighting feels slightly more studio-lit than natural window light
Wan 2.6
- + Highly realistic textures on the wooden table and weathered red book
- + Strong adherence to the lighting instruction with realistic shadows
- + Excellent depiction of the plant seen through the glass
- − The glass cube has some geometric inconsistencies on the right vertical edge
Verdict: Both models followed the complex spatial instructions perfectly. LongCat-Image yields a cleaner, more sterile image with very high clarity, while Wan 2.6 provides a much more convincing and realistic atmosphere through textured surfaces and natural lighting. Wan 2.6 is preferred for its superior materiality and lighting consistency.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
LongCat-Image
- + Excellent depiction of raining environment and wet pavement reflections
- + Includes clear motion blur from passing cars as requested
- + Captures an 'imperfect framing' look characteristic of street photography
- − Anomalous bicycle structure with an extra wheel appearing behind the front
- − Skin texture on the man's face feels slightly smoothed and less 'natural' than requested
Wan 2.6
- + Exceptional skin texture and realistic age details on the man
- + High attention to detail with rain droplets clinging to the man's jacket
- + Correct bicycle anatomy and logical interaction with a tool
- − Failed to include the requested motion blur on the passing car
- − The rain effect on the jacket looks slightly static/frozen like pearls rather than a candid shot
Verdict: Wan 2.6 provides a much more convincing and realistic portrait of the elderly man with impressive skin textures and realistic clothing, whereas LongCat-Image suffers from a major anatomical failure in the bicycle's design. While LongCat-Image followed the motion blur instruction better, Wan 2.6 is the superior image due to its overall coherence and high level of detail.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
LongCat-Image
- + Excellent depiction of ornate engraved plate armor
- + Clear and symmetrical composition
- + Vibrant warm colors and effective bokeh sparks
- − The character looks a bit too clean and youthful for 'battle-worn'
- − The hair beads look somewhat like modern plastic jewelry
Wan 2.6
- + Superb adherence to the 'battle-worn' and 'dirt on skin' prompts
- + Highly lifelike and emotional eyes with wetness texture
- + Excellent clothing texture with frayed cloth and rugged leather straps
- − The composition is a bit more crowded than Model A
- − Slightly less emphasis on the 'ornate' aspect of the armor engraving compared to A
Verdict: Wan 2.6 is the clear winner for its superior interpretation of 'battle-worn' and its exceptional texture work on the skin, eyes, and clothing. While LongCat-Image provides a beautiful and clean aesthetic, Wan 2.6 captures the grit and realism requested in the prompt more effectively.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
LongCat-Image
- + Includes all requested food categories including mains and appetizers
- + Dynamic layout with vibrant color blocks
- − Text is largely gibberish and poorly rendered
- − Graphic design feels cluttered and lacks whitespace
Wan 2.6
- + Excellent clean and professional minimalist layout
- + High-quality, realistic food photography in a neat grid
- + Text is legible and uses appropriate sans-serif fonts
- − The grid relies heavily on pizza photos, showing less variety for mains
- − The vibrant accents are limited to the corners rather than throughout the design
Verdict: Wan 2.6 is the clear winner as it produces a professional, usable menu design with high-quality photography and legible typography. LongCat-Image fails on the basic requirement of text rendering, resulting in a cluttered and illegible layout that feels amateurish by comparison.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
LongCat-Image
- + Excellent graphic design layout with clean typography
- + High-quality, appetizing burger rendering
- + Vibrant and well-executed starburst graphic
- − Fails the 'exploded view' requirement, showing a mostly assembled burger
- − Lacks the sense of motion requested in the prompt
Wan 2.6
- + Successfully captures the 'exploded' concept with ingredients suspended in air
- + Strong sense of motion with sauce drips and smoke
- + Fiery text effects are highly integrated into the background theme
- − Starburst graphic is a bit cluttered in the corner
- − Text at the bottom is slightly less legible against the fire
Verdict: While LongCat-Image produced a very clean and professional advertisement, it failed the primary core architectural instruction of an 'exploded burger'. Wan 2.6 followed all prompt instructions perfectly, creating a dynamic, mid-air explosion of ingredients with impressively themed fiery typography.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
LongCat-Image
- + Features a clear, centered composition of a chalkboard stand in a café environment.
- + Captures a stylized chalk-like weight to the text.
- − Text is largely unintelligible with major spelling errors like 'ToAYS GYAYS' and 'ARLIL'.
- − Failed to render the requested specific menu items correctly.
- − Text rendering looks like a digital font attempt rather than natural handwriting.
Wan 2.6
- + Excellent prompt adherence with near-perfect spelling of all requested menu items.
- + Realistic chalk texture including smudges and dust on the board.
- + Authentic handwritten cursive and print styles that look human-made.
- − The perspective is slightly angled rather than a direct front-on shot.
- − One small character overlap in the title 'SPECIALS'.
Verdict: Wan 2.6 is the clear winner as it successfully rendered almost all the complex text requested in the prompt with high accuracy and a very realistic chalk texture. In contrast, LongCat-Image failed significantly on the text rendering, producing gibberish and failing to follow the specific menu item instructions.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
LongCat-Image
- + Excellent anatomical clarity for the horse's legs and structure.
- + Includes detailed lunar landscape and multiple background celestial bodies.
- − Failed the inverse prompt instruction (astronaut is riding the horse).
- − The randomly floating aircraft in the top left feels out of place and low resolution.
Wan 2.6
- + Impressive cinematic lighting and color palette with vibrant nebulae.
- + High level of detail in the horse's mane and the reflections on the visor.
- − Failed the inverse prompt instruction (astronaut is riding the horse).
- − Anatomical issues with the horse's front legs appearing merged or confusingly posed.
Verdict: Both models failed the negative constraint to have the horse on top of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. Wan 2.6 is preferred for its superior cinematic qualities, lighting, and textures, whereas LongCat-Image feels more like a static collage of separate elements.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
LongCat-Image
- + Excellent fur texture on the capybara and clean lighting.
- + Dynamic perspective and sharp focus on the subjects.
- − Internal car perspective shows the capybara partially emerging from the windowsill.
- − Includes two passengers instead of the requested single businesswoman.
- − The capybara's hand/paw looks like a bird talon/human hybrid.
Wan 2.6
- + Perfect adherence to the single businesswoman requirement.
- + Captures the 'bored' expression of the passenger perfectly.
- + Great environmental atmosphere with the rain and Times Square lights.
- − The passenger is seated in the front seat instead of the back seat as requested.
- − The capybara's head is slightly centered rather than in the driver's seat position.
Verdict: LongCat-Image provides higher textural detail and a more traditional 'professional' attire for the driver, but it fails to follow the character count and anatomy constraints. Wan 2.6 captures the specific mood and narrative requested ('normal, bored expression') much more effectively, and while it places the passenger in the front seat, the overall composition and realism are superior for the requested prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
LongCat-Image
- + Includes all requested elements like thorns and webs clearly
- + Title text is very sharp and readable
- − Hallucinates names like 'Julie' and misspells the location as 'The Armiees'
- − Visual composition looks a bit cluttered and less cinematic
Wan 2.6
- + Excellent text accuracy for all details including the location
- + Cinematic lighting and high-quality artistic composition
- + Better integration of the 'parchment' aesthetic within the background
- − The 'webs and thorns' border is a bit messy and overlaps with the background trees
- − Scroll banner is slightly less detailed than the rest of the image
Verdict: Wan 2.6 is the clear winner because it correctly renders the specific event details requested (date, time, and location) whereas LongCat-Image hallucinates names and misspells the location. Wan 2.6 also offers a much more cohesive, polished, and atmospheric visual style that better fits the 'cinematic' requirement of the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
LongCat-Image
- + Excellent 3D miniature toy-like texture and rendering
- + Clean and readable text with professional typography
- + Vibrant colors and appealing stylization
- − The flag is a waving icon rather than a flat graphic as requested
- − Composition is slightly off-center vertically
Wan 2.6
- + Perfect 45-degree isometric perspective and composition
- + Strict adherence to the 'raised diorama base' and 'solid background' prompts
- + Excellent placement and scaling of text and flag
- − The textures look slightly flatter and less 'refined' than Image A
- − Lighting is a bit more clinical compared to the soft warmth of the competitor
Verdict: Both models followed the prompt exceptionally well, but Wan 2.6 is the winner for its superior adherence to the isometric perspective and the layout of the text elements. While LongCat-Image has better individual textures on the sushi, Wan 2.6 provided the requested diorama base and a more balanced square composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
LongCat-Image
- + Strong god ray effect and vibrant colors
- + Clear subject focus
- − Failed to include the rabbit as a separate animal, merging it into a 'cat-rabbit' hybrid
- − Butterflies appear flat and lack realistic integration with the lighting
- − Anatomical issues with the animals' paws and ears
Wan 2.6
- + Successfully included all four distinct animals requested in the prompt
- + Dynamic and playful composition that matches the 'tumbling' request
- + Realistic lighting integration with beautiful fur texture and dew sparkles
- − The fox's facial features are slightly distorted
- − Some floating dandelion seeds appear overly sharp or disconnected
Verdict: Wan 2.6 is the clear winner because it correctly interprets the prompt by including four distinct baby animals, whereas LongCat-Image creates a bizarre cat-rabbit hybrid. Wan 2.6 also captures the chaotic, playful energy of the prompt with much more realism and better lighting integration.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
LongCat-Image
- + Excellent adherence to the vintage woodcut/subtle texture request
- + Good rendering of the banner and date
- + Accurate text spelling
- − Repetitive text ('Caffè' appears twice)
- − Composition feels a bit cluttered and cramped
- − Steam lines are somewhat chaotic
Wan 2.6
- + Strong minimalist vector aesthetic
- + Clean and balanced composition
- + Good color palette adherence with warm brown and cream tones
- − The banner is very small and tucked to the side compared to the central placement requested
- − Less emphasis on the vintage texture
Verdict: LongCat-Image captures the 'vintage' and 'textured' aspect of the prompt much more effectively with its etched style, though it suffers from some text repetition. Wan 2.6 provides a cleaner, more professional vector logo that feels more 'minimalist', but the banner is less prominent than requested. Overall, LongCat-Image is the preferred choice for its stylistic character and texture.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
LongCat-Image
- + Strong adherence to the vector icon style and layout requested.
- + Correctly follows the NASA-inspired color palette.
- + Includes multiple specific stages of the mission visually.
- − Text rendering is garbled and contains nonsensical characters.
- − Includes a space shuttle icon which is historically inaccurate for the Apollo missions.
Wan 2.6
- + Legible and accurate text for the mission and crew names.
- + Clean, minimalist aesthetic that follows the requested color palette.
- − Completely fails to include the requested 6-step infographic steps.
- − The background contains digital textures/noise that deviate from the 'clean' and 'flat-vector' instruction.
- − Composition is very sparse and lacks the requested iconography.
Verdict: LongCat-Image succeeds in creating the layout and iconographic flow of an infographic as requested, despite historical inaccuracies and poor text quality. Wan 2.6 fails the core prompt instruction to visualize the six mission steps, providing only a simple title and names. LongCat-Image is the preferred choice for following the complex structural requirements of the prompt.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English