6B parameter image generation model excelling at rendering multilingual text directly in generated images
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
LongCat-Image
#62 of 62 in Text-to-Image
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Where the votes landed
LongCat-Image
0.0%
win rate
Ties
0.0%
Wan 2.5 (Preview)
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
LongCat-Image
- + Excellent realism in the glass texture and reflections.
- + Accurate interpretation of 'small' regarding the blue sphere.
- + Clean and modern composition with high clarity.
- − The lighting is a bit flat compared to the other model.
Wan 2.5 (Preview)
- + Beautiful warm lighting and dynamic shadows.
- + Good realistic texture on the red book cover.
- + Creative use of atmospheric dust particles in the light.
- − The blue sphere's scale is quite large relative to the cube.
- − Minor clipping where the sphere appears to be touching the top of the glass cube.
Verdict: Both models followed all spatial instructions perfectly. LongCat-Image is preferred as it better captured the 'small' size of the sphere requested and provided a more physically accurate glass container without clipping issues. Wan 2.5 (Preview) had superior lighting and atmosphere but made the sphere significantly larger and less central to the composition.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
LongCat-Image
- + Captures motion blur from passing cars very effectively.
- + Includes realistic rain splashes and wet pavement reflections.
- + Excellent mood and cinematic lighting that fits the street photography aesthetic.
- − Physical logic errors in the bicycle structure, showing an extra third wheel/frame component.
- − Visible rain streaks are overly uniform and look a bit like a filter.
Wan 2.5 (Preview)
- + Exceptional skin texture and anatomical detail on the man's hands and face.
- + The bicycle structure is much more coherent and realistic.
- + Very soft, natural depth of field that feels professional.
- − Misses the 'motion blur from passing cars' requirement almost entirely.
- − The lighting is a bit flat compared to the requested cinematic look.
Verdict: Wan 2.5 (Preview) produces a much more anatomically and structurally sound image, with impressive skin textures and a logical bicycle. However, LongCat-Image better captures the specific atmospheric requests for motion blur and cinematic rain, despite the significant anatomical failure of generating a nonsensical three-wheeled bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
LongCat-Image
- + Excellent rendition of ornate engraved plate armor with high-quality metallic reflections.
- + Very colorful and detailed beads in the hair braids, effectively meeting the prompt.
- + Strong bokeh effect and lighting that creates a heroic, cinematic atmosphere.
- − The facial scars look a bit like flat textures or stickers rather than physical wounds.
- − The facial hair and skin texture lack the micro-detail found in the competitor.
Wan 2.5 (Preview)
- + Incredible skin texture and facial realism, with lifelike eyes and convincing dirt smudges.
- + Highly detailed texture on the tattered cloth and leather straps, appearing very tactile.
- + Excellent hair braiding with silver-toned beads integrated naturally into the style.
- − Lower lighting contrast on the armor compared to Model A, making the 'battle-worn' aspect feel more like dirt than metal fatigue.
- − The secondary hair braid on the left side appears to merge into the background or shoulder awkwardly.
Verdict: Wan 2.5 (Preview) produces a more lifelike and gritty portrait with superior skin and fabric textures, truly capturing the 'battle-worn' aesthetic. LongCat-Image excels in the depiction of the ornate plate armor and vibrant cinematic lighting, but its facial details feel slightly more artificial by comparison. Wan 2.5 (Preview) is the winner for its impressive realism and adherence to the fine texture details requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
LongCat-Image
- + Includes a wider variety of realistic food photography styles
- + Uses vibrant color blocking to define sections
- − Lacks the requested minimalist design aesthetic
- − Text is extremely garbled and messy compared to the competitor
- − Layout feels cramped and lacks white space
Wan 2.5 (Preview)
- + Perfect adherence to the minimalist design requested
- + Clean grid layout with excellent use of white space
- + Captions and headings are much more legible and logically placed
- − Food photos are more repetitive and look somewhat synthetic
- − Limited text in the side column makes it look like a placeholder
Verdict: LongCat-Image provides some good individual food photos but fails the 'minimalist' requirement, resulting in a cluttered and disorganized layout. Wan 2.5 (Preview) perfectly captures the clean, modern aesthetic of a professional menu with a consistent grid and superior typography, making it the clear winner for this design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
LongCat-Image
- + Excellent typography with a professional 3D liquid-metal and fire effect.
- + High-quality photorealistic rendering of the burger textures, especially the bun and meat.
- + Strict adherence to the starburst and text placement requirements.
- − Fails to follow the 'exploded burger' instruction, showing a mostly assembled burger instead.
- − The background fire and coals look a bit static compared to the dynamic request.
Wan 2.5 (Preview)
- + Perfectly captures the 'exploded' view with components suspended in mid-air.
- + Highly dynamic composition with sauce droplets and floating ingredients enhancing the sense of motion.
- + Excellent integration of the fiery glowing effect on the text and starburst.
- − The 'MAGIC BURGER' text has a dripping effect that looks slightly more like cheese/honey than fire.
- − The bottom bun's angle is a bit awkward relative to the other ingredients.
Verdict: While LongCat-Image produces a very polished and professional-looking advertisement, it failed the primary prompt instruction to create an 'exploded' burger. Wan 2.5 (Preview) followed all instructions perfectly, creating a dynamic, high-energy composition with all burger components suspended as requested, making it the superior choice for this specific prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
LongCat-Image
- + Excellent chalk texture including realistic dust and smudges at the bottom of the board.
- + Strong cafe background bokeh and warm lighting.
- − Text rendering is mostly gibberish with severe spelling errors.
- − Failed to follow the specific menu items requested in the prompt.
Wan 2.5 (Preview)
- + Perfect text rendering with near-flawless spelling of all complex menu items.
- + Consistent and attractive handwritten chalk style that feels authentic.
- + Followed all instructions, including the specific date and price points.
- − Slight repetition at the end of the menu where it lists the cookies twice.
- − The perspective is slightly angled rather than a straight-on shot.
Verdict: Wan 2.5 (Preview) significantly outperformed LongCat-Image by accurately rendering the requested text, including complex items like 'Truffle Mushroom Risotto'. LongCat-Image failed the primary task of text generation, resulting in unreadable gibberish despite having a very realistic chalk texture and environment.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
LongCat-Image
- + Successfully placed the horse on top of the astronaut as requested by the specific reversal prompt.
- + High degree of surrealism in the juxtaposition of figures.
- + Clear and consistent lighting on the lunar surface environment.
- − Anatomical issues with the horse's legs, particularly the extra joint appearing in the back leg.
- − Low-quality rendering of secondary elements like planes and satellites.
Wan 2.5 (Preview)
- + Excellent visual quality and resolution.
- + Superior composition with a cinematic sense of movement and depth.
- + Highly detailed rendering of the astronaut suit and horse's muscles.
- − Failed the negative constraint; the astronaut is riding the horse instead of the horse riding the astronaut.
- − The scene is more conventional than the surreal concept requested.
Verdict: LongCat-Image is the superior model for this prompt because it successfully followed the complex logical instruction to have the horse ride the astronaut. While Wan 2.5 (Preview) produced a significantly more beautiful and technically polished image, it completely failed to interpret the specific 'horse on top' requirement, resulting in a generic (though high-quality) astronaut-on-horse image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
LongCat-Image
- + Excellent texture on the capybara's fur
- + Sharp, clear focus on the passenger characters
- + Vibrant colors and high-quality lens flare effects
- − The capybara only has one paw on the steering wheel despite the prompt request
- − The composition feels slightly crowded with two passengers instead of one
- − The capybara's paw looks more like a human hand with claws
Wan 2.5 (Preview)
- + Perfect adherence to 'both front paws on the steering wheel'
- + Captures the bored, unaware expression of the passenger perfectly
- + Great atmospheric lighting and rain effects on the car exterior
- − The capybara's face is a bit symmetrical and less expressive than Model A
- − The steering wheel scale looks slightly off compared to the driver
Verdict: Wan 2.5 (Preview) followed the prompt instructions more accurately, specifically regarding the capybara's positioning with both paws on the wheel and the passenger's expression. While LongCat-Image has higher texture detail and sharper focus on the people, it failed the specific paw placement requirement and included an extra passenger.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
LongCat-Image
- + Excellent typography style for the main title
- + Includes the thorn and web border as requested
- + Cinematic lighting on the pumpkin
- − Significant spelling errors in the bottom event details
- − The torn parchment layer is a bit distracting from the background elements
- − The scroll banner is small and lacks detail
Wan 2.5 (Preview)
- + Perfect text rendering for both titles and event details
- + Superior visual quality and artistic composition
- + Beautifully integrated jack-o-lantern and twisted tree elements
- − The thorns are slightly less prominent compared to the spider webs
- − Slightly less 'parchment' texture on the paper compared to the other model
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully rendered every piece of text correctly, including the date, time, and specific location. While LongCat-Image had a strong aesthetic, it failed on the legibility of the bottom text and suffered from minor artifacts in the lettering.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
LongCat-Image
- + Excellent text rendering and alignment of both 'JAPAN' and 'SUSHI'.
- + Accurate representation of the requested flag icon.
- + Higher complexity in the 3D model with realistic textures for the rice and fish.
- − The perspective is slightly lower than the requested 45° isometric angle.
- − The base feels more like a standard wooden tray than a specific diorama base.
Wan 2.5 (Preview)
- + Perfectly captures the 45° top-down isometric perspective.
- + Very clean, minimal diorama base as requested.
- + Lighting is soft and aesthetically pleasing, enhancing the cartoon 3D look.
- − The flag icon is less integrated and looks like a simple flat graphic.
- − Text is slightly less refined compared to the first model.
Verdict: Both models followed the prompt exceptionally well, but LongCat-Image edges ahead slightly due to the superior quality of the typography and the subtle textures on the sushi ingredients. While Wan 2.5 (Preview) better captured the specific isometric viewing angle and diorama base, LongCat-Image produced a more professional-looking graphic design with better detail on the flag icon.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
LongCat-Image
- + Strong composition with a clear, happy focal point on the golden retriever.
- + Rich golden sunrise lighting and prominent god rays that enhance the mood.
- + Vibrant colors in the wildflowers.
- − Failed to generate a bunny, instead creating a 'cat-bunny' hybrid with elongated ears.
- − The fox's tail transition looks anatomicaly awkward.
Wan 2.5 (Preview)
- + Successfully included all four requested animals (dog, cat, bunny, fox) as distinct species.
- + Better sense of movement and 'playful chasing' through varied poses.
- + Excellent rendering of dew sparkles and fine fur textures.
- − The fox has strangely glowing, human-like blue eyes which looks unnatural.
- − The butterflies appear somewhat flat compared to the surrounding 3D environment.
Verdict: Wan 2.5 (Preview) is the winner because it successfully followed the prompt to include all four requested animals, including a distinct baby bunny which LongCat-Image failed to render correctly (instead merging it with a cat). While LongCat-Image has slightly better lighting aesthetics, Wan 2.5 captured the chaotic, playful energy and specific subject count required by the prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
LongCat-Image
- + Strong hand-drawn woodcut aesthetic
- + Highly detailed line work on the banner and steam
- + Good application of the requested subtle texture on the background
- − Text is repetitive and messy with 'Caffè' appearing twice
- − Styling of the top 'Caffè' is distorted and difficult to read
- − Composition is a bit cluttered compared to a minimalist logo
Wan 2.5 (Preview)
- + Excellent vector emblem style with clean lines
- + Perfect spelling and typography layout
- + Professional composition that feels like a real commercial logo
- − The 'steam' is a bit small relative to the cloche
- − Slightly less 'subtle' texture on the edges, appearing more like crumpled paper
Verdict: Wan 2.5 (Preview) produced a superior logo that accurately follows the typography and minimalist vector requirements. While LongCat-Image has a charming hand-etched style, it failed on the text rendering by repeating words and creating a cluttered, less legible design.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
LongCat-Image
- + Clean vector illustration style.
- + Follows the specified NASA-inspired color palette effectively.
- − Fails to follow the requested 6-step sequence.
- − Text and typography are garbled and illegible.
- − Logical flow of the infographic is poor.
Wan 2.5 (Preview)
- + Successfully includes almost all requested steps with legible labels.
- + Highly accurate text rendering for the astronaut names.
- + Professional layout with a clear hierarchical flow.
- − The 'Saturn V' icon looks more like a shuttle mashup than a rocket.
- − The trajectory lines for 'Translunar' and 'Lunar Orbit' are slightly messy/overlapping.
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully follows the complex instructions for a multi-step infographic with readable labels and specific astronaut names. LongCat-Image provides some nice vector elements but fails to structure them into the requested sequence and suffers from significant text corruption.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities