Head to head
Esc

Models · slot A

to navigate to pick

LongCat-Image Meituan Qwen Image 2512 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

LongCat-Image

9.8 arena score

#62 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2512

22.9 arena score

#30 of 62 in Text-to-Image

Vote tally

Where the votes landed

LongCat-Image

0%

win rate

Ties

0%

Qwen Image 2512

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent adherence to lighting instructions with a clear soft glow from the left.
  • + Highly realistic glass materials with accurate refractions and reflections.
  • + Beautiful, clean composition with a professional photographic feel.
  • The glass cube has double edges at the bottom which look slightly like a base rather than a single panel.

Qwen Image 2512

  • + Successfully includes all requested elements in the correct positions.
  • + Good textural detail on the red book cover.
  • + Solid spatial arrangement with the plant visible through the glass.
  • Physics artifacts on the right side of the cube where a second blue sphere appears to exist or be reflected incorrectly.
  • The lighting is flatter and less directional than requested.
  • The glass panels appear slightly inconsistent in thickness.

Verdict: LongCat-Image provides a superior result with more realistic material rendering and much more accurate lighting that follows the prompt. Qwen Image 2512 includes all elements but suffers from confusing reflections and artifacts that make the glass cube look cluttered and less transparent.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent handling of the rainy atmosphere and water reflections on the ground.
  • + Strong cinematic composition with a meaningful use of leading lines.
  • + Accurate representation of the 'motion blur from passing cars' request.
  • The red bicycle has significant structural errors, appearing to have three wheels or a detached frame.
  • The man's hands while repairing the bike look mangled and anatomically incorrect.

Qwen Image 2512

  • + Exceptional skin texture and facial realism, capturing the 'natural skin texture' perfectly.
  • + The bicycle is much more coherent and realistic in its engineering.
  • + Captures the 'candid' and 'imperfect framing' feel effectively with a direct gaze.
  • Does not show the man actually 'repairing' the bike; he is simply posing with it.
  • The background cars are static rather than showing the requested motion blur.

Verdict: LongCat-Image excels at environmental storytelling and atmosphere, following the 'motion blur' and 'rain' cues more effectively, but suffers from severe anatomical and mechanical distortions in the subject. Qwen Image 2512 produces a much more realistic person and a coherent bicycle, but fails the action requirement ('repairing') and specific background motion requests. Qwen Image 2512 is the overall winner because its technical flaws are less distracting than the mangled limbs and impossible bicycle in LongCat-Image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent depiction of ornate engraved plate armor with high-contrast reflections
  • + Clean and vibrant color palette with strong bokeh effects
  • + Accurate representation of the requested beads in the braids
  • The facial features look a bit too clean and youthful for 'battle-worn'
  • The scars look like minor surface injuries or stains rather than textured skin damage

Qwen Image 2512

  • + Performs better on the 'battle-worn' aspect with more convincing skin texture, dirt, and deep scars
  • + Superior detail on the leather straps and the transition between different armor layers
  • + The lighting is more atmospheric and integrates better with the subject's face
  • The armor engraving is slightly less sharp than in Model A
  • The torch in the background is a bit blurry/distorted compared to the rest of the scene

Verdict: While LongCat-Image produces a very clean and aesthetically pleasing image with beautiful armor reflections, Qwen Image 2512 better captures the 'battle-worn' essence of the prompt through more realistic skin textures and scars. Qwen Image 2512 also provides the requested close portrait framing more effectively, whereas LongCat-Image feels more like a medium shot.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Strong use of vibrant primary colors and bold accents
  • + Higher resolution food photography with more appetizing textures
  • + Clever use of negative space for a modern feel
  • Text is largely nonsensical and poorly rendered
  • Layout feels a bit cluttered with overlapping elements

Qwen Image 2512

  • + Excellent adherence to the 'grid' requirement for food photos
  • + Clean, professional typography that follows the requested sections (Appetizers/Mains)
  • + Consistent color palette and icon usage
  • The food photos are repetitive and lack distinct variety
  • The overall image size/resolution feels slightly lower than Model A

Verdict: Qwen Image 2512 is the superior choice because it strictly adhered to the design layout requested, including the specified sections and the grid format for photos. While LongCat-Image produced higher quality individual food images, its layout lacks the professional structure and legibility typical of a restaurant menu.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent text legibility and clean graphic design for the logo.
  • + Vibrant lighting and high-quality rendering of the charcoal and embers.
  • Failed the core prompt instruction for an 'exploded burger' layout, as the burger is mostly assembled.
  • The composition feels more like a static menu item than a dynamic action shot.

Qwen Image 2512

  • + Perfectly captured the 'exploded' instruction with components suspended in mid-air.
  • + Detailed textures on the patty and vegetables enhance the photorealistic quality.
  • + Typography is well-integrated into the fiery theme.
  • Missed the word 'TIME' in the 'LIMITED TIME ONLY' phrase.
  • The starburst element is slightly less polished than the rest of the graphics.

Verdict: Qwen Image 2512 is the clear winner because it correctly followed the primary creative instruction to depict an 'exploded' burger with suspended components, whereas LongCat-Image generated a standard assembled burger. While LongCat-Image had slightly cleaner text rendering, Qwen Image 2512 captured the dynamic motion and complex composition requested by the prompt.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + The chalk texture on the board and board edges looks highly realistic.
  • The text is largely illegible with numerous spelling errors like 'ToAYS STAYS'.
  • The prices and items are mismatched and disorganized compared to the prompt.
  • The 'elegant cursive' requirement for the title was not met.

Qwen Image 2512

  • + Excellent text adherence with almost perfect spelling of all requested items.
  • + Beautifully rendered elegant cursive chalk handwriting.
  • + The composition is clean and follows the logical hierarchy of a real menu.
  • Minor spelling error in 'Risitto' instead of 'Risotto'.
  • The 'Brown Butter' item was cut off in the prompt but the model correctly inferred the full common name, which might be a slight overreach if strict literalism was desired.

Verdict: Qwen Image 2512 is the clear winner as it successfully rendered almost all the complex text requested with an elegant handwritten aesthetic. LongCat-Image failed significantly on the text rendering, producing garbled letters and nonsensical words that did not follow the prompt's specific menu items.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent environment and cinematic composition with planets and spacecraft.
  • + Good clarity on both the horse and the astronaut's equipment.
  • + High level of detail in the lunar surface and stars.
  • The horse has an extra leg (five legs are visible).
  • Several visible artifacts in the sky and near the spacecraft objects.

Qwen Image 2512

  • + Stronger visual impact with a more realistic horse texture and lighting.
  • + The astronaut's face is visible and well-rendered within the helmet.
  • + Cleaner overall image quality with fewer glitches/artifacts in the background.
  • The composition is a bit more generic with less variety in the space background.
  • The floating horse posture feels slightly less 'cinematic' than the grounded action of the other model.

Verdict: Both models failed the negative constraint to have the 'horse on top' of the astronaut, instead providing the standard astronaut-on-horse interpretation. LongCat-Image provides a more detailed world with nice environmental storytelling, but it suffers from a significant anatomical error (five legs). Qwen Image 2512 is the preferred choice as it is technically cleaner, has better lighting, and avoids the major anatomical failures seen in the competitor.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent photo-realism for the capybara's fur texture.
  • + High-quality lighting that feels natural for a city at night.
  • + Good background bokeh with recognizable city lights.
  • Anatomy error with the capybara's paws being depicted with human-like fingers.
  • The passenger is looking at her phone but appears duplicated or has a ghostly double behind her.
  • Perspective on the taxi roof sign is slightly distorted given the camera angle.

Qwen Image 2512

  • + Perfect adherence to the requirement of both front paws on the steering wheel.
  • + The passenger's expression perfectly captures the 'bored and normal' look requested.
  • + Better composition for a 'driver perspective' shot through the windshield.
  • The capybara's face is slightly less detailed than in the other model.
  • The steering wheel appears to be floating or not connected to a dash correctly.
  • The hands/paws have an odd, leathery texture.

Verdict: LongCat-Image provides superior realism in lighting and texture, but suffers from significant hallucination issues including a duplicated passenger and incorrect animal anatomy. Qwen Image 2512 follows the specific prompt instructions more accurately, particularly regarding the capybara's paws on the wheel and the passenger's bored expression, resulting in a more coherent scene.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent typography style that matches a vintage invitation
  • + Good conceptual use of the parchment paper background
  • + Includes all requested elements like thorns and webs
  • Several typos in the small text including the location
  • Composition feels a bit fragmented with the die-cut center
  • Unnecessary artifacts like the random '3uulie' text

Qwen Image 2512

  • + Accurate rendering of the event details text
  • + Atmospheric and cohesive cinematic lighting
  • + Elegant layout with well-integrated gothic elements
  • Spelling error in the main title ('Hallowern')
  • Thorn border is slightly less distinct than the other version

Verdict: Qwen Image 2512 produces a much more polished and atmospherically cohesive image with accurate event details, despite a visible typo in the word 'Hallowern'. LongCat-Image suffers from significant legibility issues and hallucinations in its text, making parts of the invitation nonsensical.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent typography rendering with clean, professional fonts
  • + Superior 3D modeling textures that feel tactile and high-quality
  • + Consistent lighting and shadows that enhance the isometric feel
  • The 'diorama base' is a simple wooden tray rather than a themed environment

Qwen Image 2512

  • + Successfully creates a more complex diorama base with organic elements
  • + Includes a wider variety of sushi types
  • + Follows the isometric perspective accurately
  • Text rendering is slightly less refined with inconsistent kerning and styling
  • The flag icon is a simple rectangle rather than a graphic element

Verdict: LongCat-Image provides a much cleaner, more professional look with high-quality textures and superior typography that matches the 'ultra-clean' request. Qwen Image 2512 does a better job interpreting the 'diorama base' part of the prompt by adding environmental details, but the overall visual fidelity and text clarity of LongCat-Image make it the more appealing and higher-quality result.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Excellent light ray effects and atmospheric dew sparkles
  • + Separates the subjects well with a sense of motion
  • + Good representation of the puppy and fox kit
  • Failed the prompt by merging the kitten and bunny into a single hybrid creature with rabbit ears on a cat's head
  • The butterflies look like flat 2D decals with inconsistent lighting

Qwen Image 2512

  • + Successfully includes all four distinct animals (puppy, kitten, bunny, and fox) accurately
  • + Complex composition with realistic interactions between the animals
  • + Butterflies have more natural positioning and integration into the scene
  • The fox kit's ears are slightly oversized and less 'kit-like'
  • Lacks the distinct dew drop 'sparkles' requested compared to the other model

Verdict: Qwen Image 2512 is the clear winner because it correctly renders all four specified animals, whereas LongCat-Image creates a bizarre cat-rabbit hybrid. Qwen Image 2512 also manages deep, soft fur textures and golden hour lighting while maintaining better anatomical accuracy across the entire group.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Successfully included all text elements including the year banner
  • + Followed the color palette and texture requirements well
  • Redundant text rendering with 'Caffè' appearing twice
  • Top text is poorly integrated and clutters the cloche illustration
  • Poor alignment of the banner and text

Qwen Image 2512

  • + High visual quality with professional shading and line work
  • + Excellent typography and clear brand name presentation
  • + Superior composition with a balanced layout for a logo
  • Steam is a bit more illustrative and less minimalist than requested
  • Text color on the banner is slightly inconsistent with the main heading

Verdict: Qwen Image 2512 produces a much more professional and coherent logo with clean typography and a well-centered cloche dome. LongCat-Image suffers from text repetition and a cluttered, poorly aligned layout that fails to capture the 'minimalist' aspect of the prompt.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

LongCat-Image
Qwen Image 2512

AI Judge Analysis

LongCat-Image

  • + Strong NASA-inspired color palette
  • + Clean vector iconography style
  • Fails to follow the 6-step logical flow requested
  • Nonsense text and glitched typography for the title
  • Includes space shuttles instead of Saturn V and Apollo modules

Qwen Image 2512

  • + Accurately follows the 6-step mission sequence requested
  • + Correct iconography for Saturn V and the Lunar Module
  • + Legible main titles and logical flow of information
  • Text duplication error with '2. Earth Orbit' appearing twice
  • Slightly more complex shading than the requested flat-vector style

Verdict: Qwen Image 2512 is the clear winner as it successfully follows the logical steps of the mission and uses correct historical hardware (Saturn V and Lunar Module), whereas LongCat-Image uses generic space shuttles and fails to include the requested steps. Qwen Image 2512 also features much more legible text, despite a minor duplication error.

Next steps

Explore each model