Head to head
Esc

Models · slot A

to navigate to pick

LongCat-Image Meituan Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

LongCat-Image

9.6 arena score

#62 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

LongCat-Image

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent adherence to lighting instructions with realistic soft light from the left.
  • + Clean, minimalist composition that feels realistic.
  • + Correct spatial orientation of all requested objects.
  • The glass cube walls look a bit thick, resembling a container rather than a solid or thin-walled cube.

Qwen Image Max

  • + Strong visual quality and sharp textures on the wooden table.
  • + Good reflection and transparency effects on the cube.
  • The blue sphere is awkwardly floating in the center of the cube without support.
  • The black pattern on the red book was not requested.

Verdict: Both models followed the complex spatial instructions well. LongCat-Image is the winner because it adhered more strictly to the prompt without adding unnecessary details like the book pattern, and it handled the physics more realistically by placing the ball on the bottom of the cube rather than having it float. Qwen Image Max had slightly sharper textures but felt less natural in its composition.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent depiction of shiny, engraved plate armor with high-contrast reflections.
  • + Accurately represents multiple braided sections with colorful beads as requested.
  • + Very clean and aesthetic facial features with specific dirt and wound details.
  • The character looks a bit too pristine and youthful for a 'battle-worn' paladin.
  • The leather strap lacks high-frequency texture compared to the other model.

Qwen Image Max

  • + Superb realism in the aged, weathered skin and deep scars which fits 'battle-worn' perfectly.
  • + Highly detailed texture on the leather straps and frayed cloth underlayer.
  • + More natural lighting integration with the torchlight and bokeh sparks.
  • The armor looks slightly more worn/rusted rather than 'ornate plate', though engraving is present.
  • Focus is a bit softer on the lower half of the image.

Verdict: Both models followed the prompt exceptionally well, but Qwen Image Max is the winner for its superior interpretation of 'battle-worn'. LongCat-Image produced a very clean, 'fantasy hero' look with highly reflective armor, whereas Qwen Image Max captured the gritty, lifelike textures of the skin, leather, and fabric much more convincingly.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent, clean text rendering with a high-end commercial aesthetic.
  • + Very high photorealistic detail in the ingredients and charcoal embers.
  • + Composition is well-balanced and resembles a professional billboard advertisement.
  • Failed to deliver an 'exploded' view; the burger is mostly assembled and static.
  • The secondary text 'LIMITED TIME ONLY' is cramped within the starburst.

Qwen Image Max

  • + Captured the 'exploded' and 'mid-air motion' aspect much more effectively with flying ingredients.
  • + The fiery text effect is very dynamic and matches the background perfectly.
  • + Great sense of energy and better adherence to the 'exploded' instruction.
  • The lettuce placement looks a bit awkward and detached near the back.
  • Slightly more 'digital art' feel compared to the clean commercial photorealism of the competitor.

Verdict: LongCat-Image produces a very polished, professional-looking advertisement with superior text clarity, but it fails to provide the requested 'exploded' view. Qwen Image Max follows the prompt's layout instructions much better by creating a dynamic, disassembled look with flying ingredients, making it the better interpretation of the creative brief despite slightly messier composition.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent photorealism in the capybara's fur and the taxi's exterior lighting.
  • + Matches the specific request for a 'yellow taxi driver cap' with insignia.
  • + Captures the bored expression of the passengers perfectly.
  • The perspective is more 'beside' the car than 'inside' the car.
  • The capybara's hands look more like claws and are not naturally gripping the wheel.
  • Includes two passengers when the prompt implied a single businesswoman.

Qwen Image Max

  • + Successfully positions the camera 'inside' the taxi as requested, creating a more immersion POV.
  • + Features very realistic interior details like the dashboard and cracked car ceiling.
  • + The capybara's hands look remarkably natural while gripping the steering wheel.
  • The cap is a plain baseball cap rather than a traditional taxi driver cap.
  • The capybara appears to have one anthropomorphic hand and one more animal-like paw.
  • The passenger in the back is slightly cut off by the frame.

Verdict: Qwen Image Max followed the spatial instructions more accurately by placing the camera inside the vehicle, providing a much more cohesive interior scene. While LongCat-Image produced a cleaner exterior aesthetic and more accurate headwear, its 'claw-like' hands and external perspective made it less successful at fulfilling the specific prompt details.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent typography style for the main title
  • + Includes all elements like thorns, webs, and bats
  • + Vibrant and cinematic jack-o-lantern lighting
  • Several typos in the event details at the bottom
  • The parchment looks a bit digital and clean
  • The location name 'The Arches' is misspelled

Qwen Image Max

  • + Flawless text rendering for both the title and event details
  • + Beautifully integrated scroll banner with accurate text
  • + Superior vintage aesthetic with better composition and atmospheric depth
  • The thorns in the border are slightly less sharp than in the other model

Verdict: Qwen Image Max is the clear winner because it correctly rendered every piece of text requested, including the specific date and location, while maintaining a cohesive vintage gothic aesthetic. LongCat-Image struggled significantly with the text at the bottom, resulting in several typos and nonsensical words despite having a strong central illustration.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent miniature cartoon aesthetic with soft, clay-like textures.
  • + Clean, minimalist composition that adheres well to the small garnish request.
  • + High quality decorative text rendering and flag icon.
  • The wooden base feels a bit simple compared to the 'diorama' prompt.
  • Rice texture is a bit large and stylized, leaning more into cartoon than PBR realism.

Qwen Image Max

  • + Superior PBR material rendering, especially on the glass base and moist sushi toppings.
  • + Accurate variety of sushi types with realistic rice grain detail.
  • + Very clean, modern typography and professional layout.
  • The composition feels slightly crowded for a 'minimalist' prompt.
  • The green onion garnish in the center looks a bit flat compared to the rest of the 3D elements.

Verdict: Qwen Image Max is the winner for its superior material qualities and realistic rice textures which better capture the 'realistic PBR materials' aspect of the prompt. While LongCat-Image delivers a charming cartoon style, Qwen Image Max's glass diorama base and sophisticated lighting provide a more professional aesthetic that fits the high-clarity requirement.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent rendition of golden sunrise lighting and god rays.
  • + Vibrant and sharp focal point on the golden retriever puppy.
  • + Included floating dew drops as requested in the prompt.
  • Anatomical failure: the kitten has rabbit ears instead of cat ears.
  • The bunny specified in the prompt is missing completely.
  • The butterflies appear flat and lack 3D integration with the lighting.

Qwen Image Max

  • + Successfully captured the 'tumbling together' action with a more dynamic composition.
  • + Includes all requested animals (puppies, kitten, and fox), though the bunny is arguably another small golden mammal here.
  • + Better integration of butterflies within the scene's depth.
  • Missed the specific requested 'bunny' in favor of a second golden retriever puppy.
  • The fox's front leg anatomy is slightly disjointed where it touches the puppy.

Verdict: While LongCat-Image has beautiful lighting and texture, it fails significantly on prompt adherence by merging the cat and bunny into a 'cat with rabbit ears' and missing the bunny entirely. Qwen Image Max captures the 'tumbling' energy of the prompt much better and, despite substituting a second puppy for a bunny, provides a more coherent and anatomically believable set of animals.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

LongCat-Image
Qwen Image Max

AI Judge Analysis

LongCat-Image

  • + Excellent typography rendering for the 'Est. 1720' banner
  • + Provides a distinct vintage illustration style with nice line work
  • Redundant text with 'Caffè' appearing twice
  • Text layout is cluttered and covers the cloche dome incorrectly
  • The primary name 'Florian' feels disconnected from 'Caffè'

Qwen Image Max

  • + Perfect adherence to the full prompt including all text elements
  • + Clean, balanced composition suitable for a professional logo
  • + Superior vector-style shading and elegant steam illustration
  • The background texture is a bit heavier on the edges than 'subtle'

Verdict: Qwen Image Max produced a much more professional and coherent logo that followed all text instructions perfectly, whereas LongCat-Image suffered from redundant text and a cluttered layout. Qwen Image Max's interpretation of the cloche dome and the banner was more aesthetically pleasing and better aligned with the 'minimalist vector' request.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.