6B parameter image generation model excelling at rendering multilingual text directly in generated images
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
LongCat-Image
#61 of 62 in Text-to-Image
Qwen Image
#35 of 62 in Text-to-Image
Where the votes landed
LongCat-Image
0.0%
win rate
Ties
0.0%
Qwen Image
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
LongCat-Image
- + Excellent photographic quality with realistic glass caustics and reflections.
- + Strong adherence to all spatial instructions and lighting direction.
- + The plant is clearly visible through the glass cube as requested.
- − The glass cube walls appear slightly inconsistent in thickness at the corners.
Qwen Image
- + Satisfies all prompt requirements including object placement and Lighting.
- + Clean composition with a nice shallow depth of field.
- − The glass rendering is less realistic, appearing more like a mirror on the bottom surface.
- − Perspective of the cube is slightly skewed, particularly the vertical corner lines.
Verdict: Both models followed the complex spatial prompt perfectly. LongCat-Image is the winner because it exhibits superior realism in how light interacts with the glass, showing convincing refraction and transparency, whereas Qwen Image has some minor perspective issues and less detailed textures.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
LongCat-Image
- + Excellent atmospheric lighting with vibrant red color pops and high-contrast reflections.
- + Captures the requested 'motion blur' on passing cars very effectively.
- + Cinematic framing and a more dynamic composition with depth.
- − The bicycle structure is physically impossible, appearing to have three wheels or overlapping frames.
- − The subject's hands are mangled and fuse with the metal of the bike.
Qwen Image
- + More realistic and coherent bicycle geometry compared to Image A.
- + Good adherence to the 'candid' street photography style with natural skin tones.
- + Better depiction of the man's hands interacting with the seat/frame.
- − Failed to include 'motion blur' from passing cars as requested.
- − The lighting is somewhat flat and lacks the 'cinematic' quality found in the competing image.
- − The man's feet appear to be floating or poorly integrated into the wet pavement.
Verdict: LongCat-Image delivers a much more cinematic and atmospheric result that perfectly captures the motion blur and lighting requested, but it suffers from severe structural hallucinations regarding the bicycle and human anatomy. Qwen Image provides a technically more accurate object (the bike) and cleaner anatomy, but misses key prompt elements like motion blur and feels less visually engaging. LongCat-Image is the preferred choice for its superior mood and adherence to the photographic style, despite the AI artifacts.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
LongCat-Image
- + Excellent depiction of ornate, high-shine plate armor with realistic thickness and engraving
- + Accurate focus and shallow depth of field as requested
- + Clean rendering of multiple braids with color-consistent beads
- − The 'scars' look more like fresh, deep puncture wounds or skin sores rather than faint battle scars
- − The subject's expression is somewhat passive for a 'battle-worn' warrior
Qwen Image
- + Stronger 'battle-worn' characterization with more realistic facial scarring and a gritty expression
- + High detail on the cross-over leather straps and underlayer textures
- + Dynamic lighting with the torch visible in the frame providing clear directionality
- − The sparks have a stylized, 'star-filter' look that feels less natural than standard bokeh
- − Some floating beads on the left side and anatomical awkwardness in the hair braiding
Verdict: Both models followed the prompt well, but LongCat-Image excels in technical visual quality and the rendering of the plate armor's metallic luster. Qwen Image provides a more compelling and thematic interpretation of a battle-hardened character, though it suffers from minor artifacts in the hair and less natural spark effects.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
LongCat-Image
- + Excellent food photography quality with realistic textures
- + Good use of vibrant pink accents for a modern feel
- + Clear grid-based layout for images
- − Text rendering is completely garbled and illegible
- − Typography lacks the requested professional sans-serif feel
- − Sections are poorly defined and chaotic
Qwen Image
- + Stronger adherence to 'minimalist' prompt with clean lines
- + Readable category headings ('Appetizers/', 'Pizza/mains')
- + Consistent grid layout for food photography on the left
- − Food images look somewhat artificial/plastic compared to Model A
- − Large empty red square in the center of the grid serves no purpose
- − Repeated use of the same salad/pizza images reduces variety
Verdict: Qwen Image is the superior choice because it provides a functional menu layout with legible headings and organized sections, directly fulfilling the prompt's structural requirements. While LongCat-Image has higher quality individual food photos, its typography and overall layout are too cluttered and the text is entirely nonsensical, making it unusable as a design mockup.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
LongCat-Image
- + Excellent typography with brilliant fire and glass effects
- + High-quality photorealistic rendering of the main burger
- + Great use of lighting and glowing embers
- − The burger is not truly 'exploded' as requested, but rather floating as a whole unit
- − Small artifacts in the starburst shape
Qwen Image
- + Better adherence to the 'exploded' prompt with components flying apart
- + Very dynamic composition and sense of motion
- + Good rendering of the fiery background
- − Typography is slightly less polished than Model A
- − Some ingredients in the air look like generic plastic chunks (e.g., the sauce blobs)
Verdict: LongCat-Image produces a more professional and visually appealing advertisement with superior text rendering, but fails to 'explode' the burger. Qwen Image follows the 'exploded burger' instruction much more accurately, creating a more dynamic sense of motion even if the individual textures are slightly less realistic. Qwen Image is the winner for following the specific creative direction of the prompt more closely.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
LongCat-Image
- + Excellent realistic chalk texture with dusty residue at the bottom
- + Realistic framing and lighting within a café setting
- − Text is largely illegible and contains numerous spelling errors
- − Failed to include the third specific menu item
- − Layout is confusing with prices separated from items by lines
Qwen Image
- + Text is highly legible and follows the prompt's specific menu items
- + Good composition with clear hierarchy between title and items
- + Successfully rendered the 'Brown Butter...' item that was cut off in the prompt
- − Year is incorrectly rendered as '20026'
- − Handwriting looks slightly more like a digital font than natural chalk variations
- − The chalk texture is very clean, missing some of the grittiness requested
Verdict: Qwen Image is the superior choice because it successfully followed the complex text instructions and rendered the specific menu items requested, whereas LongCat-Image produced mostly gibberish text. While LongCat-Image had a more realistic chalk texture and artistic board appearance, Qwen Image's ability to maintain legibility and accuracy (despite an extra zero in the date) makes it far more useful.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
LongCat-Image
- + Strong cinematic lighting and high-contrast color palette.
- + Includes interesting background elements like lunar landers and distant planes.
- + High level of detail on the astronaut suit and horse's mane.
- − Completely failed the semantic prompt to have the horse 'on top' of the astronaut.
- − Anatomical issues with the horse's legs, particularly the front right leg.
Qwen Image
- + Features a cleaner, more coherent composition with a large planetary backdrop.
- + Good anatomical consistency for the horse in a zero-gravity pose.
- + Accurate rendering of the astronaut's gear and equestrian tack.
- − Failed to follow the specific instruction to have the horse riding the astronaut.
- − The surreal elements are limited to the setting rather than the subjects.
Verdict: Both models failed the specific logic test of the prompt, which requested the horse to be on top of the astronaut ('horse riding astronaut'). Instead, both LongCat-Image and Qwen Image produced conventional 'astronaut riding a horse' images. Qwen Image is slightly better overall as it avoids the anatomical distortions seen in the horse's legs in the LongCat-Image version, though both are technically failures of prompt adherence.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
LongCat-Image
- + Features a highly realistic fur texture for the capybara.
- + Includes the requested 'yellow taxi driver cap' specifically.
- + Captures the bored expression of the passengers well.
- − The capybara's paw looks more like a human hand with long claws.
- − The composition feels a bit cramped with the taxi light appearing inside the door frame.
Qwen Image
- + Excellent anatomical rendering of the capybara's paws on the steering wheel.
- + The lighting and depth of field provide a more cinematic feel.
- + Successfully places both paws on the steering wheel as requested.
- − The capybara is not wearing a shirt/jacket combination as clearly as the prompt suggested (looks more like a single unit jacket).
- − The passenger is looking away from her phone rather than at it.
Verdict: Qwen Image is the winner because it successfully renders both paws on the steering wheel with convincing anatomy, whereas LongCat-Image generates an eerie human-like hand. Qwen Image also has a better overall composition and lighting that feels more like a professional photograph.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
LongCat-Image
- + Excellent typography rendering for the main title and scroll text.
- + Great cinematic lighting on the central Jack-o'-lantern.
- + Very strong adherence to the 'border with webs and thorns' prompt requirement.
- − The event details at the bottom are poorly rendered with severe spelling hallucinations ('The Armiees').
- − The transition between the parchment and the central scene is a bit harsh.
Qwen Image
- + Excellent layout with a more atmospheric and integrated illustration.
- + Perfectly accurate event details at the bottom with clear legibility.
- + Sophisticated composition that balances the trees, pumpkin, and text well.
- − The main title text contains a typo ('Hallo Party' instead of 'Halloween Party').
- − The larger parchment background is slightly plain compared to the internal illustration.
Verdict: LongCat-Image excels at rendering the main title text correctly and features a very aggressive, prompt-accurate thorn border, but it fails significantly on the smaller event details. Qwen Image provides a much more cohesive and atmospheric piece of art with perfectly legible event details, though it unfortunately leaves out a few letters in the main header. Qwen Image is the preferred choice for a functional invitation due to its overall layout and clarity of logistical information.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
LongCat-Image
- + Excellent rendering of PBR materials with realistic subsurface scattering on the fish.
- + Perfectly clean typography and a well-placed flag icon as requested.
- + High visual clarity and an appealing miniature aesthetic.
- − The raised base is a simple wooden board rather than a distinct diorama landscape.
- − The placement of the flag icon is slightly crowded next to the text.
Qwen Image
- + Accurately interprets the 'raised diorama base' with a multi-layered platform.
- + Great layout with the addition of chopsticks and a physical flag pole that adds to the scene.
- + Strong adherence to the isometric perspective and cartoon style.
- − The text 'JAPAN' has slightly irregular letter shapes and spacing.
- − The materials look more like plastic/clay and less like the requested 'realistic PBR' sushi textures.
Verdict: LongCat-Image provides superior texture work and typography, making the food look more like a high-end 3D render. However, Qwen Image better understood the 'diorama' aspect of the prompt, creating a more cohesive miniature world. LongCat-Image is preferred for its exceptional clarity and professional finish.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
LongCat-Image
- + Strong god rays and vibrant lighting
- + Clear, high-contrast details on the golden retriever
- − Major prompt failure: merged a cat and bunny into a 'cabunny' hybrid
- − Butterflies appear flat and poorly integrated
- − The fox has an anatomical issue with its front legs
Qwen Image
- + Successfully rendered all four requested animals as distinct species
- + Effective use of bokeh and dew sparkles in the foreground
- + Better sense of movement and interaction between the animals
- − The bunny's face is slightly blurry compared to the other animals
- − Background lighting is a bit softer, reducing the '8K masterpiece' sharpness slightly
Verdict: Qwen Image is the clear winner because it correctly depicts all four requested animals (dog, cat, bunny, fox), whereas LongCat-Image failed significantly by merging the kitten and bunny into a single hybrid creature. Qwen Image also better captures the 'playful' aspect of the prompt with more natural posing and interaction.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
LongCat-Image
- + Excellent typography with correct spelling and accented characters.
- + Complex and visually appealing engraving-style textures.
- + Perfect adherence to all requested elements including the cloche and banner.
- − The word 'Caffè' is repeated twice, which was not explicitly requested.
- − The 'Est. 1720' text is slightly skewed to fit the ribbon curve.
Qwen Image
- + Successfully achieves a clean, minimalist vector aesthetic.
- + Good color palette adhering to the warm brown and cream tones.
- − Major typographical errors with jumbled letters in 'Florian'.
- − The steam effect is very basic compared to the artistic request.
- − Missed the accent on 'Caffè' in the main text block.
Verdict: LongCat-Image is the clear winner as it successfully rendered all text accurately and captured the 'vintage' and 'cloche' themes with professional-grade detail. Qwen Image struggled significantly with the text rendering, resulting in an unreadable brand name, and had a much simpler composition that lacked the requested 'classic' feel.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
LongCat-Image
- + Stronger vector aesthetic with classic comic-style line work
- + Balanced layout with clear sections and supporting icons
- + Adheres well to the muted red and navy color palette
- − Nonsense text rendering for all labels
- − Fails to follow the specific 6-step sequence outlined in the prompt
- − Rocket icon is more of a generic space shuttle than a Saturn V
Qwen Image
- + Attempts to follow the sequential numbering requested in the prompt
- + Text rendering is significantly more legible, correctly identifying names like Collins and Aldrin
- + Clean flat-vector style with professional-looking gradients
- − Includes instructional text '(Stop at landing)' as part of the visual header
- − Logical errors in labels (e.g., '1. Sarth Orbit' and '3. Lunar d'ólle (Earth Orbit)')
- − Mixed up the execution of the requested icons within the requested sequence
Verdict: Qwen Image is the preferred choice because it successfully interprets the requested sequential infographic format and produces legible, recognizable text, whereas LongCat-Image generates complete gibberish. While Qwen Image included the prompt's instructional parenthetical text in the design, its adherence to the specific mission details like the crew names and the numbered steps makes it a more useful output.
Explore each model
Alibaba's Qwen image model