Distilled version of HiDream AI's 17B parameter text-to-image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
HiDream I1 Fast
#50 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
HiDream I1 Fast
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent reflection work, especially showing the sphere and plant distorted in the glass
- + Very clean, minimal aesthetic that matches the glass cube concept well
- + Perfect adherence to lighting and placement instructions
- − The sphere is quite large relative to the cube, pushing the definition of 'small'
- − The book lacks paper texture on its side, looking more like a solid block
Qwen Image 2.0
- + Great material texture on the book's cover and pages
- + Realistic wood grain on the table
- + Includes the plant behind the glass as requested
- − The glass refractive logic is messy, with strange duplicate spheres appearing on the sides
- − The sphere appears to be floating mid-air inside the cube without support
- − The lighting feels a bit more washed out compared to the soft window light in Model A
Verdict: HiDream I1 Fast produces a more cohesive and visually pleasing image with impressive handling of glass reflections and transparency. While Qwen Image 2.0 has superior texture detailing on the book and table, its physics and refractive logic are confusing, resulting in distracting ghosting/duplication of the blue sphere within the glass panes.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
HiDream I1 Fast
- + Strong cinematic atmosphere with excellent bokeh and reflections.
- + Perfectly captures the sense of a wet, rainy urban environment.
- − The man is sitting on the bike rather than repairing it.
- − The car behind him lacks the requested motion blur, appearing static.
Qwen Image 2.0
- + Successfully captures the 'imperfect framing' and 'candid' street photography style.
- + Realistic skin texture and age spots on the subject's face.
- + Accurately depicts the act of 'repairing' the bicycle chain.
- − The motion blur on the car is subtle and looks slightly artificial.
- − Depth of field is a bit deep compared to the requested 50mm shallow look.
Verdict: Qwen Image 2.0 is the winner because it adhered much better to the specific action of 'repairing' the bicycle and captured the requested 'imperfect framing' which gave it a more authentic street photography feel. While HiDream I1 Fast created a very beautiful, cinematic image, the subject was simply posing on the bike, failing the primary action prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent engraving details on the plate armor.
- + Strong execution of bokeh sparks and warm torchlight atmosphere.
- + Perfectly follows the hairstyle prompt with braided hair and colorful beads.
- − The facial scars look a bit like face paint rather than physical wounds.
- − Skin texture is slightly too smooth and clean for a 'battle-worn' character.
Qwen Image 2.0
- + Superb 'battle-worn' aesthetic with realistic dirt, grit, and believable scars.
- + Highly detailed texture on the leather pauldron and cloth underlayer.
- + Lifelike, weary expression that fits the paladin archetype well.
- − The hand resting on the sword has anatomical issues with finger placement and length.
- − The lighting is a bit harsh, blowing out some details in the background sparks.
Verdict: HiDream I1 Fast produces a more 'heroic' and polished image with superior armor engravings and lighting, whereas Qwen Image 2.0 captures the 'battle-worn' grit and texture of the skin and fabric much more effectively. While HiDream I1 Fast is more visually pleasing, Qwen Image 2.0 wins on raw realism and character depth, despite a slight anatomical flaw in the hand.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
HiDream I1 Fast
- + Strong bold header section
- + Clear separation of pricing and text
- − Repetitive food photos of mostly just pizza
- − Text rendering is very poor with significant artifacts
- − Header text contains major spelling errors
Qwen Image 2.0
- + Excellent grid layout following all prompt instructions
- + High-quality, varied food photography for different sections
- + Cleaner font rendering and logical organization of menu items
- − Nonsense 'gibberish' text for dish names
- − Pricing values are repetitive and unrealistic
Verdict: Qwen Image 2.0 followed the prompt instructions much more accurately, providing a clear grid layout with distinct sections for different types of food. HiDream I1 Fast struggled with variety, showing almost exclusively pizza, and had significantly more visual artifacts in the text rendering. Qwen Image 2.0 is the clear winner for its professional composition and high-quality imagery.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent rendering of the 'MAGIC BURGER' title with a soft neon glow
- + Clear and accurate price text within a distinct starburst graphic
- + High-quality textures on the bun and vegetables
- − The burger is not truly 'exploded,' appearing mostly assembled and static
- − Secondary text 'LIMITED TIME ONLY' is garbled and barely legible
Qwen Image 2.0
- + Perfect adherence to all text requirements including 'LIMITED TIME ONLY'
- + Dynamic 'exploded' composition with ingredients separated and sauce dripping
- + Stunning fiery effect integrated into the title typography
- − The price starburst looks a bit more like a generic sticker than an integrated fiery element
Verdict: Qwen Image 2.0 outperformed HiDream I1 Fast by following every part of the prompt, specifically the 'exploded' layout and the requirement for secondary text. While HiDream I1 Fast produced a very clean burger image, the burger itself was not separated into components as requested, and it failed to render the secondary text legibly.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent photographic quality and depth of field
- + Charming café environment with realistic lighting
- − Significant text rendering failures and overlaps
- − Incorrect spelling of 'Butter' as 'Buter'
- − Messy layout with numbers floating over words
Qwen Image 2.0
- + Near-perfect text rendering and spelling
- + Realistic chalk texture with smudges and erasure marks
- + Consistent handwriting style across all menu items
- − Background is slightly more generic than model A
- − Composition is a bit tightly cropped at the top
Verdict: Qwen Image 2.0 followed the complex text-heavy prompt almost perfectly, correctly spelling difficult items and maintaining a consistent, realistic chalkboard aesthetic. HiDream I1 Fast produced a beautiful photograph but failed significantly on the text generation, with numerous overlaps, garbled letters, and spelling errors. Qwen Image 2.0 is the clear winner for its superior ability to handle precise text elements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
HiDream I1 Fast
- + High resolution and realistic textures on the spacesuit and horse.
- + Good lighting and shadows following the outdoor desert sun.
- − Completely failed the environment prompt by placing the subject in a desert instead of space.
- − Failed the specific positional instruction 'horse on top' (it is a standard astronaut riding a horse).
- − Anatomical issues with the horse's legs, specifically the floating/detached rear legs.
Qwen Image 2.0
- + Successfully captured the 'in space' environment prompt.
- + Included surreal elements like the scale-textured horse and floating droplets.
- + Higher level of detail in the starfield and cinematic composition.
- − Failed the specific positional instruction 'horse on top, not vice versa'.
- − Minor anatomical blending issues where the astronaut's leg meets the horse.
Verdict: Both models failed the complex spatial instruction to place the horse on top of the astronaut; however, Qwen Image 2.0 is the superior choice as it adhered to the 'in space' setting and 'surreal' style. HiDream I1 Fast ignored the environment prompt entirely, rendering a literal horse in a standard desert landscape with significant anatomical glitches.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent high-resolution fur texture and lighting
- + Captures the cinematic blurred city lights through the window perfectly
- − Anatomical failure with human hands attached to the capybara body
- − The woman is seated in the front passenger seat instead of the back seat
Qwen Image 2.0
- + Correctly depicts the capybara with its own paws on the wheel
- + Accurately places the passenger in the back seat as requested
- + Professional taxi driver cap style is more authentic
- − Slightly lower resolution and more noise compared to Model A
- − The capybara's head shape is slightly distorted
Verdict: While HiDream I1 Fast has superior lighting and textures, it fails significantly on prompt adherence by giving the capybara human hands and placing the passenger in the wrong seat. Qwen Image 2.0 correctly follows all spatial and anatomical instructions, making it the more successful interpretation of the prompt despite slightly lower image clarity.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
HiDream I1 Fast
- + Includes all primary elements like bats, trees, and jack-o-lantern.
- + Effective use of high-contrast colors (orange and teal).
- − Text rendering is poor with several typos and garbled words.
- − The 'scroll banner' text is cut off and illegible.
- − The placement of the date/time text looks messy and repetitive.
Qwen Image 2.0
- + Excellent text rendering throughout the entire poster.
- + Sophisticated cinematic lighting and atmospheric foggy background.
- + Followed all instructions including the specific thorn and web border.
- − The parchment texture has some dark spots that look slightly like digital artifacts.
- − Composition is a bit crowded around the center scroll.
Verdict: Qwen Image 2.0 is the clear winner as it successfully rendered every piece of requested text accurately and legibly, whereas HiDream I1 Fast struggled with significant typos. Qwen Image 2.0 also produced a more cohesive and artistic composition with a detailed thorn border and moody atmosphere that matched the 'cinematic' prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent adherence to the '3D cartoon' and 'isometric miniature' aesthetic.
- + Clean text rendering and well-placed graphics.
- + Consistent PBR material feel with soft, rounded corners.
- − The sushi composition is slightly nonsensical with a nigiri topping placed over a maki roll.
- − Chopsticks are slightly distorted at the tips.
Qwen Image 2.0
- + High photographic realism in the sushi textures.
- + Accurate sushi types (nigiri, eel, maki) with realistic detailing.
- + Correct text and flag inclusion.
- − Ignored the '3D cartoon' style requested in the prompt.
- − Text layout is less integrated with the overall design compared to Model A.
- − Lighting is a bit harsh, creating some specular highlights that clash with 'gentle lighting'.
Verdict: HiDream I1 Fast better followed the stylistic instructions, creating a cohesive 3D cartoon diorama that matches the requested isometric perspective and aesthetic. While Qwen Image 2.0 produced more realistic food, it failed to capture the 'cartoon' and 'miniature' aspects of the prompt, resulting in a standard product photo.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
HiDream I1 Fast
- + Excellent lighting and soft focus
- + High visual clarity and fur detail
- + Characters have very expressive, endearing eyes
- − Failed to include the baby bunny
- − Includes two kittens instead of one
- − Animals are mostly static rather than playful or tumbling
Qwen Image 2.0
- + Successfully included all four requested animals (dog, cat, fox, bunny)
- + Captures the 'playfully chasing' and 'tumbling' aspect of the prompt well
- + Strong execution of god rays and morning dew sparkling on flowers
- − The fox kit has a slightly distorted facial structure while tumbling
- − Lower contrast in the foreground compared to Model A
Verdict: Qwen Image 2.0 is the clear winner as it successfully included all four requested animals and accurately depicted the 'tumbling' and 'chasing' action described in the prompt. HiDream I1 Fast failed to include the bunny and doubled the kitten count, providing a beautiful but more static portrait that didn't follow the scene's movement instructions.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
HiDream I1 Fast
- + Strong composition with a cohesive circular frame
- + Accurate typography for the main restaurant name
- + Captures the vintage minimalist aesthetic well
- − The date in the banner contains a typo '17210' instead of '1720'
- − The steam effect is a bit abstract and disconnected
- − Visible brush stroke artifacts on the right side of the dome
Qwen Image 2.0
- + Perfect text rendering for both name and date
- + Excellent use of texture and shading on the cloche dome
- + Clean, high-quality vector-style lines and professional layout
- − The steam looks more like a flame or a logo inside the dome than rising steam
- − Slightly less 'minimalist' than the other version due to heavy shading
Verdict: Qwen Image 2.0 is the winner because it successfully followed all text instructions, including the correct date '1720', whereas HiDream I1 Fast included an extra digit. Qwen Image 2.0 also demonstrated superior technical execution with cleaner lines and more sophisticated shading that better fits a restaurant branding context.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
HiDream I1 Fast
- + Successfully uses the requested NASA-inspired color palette.
- + Includes a large central rocket illustration that anchors the design.
- − Text is heavily garbled and containing numerous spelling errors like 'DESCENG' and 'EARH OPBIT'.
- − Infographic layout is nonsensical, with icons and lines that do not represent a clear progression of steps.
- − Fails to follow the logical order of mission steps requested in the prompt.
Qwen Image 2.0
- + Excellent layout that follows the logical chronological progression of the mission steps.
- + Text is highly legible and correctly spelled, including additional context like 'Tranquility' and astronaut names.
- + Iconography matches the prompt perfectly, including the trajectory arc and lunar module icons.
- − Includes a minor typo in 'Translunjar'.
- − The vertical centering of some labels relative to their icons could be slightly more precise.
Verdict: Qwen Image 2.0 is the clear winner as it successfully creates a functional, logical infographic that follows all the specific steps requested in the prompt. While HiDream I1 Fast produces a visually colorful image, its text is illegible and it fails to organize the content into a coherent sequence, whereas Qwen Image 2.0 delivers clean vector-style icons and meaningful data visualization.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request