HiDream AI's 17B parameter text-to-image model using sparse diffusion transformer with mixture of experts, achieving state-of-the-art image generation quality with strong prompt following
Settled by community votes across 12 shared challenges, with an AI judge weighing in on each.
HiDream I1 Full
#60 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
HiDream I1 Full
0%
win rate
Ties
0%
Qwen Image 2512
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
HiDream I1 Full
- + Excellent handling of light and shadows, capturing the soft window light perfectly.
- + Very realistic glass reflections and refraction of the sphere.
- + Great material textures on the wooden table and red book.
- − The sphere appears to be floating inside the cube rather than sitting on a surface.
- − The glass cube has slightly irregular edges on the left side.
Qwen Image 2512
- + Strong composition with a clear, sharp focus on all elements.
- + Excellent adherence to spatial logic, with the sphere resting on a reflective base.
- + Very clean geometry on the glass cube.
- − The lighting is a bit flat compared to Model A.
- − The glass appears slightly more like plastic due to the heavy green tint.
Verdict: Both models followed the prompt perfectly. HiDream I1 Full produced a more cinematic and atmospherically lit image with realistic lighting effects, though the sphere appears to be floating. Qwen Image 2512 created a cleaner, more logically grounded scene with the sphere resting on the bottom of the cube, although the lighting is less dynamic.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
HiDream I1 Full
- + Excellent reflection and wet pavement texture
- + Realistic 50mm shallow depth of field effect
- + Great candidate for cinematic atmosphere
- − Physical logic errors with the bicycle front wheel and frame
- − Man is sitting on the bike in a way that doesn't suggest 'repairing'
- − Background cars lack the requested motion blur
Qwen Image 2512
- + Natural skin texture and facial expression
- + Better adherence to 'imperfect framing' with a tighter street-level view
- + Higher level of detail in the bicycle mechanical parts
- − The man is posing for the camera rather than being 'candid' or 'repairing'
- − Cars in the background are too sharp, missing the requested motion blur
- − Hands appear slightly mangled around the bicycle seat
Verdict: Both models failed to incorporate the requested motion blur for background traffic. HiDream I1 Full captures a more cinematic atmosphere with superior pavement reflections, but suffers from significant anatomical and mechanical distortions in the bike. Qwen Image 2512 provides a more realistic subject and better textures, though it fails on the 'candid' requirement as the subject looks directly into the lens.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
HiDream I1 Full
- + Strong prompt adherence regarding the beaded braids and ornate engraving.
- + High contrast and vibrant lighting and reflections on the armor.
- + Clear, sharp focus on the central facial features.
- − The scars look a bit like digital paint strokes rather than realistic skin texture.
- − The bokeh sparks in the background are somewhat simplified and look like yellow dots.
Qwen Image 2512
- + Exceptional skin texture with realistic dirt and weathering effects.
- + Very natural lighting with convincing golden hour/torchlight warmth and authentic bokeh sparks.
- + Superior details on the leather straps and the fine floral engraving of the armor.
- − Slightly less symmetrical composition compared to model A.
- − Small bead details in the hair are less prominent than in model A.
Verdict: Qwen Image 2512 is the winner due to its superior realism and material textures, particularly the skin, leather, and fine metal engravings which look more lifelike. While HiDream I1 Full delivers a very striking and clean image, it has a slightly more 'digital' feel to the scars and lighting compared to the cinematic quality achieved by Qwen Image 2512.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
HiDream I1 Full
- + Excellent text legibility with clear sans-serif fonts
- + Professional clean layout with good white space management
- + High-quality, realistic food photography
- − Repetitive food photos, showing four very similar pizzas instead of diverse menu items
- − Lacks the requested 'colorful food photos in grid' at the top, opting for a list-like distribution
Qwen Image 2512
- + Successfully implements a colorful grid of varied food photos
- + Uses vibrant color accents as requested in the prompt
- + Captures the 'casual dining' vibe well with diverse item representation
- − Text contains significant spelling errors and gibberish ('RESSAGRENT', 'APPETIIZIZERS')
- − The layout feels slightly cramped with too many small elements
Verdict: HiDream I1 Full produces a much more professional and usable menu with legible text and clean formatting, though it fails to diversify the food imagery. Qwen Image 2512 follows the specific layout request for a colorful grid and vibrant accents much better, but is undermined by poor text rendering and spelling. HiDream I1 Full is the winner for overall quality and feasibility for actual design use.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
HiDream I1 Full
- + The date is rendered accurately with the correct year.
- + The chalkboard texture and wooden frame are clean and well-defined.
- − Failed almost entirely to follow the specific menu item text requested, producing gibberish like 'Fart Sanl'.
- − The text looks like a digital font overlay rather than realistic chalk handwriting.
Qwen Image 2512
- + Excellent adherence to the specific menu text requested, including complex items like 'Grilled Octopus'.
- + The handwriting looks authentic with realistic chalk textures, smudges, and variations in stroke.
- + Elegant cursive title as requested in the prompt.
- − Minor spelling error in 'Risitto' (missing the second 'o').
- − The third item text was cut off in the prompt, leading the model to hallucinate the end of the sentence.
Verdict: Qwen Image 2512 followed the instructions much better than HiDream I1 Full, accurately rendering the specific menu items and capturing the authentic chalk texture requested. While HiDream I1 Full produced sharp graphics, it failed the core text-to-image challenge by filling the menu with nonsense words and using a font that looked digital rather than handwritten.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
HiDream I1 Full
- + Excellent cinematic lighting and atmosphere
- + Strong dynamic composition with a sense of motion
- + Good integration of elements within the space environment
- − Anatomical issues with the horse's legs appearing distorted or blending into the background
Qwen Image 2512
- + High level of realistic detail on the spacesuit and horse's coat
- + Better clarity and resolution of the subjects
- + Anatomically more distinct horse legs
- − The astronaut's face inside the helmet looks slightly unnatural and doll-like
- − The lighting on the horse feels a bit flat compared to the dramatic background
Verdict: Both models failed the negative constraint to have the horse on top of the astronaut, instead providing the more common 'astronaut riding a horse' interpretation. Qwen Image 2512 is the winner due to significantly higher detail in the materials and a more coherent structure, whereas HiDream I1 Full has notable anatomical stretching and blurring in the horse's lower body.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
HiDream I1 Full
- + Excellent text rendering on the hat.
- + High-quality fur texture and lighting on the capybara.
- + Dynamic side-profile composition that feels more cinematic.
- − The capybara's hands look more like human-monkey hybrids than capybara paws.
- − The proportions of the car door and window frame are slightly warped.
Qwen Image 2512
- + Stronger adherence to the 'bored' expression requested for the passenger.
- + The capybara's hands/paws are more anatomically grounded for the animal.
- + Symmetric composition gives a clear view of the interior.
- − The capybara's head is strangely flat on top to accommodate the hat.
- − The lighting on the passenger is somewhat flat compared to the driver.
Verdict: HiDream I1 Full produces a much more visually striking and polished image with superior textures and better hat rendering. While Qwen Image 2512 captures the 'bored' expression of the passenger more accurately, the overall aesthetic quality and lighting of HiDream I1 Full give it the edge.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
HiDream I1 Full
- + Successfully included a parchment-style paper background.
- + Mostly accurate rendering of the small scroll banner text.
- + Clean, vector-like illustrations for the bats and trees.
- − Strongly failed with the event details (date, time, and location).
- − The 'Halloween Party' text is missing the word 'Invitation'.
- − The composition feels like a digital collage rather than a cinematic scene.
Qwen Image 2512
- + Excellent adherence to specific text instructions including the date, time, and location.
- + Cinematic lighting and atmosphere that matches the 'moody night sky' prompt.
- + Detailed and high-quality artistic rendering of the pumpkin and twisted trees.
- − One character typo in the headline ('Hallowern' instead of 'Halloween').
- − The border thorns are a bit repetitive and busy.
Verdict: Qwen Image 2512 is the clear winner despite a small typo in the header, as it correctly followed the instructions for the event details (date, time, and location) which HiDream I1 Full completely hallucinated. Qwen Image 2512 also delivered a much more premium, cinematic aesthetic with superior lighting and depth.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
HiDream I1 Full
- + Excellent text rendering with clean, professional-looking fonts.
- + Highly accurate 45° isometric perspective and diorama framing.
- + Superior material depiction, especially the translucent salmon and fish eggs.
- − Missed the request for a small flag icon next to the text.
- − Shadows on the plate feel a bit sharp compared to the soft lighting requested.
Qwen Image 2512
- + Includes all elements including the small flag icon.
- + Very cute, high-quality cartoon miniature aesthetic.
- + Good balance of garnish and environmental details on the diorama base.
- − The text has slightly irregular spacing and inconsistent outlines.
- − The perspective on the diorama base is slightly distorted rather than a true 45° isometric.
Verdict: Both models performed exceptionally well on this task. HiDream I1 Full produced a cleaner, more professional graphic with better material realism, but it missed the flag icon. Qwen Image 2512 followed every part of the prompt including the flag and decorative elements, though its text and perspective were slightly less polished than Model A.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
HiDream I1 Full
- + Great golden lighting and bokeh effect.
- + Excellent 'big expressive eyes' that look cute and glossy.
- − Failed to include the requested baby bunny.
- − Included two kittens instead of one, deviating from the list.
- − Fur texture looks a bit generated and smooth, lacking hyper-photorealistic detail.
Qwen Image 2512
- + Successfully included all four specific animals: retriever, kitten, bunny, and fox.
- + Higher level of realism in the fur textures and butterfly details.
- + Better adherence to the 'god rays' and 'dew sparkles' atmospheric prompts.
- − The fox's eyes appear slightly asymmetrical.
- − Composition is a bit crowded as the animals are squeezed together.
Verdict: Qwen Image 2512 is the clear winner as it successfully included all four requested animals, including the bunny which HiDream I1 Full missed entirely. Furthermore, Qwen Image 2512 achieved a more realistic textures for the fur and environment, whereas HiDream I1 Full felt more like a digital illustration despite the 'hyper-photorealistic' instruction.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
HiDream I1 Full
- + Clean vector aesthetic with high minimalism
- + Accurate text rendering for the Est. 1720 banner
- + Balanced brown and cream color palette
- − Completely missing the primary brand name 'Caffè Florian'
- − The cloche dome design is oversimplified and lacks detail
- − The texture is very flat compared to the prompt request
Qwen Image 2512
- + Perfect text rendering for both the brand name and the establishment date
- + Beautiful vintage illustration style with high-quality shading and steam effects
- + Strongest adherence to the 'subtle texture' and 'vintage' aesthetic marks
- − Slightly less 'minimalist' than Model A
- − The cloche dome is highly detailed, bordering on a complex illustration rather than a simple vector emblem
Verdict: Qwen Image 2512 is the clear winner as it successfully included all requested text elements, including the brand name 'Caffè Florian' which HiDream I1 Full omitted entirely. Qwen Image 2512 also provided a much richer interpretation of the 'vintage' and 'texture' prompts, creating a cohesive and professional logo design.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
HiDream I1 Full
- + Excellent typography rendering for the main title
- + Perfectly captures the requested palette and clean flat-vector aesthetic
- − Missed the step-by-step sequential logic of the specific 6 steps
- − Text for the icons includes meta-commentary like 'EARTH V + ICON'
Qwen Image 2512
- + Followed the instructional logic of the 6 steps much more closely
- + Includes the Lunar Module and surface landing visuals requested
- + Detailed vector icons for Earth and the Moon
- − Poor text spelling and legibility for the descriptive steps
- − Visual layout is a bit cluttered with redundant labels
Verdict: Qwen Image 2512 followed the complex prompting instructions for a 6-step sequence much better than HiDream I1 Full, even including the astronaut names. However, HiDream I1 Full produced a much cleaner, professional-looking graphic with perfect title text, whereas Qwen Image 2512 suffered from significant spelling errors and redundant labels.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.