HiDream AI's 17B parameter text-to-image model using sparse diffusion transformer with mixture of experts, achieving state-of-the-art image generation quality with strong prompt following
Settled by community votes across 12 shared challenges, with an AI judge weighing in on each.
HiDream I1 Full
#60 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
HiDream I1 Full
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
HiDream I1 Full
- + Perfectly follows all spatial instructions including the book on top.
- + Excellent use of soft, cinematic lighting from the window.
- + Highly realistic glass textures with convincing reflections and refractions.
- − The blue sphere appears to be floating unnaturally in the center of the cube.
Stable Diffusion 3.5 Large
- + Clear and sharp detail on the wooden table and book textures.
- + Accurately places a sphere and a red book within the frame.
- − Fails the spatial prompt by putting the cube on top of the book rather than the book on top of the cube.
- − The sphere is sitting on the book, which is also inside the cube, contradicting the specified arrangement.
Verdict: HiDream I1 Full perfectly followed the complex spatial instructions, accurately placing the red book on top of the cube and the plant behind it. Stable Diffusion 3.5 Large failed to follow the positional logic, placing the cube on top of the book. HiDream I1 Full also produced a more aesthetically pleasing image with superior lighting and depth of field.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
HiDream I1 Full
- + Excellent shallow depth of field and bokeh
- + Subtle, realistic reflections on the wet pavement
- + The man appears more organically integrated into the scene
- − The bicycle frame is physically impossible (seat post and down tube alignment)
- − The man's hands are mangled and blending into the bike frame
- − Missing requested motion blur on cars
Stable Diffusion 3.5 Large
- + More dynamic lighting and convincing rain streaks
- + Better posture showing the 'repairing' action more clearly
- + Captures the sense of a busy street with large vehicles in the background
- − The bicycle structure is very primitive and lacks spokes
- − The man's right arm is oddly elongated and lacks proper anatomy
- − Shadows/reflections on the ground are less convincing than Model A
Verdict: Both models struggle significantly with the complex geometry of a bicycle. HiDream I1 Full produces a more cinematic atmosphere with superior wet-road reflections and skin texture, though the bike frame is a mess of pipes. Stable Diffusion 3.5 Large has better rain effects and highlights, but the lack of spokes on the wheels makes the image feel less realistic. HiDream I1 Full is the winner for a more balanced photographic quality despite the structural errors.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
HiDream I1 Full
- + Excellent adherence to the 'beads in hair' prompt with clear, visible beads.
- + Strong, dramatic lighting that creates a high-contrast cinematic feel.
- + Intense, lifelike eye detail and distinct battle scars.
- − The armor engraving looks a bit more generic compared to the intricate patterns in Image B.
- − Slightly more 'over-sharpened' digital look on the skin textures.
Stable Diffusion 3.5 Large
- + Exquisite, highly detailed engraving on the plate armor that feels very authentic.
- + Exceptional texture work on the cloth scarf/underlayer and gambeson.
- + More naturalistic skin tones and subtle dirt/blood smearing.
- − Failed to include the specific request for 'small beads' in the braided hair.
- − Lighting is slightly flatter compared to Model A, losing some of the 'warm torchlight' atmosphere.
Verdict: HiDream I1 Full delivers a more cinematic and punchy portrait that follows the specific hair detail requirements perfectly. However, Stable Diffusion 3.5 Large showcases superior technical skill in rendering the ornate engravings of the armor and the complex textures of the undergarments, despite missing the hair beads. Image A is preferred for its better adherence to all prompt elements and striking eye detail.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
HiDream I1 Full
- + Excellent text legibility and font choice.
- + Clear section headers that adhere to the prompt.
- + Professional and clean white-space management.
- − Lack of variety in food photos, showing only pizza regardless of the section label.
- − Layout feels slightly more like a flyer than a multi-page menu design.
Stable Diffusion 3.5 Large
- + Strong aesthetic composition with a grid of diverse food photos.
- + Captures the 'vibrant accents' and 'casual dining' atmosphere well.
- + Higher creative merit in the layout design.
- − Very poor text rendering with significant spelling errors in headers.
- − The text content becomes unreadable and messy at the smaller scale.
Verdict: HiDream I1 Full produces a much more functional and professional menu design with legible text and a clean minimalist aesthetic, though it lacks variety in its food imagery. Stable Diffusion 3.5 Large offers a superior artistic layout with a great grid of food, but is ultimately undermined by significant spelling errors and incoherent text rendering. HiDream I1 Full is the likely winner for its practical utility and adherence to typography requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
HiDream I1 Full
- + The title text is mostly legible and follows the requested date.
- + Strong chalk texture and consistent handwriting style across the board.
- + Good layout with a charming illustration of a coffee cup.
- − Completely failed to follow the specific menu item text requested.
- − Text below the title is nonsensical gibberish.
- − The handwriting looks somewhat like a digital font despite the chalk texture.
Stable Diffusion 3.5 Large
- + Attempts to include the specific menu items requested in the prompt.
- + Excellent composition showing the chalkboard within a realistic, cozy café environment.
- + Text has a very authentic, hand-drawn look with varying sizes.
- − Significant spelling errors throughout the board (e.g., 'TODAAY', 'Ottpups', 'Cholcalte').
- − Incorrect date rendered ('2024' instead of '2026').
- − The handwriting is less legible than Model A's title.
Verdict: Both models struggled with the complex text requirements. HiDream I1 Full produced much cleaner, more legible title text and better chalk textures but completely ignored the specific food items. Stable Diffusion 3.5 Large attempted the specific food items and provided a much better overall scene composition, but suffered from numerous spelling errors and failed the specific date requirement.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
HiDream I1 Full
- + Strong cinematic lighting and high contrast.
- + Sharp rendering of the astronaut suit and horse's mane.
- + Clean composition with a clear view of the earth in background.
- − Completely failed to follow the instruction for the horse to be on top.
- − The horse has an extra, distorted fifth leg emerging from its belly.
Stable Diffusion 3.5 Large
- + Captured a more ethereal, surreal atmosphere with stardust effects.
- + High level of detail on the space suit textures.
- + Composition feels more integrated with the space environment.
- − Completely failed to follow the instruction for the horse to be on top.
- − Visible artifact on the horse's muzzle resembling a black mask or distortion.
Verdict: Both models failed the core spatial reasoning challenge of placing the horse on top of the astronaut, instead producing the common trope of an astronaut riding a horse. HiDream I1 Full has better lighting but suffers from a significant anatomical error with a fifth leg, while Stable Diffusion 3.5 Large offers a more cohesive surreal atmosphere despite the muzzle artifact.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
HiDream I1 Full
- + Successfully includes the businesswoman in the back seat as requested
- + Excellent capybara face and fur texture
- + Clear and readable 'TAXI' text on the cap
- − The paws on the steering wheel look more like human-primate hybrid hands than capybara paws
- − The interior perspective is slightly cramped
Stable Diffusion 3.5 Large
- + Dynamic lighting and higher color contrast
- + The capybara's anatomy and paws look more natural for the animal
- + Good rendering of the leather jacket and clothing
- − Completely failed to include the human businesswoman in the back seat
- − The capybara appears to have human legs and pants, which was not requested
- − The steering wheel is floating/not correctly attached to the dashboard
Verdict: HiDream I1 Full followed the complex prompt requirements much better than Stable Diffusion 3.5 Large, which completely ignored the passenger and the 'paws on the wheel' instruction. HiDream I1 Full correctly captures the narrative of the scene including the bored businesswoman, making it the superior choice for prompt adherence.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
HiDream I1 Full
- + Strong composition with a balanced layout
- + Excellent high-contrast glowing jack-o-lantern
- + Clean, legible title text and banner phrasing
- − Failed completely to include the specific event details (date, time, location)
- − Gibberish text in the bottom half of the invitation
Stable Diffusion 3.5 Large
- + Rich, atmospheric lighting and artistic detail
- + Beautifully textured parchment paper effect
- + Creative use of layered depth with the trees and moon
- − Missing the specific date, time, and location details entirely
- − The central element is a moon rather than a dominant central jack-o-lantern as requested
- − Text is somewhat squashed at the bottom and contains some minor artifacts
Verdict: HiDream I1 Full provides a cleaner, more traditional invitation layout with highly legible main titles, though its bottom text is nonsensical. Stable Diffusion 3.5 Large offers far superior artistic quality and atmosphere, but it ignored the 'central jack-o-lantern' instruction in favor of a moon and failed to include the event details. Both models struggled with the specific text fields (date/time/location), but HiDream I1 Full is slightly more functional as an invitation template.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
HiDream I1 Full
- + Excellent adherence to the '3D cartoon' and 'soft refined texture' style requirements.
- + Perfectly follows the typography request with large bold 'JAPAN' and 'SUSHI' correctly placed.
- + Ultra-clean composition on a high-quality diorama base.
- − Missed the request for a small flag icon.
Stable Diffusion 3.5 Large
- + Includes the flag icon requested in the prompt.
- + High level of detail on individual sushi grains and ingredients.
- + Good implementation of the isometric 45° angle.
- − Fails to place the text correctly as part of the image overlay, instead putting it on a small sign.
- − Scene is cluttered with chopsticks and extra bowls, ignoring the 'minimal garnish' instruction.
- − The texture is more realistic than the 'cartoon scene' requested.
Verdict: HiDream I1 Full followed the aesthetic and layout instructions much more closely, delivering a clean, stylized 3D diorama that matches the minimalist cartoon prompt. While Stable Diffusion 3.5 Large captured more small details like the flag, it ignored the typography placement instructions and included several unrequested objects that cluttered the 'minimal' scene.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
HiDream I1 Full
- + Excellent fur texture and clarity on the animals.
- + Vibrant colors and high-contrast lighting that feels magical.
- + Strong presence of butterflies and flowers.
- − Failed to include the baby bunny entirely.
- − Included two kittens instead of one.
- − The pose is static and does not show the requested 'chasing' or 'tumbling' action.
Stable Diffusion 3.5 Large
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox kit.
- + Captured the sense of movement with an action-oriented 'chasing' composition.
- + Beautiful use of bokeh, dew sparkles, and god rays.
- − The fox kit in the background is slightly blurry and lacks facial detail.
- − The kitten's ears and head shape are slightly distorted.
Verdict: Stable Diffusion 3.5 Large is the winner as it accurately followed the prompt by including all four distinct animals and capturing the requested action of chasing butterflies. HiDream I1 Full produced a higher-quality single-point focus and better fur texture, but it failed the prompt instructions by omitting the bunny and providing two kittens in a static pose.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
HiDream I1 Full
- + Clean vector-style execution
- + Excellent center-aligned composition
- + Correct 'Est. 1720' text rendering
- − Completely failed to include the primary name 'Caffè Florian'
- − Unusual interpretation of a cloche which looks more like a bell dome over a cup
Stable Diffusion 3.5 Large
- + Successfully rendered the requested name 'Caffé Florian'
- + Captured the 'subtle texture' and vintage paper aesthetic well
- + Correct cloche dome shape with rising steam
- − Added an extra accent on the primary text ('Cafféé')
- − Graphic elements in the center are slightly disjointed with steam appearing in two different places
Verdict: Stable Diffusion 3.5 Large is the clear winner because it included the primary brand name 'Caffè Florian', whereas HiDream I1 Full completely ignored it. While HiDream I1 Full had a cleaner graphic style, Stable Diffusion 3.5 Large better captured the vintage texture and complete text requirements of the prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
HiDream I1 Full
- + Successfully rendered readable main titles and labels.
- + Excellent adherence to the clean, flat vector aesthetic.
- + Followed the color palette instructions perfectly.
- − Confused aspects of the prompt, including literal prompt text like '+ ICON' in the labels.
- − Incorrect iconography, using a Space Shuttle-style craft instead of a Saturn V.
- − Failed to include all 6 requested steps, stopping after 3 or 4 vague sections.
Stable Diffusion 3.5 Large
- + Includes a high density of visual elements and technical-looking data.
- + Better captures the complex 'infographic' layout feel with various charts and planets.
- + Captures the NASA-inspired color scheme well.
- − Text is entirely illegible gibberish compared to Model A.
- − Lines and icons are cluttered and lack the 'clean, modern' look requested.
- − Fails to clearly delineate the 6 specific steps requested in the prompt.
Verdict: HiDream I1 Full is the winner because it successfully produces legible text and a clean vector style, whereas Stable Diffusion 3.5 Large generates messy, unreadable graphics. Although HiDream I1 Full struggled with the technical accuracy of the Saturn V and literal interpretations of the prompt text, its output is a functional poster design compared to the cluttered noise of the competitor.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency