Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
OmniGen v2
#57 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
OmniGen v2
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
OmniGen v2
- + Perfect adherence to spatial instructions including the sphere inside the cube and the book on top.
- + High visual quality with realistic light refraction and shadows on the wooden surface.
- + Accurately renders the plant behind the cube visible through the glass.
- − The plant in the background is quite large and dominant, though still follows the prompt.
Stable Diffusion 3.5 Medium
- + Clean photographic aesthetic with nice depth of field.
- + The glass material has realistic thickness and edge highlights.
- − Failed spatial relationships by putting the book inside the cube and the sphere on top of the book.
- − The plant is hanging from the top instead of being positioned behind the cube.
- − The book doesn't clearly look like a book, appearing more like a red platform.
Verdict: OmniGen v2 followed all spatial and logical requirements of the prompt, correctly placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Medium struggled with the arrangement, incorrectly placing the red object inside the cube and positioning the plant oddly at the top of the frame. OmniGen v2 is the clear winner for its superior prompt adherence and realistic rendering of the scene.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
OmniGen v2
- + Excellent clarity and crisp rendering of the subject and bicycle.
- + Strong adherence to the 'red bicycle' and 'reflections' prompts.
- + Clean composition with a professional cinematic look.
- − Lacks the requested 'imperfect framing' and 'motion blur' from cars.
- − The lighting and skin look slightly too clean/smooth for a candid street photo.
- − The bike is being held rather than repaired.
Stable Diffusion 3.5 Medium
- + Successfully captures the 'candid', 'imperfect framing', and gritty street atmosphere.
- + Includes realistic motion blur on the background vehicles.
- + Film-like texture matches the cinematic but realistic request well.
- − Proportions of the bicycle and the man's hands are warped and anatomically incorrect.
- − The face is obscured and lacks the 'natural skin texture' requested.
- − Poor image coherence with the bike frame blending into the man's leg.
Verdict: OmniGen v2 produces a much cleaner and more aesthetically pleasing image with high technical quality, though it ignores the specific 'candid' and 'motion blur' stylistic prompts. Stable Diffusion 3.5 Medium captures the intended mood and secondary details better but fails significantly on basic anatomy and object coherence, resulting in a distorted bicycle and person. OmniGen v2 is preferred for its overall visual integrity.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
OmniGen v2
- + Excellent depiction of warm torchlight reflecting off the plate armor.
- + Beautifully detailed engraving on the metalwork.
- + Smooth, aesthetic skin textures and lifelike eyes.
- − The 'battle-worn' aspect is very light, with dirt appearing more like decorative freckles.
- − The beads in the hair are present but look more like modern hair clips.
Stable Diffusion 3.5 Medium
- + Successfully captures a grittier 'battle-worn' aesthetic with more visible scars and skin texture.
- + Highly detailed engraving on the armor which feels authentically weathered.
- + Excellent implementation of the braids and bokeh spark effects.
- − The lighting feels slightly colder than the requested 'warm torchlight' despite the sparks.
- − Minor anatomical oddities in how the hair braids emerge from the head.
Verdict: Stable Diffusion 3.5 Medium is the likely winner as it better captures the 'battle-worn' intensity and fine details of the prompt, whereas OmniGen v2 produces a much cleaner, more 'glamorized' portrait. While OmniGen v2 handles the warm lighting better, Stable Diffusion 3.5 Medium delivers more on the specific textures of scars, leather, and weathered metal.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
OmniGen v2
- + Features high-quality, professional food photography with vibrant color accents.
- + Effective use of grid-based layout and negative space for a clean aesthetic.
- + Includes the requested sections for Appetizers, Pizzas, and Mains.
- − Numerous spelling errors in headings like 'RESTAURATED MENTS' and 'PIZZZZAN'.
- − The text blocks for menu items are mostly illegible placeholder lines.
Stable Diffusion 3.5 Medium
- + More realistic density of text for a menu, including prices and item descriptions.
- + Excellent grid alignment and consistent photo styling.
- + Better text rendering for smaller fonts compared to the other model.
- − Headings are less legible and do not strictly follow the prompt's section requests.
- − The overall color palette is a bit muted compared to Model A's vibrant accents.
Verdict: OmniGen v2 excels in visual presentation and high-quality food photography that fits the 'vibrant' and 'modern' prompt, despite significant spelling errors in the headers. Stable Diffusion 3.5 Medium produces a more functional-looking menu with pricing and detailed text, but it feels slightly more cluttered and less stylistically bold.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
OmniGen v2
- + Excellent typography rendering with clean, professional-looking fonts.
- + Highly legible and well-balanced advertisement layout.
- + Colors are vibrant and fit the 'fiery' theme perfectly.
- − Completely failed the 'exploded burger' and 'suspended components' requirement, showing a standard stacked burger.
- − Missing the Euro symbol (€) in the price tag.
- − Text rendering for 'LIMITED TIME ONLY' is slightly cut off and lacks the 'fiery' effect.
Stable Diffusion 3.5 Medium
- + Successfully captured the Euro symbol (€) and the price accurately.
- + Better background texture with more realistic fire and glowing embers.
- + The burger appears to be floating, capturing some sense of suspension.
- − Failed the 'exploded burger' requirement, keeping the ingredients mostly assembled.
- − Text lacks the glowing, fiery effect requested and is quite small/plain for a primary title.
- − The 'starburst' for the price is more of a line-art sunburst and lacks the advertisement style and fire requested.
Verdict: Both models failed to deliver the 'exploded' nature of the burger where components should be separated in mid-air. OmniGen v2 produced a much better graphic design layout for an advertisement, whereas Stable Diffusion 3.5 Medium provided a more photorealistic burger and better background elements, including the correct currency symbol. OmniGen v2 is the narrow winner for its superior font rendering and overall ad-like composition, despite the assembly error.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
OmniGen v2
- + The date 'APRIL 30, 2026' is rendered clearly and accurately
- + Follows the menu items more closely with legible pricing
- + Good simulated chalk texture with smudges
- − Spelling error in the main title ('SPECALS')
- − Significant overlapping of lines and jumbled text in the center section
Stable Diffusion 3.5 Medium
- + Features a very authentic chalk aesthetic with realistic smudges and dusty background
- + Uses an artistic chalkboard layout that feels more like a real café
- − Total failure on text rendering and date accuracy
- − Failed to follow the specific 'elegant cursive' requirement for the title
- − Illegible scribbles for most of the menu items
Verdict: OmniGen v2 performed significantly better at prompt adherence, specifically regarding text accuracy and capturing the requested menu items and date. While Stable Diffusion 3.5 Medium captured the 'cozy café' atmosphere and realistic chalk texture more effectively, its inability to render legible text makes it fail the core requirements of the challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
OmniGen v2
- + Clean, vector-like aesthetic with crisp lines
- + Good lighting contrast on the horse and suit
- + Solid anatomical structures for the horse
- − Failed the negative constraint; the astronaut is riding the horse instead of the horse riding the astronaut
- − Lacks complex textures, appearing somewhat plastic or artificial
- − Simple background lacks cinematic depth
Stable Diffusion 3.5 Medium
- + Atmospheric cinematic lighting with a detailed planet horizon
- + Highly detailed star field and nebula effects
- + Good sense of scale and spatial composition
- − Failed the negative constraint; the astronaut is riding the horse
- − Deformed horse legs with extra joints and hooves
- − Poor physical integration where the rider's legs meet the horse
Verdict: Both models completely failed the specific logic constraint 'horse on top, not vice versa,' instead producing standard images of astronauts riding horses. OmniGen v2 produced a cleaner, more illustrative image with better horse anatomy, while Stable Diffusion 3.5 Medium offered a more cinematic background but suffered from significant anatomical defects in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
OmniGen v2
- + Successfully places the capybara in a dark jacket and yellow hat.
- + The businesswoman clearly shows the bored, phone-using behavior requested.
- − Anatomical failure with human hands coming out of the capybara's sleeves.
- − Perspective is awkward, making it look as though the passenger is sitting in the passenger seat rather than the back seat.
Stable Diffusion 3.5 Medium
- + Excellent prompt adherence for the capybara, featuring realistic paws on a steering wheel.
- + Distinct and accurate separation between the driver's seat and the back seat.
- + High visual quality and professional lighting that matches a cinematic photography style.
- − The businesswoman in the background is not looking at a phone as requested.
- − The 'TAXI' sign on top is partially cut off and reversed.
Verdict: Stable Diffusion 3.5 Medium is the winner because it successfully renders capybara paws operating the vehicle, whereas OmniGen v2 mistakenly attached human hands to the animal. While OmniGen v2 followed the passenger's instructions better, Stable Diffusion 3.5 Medium produced a more coherent and anatomically logical scene with superior lighting.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent typography for the main title with the correct spelling.
- + Strong visual composition with a central, glowing focal point.
- + Successfully includes almost all requested text elements, including date and location.
- − The text on the scroll banner is gibberish.
- − The bottom text section contains some repetition and slight misspellings like 'Frigts' and 'Arcas'.
Stable Diffusion 3.5 Medium
- + Features a highly detailed parchment texture and intricate thorny border.
- + Good use of space with twisted trees and multiple Jack-o-lanterns.
- − Very poor spelling in the main heading ('Halloweeen Inviloween').
- − Incorrect date '226' and multiple typos in smaller text strings.
- − Lacks the 'scroll banner' requested in the prompt.
Verdict: OmniGen v2 is the clear winner as it successfully rendered the complex title and most of the event details with legible, attractive typography. While Stable Diffusion 3.5 Medium captured the 'vintage parchment' aesthetic well, its failure to spell the primary title correctly and its nonsensical date rendering make it unusable as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent text rendering with clean, bold typography.
- + Strong adherence to the isometric miniature diorama request with a raised base.
- + Vibrant lighting and clean cartoon aesthetic.
- − The flag icon is a generic red/yellow flag rather than the requested Japan flag.
- − Anatomical oddity in the sushi design where pieces appear to be a mix of nigiri and rolls with a tail.
Stable Diffusion 3.5 Medium
- + Realistic PBR-style textures for the roe and nori seaweed.
- + Good centered composition and soft lighting.
- − Failed to include the 'raised diorama base' component of the prompt.
- − Text rendering is messy, with 'JAPAN' being too small and 'SUSHI' having typographical artifacts.
- − Missing the flag icon entirely.
Verdict: OmniGen v2 followed the architectural requirements of the prompt much better, successfully creating a miniature diorama on a raised base with clear, bold text. While Stable Diffusion 3.5 Medium had nice material textures, it failed on the text rendering and the 'diorama' structure, making it a weaker overall match for the specific prompt instructions.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
OmniGen v2
- + Strong adherence to the requested sunset lighting and god rays
- + Extremely clean, vibrant characters with expressive faces
- − Has a very stylized 3D animation look rather than the requested hyper-photorealistic style
- − Fails to include all requested animals, missing the baby bunny and featuring a generic hybrid instead of a distinct fox
- − Characters are static and posing rather than tumbling/playing
Stable Diffusion 3.5 Medium
- + Closer to a photographic style with more natural fur textures
- + Includes more variety in the animals present
- + Better depiction of the meadow environment with deeper layering of flowers
- − Failed to include the requested baby bunny
- − Anatomy issues on the central kitten, which looks like a cat/fox hybrid with elongated paws
- − Less distinct 'god rays' compared to Model A
Verdict: Both models failed to include all four requested animals, with neither producing the baby bunny. OmniGen v2 produced a high-quality but overly stylized image that resembles a Pixar movie poster rather than realism, while Stable Diffusion 3.5 Medium attempted a more realistic texture but struggled with coherent animal anatomy. Stable Diffusion 3.5 Medium is the slight winner for better texture detail and for at least attempting a more grounded photographic aesthetic.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
OmniGen v2
- + Perfect vector logo aesthetic with clean, minimalist lines.
- + Includes all requested prompt elements like the steam and specific dates.
- + Composition is balanced and professionally centered.
- − Misspelled the name as 'CAFFFLORIN'.
- − The minimalist style might feel a bit too generic for a vintage logo.
Stable Diffusion 3.5 Medium
- + Excellent hand-drawn vintage illustrative style.
- + Great texture and sophisticated choice of vintage typography.
- − Multiple spelling errors in the main name and the date ('Est 170').
- − Fails to include the cloche dome, instead using a shape resembling a tea pot or bowl.
- − Too cluttered for a 'minimalist' logo request.
Verdict: OmniGen v2 better captured the requested minimalist vector style and included the specific icon (cloche) and date accurately, despite a slight typo in the brand name. Stable Diffusion 3.5 Medium produced a more artistic vintage illustration, but it failed on core prompt instructions like the 1720 date and the minimalist requirement.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
OmniGen v2
- + Excellent adherence to the clean, flat vector aesthetic.
- + Follows the specified NASA color palette accurately.
- + Strong, organized layout that feels like a modern infographic.
- − Failed the mission number, displaying 'Apolo 17' instead of Apollo 11.
- − Text consists of garbled nonsense words.
- − Failed to include specific icon requests like the Saturn V or the trajectory arc.
Stable Diffusion 3.5 Medium
- + Successfully captured the requested steps in a sequential vertical list.
- + Good use of the color palette and cosmic texture.
- + Text is slightly more legible, though still containing many typos.
- − Failed to provide the 'flat vector' style, leaning more towards an illustrative/painterly look.
- − Composition feels cluttered and less like a professional infographic.
- − Icons are inconsistent and do not accurately represent the mission phases requested.
Verdict: OmniGen v2 followed the stylistic constraints much better, delivering a crisp flat vector design that looks like a real poster, despite getting the mission number wrong and having gibberish text. Stable Diffusion 3.5 Medium managed to include all six steps sequentially, but the visual style was too rough and failed to meet the 'clean, modern vector' requirement. Overall, OmniGen v2 is preferred for its superior layout and aesthetic coherence.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding