Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
OmniGen v2
#56 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
OmniGen v2
0.0%
win rate
Ties
0.0%
Z-Image Turbo
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
OmniGen v2
- + Excellent adherence to the glass cube geometry with clean edges.
- + Superior lighting and reflections on the blue sphere and glass surfaces.
- + Very high resolution and photographic clarity throughout the image.
- − The plant is behind the cube but doesn't show significant refraction through the glass body compared to the edges.
Z-Image Turbo
- + Accurately captures all prompt elements including the small sphere and plant placement.
- + Realistic wood texture on the table surface.
- − Lower overall resolution and slight blurriness compared to OmniGen v2.
- − The glass cube has a mirrored base which wasn't requested, creating an extra reflection.
- − The red book has some minor edge artifacts.
Verdict: OmniGen v2 produces a much sharper, high-fidelity image with more sophisticated lighting and material rendering. While Z-Image Turbo captures all requested elements, it suffers from lower clarity and introduces an unrequested mirrored bottom for the cube. OmniGen v2 is the winner due to its superior visual quality and clean execution of the spatial requirements.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
OmniGen v2
- + Excellent reflections on the wet pavement
- + Strong 50mm lens look with smooth bokeh
- + Vibrant colors and high visual clarity
- − Anatomical errors in the bicycle structure, specifically the pedals and chain guard area
- − The rain effect looks like a simple Photoshop overlay rather than a natural part of the scene
Z-Image Turbo
- + Successfully captures the requested 'motion blur' from passing cars
- + Realistic skin texture and age-appropriate clothing
- + More authentic 'candid street photo' composition
- − Muddled details in the bicycle spokes and frame
- − Lacks the distinct reflections on the pavement mentioned in the prompt
Verdict: Z-Image Turbo delivers a much more authentic candid feeling with realistic motion blur and skin textures, successfully hitting the 'no stylization' requirement. OmniGen v2 produces a more aesthetically pleasing image with beautiful reflections, but it fails on the motion blur prompt and has significant structural flaws in the bicycle's design.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
OmniGen v2
- + Excellent engraving detail on the pauldrons
- + High-contrast lighting with beautiful warm tones
- + Clean and aesthetically pleasing character design
- − Character looks too clean and 'pristine' for a battle-worn description
- − The dirt/scars look like superficial makeup dots rather than grit
- − Missing the requested bokeh sparks effect
Z-Image Turbo
- + Perfectly captures the 'battle-worn' aesthetic with realistic scars and dirt
- + Excellent adherence to all prompt elements including bokeh sparks and torchlight source
- + Highly detailed textures on the gambeson and chainmail underlayers
- − The engraving on the armor is slightly less sharp than Model A
- − Composition is a bit tighter on the right side
Verdict: While OmniGen v2 produced a beautiful and clean portrait, it failed to capture the 'battle-worn' grit and specific environmental effects like bokeh sparks. Z-Image Turbo followed the prompt much more accurately, delivering realistic skin textures, believable scarring, and a superior implementation of the torchlight and cloth textures.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
OmniGen v2
- + Excellent grid layout with a mix of text and imagery.
- + Includes high-quality, vibrant food photos and colored backgrounds.
- + Includes a wide variety of food types matching the prompt categories.
- − Significant spelling errors in large headings (e.g., 'RESTAURATED MENTS', 'APPTETIZES').
- − The font choice feels slightly more rustic than the requested 'bold sans-serif'.
Z-Image Turbo
- + Features a very clean, professional, and modern grid structure.
- + Adheres better to the text requirements with legible 'APPETIZERS' and 'PIZZA' text.
- + High visual quality and consistency across all food photos.
- − Slight spelling error in 'PIZZA MANS' and 'SE TIIION'.
- − The layout is a bit more repetitive than the diverse design of the first model.
Verdict: OmniGen v2 creates a more visually interesting and artistic menu spread, but it suffers from severe spelling issues on nearly every major heading. Z-Image Turbo provides a much more functional and professional layout with cleaner sans-serif typography and better adherence to requested section titles, making it the more usable design overall.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
OmniGen v2
- + Text is clear and prominent with a bold graphic style.
- + Vibrant lighting effects create a high-energy ad feel.
- + Includes all requested text elements including the price and starburst.
- − Failed the 'exploded burger' requirement, showing a fully assembled burger.
- − Textures look somewhat artificial and more like a 3D render than a photorealistic photo.
- − The 'limited time only' text is partially cut off on the left.
Z-Image Turbo
- + Excellent adherence to the 'exploded' and 'suspended' burger prompt.
- + Superior photorealistic textures on the meat, bun, and vegetables.
- + Successfully integrated all text with the requested glowing/fiery effect and the correct currency symbol.
- − The 'exploded' effect is slightly subtle compared to a full deconstruction.
- − Some minor artifacting in the smoke trails at the bottom.
Verdict: Z-Image Turbo is the clear winner as it followed the complex 'exploded burger' prompt while maintaining a much higher level of photorealistic detail. OmniGen v2 failed to explode the burger components and provided a more generic, plastic-looking 3D render with cut-off text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
OmniGen v2
- + The frame has a clean and professional look.
- + Attempts to follow the chalk texture requirement.
- − Contains numerous spelling errors like 'Specals' and 'glte free'.
- − Text layout is cluttered and overlapping significantly.
- − Failed to render the text in elegant cursive as requested.
Z-Image Turbo
- + Excellent text legibility and significantly better spelling accuracy.
- + The chalk texture and smudge marks on the board look highly realistic.
- + Well-organized layout that follows the hierarchy of the prompt.
- − One minor spelling error in 'Mustroom' instead of 'Mushroom'.
- − Did not use an 'elegant cursive' font for the title as specified.
Verdict: Z-Image Turbo is the clear winner as it produced a highly legible, realistic chalkboard with only one minor spelling error and excellent composition. OmniGen v2 failed significantly on text rendering, resulting in misspelled words, overlapping lines, and poor overall aesthetics.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
OmniGen v2
- + Successfully captures a surreal, clean aesthetic.
- + High contrast and vibrant colors create a cinematic feel.
- + Good preservation of the horse's anatomy despite the unusual setting.
- − Failed the primary logic constraint of having the horse on top of the astronaut.
- − The composition is quite static with the horse seemingly standing on an invisible floor.
Z-Image Turbo
- + Dynamic posing makes the 'riding in space' concept feel more active.
- + Realistic lighting and texture on the space suit.
- + Detailed rendering of the horse's coat and mane.
- − Failed the primary logic constraint of having the horse on top of the astronaut.
- − A bit of artifacting near the horse's hooves against the starfield.
Verdict: Both models failed the specific prompt constraint to have the horse on top of the astronaut, instead defaulting to the typical 'astronaut riding horse' trope. OmniGen v2 produced a cleaner, more stylized illustration, while Z-Image Turbo opted for a more realistic and dynamic composition that feels more cinematic, though it still suffers from the same prompt adherence failure.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
OmniGen v2
- + Excellent high-contrast colors and clean lighting
- + The capybara's facial texture is very sharp and realistic
- − Major logic failure with human hands emerging from the capybara's sleeves to steer
- − The human passenger is placed in the front passenger seat rather than the back seat as requested
Z-Image Turbo
- + Accurately depicts paws on the steering wheel instead of human hands
- + Correctly places the businesswoman in the back seat as specified in the prompt
- + Superior adherence to the specific 'taxi driver cap' style
- − The transition between the capybara's neck and the jacket is slightly awkward
- − The blurred background lights are less vibrant than Model A
Verdict: Z-Image Turbo is the clear winner because it correctly follows the complex spatial and anatomical requirements of the prompt. While OmniGen v2 has high visual clarity, it fails significantly by giving the capybara human hands and placing the passenger in the front seat, whereas Z-Image Turbo realistically depicts paws and the correct seating arrangement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
OmniGen v2
- + Legible and stylish main title text
- + Clean graphic illustration style
- − Garbled text in the banner and event details
- − Incorrect spelling of the location 'The Arches'
- − Lacks the requested 'thorns' in the border
Z-Image Turbo
- + Excellent adherence to all visual elements including thorns and webs
- + Highly accurate text for the date, time, and location
- + Atmospheric, cinematic lighting on a 3D-style central jack-o-lantern
- − Minor typo in the location ('Archves' instead of 'Arches')
- − The scroll banner is split and looks slightly awkward behind the pumpkin
Verdict: Z-Image Turbo is the clear winner as it successfully incorporated the specific border elements like thorns and webs while maintaining high text accuracy for the event details. OmniGen v2 produced a clean layout but failed significantly on the text rendering for the lower half of the invitation and missed the thorn requirement.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
OmniGen v2
- + Excellent text rendering with clean 3D drop shadows.
- + Vibrant lighting and appealing PBR material representation for the fish.
- + Superior composition that feels like a professional 3D render.
- − The flag icon is generic and does not represent the Japanese flag.
- − The sushi pieces contain odd black squares in the rice centers.
Z-Image Turbo
- + Successfully included a recognizable flag icon, though it is the incorrect country.
- + Clean isometric base and smooth cartoon textures.
- − Incorrectly used the flag of China for a prompt specifically requesting Japan.
- − The text rendering is flatter and less stylistically cohesive than the subject matter.
- − The sushi roll/nigiri hybrid looks less anatomically correct than Model A.
Verdict: OmniGen v2 is the winner due to its superior 3D stylization, lighting, and high-quality text rendering, which perfectly matches the 'bold' and 'miniature 3D cartoon' aesthetic requested. While Z-Image Turbo followed the flag icon instructions more literally, it mistakenly used the flag of China for a Japanese-themed prompt and produced a less polished visual result.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
OmniGen v2
- + Strong implementation of god rays and sunrise lighting
- + Vibrant colors that enhance the joyful theme
- − Failed to include all four requested animals, missing the bunny entirely
- − Highly stylized, 3D-render aesthetic rather than the requested hyper-photorealistic style
- − Animals are sitting static instead of 'playfully chasing' and 'tumbling'
Z-Image Turbo
- + Successfully included all four animals: golden retriever, kitten, bunny, and fox kit
- + Achieved a much higher level of photorealism in fur texture and lighting
- + Captures an active, playful composition with animals actually interacting
- − The fox kit has a slightly distorted eye structure
- − Small artifacting around the kitten's paws
Verdict: Z-Image Turbo is the clear winner as it followed the prompt's count requirement (4 animals) and style (photorealistic), whereas OmniGen v2 missed one animal and produced a cartoonish, illustrated look. Z-Image Turbo also captured the requested 'tumbling' action successfully, providing a much more dynamic and realistic scene.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
OmniGen v2
- + Successfully converts the image into a clear anime style.
- + Preserves the iconic composition and colors while adding the requested 'warm' mood.
- + Characters' facial expressions and clothing patterns translated well to the new aesthetic.
- − The facial expressions are too friendly, losing the humorous 'jealousy' and 'shame' dynamic of the original meme.
- − The style is more generic modern anime than the specific hand-painted Ghibli texture requested.
Z-Image Turbo
- + Matches the 'soft pastel colors' and 'gentle lighting' instruction very effectively.
- + Maintains the exact facial expressions and poses of the original photography.
- − Completely fails the primary instruction to transform the photo into an illustration.
- − The output is just a slightly filtered photograph rather than an artistic rendition.
- − Lacks the 'hand-painted textures' requested in the prompt.
Verdict: OmniGen v2 successfully followed the core instruction to transform the source image into an illustration, capturing the composition and colors of the original meme perfectly in a new medium. Z-Image Turbo merely applied a soft color grade to the photo, failing to create an illustration at all. While OmniGen v2 lost some of the narrative nuance in the characters' faces, it is the clear winner for actually executing the requested transformation.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
OmniGen v2
- + Excellent adherence to the 'windy' hair instruction with clear directional flow.
- + Added leaves are bright and stylistically fit the energetic vibrance.
- + The overall image is high-contrast and very lively.
- − Failed to preserve the source image, essentially regenerating the scene with a different woman and dog.
- − Loss of environmental detail from the original (e.g., the bridge and water have been smoothed/simplified).
Z-Image Turbo
- + Excellent source preservation, maintaining the original woman, dog, and background layout.
- + Successfully added flying leaves while keeping the realism of the original scene.
- + Subtle but effective hair movement that looks natural.
- − The motion in the hair is less 'dynamic' than Model A, appearing as a slight breeze rather than strong wind.
- − A few leaves look slightly like small artifacts due to their size.
Verdict: OmniGen v2 produces a more energetic result but fails the core task of an edit by completely regenerating the subject and dog, losing the original's identity. Z-Image Turbo successfully applies the requested motion and adds leaves while preserving the specific people and locations from the source image, making it the superior editing tool.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
OmniGen v2
- + Clean vector emblem style
- + Includes the banner element requested
- + Correctly identifies the date establishment
- − Misspelled the name as 'CAFFFLORIN'
- − Line work on the cloche is slightly asymmetrical
Z-Image Turbo
- + Perfect spelling of 'Caffè Florian' including the accent
- + Professional minimalist vector aesthetic
- + High visual clarity and balance
- − Omitted the requested banner element
- − Text is centered below rather than within a cohesive emblem
Verdict: OmniGen v2 followed the layout instructions better by including the banner, but failed significantly on spelling. Z-Image Turbo produced a much more professional and correctly spelled logo, although it replaced the requested banner with simple lines. Z-Image Turbo is the winner for accurate text rendering and superior aesthetic polish.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
OmniGen v2
- + Matches the requested NASA-inspired color palette perfectly.
- + Clean professional graphic design layout with grid structure.
- + Adheres strictly to the flat-vector style with crisp lines.
- − Significant spelling errors in every label.
- − Incorrect mission number (Apollo 17 instead of 11).
- − Icons are abstract and don't clearly represent the specific steps requested.
Z-Image Turbo
- + Features a much better Saturn V and Lunar Module illustration.
- + Contains fewer spelling errors and correctly identifies the mission as Apollo 11.
- + Composition clearly represents the specific steps of the mission.
- − Fails to include the 'muted red' and 'navy' primary background aesthetic requested.
- − Icons are inconsistent in style compared to Model A.
- − Includes a strange yellow/orange color on the lander that deviates from the palette.
Verdict: OmniGen v2 produces a much better 'infographic' layout with a professional aesthetic, but fails significantly on text accuracy and icon specificity. Z-Image Turbo creates better individual illustrations for the requested steps and gets the mission name right, but the overall design feels less like a cohesive poster. Z-Image Turbo is the likely winner for actually attempting all instructions, despite the weaker layout.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering