Head to head
Esc

Models · slot A

to navigate to pick

Grok Imagine Image xAI OmniGen v2 VectorSpaceLab

Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.

Grok Imagine Image

23.4 arena score

#26 of 62 in Text-to-Image

Skill signature · Text-to-Image

OmniGen v2

16.8 arena score

#57 of 62 in Text-to-Image

Vote tally

Where the votes landed

Grok Imagine Image

100.0%

win rate

Ties

0.0%

OmniGen v2

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 15

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent photorealistic texture on the wooden table and book.
  • + Realistic glass refractions and caustics on the table surface.
  • + Accurately places the plant behind the glass with realistic distortion.
  • The blue sphere is levitating unnaturally in the center of the cube.
  • The cube is more of a hollow rectangular prism than a solid geometric cube.

OmniGen v2

  • + Solid cube geometry with a logical placement of the blue sphere on the bottom surface.
  • + Clean, vibrant colors and clear adherence to all prompt elements.
  • + Strong lighting consistency and soft shadows.
  • The texture of the wooden table is a bit flat and looks more like laminate.
  • The plant behind the glass lacks the realistic refractive distortion seen in the other model.

Verdict: Grok Imagine produces a more convincing photographic result with impressive material physics, particularly the way the glass refracts the plant and table. However, OmniGen v2 provides a more logical spatial arrangement of the objects, as the sphere in Grok Imagine appears to be floating. Grok Imagine is the likely winner for its superior visual fidelity and realistic rendering of the specific lighting request.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 'imperfect framing' and 'candid' street photography style.
  • + Realistic motion blur from the passing vehicle and natural film-like grain.
  • + Highly authentic bicycle details and believable lighting.
  • The man's face is mostly hidden, though this fits the 'candid' instruction.
  • Slightly less clarity on the man's hands.

OmniGen v2

  • + Clear facial details and identifiable features.
  • + Nice reflections on the wet pavement.
  • Fails on the 'motion blur from passing cars' instruction as the car is stationary.
  • The rain effect looks like a static overlay rather than a natural environmental element.
  • The bicycle geometry is incorrect, particularly where the frame meets the rear wheel and pedal area.

Verdict: Grok Imagine produced a far more convincing 'candid' photograph that feels like a real moment captured on a 50mm lens, successfully incorporating the motion blur and imperfect framing requested. OmniGen v2 created a more staged-looking image with significant anatomical errors in the bicycle and failed to include the requested motion blur.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Extremely detailed and intricate ornate engraving on the plate armor
  • + More realistic skin texture and lifelike eyes
  • + Successful execution of the cloth underlayer and leather straps with high-resolution texture
  • The sparks look slightly like static dots in some areas

OmniGen v2

  • + Good interpretation of hair braids and beads
  • + Clean composition with clear warm lighting focus
  • Armor engravings are much simpler and lack the 'ornate' quality of the prompt
  • Leather straps and cloth layers are missing or very basic in detail
  • The dirt on the face looks specifically like painted dots rather than 'battle-worn' grime

Verdict: Both models captured the essence of the prompt, but Grok Imagine Image is significantly more detailed, particularly in the intricate filigree of the armor and the realistic texture of the skin and leather. OmniGen v2 produced a cleaner, simpler image that felt less 'battle-worn' and lacked the requested texture in the underlayers.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography with mostly legible section headers and item names
  • + Clean, professional layout that closely mimics a real-world restaurant menu
  • + The food photography is integrated naturally with the text, creating high visual appeal
  • Contains several repeating items like 'Steak Frites' and 'Grilled Salmon'
  • Some minor spelling errors in smaller text and headers (e.g., 'Veggie Deasiy')

OmniGen v2

  • + Strong use of vibrant color blocks as accents
  • + Clear grid structure for the top section photography
  • Severe spelling errors in headers (e.g., 'Apptetizes', 'Restaurated Ments')
  • The layout feels more like a flyer or a rough template than a functional professional menu
  • Lower resolution/clarity in the food imagery compared to the other model

Verdict: Grok Imagine produces a much more realistic and usable menu design with professional typography and logical item placement. While it has some repetitive entries, OmniGen v2 fails significantly on text rendering and professional aesthetic, featuring numerous nonsensical spelling errors and a disjointed layout.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Grok Imagine Image
OmniGen v2
100% wins 0% ties 0% wins

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 'exploded burger' requirement with dynamic component suspension.
  • + Text is perfectly rendered, including the specific currency symbol and fiery effect.
  • + Photorealistic texture on the meat patty, lettuce, and bun.
  • The starburst element for the price is a bit busy compared to the rest of the professional composition.

OmniGen v2

  • + Bold, clean graphic design that feels like a professional fast-food advertisement.
  • + Good integration of the logo with the fiery background glow.
  • Completely failed the 'exploded burger' request, showing a static, assembled burger.
  • Text is cut off on the left side ('LIMITED TIME ONLY').
  • Missing the Euro symbol (€) requested in the price starburst.

Verdict: Grok Imagine followed every detail of the prompt, particularly the challenging 'exploded view' and specific text requirements. OmniGen v2 failed to explode the burger components and had significant text clipping issues, making it unsuitable for the requested advertisement style.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent text rendering with perfect spelling and grammar.
  • + Realistic chalk texture with convincing dust and smudges on the board.
  • + Consistent handwriting style that maintains a natural, non-digital look.
  • The title is in print-style handwriting rather than the requested 'elegant cursive'.

OmniGen v2

  • + Good contrast between the chalkboard and the frame.
  • + Layout follows a traditional menu structure.
  • Significant spelling errors throughout the text (e.g., 'SPECALS', 'Lemont', 'Octapus').
  • Text becomes unintelligible gibberish in the middle sections.
  • The handwriting looks more like a digital brush than natural chalk on a textured surface.

Verdict: Grok Imagine is the clear winner as it successfully rendered all of the requested text with perfect spelling and a highly realistic chalk texture. OmniGen v2 struggled significantly with text adherence, producing numerous spelling errors and garbled characters that make the menu unreadable.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Successfully follows the logical constraint of the horse being 'on top' through a surreal interpretation.
  • + Excellent use of lighting, color, and cinematic nebula effects.
  • + High level of detail in the astronaut's suit and the horse's anatomy.
  • The horse has an extra leg merging into its front right area.
  • The relationship between the horse's hoof and the astronaut's hand is slightly ambiguous.

OmniGen v2

  • + Clean, simple composition with clear subjects.
  • + Solid rendering of the astronaut's gear and the horse's coat.
  • Completely fails the negative constraint of 'horse on top, not vice versa'.
  • Flat, less detailed background compared to the 'cinematic' request.
  • Anatomical errors in the horse's legs and hooves.

Verdict: Grok Imagine is the clear winner as it successfully interpreted the surreal and difficult prompt requirement of the horse being on top of the astronaut. OmniGen v2 defaulted to a standard 'astronaut riding a horse' image, failing the primary logical constraint of the text. Grok Imagine also provided a much more detailed and atmospheric background in line with the 'cinematic' keyword.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent photorealism with realistic lighting and textures.
  • + Accurately follows the prompt by placing the capybara in the driver's seat and the passenger in the back.
  • + Capture the requested 'bored' expression of the passenger perfectly.
  • The capybara's paws are rendered more like claws/fingers than natural capybara anatomy.

OmniGen v2

  • + Features a sharp, high-contrast image.
  • + The capybara's face is well-rendered with clear textures.
  • Significant failure in composition: the passenger is sitting in the driver's seat with the capybara on her lap or merged through her.
  • Human hands are shown steering the wheel instead of the capybara's paws.
  • The passenger is in the front seat instead of the requested back seat.

Verdict: Grok Imagine is the clear winner as it successfully interprets the spatial requirements of the prompt, placing the capybara and passenger in their correct respective seats with highly realistic lighting. OmniGen v2 fails fundamentally in its composition, merging the passenger and driver into the same seat and using human hands to steer, which directly contradicts the prompt.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent text rendering with perfect spelling and placement.
  • + Strong adherence to all visual elements including thorns, webs, and scroll banner.
  • + Sophisticated cinematic lighting and authentic vintage parchment texture.
  • The transition between the top thorns and the moon feels slightly cluttered.

OmniGen v2

  • + Bold, high-contrast color palette with a vibrant glow.
  • + Clear and legible title text.
  • Failed to render the event details accurately, with numerous misspellings and garbled text.
  • The scroll banner contains illegible gibberish.
  • The composition feels flatter and more like a modern digital graphic than a vintage gothic invitation.

Verdict: Grok Imagine followed the prompt with high precision, successfully rendering all required text perfectly while maintaining a cohesive, atmospheric gothic aesthetic. In contrast, OmniGen v2 struggled significantly with text legibility and failed to include the specific details and scroll content requested.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent text rendering and placement that perfectly follows the prompt.
  • + Realistic 3D material feel with good wood grain and subsurface scattering on the fish.
  • + Captures the Japanese flag correctly above the text.
  • The sushi rice grains look slightly like bubbles or pearls rather than traditional rice.

OmniGen v2

  • + Clean isometric diorama base with distinct layers.
  • + Good 3D drop shadow effect on the typography.
  • Fails to include the Japanese flag correctly, showing a generic yellow/red flag instead.
  • The sushi design is illogical, featuring a 'tail' on one piece that doesn't match standard sushi types.
  • Less variety in the sushi types compared to the other model.

Verdict: Grok Imagine is the clear winner as it accurately rendered the Japanese flag and provided a much more realistic and appetizing variety of sushi. While OmniGen v2 followed the isometric diorama layout well, it failed on the specific flag request and the sushi models felt more like generic plastic toys than high-quality 3D renders.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Successfully included all four requested animals: puppy, kitten, bunny, and fox.
  • + Dynamic and energetic composition showing the animals actually tumbling and chasing.
  • + Better rendering of textures with more realistic lighting and dew sparkles.
  • The fox kit has somewhat stylized, oversized features relative to the others.
  • The background is quite blurry, though it fits the bokeh request.

OmniGen v2

  • + Clearly includes the butterflies mentioned in the prompt.
  • + Bright, vibrant colors that fit a 'wholesome' vibe.
  • Failed to include the baby bunny, showing only three animals.
  • The image has a very flat, digital cartoon-like aesthetic rather than 'hyper-photorealistic'.
  • The pose is static and misses the 'tumbling together' action requested.

Verdict: Grok Imagine followed the complex prompt much more accurately by including all four specific animals and capturing the sense of movement and 'tumbling'. While OmniGen v2 included the butterflies, it failed to generate the bunny and resulted in a stylistically flat, illustrated look that ignored the 'hyper-photorealistic' requirement.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellently captures the painterly, watercolor-textured background typical of Ghibli films
  • + Maintains the specific facial expressions and character features of the original people
  • + Preserves the structure and architectural detail of the background faithfully

OmniGen v2

  • + Successfully translates the scene into a clean anime art style
  • + Matches the character poses and clothing colors accurately
  • Faces look generic and have lost the distinct 'distracted' and 'indignant' expressions of the original meme
  • Colors are too saturated and flat, failing the 'soft pastel' and 'hand-painted' requirements
  • Lighting is flat and lacks the requested warm, nostalgic mood

Verdict: Grok Imagine Image provides a superior edit by perfectly capturing the atmospheric qualities of a Studio Ghibli film, including the hand-painted textures and soft lighting. While OmniGen v2 produces a clean anime illustration, it loses the specific emotional nuance of the source image and ignores the prompt's request for soft pastel colors and a nostalgic mood. Grok manages to balance the stylization with an impressive preservation of the original subjects' likenesses and expressions.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
Grok Imagine Image
Before After
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent prompt adherence with numerous flying leaves throughout the scene.
  • + Strong preservation of the original subject's face, clothing, and background details.
  • + Dynamic hair and canine ear motion that feels natural to the source.
  • The sheer number of leaves can feel a bit cluttered and overlayed.
  • Slight loss in texture sharpness compared to the original image.

OmniGen v2

  • + Natural and elegant hair motion blowing to the left.
  • + Clean composition with few distractions.
  • Failed to include sufficient 'flying leaves,' with only a few stylized orange shapes present.
  • Significant changes to the subject's face and the dog's appearance, moving away from being a faithful edit.
  • The environment textures (grass and flowers) have been smoothed out into a more illustrative style.

Verdict: Grok Imagine is the clear winner as it successfully follows the editing instructions while maintaining the identity of the woman and the dog from the source image. In contrast, OmniGen v2 significantly alters the subject's facial features and simplifies the background textures, making it more of a recreation than an edit, and largely fails to add the requested flying leaves.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography with correct spelling and accent marks
  • + Sophisticated vector style with subtle grain texture
  • + Includes a creative 'caffe' cup hidden within the cloche design
  • Repeats the 'Est. 1720' text twice which creates layout redundancy

OmniGen v2

  • + Follows the 'banner' requirement for the typography effectively
  • + Clean minimalist vector lines
  • + Good use of warm cream and brown tones
  • Major spelling errors in the brand name ('CAFFFLORIN')
  • Steam effect is very thin and lacks the requested vintage feel

Verdict: Grok Imagine Image is the clear winner as it successfully renders the text 'Caffè Florian' with perfect spelling and professional typography. While OmniGen v2 has a nice banner layout, its failure to spell the primary brand name correctly makes it unusable as a logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Grok Imagine Image
OmniGen v2

AI Judge Analysis

Grok Imagine Image

  • + Successfully included all six requested steps in the correct order.
  • + Captures the flat-vector infographic style perfectly with a cohesive NASA color palette.
  • + Excellent text legibility and accuracy for main headers and crew names.
  • Minor spelling glitches in secondary text like '3rajcoory' and 'Transluiory'.
  • The icons for Earth Orbit and Lunar Orbit are very similar, both featuring rings.

OmniGen v2

  • + Features a clean, modern layout with a strong grid system.
  • + Adheres strictly to the requested muted red and navy color palette.
  • Fails to follow the sequential 6-step prompt, showing only 3 distinct iconography areas.
  • Severe spelling errors, including referring to the mission as 'Apolo 17' instead of Apollo 11.
  • The iconography is abstract and does not represent the requested Saturn V or Lunar Module effectively.

Verdict: Grok Imagine is the clear winner as it followed every step of the complex 6-part instructional prompt and maintained a high level of text accuracy for a generative model. OmniGen v2 failed significantly on prompt adherence, providing the wrong mission number, missing half the steps, and producing mostly illegible text.

Next steps

Explore each model