An image generation model by xAI designed to generate highly aesthetic images from text descriptions.
Settled by community votes across 7 shared challenges, with an AI judge weighing in on each.
Grok Imagine Image
#25 of 62 in Text-to-Image
Qwen Image Edit 2511
#16 of 32 in Image Editing
Where the votes landed
Grok Imagine Image
38.1%
win rate
Ties
9.5%
Qwen Image Edit 2511
52.4%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
Grok Imagine Image
- + Excellent preservation of the white Rolls-Royce Phantom Drophead Coupe's external design and details.
- + Highly realistic integration of the car onto a coastal road with motion blur.
- + Accurate depiction of the California coastline environment.
- − Completely failed to use the specific man provided in the second source image, substituting a generic older white male driver.
Qwen Image Edit 2511
- + Successfully preserved the identity, hairstyle, and clothing of the man from the source image.
- + Accurately places the subjects in a California coastline setting as requested.
- + Good composition that focuses on the specific characters provided in the task.
- − Failed to preserve the specific car model, changing the luxury Rolls-Royce to a different convertible with a brown wood interior.
- − The steering wheel and dashboard have some warped geometry/perspective.
Verdict: This comparison highlights a classic trade-off in image editing: Grok Imagine Image perfectly preserved the car but ignored the specific person provided, while Qwen Image Edit 2511 perfectly preserved the person but changed the car. Qwen is the preferred model for this specific task because identifying and placing the person from the second source image is a more complex multi-modal requirement than simply keeping the car, even though it struggled with the interior details of the vehicle.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
Grok Imagine Image
- + Preserves the composition and background of Image 1 perfectly.
- + Maintains high resolution and realistic lighting.
- − Completely ignored the instruction to use Image 2 as the character reference.
- − Failed to change the person's face, hair, or gender.
Qwen Image Edit 2511
- + Successfully included the character from Image 2 in the frame.
- + Matches the background texture of Image 2.
- − Failed the primary task of replacing the person in the pose with the character from Image 2.
- − Overlayed the two images awkwardly, resulting in a physical impossibility where the man stands behind the original person.
- − The original woman from Image 1 remains in the foreground unchanged.
Verdict: Both models failed significantly on this complex character swap and pose-matching task. Grok Imagine essentially ignored the character reference (Image 2) entirely, providing only a slight variation of the original Image 1, while Qwen Image Edit 2511 simply composited the character from Image 2 behind the original person, failing to integrate them into the requested pose. Qwen is slightly more aware of the visual elements requested, but neither achieved the requested edit.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
Grok Imagine Image
- + Successfully preserved the exact face and hair of the subject from Image 1.
- + High resolution and sharp detail on the clothing textures.
- + Good lighting and shadow integration for the garment.
- − Completely failed to use the clothing from Image 2, generating a random regal outfit instead.
- − One hand is partially merged with the fabric of the robe.
Qwen Image Edit 2511
- + Successfully identified and utilized the correct outfit, scarf, and sunglasses from Image 2.
- + Maintained the beach background from Image 1.
- + Accurately positioned the clothing onto the subject's pose.
- − Failed to preserve the model's face and hair, replacing the subject with the man from Image 2.
- − Loss of sharp detail and clarity compared to the original images.
- − Subject appears poorly blended with the sand at the feet.
Verdict: Both models failed significantly on specific instructions. Grok Imagine excelled at maintaining the original person's identity but completely ignored the reference clothing provided in Image 2. Qwen Image Edit 2511 followed the clothing instructions much better but failed to keep the base person's face, essentially performing a background swap rather than an outfit swap. Grok is slightly preferred for maintaining the core identity and higher visual quality, despite the incorrect outfit.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Grok Imagine Image
- + Excellent source preservation, keeping the face and lighting identical
- + Very realistic and natural-looking hair texture and density
- + Perfectly integrated hairline that matches the sideburns
- − The hair is quite short, bordering on a buzz cut rather than 'full'
Qwen Image Edit 2511
- + Successfully added a very thick and full head of hair
- + Maintained the original facial features and background correctly
- − Hair texture and curls look oily or plastic-like
- − The lighting on the hair doesn't perfectly match the environment
- − The hair volume is slightly exaggerated, looking like a wig
Verdict: Grok Imagine produced a much more realistic result by seamlessly integrating the new hair with the existing beard and lighting, though the style is shorter than expected. Qwen Image Edit 2511 followed the 'full and thick' instruction more literally, but the hair texture looks artificial and poorly blended with the rest of the image.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Grok Imagine Image
- + Successfully incorporates all elements: tv anchor desk, hockey rink/pucks/skates, and dogs.
- + Excellent preservation of the subject's facial features while achieving a caricature effect.
- + High visual quality with clean rendering and legible text.
Qwen Image Edit 2511
- + Bold, classic comic-book caricature style with vibrant colors.
- + Captures the 'tv anchor' and 'dog' aspect of the prompt well.
- + The face is very expressive and matches the humorous tone requested.
- − Completely misses the 'hockey' requirement of the prompt.
- − Contains gibberish text in the speech bubble and nonsensical elements like a candle in front of a laptop.
- − Hand and microphone rendering is messy and anatomically incorrect.
Verdict: Grok Imagine followed the prompt much more accurately, successfully integrating the hockey theme which Qwen Image Edit 2511 completely omitted. While both models did a good job capturing the subject's likeness in a caricature style, Grok Imagine's composition is more coherent and lacks the nonsensical artifacts found in Qwen's output.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Grok Imagine Image
- + Excellent capture of the Ghibli art style with hand-painted textures and soft line work.
- + Preserves the composition and poses of the iconic source image perfectly.
- + Creates a beautiful, dreamy watercolor-style background that fits the nostalgic theme.
- − The facial features are simplified significantly compared to the original subjects.
Qwen Image Edit 2511
- + Successfully translates the image into a clean anime illustration style.
- + Retains more recognizable facial likeness to the original people in the photo.
- + Maintains high clarity and clean linework throughout the composition.
- − The style leans more toward generic modern digital anime rather than the specific 'Ghibli' hand-painted look requested.
- − The lighting is flat compared to the soft, atmospheric glow in Model A.
Verdict: Both models successfully interpreted the scene into an illustration, but Grok Imagine Image (Model A) followed the specific 'Studio Ghibli' and 'hand-painted' instructions much more accurately with its watercolor textures and soft edges. While Qwen Image Edit 2511 (Model B) preserved the likenesses of the subjects better, it lacked the specific artistic charm and nostalgic mood requested in the prompt.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Grok Imagine Image
- + Successfully added a large volume of falling leaves
- + Effectively modified the hair and dog's ears to show wind direction
- + High source preservation with minimal facial distortion
- − The wind effect on the hair is slightly less dramatic than the competitor
- − Added leaves appear slightly flat and repetitive in texture
Qwen Image Edit 2511
- + Excellent dynamic hair movement with a more realistic 'blown' appearance
- + Leaves show varied colors (green and orange) and motion blur for energy
- + Perfect preservation of the original subjects and background
- − Fewer leaves overall compared to Model A
- − A few digital artifacts where the hair strands meet the background
Verdict: Both models followed the instructions well, preserving the source image while adding motion. Qwen Image Edit 2511 is the winner due to the superior quality of the 'hair blowing' effect and the use of motion blur on the leaves, which creates a much more energetic feel than Grok Imagine Image's static-looking leaves.
Explore each model
Alibaba's Qwen image editing model for instruction-based image modifications and transformations