OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Recraft V4.1
#53 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
100.0%
win rate
Ties
0.0%
Recraft V4.1
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Accurate rendering of the glass cube structure
- + Clear visibility of the plant through the glass
- − The sphere appears slightly large relative to the cube
Recraft V4.1
- + Natural wood grain and reflections on the table
- + Good use of light and shadow
- − The sphere is unrealistically levitating in the center
- − The plant's occlusion through the glass is less distinct
Verdict: Both models followed the complex spatial instructions well. GPT Image 1 is slightly better because it places the sphere on the floor of the cube naturally, whereas Recraft V4.1 features a levitating sphere that looks like a physics error.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Natural skin texture and detailed facial features
- + Excellent use of shallow depth of field and bokeh
- − Anatomical issues with the subject's left hand
- − Missing the requested motion blur on the passing vehicles
Recraft V4.1
- + Better capture of the requested motion blur on background cars
- + Authentic feeling rain texture and atmosphere
- − Unnatural posture and neck anatomy of the subject
- − Lower resolution and visible noise in dark areas
Verdict: GPT Image 1 produces a more coherent and visually appealing portrait with superior details on the man's face and hands. Recraft V4.1 succeeds better at technical requests like motion blur and rain intensity, but suffers from distorted human proportions.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture with realistic pores and subtle imperfections.
- + Highly intricate engraving detail on the pauldrons and gorget.
- + Effective use of warm rim lighting that enhances the atmosphere.
- − The braids are less distinct compared to the other model.
- − The background is very dark, losing some potential for environmental storytelling.
Recraft V4.1
- + The hair braids with beads are rendered with high clarity and detail.
- + Stronger adherence to the 'battle-worn' prompt with a prominent facial scar.
- + Excellent texture on the leather strap and cloth underlayer.
- − The skin rendering looks slightly smoother and more digital than Image A.
- − The facial expression is a bit vacant compared to the intensity in Image A.
Verdict: Both models followed the prompt exceptionally well, but Recraft V4.1 succeeded more effectively in showcasing all the requested decorative details, such as the beads in the braids and the leather straps. While GPT Image 1 has superior, more lifelike skin textures and a more compelling character expression, Recraft V4.1's overall composition and clarity in the engraving make it a more well-rounded response to the specific prompt items.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Features large, high-quality, appetising food photography.
- + Excellent use of bold sans-serif typography and vibrant colored icons.
- + Very clean and professional layout that feels balanced and modern.
- − Nonsense filler text for descriptions ('Apperoiation descrigion').
- − Includes a 'Main' item under the 'Pizza' heading, showing a minor logical error.
- − The grid of photos obscures some potential room for more text-heavy menu items.
Recraft V4.1
- + Excellent text rendering with real ingredients and dish names.
- + Clear logical separation of categories including a specific 'Mains' section.
- + Good use of professional design elements like colored underlines and status tags like 'GLUTEN FREE'.
- − The food photos are much smaller and less high-fidelity than those in Image A.
- − The right-aligned image column creates a lot of white space in the center that feels slightly unpolished.
- − Composition is a bit bottom-heavy and lacks the 'wow' factor of the large hero images.
Verdict: GPT Image 1 produces a more visually striking 'hero' style menu with superior image quality, but it fails on text legibility by using gibberish. Recraft V4.1 functions much better as an actual menu design with perfectly rendered English text and appropriate categories, making it more useful for practical design work despite smaller photo assets. GPT Image 1 is preferred for pure aesthetic inspiration, but Recraft V4.1 wins for utility and professional execution.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the meat and bun.
- + Very clean typography that perfectly matches the 'glow' request.
- + Superior lighting consistency where the burger appears illuminated by the surrounding embers.
- − Missed a digit in the price starburst, displaying '.99' instead of '6.99'.
- − The 'exploded' effect is quite vertical and stiff, lacking some dynamic motion.
Recraft V4.1
- + Successfully included all requested text and the correct price.
- + Dynamic, diagonal composition creates a stronger sense of movement and suspension.
- + The fiery font style is very creative and fits the theme well.
- − The starburst element looks like a flat vector graphic, clashing with the photorealistic burger.
- − The lighting on the burger components is a bit cool and disconnected from the fiery background.
Verdict: GPT Image 1 offers significantly better visual quality and lighting, making the food look appetising and the text look integrated into the scene, though it failed to correctly render the price digits. Recraft V4.1 followed the text instructions more accurately and had a more dynamic layout, but the inconsistent art styles between the starburst and the food make it look like a low-quality composite.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and accuracy
- + Rich chalk texture and realistic lighting on the board
- − The handwriting looks a bit too uniform/perfect, bordering on a font-like appearance
- − Failed to render the title in the requested 'elegant cursive' style
Recraft V4.1
- + Successfully captured the 'handwritten' feel with varying pressure and organic imperfections
- + Better depiction of the chalkboard surface with smudges and dust
- + Includes the complete third menu item context
- − The letters and numbers are slightly less clean and harder to read at a glance
- − Also failed to render the title in 'elegant cursive'
Verdict: Both models followed the complex prompt with nearly perfect spelling, which is impressive. Recraft V4.1 is the winner because its handwriting feels more authentic to a real person writing on a chalkboard, whereas GPT Image 1 looks slightly more like a digital font overlay due to its extreme spacing and letter consistency.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent anatomical realism for both the horse and the spacesuit texture.
- + Moody and cinematic lighting with high contrast against the dark space background.
- + Clear composition that balances the subject with the curve of the planet below.
- − Failed the specific spatial instruction for the horse to be on top of the astronaut.
Recraft V4.1
- + Vibrant and colorful interpretation of space with detailed nebulae and asteroid field.
- + Dynamic lighting with a strong back-light source creating a glowing mane effect.
- + High level of detail in the background environment.
- − Failed the complex instruction to have the horse on top of the astronaut.
- − Minor anatomical wonkiness where the astronaut's leg meets the saddle.
Verdict: Both models failed the specific prompt constraint to have the horse on top of the astronaut, instead defaulting to the standard image of an astronaut riding a horse. GPT Image 1 is preferred because of its superior lighting, texture realism, and more grounded cinematic aesthetic compared to the slightly cluttered background of Recraft V4.1.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic quality with a shallow depth of field.
- + Accurately represents the requested 'dark jacket' and a more formal taxi cap.
- + The capybara's expression is very professional and human-like in its posture.
- − The paws on the steering wheel look more like primate hands than capybara paws.
- − The human passenger is seated directly in the front/middle rather than clearly in the back seat.
Recraft V4.1
- + Better spatial composition showing the distinction between the front and back seats.
- + The capybara's paws are more anatomically accurate while still gripping the wheel.
- + Shows more of the taxi's mechanical details like the gear shift and ignition.
- − The capybara is missing the requested dark jacket.
- − The lighting on the human passenger feels slightly flat compared to the driver.
- − The taxi sign on top of the car is oddly placed directly over the window frame.
Verdict: GPT Image 1 offers superior cinematic lighting and better adherence to the clothing descriptions, creating a very convincing atmosphere. However, Recraft V4.1 provides a much better sense of the car's interior layout and maintains more realistic capybara anatomy, despite missing the jacket detail.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent legibility and clean graphic design
- + Perfect adherence to all text requirements including date and location
- + Cohesive vintage parchment texture and moody atmosphere
- − The 'gothic title' font is a standard serif rather than decorative gothic lettering
- − Composition is very safe and centered
Recraft V4.1
- + Stylistically superior gothic typography for the main title
- + High visual detail on the jack-o-lantern and clouds
- + More creative and dynamic illustrative style
- − The scroll text is tiny and difficult to read
- − The footer text is plain and lacks the 'vintage poster' integration of Model A
- − Border details like skulls were not explicitly requested
Verdict: GPT Image 1 (Model A) is the better functional invitation, providing clear, well-spaced text that follows the prompt's structural requirements perfectly. Recraft V4.1 (Model B) has much more impressive artistic flair and superior gothic typography, but fails to balance the layout for readability, making the smaller text elements an afterthought.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent 3D cartoon lighting and soft, clay-like textures.
- + Perfect adherence to text layout and placement.
- + Clean, harmonious color palette with soft shadows.
- − The sushi toppings look a bit like plastic or playdough rather than 'realistic PBR' fish.
Recraft V4.1
- + Higher realism in the sushi textures, particularly the fish grain.
- + Sophisticated diorama base with tiny legs and textured surface.
- + Crisp text and accurate flag placement.
- − The text is placed further apart and smaller than indicated in the prompt.
- − The wasabi has a slightly odd, flat shape compared to the rest of the 3D scene.
Verdict: Both models followed the prompt exceptionally well, but GPT Image 1 (Model A) better captured the specific '3D cartoon' aesthetic with its soft, rounded lighting and bold, centered typography. While Recraft V4.1 (Model B) offers more realistic textures on the fish and a more detailed base, GPT Image 1 feels more cohesive as a single illustrated piece.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'tumbling together' prompt with a high-energy composition.
- + Superior lighting effects with strong god rays and a warm, cohesive golden glow.
- + Very expressive and cute facial features that capture the 'joyful vibe'.
- − The kitten has an extra-long, somewhat distorted paw reaching toward the rabbit.
- − The fox's front paws are rendered with a slightly unnatural, dark texture.
Recraft V4.1
- + Incredible fur detail and texture across all four animals.
- + The composition feels more spacious and realistic in its arrangement of animals.
- + Excellent rendering of dew sparkles on the wildflowers in the foreground.
- − The golden retriever puppy looks significantly larger and more adult-proportioned than the other 'baby' animals.
- − The 'dark line' or anatomical glitch where the kitten's front leg meets its body is distracting.
Verdict: Both models followed the complex prompt well, but GPT Image 1 captures the 'tumbling' and 'joyful' aspect better with a more dynamic, front-facing composition. Recraft V4.1 has more sophisticated texture work and realistic flowers, but the golden retriever appears too large, making the scene feel less like a group of tiny babies.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Perfect text accuracy including the accent mark
- + Strong vector emblem style with clear, bold lines
- + Effective use of the 'Est. 1720' banner at the bottom
- − Failed the light background prompt requirement (used black background)
- − Texture is a bit grainy rather than 'subtle texture'
Recraft V4.1
- + Correctly followed the light background and cream color prompt
- + Elegant typography that fits a vintage restaurant theme
- + Accurately rendered the cloche, steam, and banner elements
- − The accent on 'CAFÈ' is floating too high and is slightly detached
- − The 'Est. 1720' text is slightly less crisp than the main brand name
Verdict: Recraft V4.1 is the winner because it adhered to all prompt instructions, specifically the light background and cream tones which GPT Image 1 ignored by using a black background. While GPT Image 1 had slightly bolder vector lines, Recraft V4.1 captures the 'vintage minimalist' aesthetic more effectively with superior color and layout choices.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering for character names and mission phases.
- + Clean icon illustrations for the Earth and Saturn V rocket.
- + Effective use of the requested color palette.
- − Poor layout logic where text labels do not align well with their corresponding icons.
- − Spelling error in 'EARLLUNAR'.
- − Missing step 4 (Lunar Orbit) in the visual sequence.
Recraft V4.1
- + Strict adherence to all 6 requested steps with numbered markers.
- + Highly professional and consistent graphic design aesthetic.
- + Accurate text rendering and layout spacing.
- − Step numbering error where '04' is repeated twice instead of '04' and '05'.
- − The descent and landing icons are virtually identical with minimal variation.
- − Very small, repetitive sub-labels under the main titles.
Verdict: Recraft V4.1 is the winner as it successfully follows the requested 6-step structure and delivers a cohesive, professionally designed infographic layout. While GPT Image 1 has slightly better character silhouettes and larger icons, its layout is chaotic and fails to present the information in a logical or chronological order.
Explore each model
Recraft's V4.1 standard tier text-to-image model — refines V4's photorealism with more natural lighting, softer gradients, and sharper illustration styles for everyday creative work