Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
OmniGen v2
#56 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
OmniGen v2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealism and texture on the wooden table and plant.
- + High-quality glass reflections and refractions accurately showing the plant behind.
- − Failed the spatial prompt by adding an extra blue sphere on top of the book.
OmniGen v2
- + Perfect adherence to all spatial instructions in the prompt.
- + Clean, realistic composition with accurate lighting from the left.
- − The glass cube appears slightly open or like a container rather than a solid object.
- − Mild aliasing/softness on the background plant compared to the foreground.
Verdict: While FLUX.1 [schnell] produced a more visually stunning and detailed image with complex refractions, it failed the specific logic of the prompt by adding an extra sphere. OmniGen v2 followed all instructions perfectly, including the exact placement of objects, making it the more successful image for this challenge.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent realism in skin textures and lighting
- + Realistic wet pavement reflections and depth of field
- + Stronger visual storytelling through the urban Japanese backdrop
- − Anatomical issues with the man's hands grasping the handlebars
- − Missing the requested motion blur on the passing cars
OmniGen v2
- + Successfully captures visible rain droplets and distinct reflections
- + Clean composition with a clear focus on the subject
- + Better hand/handlebar interaction
- − The man's feet appear to be floating or misaligned with the ground
- − The 'rain' effect looks like a simple overlay filter rather than integrated environmental rain
- − Background cars lack the requested motion blur and look static
Verdict: FLUX.1 [schnell] provides a much more convincing cinematic and realistic image with superior skin textures and environmental lighting that feels like a genuine street photo. While OmniGen v2 captures the rain more explicitly, the integration of the man into the scene is flawed, particularly where his feet meet the pavement, making FLUX.1 [schnell] the preferred choice for adherence to the 'realistic' and 'no stylization' requirements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Exceptional level of skin texture and pore detail
- + Intense and convincing expression
- + Stronger adherence to the 'battle-worn' description
- − The braids and beads are a bit messy and indistinct
- − Lighting feels a bit flat compared to the dramatic prompt
OmniGen v2
- + Excellent depiction of the hair braids with visible beads
- + High-quality engraving detail on the plate armor
- + Atmospheric lighting and bokeh sparks in the background
- − The skin looks too clean and smooth for a battle-worn character
- − The 'dirt' on the face looks more like stylized freckles or paint than grime
Verdict: While FLUX.1 [schnell] captures the grit and texture of a battle-worn warrior significantly better, OmniGen v2 displays a superior interpretation of the 'paladin' aesthetic with more ornate armor and clearly defined hair braids. OmniGen v2 is slightly preferred for a fantasy portrait as it balances the specific decorative requirements of the prompt with a cleaner, more appealing composition.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Stronger adherence to the minimalist design aesthetic requested
- + Clearer and more readable bold sans-serif header fonts
- + Better alignment and professional white space management
- − Nonsense word 'ORFEFUS' used instead of 'MAINS'
- − The text in the body of the menu is very blurry and illegible
OmniGen v2
- + Better use of vibrant color accents as requested in the prompt
- + High-quality, distinct food photography with better lighting
- + Covers a wider range of the requested sections (Appetizers, Mains/Aliinars)
- − Significant spelling errors in large titles ('RESTAURATED MENTS', 'PIZZZZAN')
- − Composition feels cluttered compared to a modern minimalist style
Verdict: FLUX.1 [schnell] captures the 'modern minimalist' aesthetic much more effectively with its clean layout and professional use of white space, although it fails to correctly spell all section headers. OmniGen v2 provides more vibrant colors and better individual food images, but the layout is busier and the spelling in the primary headings is noticeably poor. FLUX.1 [schnell] is the preferred choice for a design-focused task where cleanliness is key.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic texture on the meat, cheese, and buns.
- + Dynamic composition with effective use of shallow depth of field for the background.
- + Strong adherence to the fiery background and glowing embers requirement.
- − Serious typographical errors including 'AGIC BURGER' and a redundant '€699' in the starburst.
- − The burger is largely assembled rather than being 'exploded' as requested in the prompt.
OmniGen v2
- + Accurate text rendering for the main title and price starburst.
- + Bright, punchy graphic design suitable for a commercial advertisement.
- + Clean layout with the text integrated into the fiery theme.
- − Failed the 'exploded' and 'suspended in mid-air' components requirement entirely.
- − Lacks the photorealistic detail requested, appearing more like a 3D digital illustration.
- − The text 'LIMITED TIME ONLY' is poorly placed and partially cut off.
Verdict: FLUX.1 [schnell] captures the requested cinematic mood and photorealistic textures much better than the competition, although it fails significantly on the text rendering and price accuracy. OmniGen v2 succeeds in spelling the brand name correctly and using a starburst, but it misses the core mechanical prompt requirements of an exploded, suspended burger, resulting in a static image. FLUX.1 [schnell] is preferred for its superior visual quality and atmosphere despite the typos.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent chalk texture and realistic handwriting variation
- + Clean composition that resembles a real cafe setting
- + Legible handwriting that captures the 'marker/chalk' aesthetic well
- − Significant spelling errors throughout the menu items
- − Failed to include the 'elegant cursive' style for the header
- − Inaccurate date formatting ('Pril' instead of 'April')
OmniGen v2
- + More convincing chalk stroke texture with varying thickness
- + Correctly included the full date 'April 30, 2026'
- + Better attempt at rendering the specific menu items requested
- − Text is cluttered and overlaps in several places
- − Significant misspelling in the title ('SPECALS')
- − Failed to use elegant cursive for the header
Verdict: Both models struggled with the specific 'elegant cursive' requirement and perfect spelling, but FLUX.1 [schnell] produced a much more professional and realistic composition. While OmniGen v2 had better chalk-like textures, the layout was messy and the title was misspelled, whereas FLUX.1 [schnell] maintained better overall visual coherence despite its own spelling struggles.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully followed the difficult prompt instruction of placing the horse on top of the astronaut.
- + High cinematic quality with professional lighting and depth.
- + Creative Interpretation of the surreal elements.
- − The horse has two heads/torsos merged together.
- − The astronaut's anatomy and gear are somewhat messy and incoherent.
OmniGen v2
- + Clean, clear resolution with vibrant colors.
- + Realistic horse and astronaut textures.
- − Completely failed the primary prompt instruction to place the horse on top.
- − Composition is very generic and lacks the requested cinematic feel.
- − Multiple moons in the background look like copy-pasted stickers.
Verdict: FLUX.1 [schnell] is the clear winner as it successfully interpreted the challenging spatial instruction of having the horse riding the astronaut, despite some anatomical merging issues. OmniGen v2 defaulted to a standard trope and completely ignored the 'horse on top' requirement, resulting in a much less creative output.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and realistic lighting
- + High adherence to the requested camera perspective and composition
- + Accurate human fingers and phone interaction
- − The capybara's eyes look a bit cartoonishly round compared to a real capybara
OmniGen v2
- + Successfully rendered the capybara in a formal dark jacket
- + Good bokeh effect on the background city lights
- − Major logic error with human hands growing out of the steering wheel and jacket
- − Human passenger is sitting in the passenger seat instead of the back seat
- − The capybara's face is unnaturally blended with the hat and seems distorted
Verdict: FLUX.1 [schnell] followed the prompt instructions perfectly, placing the human in the back seat and creating a coherent, photorealistic scene. OmniGen v2 suffered from several anatomical errors, including placing the human in the front seat and merging the human's hands with the capybara's steering wheel.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent gothic atmosphere with cinematic lighting and a dark parchment texture
- + Strong visual composition with spooky silhouettes of bats and twisted trees
- + Clearly legible decorative fonts that match the vintage theme
- − Significant text errors including redundant date/time lines and garbled words like 'firiichts' and 'butistigtion'
- − Missing the specific prompt instruction to place the title on the parchment (it overlaps a banner instead)
OmniGen v2
- + Good use of the parchment border effect requested in the prompt
- + Bold and readable main title text with high contrast
- + Vibrant central jack-o-lantern that glows effectively
- − Major spelling errors in the event details and invitation banner
- − Cluttered layout with overlapping text elements at the bottom
- − The border design looks a bit more like generic clip-art than a polished gothic poster
Verdict: Both models struggled with the complex text requirements, providing several spelling and formatting errors. FLUX.1 [schnell] produced a far superior aesthetic with its moody lighting and cohesive gothic art style, whereas OmniGen v2 felt more like a basic digital collage. FLUX.1 [schnell] is the winner for its high visual quality and atmospheric consistency, despite the text hallucinations.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean aesthetic with a very soft, high-quality light blue background.
- + Precise rendering of the Japanese flag icon.
- + Higher fidelity and more realistic textures on the sushi fish and rice.
- − Failed to include the word 'SUSHI' below the main text.
- − The sushi roll/nigiri hybrid looks a bit awkward with the cucumber placement.
OmniGen v2
- + Successfully included all requested text ('JAPAN' and 'SUSHI') in a bold style.
- + Stronger 3D cartoon/isometric diorama feel that matches the requested style well.
- + Vibrant colors that pop against the background.
- − The flag icon is incorrect, featuring red, yellow, and blue instead of the Japanese flag.
- − Overall resolution and texture quality look slightly more 'plastic' compared to Model A.
Verdict: FLUX.1 [schnell] produced a cleaner, more realistic image with better textures and a correct flag, but it missed a key text element from the prompt. OmniGen v2 adhered better to the text requirements and the cartoon aesthetic, but failed on the specific flag design and had lower texture fidelity. FLUX.1 [schnell] is the likely winner due to its superior visual polish and icon accuracy.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture realism and lighting integration
- + High level of detail in the wildflower meadow and bokeh depth
- + Includes all four requested animals with distinct features
- − The 'bunny' looks like a hybrid cat/rabbit creature
- − The scene is more static than the 'playfully chasing' prompt implies
OmniGen v2
- + Strong 'god rays' effects that match the prompt well
- + Vibrant colors and high contrast
- − Failed to include four distinct animals, missing the fox or bunny depending on interpretation
- − Art style is 3D-cartoonish rather than 'hyper-photorealistic'
- − Anatomical issues with the animals' paws and joints
Verdict: FLUX.1 [schnell] significantly outperformed OmniGen v2 by adhering closer to the 'hyper-photorealistic' instruction and including the correct number of animals. While OmniGen v2 captured the 'god rays' more literally, the resulting image looks like a stylized digital illustration rather than a masterpiece photo, and it failed to include all the requested animal types.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector aesthetic
- + Good balanced composition with the circular border
- − Severely misspelled the restaurant name as 'CAFEÉ FRAMILAN'
- − Incorrect date '7720' instead of '1720'
- − Missing the steam element
OmniGen v2
- + Includes the requested steam element
- + Correct establishment date (1720)
- + Closer text rendering to the target name
- − Small spelling error in the name ('CAFFFLORIN' instead of 'Caffè Florian')
- − Slightly bulky line weights for a minimalist logo
Verdict: OmniGen v2 is the winner because it successfully included all prompt elements, including the steam and the correct year, whereas FLUX.1 [schnell] hallucinated an entirely different name and a futuristic date. Although OmniGen v2 has a minor spelling typo, it is far more usable and adherent to the specific request than FLUX.1 [schnell].
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong composition that visually represents a celestial path
- + Higher level of detail in the rocket and moon icons
- + Follows the NASA color palette naturally across the design
- − Text is largely nonsensical gibberish
- − The 'Saturn V' icon looks more like a shuttle or generic rocket than the specified vehicle
- − The layout of the steps is confusing and difficult to follow chronologically
OmniGen v2
- + Layout clearly presents 6 distinct infographic sections
- + Text is much more legible and attempts to follow some mission terminology
- + Clean, balanced grid composition with a modern feel
- − Confuses the mission name as 'Apolo 17' and NASA as 'NSA'
- − Iconography is very generic and lacks the specific Saturn V and Lunar Module details requested
- − The red is slightly too bright compared to the requested 'muted red'
Verdict: OmniGen v2 is the winner despite small typos, as it successfully organized the infographic into the 6 requested steps with legible headings, whereas FLUX.1 [schnell] produced a beautiful but confusing diagram with unreadable text. OmniGen v2 also adhered better to the 'clean, modern grid' aesthetic typical of infographics.
Explore each model
Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on