FP8 quantized variant of Black Forest Labs' FLUX.1 [schnell] model, offering ~2x faster inference with reduced precision while maintaining high-quality image generation in 4 steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell] FP8
#48 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell] FP8
0.0%
win rate
Ties
0.0%
Vidu Q2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent handling of glass transparency and realistic internal reflections
- + High-quality soft lighting that creates a bright, clean aesthetic
- + Accurate follow-through on the 'plant behind the cube' request with natural bokeh
- − The glass object is more of a tall rectangular prism than a strict cube
- − Includes an internal shelf not requested in the prompt
Vidu Q2
- + Perfect cubic geometry for the glass container
- + Highly realistic textures on the book cover and wooden table
- + Accurate sphere placement and shadow rendering
- − The plant is behind the cube but doesn't feel as integrated into the refraction of the glass as Model A
- − The sphere looks a bit like a matte plastic toy compared to the more artistic marble look in A
Verdict: Both models followed the prompt's spatial instructions perfectly. FLUX.1 [schnell] FP8 produced a more polished, ethereal image with superior lighting, though it deviated slightly from the 'cube' shape by making it a prism. Vidu Q2 captured the geometry and material textures of the book and table with higher tactile realism, making it the more technically accurate interpretation of the prompt's objects.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent handling of shallow depth of field and bokeh
- + Accurately depicts light rain and wet pavement reflections
- + High technical clarity with professional-looking lighting
- − Man looks more Caucasian/Latino than Japanese
- − The man is holding the handlebars rather than repairing the bike
- − Lacks the requested 'motion blur' from passing cars
Vidu Q2
- + Successfully captures the 'imperfect framing' and 'candid' feel requested
- + The man is clearly engaged in the act of repairing the chain
- + Excellent skin texture and realistic, aged hand details
- − The passing car lacks the requested motion blur
- − Anatomy issues where the man's arm appears disconnected from his body
- − Minor structural artifacts in the bike frame construction
Verdict: Vidu Q2 better followed the prompt's thematic requirements, such as the specific action of 'repairing' and the requested 'imperfect framing' for a candid feel. However, FLUX.1 [schnell] FP8 produced a more aesthetically pleasing and high-quality image, despite failing on the specific ethnicity and action requested.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Intense and highly detailed skin textures and lifelike eyes
- + Dramatic and evocative use of warm lighting and shadow
- + Strong focus on the 'close portrait' aspect of the prompt
- − The hair is messy rather than having clear braids as requested
- − Lower resolution appearance with more visible digital artifacts
Vidu Q2
- + Excellent adherence to the 'braided hair with beads' prompt
- + Very high clarity and crisp rendering of armor engravings and textures
- + Perfect balance of scars, dirt, and battle-worn effects on the armor and skin
- − The composition is a medium shot rather than a 'close portrait'
- − Backdrop bokeh sparks are a bit sparse
Verdict: While FLUX.1 [schnell] FP8 captures a more intense and dramatic close-up with impressive facial textures, Vidu Q2 follows the complex prompt requirements more holistically, particularly the hair details and armor engravings. Vidu Q2 provides a much cleaner, more professional look with superior clarity across the entire image.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent structure with clearly defined sections for appetizers, pizza, and mains.
- + Clean and professional minimalist layout that feels like a real functional menu.
- + Clean white background that adheres perfectly to the prompt.
Vidu Q2
- + Features more vibrant colorful food photography.
- + Uses contemporary graphic accents and colorful design elements.
- + Photos display higher dynamic range and more appetizing lighting.
- − The layout is cluttered and difficult to navigate compared to Model A.
- − Text contains significant spelling distortions even for primary headings.
- − Source images for the grid contain non-food elements and artifacts.
Verdict: FLUX.1 [schnell] FP8 is the clear winner as it provides a highly functional, legible, and professional menu design with a logical grid layout and clear section headers that match the prompt. While Vidu Q2 has more vibrant photography, the composition is disorganized and the text generation is significantly lower quality.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean layout for the main title
- + High resolution on the burger texture
- − Significant text errors including 'LIIMITED' and 'LIMIED TIME NEEY'
- − Incorrect pricing and starburst execution
- − The burger is not 'exploded' with components suspended separately as requested
Vidu Q2
- + Perfect adherence to the 'exploded' burger concept with suspended layers
- + Accurate and glowing fiery text rendering for both titles
- + Successful integration of all requested prompt elements including the starburst price
- − The currency symbol in the price is slightly malformed
- − The fire background is a bit chaotic and busy
Verdict: Vidu Q2 is the clear winner for its superior prompt adherence, successfully depicting an exploded burger with suspended layers and perfectly rendered glowing text. FLUX.1 [schnell] FP8 failed to explode the burger components and suffered from severe typographical errors and incorrect pricing information.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean layout with sharp wood framing
- + High legibility for specific prompt keywords like 'TODAY SPECIALS' and 'Brown butter'
- − Significant text repetition and layout errors
- − Poor price accuracy compared to prompt
- − Chalk texture looks like a digital marker rather than real chalk
Vidu Q2
- + Excellent realistic chalk texture with smudges and dust
- + Better adherence to the cursive title request
- + More natural and artistic spacing for a hand-drawn board
- − Several spelling errors in the menu items
- − The price for the final item is rendered as a symbol rather than a number
Verdict: While both models struggle with spelling the specific menu items correctly, Vidu Q2 is the clear winner for its superior visual quality and prompt adherence regarding 'chalk texture' and 'cursive' styling. FLUX.1 [schnell] FP8's output appears too clean and digital, with repetitive text lines that break the realism of a cafe menu.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Attempts the specific role-reversal requested in the prompt
- + Clean, cinematic lighting and high-resolution textures
- + Includes surreal elements like the mechanical pack on the horse
- − The anatomy of the two horses is confusing and slightly mangled
- − The astronaut is represented only by a mechanical suit/structure rather than a clear figure
Vidu Q2
- + High visual appeal with vibrant colors and cosmic details
- + Great composition with excellent rendering of the horse and astronaut
- + Very detailed and 'cinematic' as requested
- − Failed the negative constraint to have the horse on top of the astronaut
- − Followed the traditional 'astronaut riding a horse' trope despite instructions to the contrary
Verdict: FLUX.1 [schnell] FP8 was the only model to attempt the difficult 'horse on top' spatial instruction, creating a surreal and unique image, though the execution of the horses' anatomy is a bit chaotic. Vidu Q2 produced a much more beautiful and polished image, but it completely ignored the specific core instruction of the prompt, opting for a standard astronaut-on-horse depiction. FLUX.1 [schnell] FP8 is the winner for following the complex prompt logic.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent texture on the capybara's fur
- + Accurately places text on the driver's cap
- + Strong photographic lighting and bokeh
- − The woman is holding two phones simultaneously
- − Anatomically incorrect placement of paws on the steering wheel
- − The perspective makes the capybara look like it is in the same seat as the woman
Vidu Q2
- + Realistic interior layout with distinct front and back seating
- + Excellent adherence to the 'bored expression' prompt for the passenger
- + Superior paw placement on the steering wheel
- − The driver's cap is more of a generic officer cap than a taxi cap
- − Slightly less 'New York' feel in the background lights compared to Model A
Verdict: Vidu Q2 is the clear winner for its superior composition and logical consistency; it correctly differentiates the front and back seats and perfectly captures the 'bored' expression of the passenger. FLUX.1 [schnell] FP8 has higher surface-level detail but suffers from significant logical errors, such as the woman holding two phones and the capybara appearing to sit in the passenger seat rather than the driver's side.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent moody and cinematic lighting.
- + High resolution with clean digital graphics.
- + Better overall balance and composition for a dark gothic theme.
- − Significant spelling errors throughout all text blocks.
- − The parchment effect is less convincing and looks more like a modern border.
Vidu Q2
- + Convincing 'vintage parchment' texture and illustration style.
- + Includes the requested webs and thorns in the border design.
- + Artistic gothic font choice that fits the theme well.
- − Text contains several typos and character mashups.
- − Technical error in the date (30.70.2025).
- − The lighting is flat compared to the requested cinematic look.
Verdict: Both models struggled with the specific text requirements, but FLUX.1 [schnell] FP8 produced a much cleaner, more professional-looking image with effective cinematic lighting. While Vidu Q2 captured the aesthetic of 'parchment and thorns' more accurately, its low resolution and text garbling (including an impossible date) make it less usable.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent 3D miniature aesthetic with soft, refined textures
- + Accurate 45-degree isometric perspective and balanced composition
- + Clean lighting and professional rendering quality
- − Failed to render the word 'SUSHI'
- − The word 'JAPAN' is repeated incorrectly with artifacts
Vidu Q2
- + Perfect adherence to text requirements, including 'JAPAN' and 'SUSHI'
- + Includes a clear flag icon as requested
- + Vibrant colors and good material separation
- − Texture quality is slightly more plasticky and less refined than Model A
- − Perspective is slightly flatter than a true 45-degree isometric view
Verdict: Vidu Q2 is the winner because it successfully followed all text and icon instructions, whereas FLUX.1 [schnell] FP8 failed significantly on the typography. While FLUX.1 [schnell] FP8 produced a more aesthetically pleasing miniature texture, its inability to render the word 'SUSHI' and its mangling of 'JAPAN' make it less useful for the specific prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent soft lighting and bokeh effect
- + High-quality fur textures and expressive eyes
- − Failed to include a rabbit in the scene
- − Included extra indistinct felines instead of specified animals
- − Anatomical issues where animals appear merged together
Vidu Q2
- + Successfully included all four specified animals: puppy, kitten, bunny, and fox kit
- + Captures the action of chasing and tumbling better than the competitor
- + Vibrant colors with visible dew sparkles and god rays
- − Included a second puppy not requested in the prompt
- − Some distortion in the fox's legs and the butterfly wings
- − Slightly less 'photorealistic' fur texture compared to Model A
Verdict: While FLUX.1 [schnell] FP8 offers superior fur texture and lighting, it completely failed the prompt requirements by omitting the baby bunny and adding extra cats. Vidu Q2 followed the prompt much more accurately, successfully depicting the specific variety of animals requested along with a more dynamic composition of them playing in the field.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Successfully captured a vector emblem style.
- + Excellent use of the requested warm brown and cream color palette.
- + Includes the requested 'Est. 1720' text accurately.
- − Failed to render the restaurant name correctly, spelling it as 'AFe FLAMILAN'.
- − The central icon looks more like an architectural dome than a kitchen cloche dome.
Vidu Q2
- + Accurately depicts a retro cloche dome with visible steam.
- + High visual clarity and nice subtle texture on the background.
- − Serious failures in text rendering and spelling across all text elements.
- − Includes 'Caffee' and 'Farmiin' instead of the requested name.
- − The composition feels cluttered with too many lines of conflicting text.
Verdict: Both models failed to correctly spell the primary name 'Caffè Florian'. However, FLUX.1 [schnell] FP8 is the superior choice because it better understands the vintage vector emblem aesthetic and maintains a cleaner, more professional secondary text. Vidu Q2 followed the 'cloche dome' and 'steam' instructions more literally but suffered from severe typographic garbling and messy layout.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Successfully used the requested navy-heavy NASA color palette.
- + Excellent flat-vector design aesthetic with professional layout and clean lines.
- + Included the specific 'Apollo 11' text correctly.
- − Icons do not always match the specific text (e.g., 'Saturn Vicon' label has a plain white circle).
- − Text blocks contain significant gibberish underneath the main headers.
Vidu Q2
- + Icons are more detailed and visually descriptive of the mission components like the Lunar Module.
- + Good use of the steps structure mentioned in the prompt.
- − Failed the main heading text ('ALFONCH' instead of 'Apollo 11').
- − The light gray background deviates from the requested 'navy' dominant palette.
- − Contains generic 'clipart' style elements that feel less modern than the requested vector infographic style.
Verdict: FLUX.1 [schnell] FP8 is the clear winner for its superior professional infographic aesthetic and adherence to the navy-based color palette, despite some gibberish text. Vidu Q2 produces more literal icons but fails significantly on the primary typography and the requested 'modern' vector design language.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing