Black Forest Labs' ultra-high resolution image generation model, an enhanced version of FLUX1.1 [pro] optimized for premium quality output
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX1.1 [pro] Ultra
#51 of 62 in Text-to-Image
LongCat-Image
#61 of 62 in Text-to-Image
Where the votes landed
FLUX1.1 [pro] Ultra
0.0%
win rate
Ties
0.0%
LongCat-Image
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent photorealism and texture on the red book
- + Highly accurate refraction and transparency in the glass
- + Perfect adherence to spatial instructions including lighting from the left
- − The blue sphere is levitating rather than sitting on the base
- − Slight dust artifacts on the glass might look like noise to some users
LongCat-Image
- + Natural placement of the blue sphere on the bottom of the cube
- + Strong composition with a clear window light source
- + Solid rendering of the glass thickness
- − The glass refractive logic is slightly inconsistent on the edges
- − The book texture is less detailed than image A
Verdict: FLUX1.1 [pro] Ultra produces a more photorealistic image with superior texture work and light handling, though the sphere is inexplicably floating inside the cube. LongCat-Image provides a more grounded physical arrangement but lacks the fine detail and professional polish seen in the lighting and materials of FLUX1.1 [pro] Ultra.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent execution of long-exposure motion blur on passing vehicles
- + Superior realism in clothing and skin textures
- + High technical clarity and sharp focus on the subject
- − The red bicycle is missing a seat
- − The man's posture is more 'holding' than 'repairing'
LongCat-Image
- + Natural, candid 'street photography' composition with 'imperfect' framing
- + The man is actively engaged in a repair task
- + Atmospheric lighting on wet pavement
- − Significant anatomical errors with the bicycle including three wheels and fused frames
- − Failed to include the requested motion blur on cars
- − Rain effects look like static white lines rather than integrated moisture
Verdict: FLUX1.1 [pro] Ultra produced a much higher quality image with impressive technical execution of motion blur and realism, though it missed the bicycle seat. LongCat-Image captured a more emotive and 'candid' scene that aligned better with the storytelling aspect of the prompt, but it suffered from severe structural failures in the bicycle geometry and failed to generate the requested motion blur.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- − The image failed to load or is a solid black square.
- − Complete failure to provide any visual content.
LongCat-Image
- + Excellent adherence to all prompt details including braided hair with beads, scars, and ornate armor.
- + Impressive material rendering on the metal engravings and leather straps.
- + Strong atmosphere with warm torchlight reflections and dynamic bokeh sparks.
- − The scarring on the face looks slightly artificial/painted on in one spot.
- − The 'close portrait' instruction could have been interpreted as a tighter framing on the face.
Verdict: FLUX1.1 [pro] Ultra failed to generate a visible image, resulting in a black frame. LongCat-Image successfully captured the narrative of a battle-worn paladin with high technical fidelity, particularly in the detailed textures of the armor and the atmospheric lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent typography with legible headers and clear hierarchy.
- + Professional layout using a realistic 3D mockup perspective.
- + High-quality, appetizing food photography that looks consistent.
- − Text includes minor spelling errors like 'Pizzzans'.
- − The pizza images are somewhat repetitive in style.
- − Perspective view makes the bottom section slightly harder to read than a flat view.
LongCat-Image
- + Strong use of vibrant color accents as requested.
- + Organized grid layout for food photos.
- + Front-facing, flat orientation is clear for reading composition.
- − Text consists of illegible gibberish characters.
- − Graphic design feels cluttered and less professional.
- − Fonts are erratic and do not follow the 'bold sans-serif' requirement consistently.
Verdict: FLUX1.1 [pro] Ultra is far superior in terms of professionalism and font rendering, producing a menu that looks like a real design asset despite a couple of minor spelling mistakes. LongCat-Image followed the prompt regarding vibrant accents but failed significantly on text quality, resulting in an unreadable and amateur-looking layout.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent photorealistic texture on the meat and buns
- + Perfectly rendered and legible text for all three requested elements
- + Strong sense of vertical explosion and motion
- − Missed the starburst element for the price
- − Bottom of the image features unnecessary AI-generated footer gibberish
LongCat-Image
- + Includes the starburst graphic as requested in the prompt
- + Dynamic lighting on the main text with a realistic fiery glow
- + Clean composition without footer artifacts
- − Failed the 'exploded burger' requirement, as components are mostly touching
- − Price text and secondary message are combined into the starburst rather than separate as implied
Verdict: FLUX1.1 [pro] Ultra created a much more impressive visual representation of an 'exploded' burger with incredible detail, though it missed the starburst requirement. LongCat-Image followed the starburst instruction but failed the primary conceptual requirement of having all burger components suspended and separated in mid-air.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent chalk texture and realistic handwriting style.
- + High level of readability for most requested text.
- + Authentic blackboard smudge and wood grain details.
- − Repeated words and phrases creating logical errors (e.g., 'fresh fresh daily').
- − Spelling errors in menu items like 'Truffle mushroomm ris'.
LongCat-Image
- + Strong aesthetic appeal with a warm café background.
- + Good chalk smudging effects on the board edges.
- − Severe spelling hallucinations and garbled text (e.g., 'ToayS STaYS').
- − Incorrect formatting of prices and menu item layout.
Verdict: FLUX1.1 [pro] Ultra is far superior in this comparison due to its ability to render legible, mostly accurate text that adheres to the prompt's specific menu requirements. While it suffers from some word repetition, LongCat-Image fails to produce coherent English words for most of the board, resulting in illegible gibberish.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent high-definition details of the horse and space suit
- + Dynamic cinematic composition with a clear view of Earth's atmosphere
- + Smooth lighting and high-quality textures
- − Failed the negative constraint; the astronaut is on top of the horse instead of vice versa
LongCat-Image
- + Vibrant colors and creative space elements like satellites and moon bases
- + Clear foreground and background separation
- − Failed the negative constraint; an astronaut is riding a horse
- − The anatomy of the horse's legs is awkward and distorted
- − Visual artifacts present around the spaceships and planet edges
Verdict: Both FLUX1.1 [pro] Ultra and LongCat-Image failed to adhere to the specific spatial constraint of placing the 'horse on top' of the astronaut. However, FLUX1.1 [pro] Ultra is the superior image due to its exceptional cinematic quality, realistic lighting, and cleaner rendering compared to the distorted anatomy and artifacts found in the LongCat-Image output.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent photographic lighting and depth of field
- + High level of texture detail in the capybara's fur and the jacket
- + Cinematic composition with a realistic night atmosphere
- − The woman is placed in the front passenger seat instead of the back seat
- − The capybara's paws are not clearly on the steering wheel as requested
- − The woman is holding the phone over the steering wheel area
LongCat-Image
- + Follows the spatial instructions perfectly with the woman in the back seat
- + Shows both front paws on the steering wheel as requested
- + Accurate depiction of a New York taxi exterior and rooftop sign
- − The woman's hands and face have some slight 'AI melting' artifacts
- − The composition feels a bit cramped compared to the other model
- − The capybara's fur texture is slightly less defined
Verdict: While FLUX1.1 [pro] Ultra produced a much more visually stunning and realistic photograph, it failed significantly on the spatial prompt instructions by putting the passenger in the front seat. LongCat-Image adhered much better to the specific layout of the prompt, correctly placing the woman in the back and the capybara's paws on the wheel, despite having slightly lower overall image quality.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent typography style that matches the gothic aesthetic
- + Clean and professional layout within a scroll design
- + Mostly accurate text rendering for the event details
- − Included a duplicate, misspelled line of text 'noickt or a night of frights'
- − Omitted 'Party' from the main title 'Halloween Party Invitation'
LongCat-Image
- + Successfully included the full title 'Halloween Party Invitation'
- + Creative use of negative space with the parchment cutout for the background
- + Stronger thorn and web border details
- − Text at the bottom is cluttered and includes hallucinations like '7um' and 'The Armiees'
- − The 'You are invited...' banner is quite small and lacks the scroll's elegance
Verdict: FLUX1.1 [pro] Ultra produces a much more polished and professional design that feels like a real invitation, despite a minor text repetition error and the missing word 'Party'. LongCat-Image follows the title prompt better but fails on the legibility and accuracy of the event details at the bottom, making it less functional as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent adherence to the 'miniature 3D cartoon scene' aesthetic with highly sophisticated lighting and textures.
- + Accurately places all requested text and icons clearly with a clean design.
- + Superior composition that feels like a complete diorama with varied elements.
- − Includes some odd artifacts like a small red bug on the wooden board.
- − The sushi rice texture is slightly less defined compared to Model B.
LongCat-Image
- + Great material representation, especially the texture of the salmon and tuna.
- + Followed the minimal garnish and 'sushi' focus very closely.
- + Clean text rendering and clear flag icon.
- − The diorama base is very simple and lacks the 'complete scene' feel of Model A.
- − Lighting is a bit flat and less cinematic than the competitor.
Verdict: FLUX1.1 [pro] Ultra produced a much more professional-looking 3D miniature scene with beautiful lighting and depth, whereas LongCat-Image felt a bit more like a basic 3D asset render. While Model A included a strange insect artifact, its overall interpretation of a 'cartoon scene' and 'diorama' was more creative and visually appealing.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Excellent soft lighting with convincing god rays and bokeh.
- + Consistent cohesive style across all animals.
- + Good depth of field that emphasizes the '8K masterpiece' aesthetic.
- − The animals look very stylized and 'Disney-like' rather than hyper-photorealistic.
- − The fur texture is a bit too smooth and painterly.
- − The fox and kitten appear almost identical in facial structure.
LongCat-Image
- + Displays more realistic fur textures and distinct features for the different species.
- + Captures the 'playfully chasing' and 'tumbling' aspect of the prompt more dynamically.
- + Stronger adherence to individual animal types, including the bushy fox tail.
- − Anatomical failure on the kitten which has large rabbit ears.
- − The rabbit mentioned in the prompt is missing as a separate entity.
- − The god rays look a bit artificial and over-sharpened compared to the background.
Verdict: FLUX1.1 [pro] Ultra produces a much more beautiful and artistic image with superior lighting, but the animals look like CGI characters rather than real babies. LongCat-Image attempts a more realistic texture and sense of motion, but fails significantly on anatomy by merging the kitten and rabbit into a single creature and missing one of the requested animals. FLUX1.1 [pro] Ultra is the winner for its professional composition and lack of glaring anatomical glitches.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + High visual quality with clean vector-style lines
- + Professional composition with an excellent emblem layout
- + Follows the warm brown and cream color palette perfectly
- − Misspelled the primary text as 'Caffè Framilian'
- − Texture on background is very subtle, almost missing
LongCat-Image
- + Correctly spelled the name 'Caffè Florian'
- + Stronger vintage texture on the background as requested
- + Excellent use of a banner for 'Est. 1720'
- − Redundant text rendering with 'Caffé' appearing three times
- − Cluttered composition where text overlaps elements awkwardly
- − The steam effect is a bit heavy and less 'minimalist'
Verdict: While FLUX1.1 [pro] Ultra has far superior technical execution and look-and-feel for a logo, it failed the simple spelling task by adding extra letters. LongCat-Image followed the text instructions more accurately and provided the requested background texture, but the overall design is cluttered and repetitive. FLUX1.1 is preferred if the text can be edited, but LongCat-Image is the literal winner for prompt adherence regarding the brand name.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX1.1 [pro] Ultra
- + Successfully incorporated all six requested steps into the design
- + Matches the requested navy, white, and muted red NASA palette perfectly
- + Features a clean, professional vector aesthetic that resembles a real infographic
- − Text contains several AI hallucinations and garbled words
- − Numerical ordering of steps is inconsistent and non-linear
LongCat-Image
- + Follows the requested color palette well
- + Clearer iconography for the lunar module and the moon
- − Failed to include the requested six-step sequence
- − Layout is more like a simple poster than a complex infographic
- − Text rendering is poor and significantly misspelled
Verdict: FLUX1.1 [pro] Ultra is the clear winner as it attempted and largely succeeded in including all six specific steps of the mission requested in the prompt. While its text and numbering are messy, LongCat-Image failed to provide the sequential infographic structure and missed multiple steps entirely.
Explore each model
6B parameter image generation model excelling at rendering multilingual text directly in generated images