spicy-mayo.com

In testing on Arena · Not officially confirmed Last updated:

Nano Banana 2.5 prompts

You cannot summon spicy-mayo on demand, so the useful thing is to arrive with prompts that reveal capability fast when it does pair. The set below probes the six areas where image models differ most: text rendering, instruction following, photoreal detail, multi-subject composition, style control, and structured layout.

How to use this set

Nothing you type changes which models Arena pairs you with. So the goal is not to summon spicy-mayo — it is to make each battle worth something when it happens.

Each prompt below isolates one capability axis and comes with a "what a good result looks like" note, so you can judge an output in the few seconds you have at the reveal panel rather than forming a vague impression. Run one prompt per battle and keep the wording identical between runs, otherwise you are comparing prompts instead of models.

Record what you see. Screenshot the output together with the reveal panel, and save the exact prompt text. During a blind-battle phase, written-down observations are the only evidence that exists.

The prompt set

Six axes, ordered roughly by how quickly they separate one model generation from another.

1Text rendering

The fastest discriminator between image model generations. Legible, correctly spelled text in-scene is still where most models fall apart.

A hand-lettered chalkboard outside a cafe that reads "TODAY: BURNT HONEY LATTE — 4.50", shot at a slight angle in overcast daylight, shallow depth of field.

What a good result looks like: Every character correct, kerning even, and the angle of the text following the board rather than floating flat on top of it.

2Instruction following

Counts, spatial relationships and negations are where models quietly substitute what is easy for what was asked.

Exactly three green ceramic mugs on a wooden table, arranged in a row from largest to smallest, with no other objects in frame and no text anywhere in the image.

What a good result looks like: Three mugs — not two, not four — correctly ordered by size, and nothing smuggled into the background.

3Photoreal detail

Skin, hair and fabric under directional light expose the difference between a sharp render and a plausible photograph.

Close-up portrait of an elderly fisherman in a knitted wool sweater, hard side light from a window on the left, visible skin texture and individual grey stubble, 85mm lens, no retouching.

What a good result looks like: Pores and stray hairs rather than airbrushed skin, fabric weave that stays consistent across folds, and light falloff that matches a single source.

4Multi-subject composition

Several interacting subjects is where anatomy, occlusion and scale go wrong at once.

Four cooks in a busy restaurant kitchen passing a single large copper pot between them, mid-motion, steam in the air, overhead fluorescent light, documentary photography style.

What a good result looks like: Correct hand and arm count, believable occlusion where bodies overlap, and consistent scale between the people and the pot.

5Style control

Tests whether a named style is actually applied or approximated with generic filter-like treatment.

A city tram stop at dusk in the style of a 1970s Japanese woodblock print: flat colour planes, visible registration offset, limited palette of indigo, ochre and off-white, no gradients.

What a good result looks like: Genuinely flat planes and a restricted palette. A soft digital painting with a colour cast means the style was approximated, not applied.

6Diagram and layout

Structured graphics need text, alignment and semantics at the same time, which is the hardest combination.

A clean infographic comparing three coffee brewing methods in three labelled columns, each with a simple line icon, a one-word title and two short bullet points, flat vector style on a cream background.

What a good result looks like: Columns actually aligned, labels matching their icons, and text that reads as words rather than as decorative glyphs.

Send us what you find

If one of these pairs you with spicy-mayo, send the prompt, the output and the reveal screenshot to hello@spicy-mayo.com. Published observations get an evidence label and credit.

We do not publish outputs that cannot be tied to a revealed battle. An unattributed image claimed to be from an anonymous model is worth nothing, and putting one up would undermine every other page here.

Prompts FAQ

Will these prompts make spicy-mayo appear?

No prompt can. Arena pairs models at random and no wording changes the odds. What these prompts do is make the most of the battle when it does pair, by testing the capabilities that actually separate image models.

Why does text rendering matter so much?

Legible text inside an image is the single hardest thing for image models and the fastest way to tell generations apart. A model that spells a five-word sign correctly is meaningfully ahead of one that produces glyph soup.

Should I record what I see?

Yes, and this is the whole point. Screenshot the reveal panel together with the output and note the exact prompt. Unreproducible observations are the only evidence anyone has during a blind-battle phase, so writing them down is what makes them worth anything.

Can I submit what I find?

Please do. Email the prompt, the output and the reveal screenshot to hello@spicy-mayo.com. Anything published gets an evidence label and credit, and we do not publish outputs we cannot attribute to a revealed battle.

Get notified at launch

Leave your email and we will send one message when spicy-mayo (or whatever it ends up being called) becomes publicly available. We do not send anything else and we do not share the list.

By submitting, you agree to our privacy policy.