In testing on Arena · Not officially confirmed Last updated:
Nano Banana 2.5 vs Nano Banana 2
No honest quality comparison is possible yet. Nano Banana 2.5 has no released model to test, and the spicy-mayo model believed to be it can only be reached at random inside Arena battles. What can be compared is status, access path, and what each name actually refers to.
What can be compared
| Nano Banana 2.5 | Nano Banana 2 | |
|---|---|---|
| What the name refers to | An unreleased model, widely assumed to be the spicy-mayo codename on Arena | The current shipped generation on this line |
| Official name | None. Google has announced nothing. | Community shorthand; the shipped products carry Gemini names |
| How you reach it | Arena blind battles only, paired at random | Gemini API, Google AI Studio, Gemini app |
| Model ID | Not disclosed | Published and stable |
| Pricing | Not disclosed | Google's published Gemini pricing |
| Reproducible testing | Not possible | Yes — fixed prompts, fixed model ID |
Why the quality comparison does not exist
A comparison worth reading needs four things, and the 2.5 side currently has none of them.
On-demand access
- Arena pairs models at random. You cannot request spicy-mayo, so you cannot run a prompt set against it.
Reproducibility
- No control over sampling settings or seeds, and no way to re-run the same generation twice.
Version stability
- A model in testing can be swapped or tuned mid-flight without notice. Two battles a week apart may not be the same weights.
Any published head-to-head score for this model was produced without those conditions. That does not make the observations worthless — it makes them anecdotes, and they should be labelled as such.
What we will publish, and when
On the day there is a released model with a stable ID, the same prompt set runs against both generations and the results go up with every prompt attached, so anyone can re-run them and disagree. Until then this page stays a status comparison rather than a quality one.
In the meantime, single-battle observations are still worth recording — see the prompt set for the capabilities that separate image models fastest, and the timeline for what has been logged so far.
2.5 vs 2 FAQ
Which is better, Nano Banana 2.5 or Nano Banana 2?
Nobody can answer that honestly right now. Benchmarking needs repeatable access to both models with the same prompts, and the 2.5 side does not have that. Any head-to-head score you find published today was not produced under controlled conditions.
Why can't you benchmark spicy-mayo?
Arena's blind battles pair models at random and do not let you request one. You cannot run a fixed prompt set against it, cannot control the sampling settings, and cannot reproduce a run. That rules out any benchmark worth publishing.
Is Nano Banana 2 itself an official name?
It is community shorthand too. The shipped Google product names on this line have been Gemini 2.5 Flash Image and its successors; the banana names come from Arena codenames and stuck because they are easier to say.
What will a fair comparison need?
A public release with a stable model ID, a documented API, and pricing. At that point the same prompt set can be run against both models and the results published with the prompts attached. That is what we intend to do here.
Get notified at launch
Leave your email and we will send one message when spicy-mayo (or whatever it ends up being called) becomes publicly available. We do not send anything else and we do not share the list.