Rank #1
gpt-image-2
BT
+80
like.photo Arena · 20260531-batch50 · Latest public view
like.photo Arena runs head-to-head battles* between image generators on real products. A panel of vision AI* judges picks winners blind, in both A/B orders. Every number is published with its uncertainty; this is diagnostic evidence, not yet a human-preference study.
Current snapshot: gpt-image-2 wins 27 of 36 decisive comparisons against nano-banana-2.
Snapshot private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50·Cite or share
How a battle works
Try it in the explorer

2/2 judgesEach vision AI judge receives exactly four inputs: the reference photo, candidate A, candidate B, and a text brief naming the product, its critical attributes, and the infographic task prompt. Model names are never included. The judge must answer with a single letter, A or B; the same pair is judged again with the order swapped, and only verdicts that survive the swap count. Here all 2 judges picked candidate B. Public demo battle; the scored battles use the same format.
Generators ranked by strict eligible-judge majority over swap-stable private replay pairs. Individual product examples are withheld.
Figure 1
bar = winrate · whisker = 95% Wilson CI · dashed = even split
Rank #1
gpt-image-2
95% CI 59% to 86%
Rank #2
nano-banana-2
95% CI 14% to 41%
Matrix
row vs column · strict-majority swap-stable pairs
| vs | gpt-image-2 | nano-banana-2 |
|---|---|---|
| gpt-image-2 | — | 27-9 (75%) |
| nano-banana-2 | 9-27 (25%) | — |
Table 1: Consensus ranking over 36 strict-majority swap-stable pairs, from 700 raw votes by 6 eligible judges.
Rank #1
gpt-image-2
BT
+80
Rank #2
nano-banana-2
BT
-80
* BT Latent 'strength' per generator from pairwise wins; higher = preferred. Reported as 100 times the natural log of the strength, averaged over eligible judges. more·* Winrate Fraction of a generator's battles that it won. A simple control metric. more
Uncertainty gpt-image-2 wins 75% of the 36 agreed pairs, 95% Wilson interval 59% to 86%. Exact two-sided binomial test against an even split: p = 0.004.
Aggregated over consensus-eligible judges using swap-stable pair wins. Product identity tracks the same physical product, while task fit and commercial preference can reward a different shot or scene.
| Axis | Leader | gpt-image-2 | nano-banana-2 |
|---|---|---|---|
| Product identity | gpt-image-2 | 67% · 22-11 | 33% · 11-22 |
| Task fit | gpt-image-2 | 80% · 28-7 | 20% · 7-28 |
| Commercial preference | gpt-image-2 | 73% · 27-10 | 27% · 10-27 |
| Overall | gpt-image-2 | 75% · 27-9 | 25% · 9-27 |
Aggregated by eligible-judge majority over swap-stable private replay pairs. Tied or non-majority quorum pairs are shown as withheld and do not assign a winner.
| Task type | Tasks | Majority | Withheld | gpt-image-2 | nano-banana-2 |
|---|---|---|---|---|---|
| Infographic | 27 | 20/21 | 1 | 85% · 17-3 | 15% · 3-17 |
| Lifestyle | 23 | 16/18 | 2 | 63% · 10-6 | 38% · 6-10 |
The same battle list is scored by every judge. Swap consistency and position bias decide whether a judge feeds the consensus.
Matrix
generator cells show winrate · wins-losses
| Judge | Eligible | Swap | B rate | gpt-image-2 | nano-banana-2 |
|---|---|---|---|---|---|
anthropic/claude-fable-5 pairwise-product-v2-multiaxis | yes | 82% | 45% | 61% · 25-16 | 39% · 16-25 |
google/gemini-3.5-flash pairwise-product-v2-multiaxis | yes | 70% | 41% | 71% · 25-10 | 29% · 10-25 |
minimax/minimax-m3 pairwise-product-v2-multiaxis | yes | 66% | 43% | 73% · 24-9 | 27% · 9-24 |
openai/gpt-5.5 pairwise-product-v2-multiaxis | yes | 88% | 46% | 91% · 40-4 | 9% · 4-40 |
qwen/qwen3.7-plus pairwise-product-v2-multiaxis | yes | 58% | 39% | 69% · 20-9 | 31% · 9-20 |
xiaomi/mimo-v2.5 pairwise-product-v2-multiaxis | yes | 58% | 35% | 52% · 15-14 | 48% · 14-15 |
z-ai/glm-4.6v pairwise-product-v2-multiaxis | no | 24% | 14% | 83% · 10-2 | 17% · 2-10 |
pairwise-product-v2-multiaxis
stable-pair filter used; 41/50 swap-consistent pairs, 9 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 35/50 swap-consistent pairs, 15 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 33/50 swap-consistent pairs, 17 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 44/50 swap-consistent pairs, 6 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 29/50 swap-consistent pairs, 21 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 29/50 swap-consistent pairs, 21 pairs withheld from stable judge ranking
Axis leaders (stable)
pairwise-product-v2-multiaxis
stable-pair filter used; 12/50 swap-consistent pairs, 38 pairs withheld from stable judge ranking; position bias: B selected 14%
Axis leaders (stable)
The private per-battle rows stay withheld. Click through the public demo battles below to see the exact A/B judge input, and watch the verdict hold when the order is swapped.

Lunara lunchbox
Infographic listing asset

Generated output

Generated output
openai/gpt-5.1 and google/gemini-3.1-pro-preview each see reference + A + B and return a single forced winner. Try swapping the order, then reveal the verdict and their reasoning.
Public demo battles, shown for input format only. The 100 private replay rows, prompts, and raw judge outputs remain withheld.
The result is a controlled VLM-judged preference signal. It is solid enough for product and model diagnostics, but not a final scientific claim about human preference.
A forced-choice* battle* between two generated images, grounded* by the same reference and target task.
This snapshot separates product identity, task fit, commercial preference, and overall winner. Product identity means the same physical product, not the same shot.
Every pair is judged in both A/B orders. Pairs that fail swap consistency* are dropped from that judge’s ranking; judges flagged for position bias* are kept for diagnostics but excluded from the consensus entirely.
The headline ranking is a Bradley-Terry* strength; winrate* is reported as a control. Judges are summarized with agreement*.
not a human-preference proof
position bias is measured
product, task, commercial
6/7 judges, 92% agreement
| Question | Current evidence | Caveat | Status |
|---|---|---|---|
| Best use | Model diagnostics and regression tracking | Not a replacement for human preference testing | Solid diagnostic |
| Sample | 700 votes, 100 battles, 50 products, 2 generators | 95% Wilson interval and exact binomial test reported with Table 1 | Preliminary |
| Judging | Multi-axis VLM decisions with product identity separated from shot matching | VLM judges can share systematic biases | Improved |
| Scope | Private replay sample: infographic (27), lifestyle (23) | Transfer to other catalog categories is unproven | Bounded |
Artifact manifest

Provenance
I am Josua, building like.photo. My goal is to take the manual decision step out of product-image creation: instead of making users inspect dozens of outputs, the system should generate, evaluate, and keep the strongest product photos automatically. This benchmark comes from a practical observation: VLMs are often noticeably ahead of image-generation models at judging whether a product is shown correctly and which image is commercially stronger. I created like.photo Arena to make those assessments persistent, reproducible, and open to critique. Feedback, missing baselines, and conversations with people working on image generation or evaluation are very welcome.
Use the snapshot URL for this exact result. The live page can advance to new suites, but private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50 keeps referring to this generated artifact.
Snapshot identity
Share actions
Plain citation
Josua Sievers. like.photo Arena: Private Replay: Nano Banana 2 vs GPT Image 2: private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50. VLM-judged pairwise product image generation benchmark. Generated Jun 12, 2026, 12:03 UTC. https://arena.like.photo/s/private-replay-nano-banana-2-vs-gpt-image-2/20260531-batch50
BibTeX
@misc{likephoto_private_replay_nano_banana_2_vs_gpt_image_2_20260531_batch50,
title = {like.photo Arena: Private Replay: Nano Banana 2 vs GPT Image 2},
author = {Sievers, Josua},
publisher = {like.photo},
year = {2026},
url = {https://arena.like.photo/s/private-replay-nano-banana-2-vs-gpt-image-2/20260531-batch50},
note = {Snapshot private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50; generated Jun 12, 2026, 12:03 UTC; 100 battles; 6/7 consensus-eligible VLM judges}
}Asterisked terms throughout the page are defined here. Hover or focus any term for the same note inline.