like.photo Arena · 20260531-batch50 · Latest public view

Which AI model makes the best product photos?

like.photo Arena runs head-to-head battles* between image generators on real products. A panel of vision AI* judges picks winners blind, in both A/B orders. Every number is published with its uncertainty; this is diagnostic evidence, not yet a human-preference study.

Current snapshot: gpt-image-2 wins 27 of 36 decisive comparisons against nano-banana-2.

Snapshot private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50·Cite or share

Top model record
27-9
gpt-image-2
Products
50
replay tasks
Judges
6/7
eligible
Agreement
92%
majority coverage

How a battle works

Try it in the explorer
Reference product photo of a stainless steel lunchbox with bamboo lid
Referencethe real product photo
Generated infographic candidate from nano-banana-2 for the lunchbox
Candidate Anano-banana-2
Generated infographic candidate from gpt-image-2 for the lunchbox2/2 judges
Candidate Bgpt-image-2

Each vision AI judge receives exactly four inputs: the reference photo, candidate A, candidate B, and a text brief naming the product, its critical attributes, and the infographic task prompt. Model names are never included. The judge must answer with a single letter, A or B; the same pair is judged again with the order swapped, and only verdicts that survive the swap count. Here all 2 judges picked candidate B. Public demo battle; the scored battles use the same format.

Generator ranking

Consensus leaderboard

Generators ranked by strict eligible-judge majority over swap-stable private replay pairs. Individual product examples are withheld.

Figure 1

Win rate with 95% confidence intervals

bar = winrate · whisker = 95% Wilson CI · dashed = even split

Rank #1

gpt-image-2

75%27-9BT +80

95% CI 59% to 86%

Rank #2

nano-banana-2

25%9-27BT -80

95% CI 14% to 41%

gpt-image-2: 100% vote coverage
nano-banana-2: 100% vote coverage

Matrix

Head-to-head record

row vs column · strict-majority swap-stable pairs

vsgpt-image-2nano-banana-2
gpt-image-2—27-9 (75%)
nano-banana-29-27 (25%)—
1Product pairs
50
judged in both orders by 7 judges (700 votes)
2Swap-stable per judge
41 / 35 / 33 / 44 / 29 / 29
order-flipped pairs dropped
3Quorum-stable pairs
39
enough stable judge votes
4Majority winner
36
basis of Table 1

Table 1: Consensus ranking over 36 strict-majority swap-stable pairs, from 700 raw votes by 6 eligible judges.

Rank #1

gpt-image-2

BT

+80

Winrate
75%
Votes
36
Wins
27
Losses
9

Rank #2

nano-banana-2

BT

-80

Winrate
25%
Votes
36
Wins
9
Losses
27

* BT Latent 'strength' per generator from pairwise wins; higher = preferred. Reported as 100 times the natural log of the strength, averaged over eligible judges. more·* Winrate Fraction of a generator's battles that it won. A simple control metric. more

Uncertainty gpt-image-2 wins 75% of the 36 agreed pairs, 95% Wilson interval 59% to 86%. Exact two-sided binomial test against an even split: p = 0.004.

Axis diagnostics

What the judges preferred by criterion

Aggregated over consensus-eligible judges using swap-stable pair wins. Product identity tracks the same physical product, while task fit and commercial preference can reward a different shot or scene.

AxisLeadergpt-image-2nano-banana-2
Product identitygpt-image-2
67% · 22-11
33% · 11-22
Task fitgpt-image-2
80% · 28-7
20% · 7-28
Commercial preferencegpt-image-2
73% · 27-10
27% · 10-27
Overallgpt-image-2
75% · 27-9
25% · 9-27
Task type split

Which model won by output type

Aggregated by eligible-judge majority over swap-stable private replay pairs. Tied or non-majority quorum pairs are shown as withheld and do not assign a winner.

Task typeTasksMajorityWithheldgpt-image-2nano-banana-2
Infographic2720/211
85% · 17-3
15% · 3-17
Lifestyle2316/182
63% · 10-6
38% · 6-10
Judge diagnostics

VLM judge panel

The same battle list is scored by every judge. Swap consistency and position bias decide whether a judge feeds the consensus.

92%
36/39 quorum stable pairs reached a strict eligible-judge majority; 3 tied or withheld
—
needs per-battle rows (withheld)
Consensus-eligible
6 / 7
4.6 stable votes / pair

Matrix

Judge reliability and preference

generator cells show winrate · wins-losses

JudgeEligibleSwapB rategpt-image-2nano-banana-2

anthropic/claude-fable-5

pairwise-product-v2-multiaxis

yes82%45%
61% · 25-16
39% · 16-25

google/gemini-3.5-flash

pairwise-product-v2-multiaxis

yes70%41%
71% · 25-10
29% · 10-25

minimax/minimax-m3

pairwise-product-v2-multiaxis

yes66%43%
73% · 24-9
27% · 9-24

openai/gpt-5.5

pairwise-product-v2-multiaxis

yes88%46%
91% · 40-4
9% · 4-40

qwen/qwen3.7-plus

pairwise-product-v2-multiaxis

yes58%39%
69% · 20-9
31% · 9-20

xiaomi/mimo-v2.5

pairwise-product-v2-multiaxis

yes58%35%
52% · 15-14
48% · 14-15

z-ai/glm-4.6v

pairwise-product-v2-multiaxis

no24%14%
83% · 10-2
17% · 2-10

pairwise-product-v2-multiaxis

anthropic/claude-fable-5

Votes
100
Cost
$13.38
Latency
9.2s
B rate* 45%Top BT* +39Consensus eligible

stable-pair filter used; 41/50 swap-consistent pairs, 9 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitygpt-image-2 24-20
Task fitgpt-image-2 27-15
Commercial preferencegpt-image-2 23-17
Overallgpt-image-2 25-16

pairwise-product-v2-multiaxis

google/gemini-3.5-flash

Votes
100
Cost
$2.00
Latency
8.5s
B rate* 41%Top BT* +80Consensus eligible

stable-pair filter used; 35/50 swap-consistent pairs, 15 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitygpt-image-2 23-9
Task fitgpt-image-2 26-8
Commercial preferencegpt-image-2 25-10
Overallgpt-image-2 25-10

pairwise-product-v2-multiaxis

minimax/minimax-m3

Votes
100
Cost
$1.04
Latency
11.3s
B rate* 43%Top BT* +85Consensus eligible

stable-pair filter used; 33/50 swap-consistent pairs, 17 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitygpt-image-2 19-13
Task fitgpt-image-2 22-6
Commercial preferencegpt-image-2 25-8
Overallgpt-image-2 24-9

pairwise-product-v2-multiaxis

openai/gpt-5.5

Votes
100
Cost
$7.19
Latency
5.5s
B rate* 46%Top BT* +200Consensus eligible

stable-pair filter used; 44/50 swap-consistent pairs, 6 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitygpt-image-2 31-7
Task fitgpt-image-2 42-4
Commercial preferencegpt-image-2 41-5
Overallgpt-image-2 40-4

pairwise-product-v2-multiaxis

qwen/qwen3.7-plus

Votes
100
Cost
$1.00
Latency
5.9s
B rate* 39%Top BT* +69Consensus eligible

stable-pair filter used; 29/50 swap-consistent pairs, 21 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitygpt-image-2 19-10
Task fitgpt-image-2 20-8
Commercial preferencegpt-image-2 20-8
Overallgpt-image-2 20-9

pairwise-product-v2-multiaxis

xiaomi/mimo-v2.5

Votes
100
Cost
$1.06
Latency
9.1s
B rate* 35%Top BT* +6Consensus eligible

stable-pair filter used; 29/50 swap-consistent pairs, 21 pairs withheld from stable judge ranking

Axis leaders (stable)

Product identitynano-banana-2 16-14
Task fitgpt-image-2 14-11
Commercial preferencegpt-image-2 14-13
Overallgpt-image-2 15-14

pairwise-product-v2-multiaxis

z-ai/glm-4.6v

Votes
100
Cost
$1.00
Latency
11.9s
B rate* 14%Diagnostic onlyPosition bias*

stable-pair filter used; 12/50 swap-consistent pairs, 38 pairs withheld from stable judge ranking; position bias: B selected 14%

Axis leaders (stable)

Product identitygpt-image-2 5-0
Task fitgpt-image-2 10-2
Commercial preferencegpt-image-2 10-2
Overallgpt-image-2 10-2
Evidence

Battle explorer

The private per-battle rows stay withheld. Click through the public demo battles below to see the exact A/B judge input, and watch the verdict hold when the order is swapped.

Original input
Reference product photo of a stainless steel lunchbox with bamboo lid

Lunara lunchbox

Infographic listing asset

Candidate A
Generated infographic candidate from nano-banana-2 for the lunchbox
nano-banana-2

Generated output

Candidate B
Generated infographic candidate from gpt-image-2 for the lunchbox
gpt-image-2

Generated output

openai/gpt-5.1 and google/gemini-3.1-pro-preview each see reference + A + B and return a single forced winner. Try swapping the order, then reveal the verdict and their reasoning.

Public demo battles, shown for input format only. The 100 private replay rows, prompts, and raw judge outputs remain withheld.

Methodology

What this benchmark claims

The result is a controlled VLM-judged preference signal. It is solid enough for product and model diagnostics, but not a final scientific claim about human preference.

01

Unit of measurement

A forced-choice* battle* between two generated images, grounded* by the same reference and target task.

02

Judging axes

This snapshot separates product identity, task fit, commercial preference, and overall winner. Product identity means the same physical product, not the same shot.

03

Bias control

Every pair is judged in both A/B orders. Pairs that fail swap consistency* are dropped from that judge’s ranking; judges flagged for position bias* are kept for diagnostics but excluded from the consensus entirely.

04

Ranking

The headline ranking is a Bradley-Terry* strength; winrate* is reported as a control. Judges are summarized with agreement*.

Claim strength
Diagnostic

not a human-preference proof

Controls
A/B swaps

position bias is measured

Judging
Multi-axis

product, task, commercial

Coverage
50 / 100

6/7 judges, 92% agreement

Scientific status matrix
QuestionCurrent evidenceCaveatStatus
Best useModel diagnostics and regression trackingNot a replacement for human preference testingSolid diagnostic
Sample700 votes, 100 battles, 50 products, 2 generators95% Wilson interval and exact binomial test reported with Table 1Preliminary
JudgingMulti-axis VLM decisions with product identity separated from shot matchingVLM judges can share systematic biasesImproved
ScopePrivate replay sample: infographic (27), lifestyle (23)Transfer to other catalog categories is unprovenBounded
Limitations
  • Judge-generator overlap. A judge can share a vendor or model family with a generator it ranks; self-preference bias is documented for LLM judges. Agreement from judges of a different vendor mitigates this but does not remove it.
  • Single draw per task. Each generator contributes one image per task, so the randomness of image generation is not sampled.
  • Asymmetric provenance. nano-banana-2 outputs are frozen historical generations while gpt-image-2 candidates were regenerated later with the same prompt and reference inputs, so generation-time pipelines differ. Tasks come from one historical batch, not a randomized sample.
  • Heuristic thresholds. Eligibility* uses a fixed 85% position-rate cut-off; Fleiss’ κ* is unstable when one generator dominates, as here.

Artifact manifest

Suite
private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50
Battle policy
round-robin-same-task-with-swaps-v1
Eligible judges
6/7
Generated
Jun 12, 2026, 12:03 UTC
Total votes
700
Products
50
Route
arena.like.photo
Storage
static public artifacts
Portrait of Josua Sievers

Provenance

Created and maintained by Josua Sievers

I am Josua, building like.photo. My goal is to take the manual decision step out of product-image creation: instead of making users inspect dozens of outputs, the system should generate, evaluate, and keep the strongest product photos automatically. This benchmark comes from a practical observation: VLMs are often noticeably ahead of image-generation models at judging whether a product is shown correctly and which image is commercially stronger. I created like.photo Arena to make those assessments persistent, reproducible, and open to critique. Feedback, missing baselines, and conversations with people working on image generation or evaluation are very welcome.

Persistent reference

Cite or share this snapshot

Use the snapshot URL for this exact result. The live page can advance to new suites, but private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50 keeps referring to this generated artifact.

Snapshot identity

Snapshot
private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50
Generated
Jun 12, 2026, 12:03 UTC
Live page
https://arena.like.photo
Permalink
https://arena.like.photo/s/private-replay-nano-banana-2-vs-gpt-image-2/20260531-batch50

Share actions

Share on X

Plain citation

Josua Sievers. like.photo Arena: Private Replay: Nano Banana 2 vs GPT Image 2: private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50. VLM-judged pairwise product image generation benchmark. Generated Jun 12, 2026, 12:03 UTC. https://arena.like.photo/s/private-replay-nano-banana-2-vs-gpt-image-2/20260531-batch50

BibTeX

@misc{likephoto_private_replay_nano_banana_2_vs_gpt_image_2_20260531_batch50,
  title = {like.photo Arena: Private Replay: Nano Banana 2 vs GPT Image 2},
  author = {Sievers, Josua},
  publisher = {like.photo},
  year = {2026},
  url = {https://arena.like.photo/s/private-replay-nano-banana-2-vs-gpt-image-2/20260531-batch50},
  note = {Snapshot private-replay-nano-banana-2-vs-gpt-image-2@20260531-batch50; generated Jun 12, 2026, 12:03 UTC; 100 battles; 6/7 consensus-eligible VLM judges}
}
Notes

Glossary

Asterisked terms throughout the page are defined here. Hover or focus any term for the same note inline.

*battle
The atomic unit of the benchmark: one ordered pair (candidate A, candidate B) for a fixed product and task, evaluated by every judge. Each unordered pair is run in both orders to expose order effects.
*battle policy
The versioned rule that determines which pairs become battles. The current policy (round-robin-same-task-with-swaps-v1) pairs candidates within the same task and emits each pair in both A/B orders.
*VLM
Vision-Language Model. A model that takes images and text together and produces text. Here it acts as an automatic judge, reading the reference and two candidates and returning a single forced choice.
*forced-choice
A comparison protocol in which the judge must return exactly one winner, A or B. No ties, abstentions, or scores are accepted as the primary signal, which keeps the unit of measurement robust and unambiguous.
*reference-grounded
Each battle is conditioned on the same reference product image(s) and the same target task. The judge is asked to prefer the candidate that preserves product identity (logos, text, shape, color, material, proportions) over the one that merely looks attractive.
*swap consistency
Also called flip consistency. The fraction of unordered pairs for which a judge names the same winner regardless of whether that candidate is shown as A or B. A value below 100% means the judge's verdict depends on presentation order.
*position bias
A systematic preference for a presentation slot rather than the image content. Measured through the B-selection rate over balanced battles; a judge is flagged when one slot is chosen in more than 85% of battles, or when swap consistency breaks down.
*B rate
The share of battles a judge resolved in favor of the B position. Because every pair is shown in both orders, an unbiased judge sits near 50%; values near 0% or 100% indicate position bias.
*consensus-eligible
A judge feeds the consensus ranking unless it is flagged for position bias or its verdicts depend on presentation order across the board. Eligibility works in two stages: pairs whose verdict flips when the A/B order is swapped are dropped from that judge's ranking, and the headline consensus keeps only the pairs on which every eligible judge names the same winner. Excluded judges are kept for diagnostics.
*winrate
The proportion of decided battles a generator won across consensus-eligible judges. Reported as a control metric alongside the Bradley-Terry score rather than as the primary ranking signal.
*Bradley-Terry
The Bradley-Terry model estimates a latent strength parameter for each generator from pairwise outcomes, so a single number summarizes many comparisons. Scores here are 100 times the natural logarithm of the estimated strength, centered around zero, and averaged over consensus-eligible judges. With only two generators the score carries the same information as the head-to-head record; it becomes informative on its own once three or more generators compete.
*observed agreement
For each battle, the share of judge pairs that name the same winner; averaged over all battles. In private replays it is computed over swap-stable pairs only, so the denominator excludes order-sensitive pairs. Unlike a chance-corrected coefficient it is easy to read and stable at small samples, so it is reported as the primary agreement figure.
*Fleiss' κ
Fleiss' kappa corrects observed agreement for the agreement expected by chance across more than two raters. When one outcome dominates and the sample is small, expected agreement approaches 1 and kappa becomes volatile or negative; at the current scale it is reported as a diagnostic only, not a quality claim.