What actually changed between generations
Every model release comes with a page of claims and no numbers you can plan with. Here is what a generation step actually changed, from the perspective of somebody who had to reprice around it.
1. The cost estimator from the previous generation does not carry over
This is the one that costs money if you miss it, and OpenAI states it explicitly: the GPT Image 2 calculator does not estimate 2.5 consumption. The older figures run roughly 4× higher at the tier most products sell most of.
So anybody who priced a 2.5 product by scaling their 2-era spreadsheet is either overcharging by a wide margin or — if they scaled the other way on a different tier — quietly losing money. Re-measure; do not scale.
2. Two models where there was one
2.5 ships as a fast model and a base model, and the interesting part is that they bill identically — same rate per million output tokens. The quality model is slower, not dearer.
Consumer chat products route between them without telling you which answered, which means "I tried it in the app and it was worse" is not a statement about the model. It is a statement about the router.
3. Token count tracks the long edge, not the area
Measured: 1024² → 2048² comes in at about 2.03×, not 4×. A wide 2K frame can bill less than a square 1K one.
Anyone pricing image size by pixel count is overcharging their larger tiers by roughly a factor of two. This is the single most useful thing to know before publishing a price list.
And one that is mostly framing
"Better instruction following" appears in essentially every release. It is usually true and almost never actionable, because it is measured on benchmarks whose prompts look nothing like the ones your users write.
The version of that claim worth testing yourself is narrow: take the ten prompts your users actually run most, run them on both generations at the same settings, and look only at the failures. Improvements in the successful cases are invisible to a user; a failure class that disappears is the thing they notice.
For us the failure class that changed was typography at small sizes — not because the release notes said so, but because it was the thing our support messages stopped being about.
The side-by-side we ended up publishing is at gptimage25.top/gpt-image-2-5-vs-gpt-image-2.