Summary
Researchers built a benchmark to evaluate AI image generation models by anchoring them against real photographs, expecting synthetic images to always score below reference photos. However, after 863 blind judgments from 13 professionals, GPT Image 2 scored higher than real photographs, revealing three critical flaws in their methodology that made their ruler unreliable for measuring AI image quality.
- The benchmark used a four-rung ladder with FLUX.2 Klein 9B (floor), Qwen Image 2 (middle), GPT Image 2 (strong), and real photographs (reference at 100), but GPT Image 2 unexpectedly surpassed the real photograph in evaluations
- The study avoided absolute rating scales and instead used blind pairwise comparisons between 192 total assets across 48 commercial scenes, with judges answering questions about brief adherence, photographic plausibility, and visual quality
- Thirteen production professionals (art directors, content specialists, and QA experts) conducted 863 blind judgments, and results were analyzed using the Bradley–Terry model, the same statistical framework used to rank chess players
- The methodology exposed fundamental flaws in how AI image benchmarks were designed, with implications for anyone running model evaluations for product or creative teams
- Lower-tier models behaved as expected with clear gaps between them, but the top of the ranking inverted, suggesting the benchmark's design needed significant revision before attempting to grade AI images again
What Are Your Ideas Today?
Tested as of August 2026. Models used: FLUX.2 Klein 9B, Qwen Image 2, GPT Image 2.
We built a ruler for AI images and placed a real photograph at the very top—the one spot we were certain no generative model could ever reach. But after 863 blind judgments from 13 production professionals, GPT Image 2 was sitting comfortably above it. Naturally, we had to figure out how that happened. Almost every answer pointed right back at the ruler.
This is the story of that benchmark: what we set out to measure, the three critical flaws we uncovered in our own methodology, and what we will change before anyone attempts to grade AI images this way again. If you regularly run model evaluations for a product or creative team, you will likely recognize at least one of these mistakes.
Why benchmark AI images against a photograph at all?
Comparing image models exclusively against one another creates a persistent shelf-life problem. Every new model release reshuffles the leaderboard, and knowing a model is “better than last quarter’s baseline” tells you nothing about whether it actually satisfies a client brief.
Thermometers solved this dilemma centuries ago. You don’t define temperature by comparing two cups of tea; you anchor the scale to fixed physical points, like the freezing and boiling points of water, and measure everything else against them.
So, we built a four-rung ladder, with every tier working from the exact same creative brief:
- Floor: FLUX.2 Klein 9B, a deliberately modest model
- Middle: Qwen Image 2
- Strong: GPT Image 2
- Reference: A real, unedited photograph of the scene

Before a single vote came in, we arbitrarily labeled the rungs 30, 55, 80, and 100. Those numbers were simply design sketches, not strict measurements. The true goal of the experiment was to see whether reality would validate them.
The logic behind our top rung felt airtight: whatever a synthetic model produces, it is not a photograph. Newer models could certainly climb closer toward 100, but nothing should ever surpass it.
How the test worked
We curated 48 scenes spanning eight commercial categories: fashion and beauty, food and beverage, interiors, layout and typography, people and lifestyle, product and e-commerce, social marketing, and street and architecture. Thirty-eight reference photos originated from Everypixel’s own commercial shoots, while ten were sourced from Unsplash.
For each photo, we drafted a detailed text description and converted it into a prompt for all three models, yielding four variations per scene and 192 total assets.
Thirteen colleagues then evaluated blind image pairs without knowing their source. Our panel included four art directors, five content distribution specialists, and four production QA experts. For every match-up, they answered three distinct questions:
- Which image follows the brief better?
- Which one is more photographically plausible?
- Which one simply looks better?
We avoided absolute 10-point scales because ratings drift wildly between individuals—and even within the same person over a long session. Forcing a choice between two side-by-side images is far more reliable. We then fed all selection data into a Bradley–Terry model—the same statistical framework used to rank chess players from match outcomes. Its superpower is placing every contender on a single continuous scale, even if certain pairs never competed head-to-head.

What came back: The bottom held, but the top flipped
The lower half of the ladder behaved precisely as expected. FLUX.2 Klein 9B landed firmly at the bottom, Qwen Image 2 claimed the middle, and clear gaps separated them. Our industry professionals can easily distinguish a weak model from a strong one.
The real surprise materialized at the absolute peak. Across the entire comparative network, the final hierarchy emerged as:
FLUX.2 Klein 9B < Qwen Image 2 < Real Photograph < GPT Image 2
Before concluding that “AI officially beats photography,” look closely at the direct head-to-head metrics. When GPT Image 2 went toe-to-toe with the actual photograph, the photograph won 46.2% of the time. With a confidence interval spanning from 39.8% to 52.9%, that safely encompasses 50%—making it a statistical coin toss.
How can a model rank higher overall yet only tie in direct matchups? Think back to chess tournaments. Two grandmasters might draw against each other, but one systematically dominates every other player in the tournament by wider margins. The rating algorithm places that dominant player on top, despite never winning the key head-to-head match. GPT Image 2 outperformed the weaker models more decisively than the photograph did, and the network model factored that in.
Thus, our honest takeaway is more nuanced than the headline: the photograph failed to hold its ground as a stable ceiling. That fatal flaw for our scaling model sent us straight back to the drawing board.
Flaw #1: The prompt was written directly from the photograph
This proved to be our most impactful misstep, hiding neatly inside the step that initially felt the most logical.
Every single text prompt originated as a literal description of its reference photo. Consider what that means for each contender: the AI model receives a clean text spec and meticulously paints every instruction. Meanwhile, the actual photograph contains everything the description accidentally omits—a stray power cable, an awkwardly cropped chair, or lighting that looks slightly flat simply due to architectural constraints.
When judges were asked which image followed the brief better, the generated image inherently won because the brief was literally an idealized summary of it. On this specific metric, the photograph won just 41.7% of direct comparisons—the only metric where the photograph’s entire confidence interval sat strictly below 50%.
In short, the photograph lost a prompt-adherence contest against a prompt written about it. Critics of AI evaluation can rightly argue that a major share of GPT Image 2’s victory stemmed from experimental design rather than pure image superiority.
Lesson learned: If your prompts derive directly from reference images, measure prompt adherence as an isolated variable, or completely exclude the reference image from that specific judgment.
Flaw #2: The tie button was pre-selected
Our evaluation interface defaulted to every question with the tie option pre-selected. If a tired judge briefly skipped a question, the system quietly recorded a default tie.
Ties ultimately accounted for 658 out of 2,589 answers, or 25.4%. Post-analysis couldn’t reliably distinguish between genuine ties and neglected interface controls because the system failed to log physical interaction tracking.
There is a silver lining: judgments resulting in triple ties actually took longer on average (a median of 57.3 seconds versus 45.0 seconds for active choices), indicating reviewers weren’t just rushing blindly through panels. Furthermore, when we completely stripped out ties and recalculated the model, GPT Image 2’s lead over the photograph actually widened from +0.302 to +0.463, leaving the overall hierarchy intact.
Even so, this bug degrades inter-rater reliability and muddies the precise gaps between rungs. Any future iteration demands a fixed interface.
Lesson learned: In a rating UI, every default choice is a vote. Keep controls entirely blank by default and log explicit user interaction.
Flaw #3: Equal screen size doesn’t equal an equitable comparison
All four images were rendered at 2048 pixels on the long edge. That sounds fair on paper—until you examine the source constraints.
The synthetic outputs were displayed close to their native generation resolutions, whereas the reference photographs were downscaled from massive, high-megapixel source files. Fine micro-details—such as natural skin texture, fabric weaves, and faint background text—represent photography’s strongest visual moat, and downscaling strips them away first.
By equalizing the display medium, we inadvertently neutralized photography’s greatest advantage. Next to our prompt bias, this stands as the most compelling counter-argument to our headline results.
Lesson learned: Equal display dimensions do not equal equivalent informational value. If you intend to benchmark against real photography, display photos at full native resolution or test multiple resolution conditions.
The part that wasn’t a bug: Expert disagreement
Each pair was reviewed by three distinct colleagues to measure consensus. When comparing images from opposite ends of the ladder, alignment was effortless (FLUX.2 Klein 9B vs. GPT Image 2 scored 64% pairwise agreement).
At the top tier, however, consensus completely broke down. GPT Image 2 versus the photograph achieved just 47% agreement. Eight of our 13 judges ranked GPT Image 2 higher overall, while five favored the real photograph.
Using Gwet’s AC1 statistic, agreement hovered between 0.254 and 0.385 across all three questions—falling below our research protocol threshold for a definitive result. Consequently, the study must be categorized strictly as directional rather than absolute.
This split insight is perhaps our most valuable takeaway: thirteen seasoned professionals whose careers depend on visual judgment could not agree on a single definition of “better” once mediocre contenders were removed. If your organization relies on a single reviewer to sign off on visual assets, that vulnerability is worth examining.
The one category where photography held firm
Breaking down results by category revealed that in seven of eight domains, the network placed GPT Image 2 above the photograph, scoring widest margins in interiors and fashion/beauty.
People and lifestyle served as the sole exception. It was the only vertical where real photography came out ahead, capturing 65.7% of direct comparisons.
We hesitate to turn this into a blanket rule. With only six scenes per category, the confidence interval for that subset still runs from 48.1% to 84.3%. It strongly suggests that scenes anchored in genuine human behavior warrant a dedicated, standalone study—nothing more, nothing less.

How to benchmark AI images without breaking your scale
Here is our blueprint for the next iteration of the ladder. If you are building an independent evaluation framework, you can adopt these guidelines directly:
- Never assume reference perfection. Empirically test whether your reference asset actually holds the top spot before building a scale on top of it.
- Implement category-specific top rungs. A reference anchor that excels for architectural interiors will fail for human portraits.
- Define references by quality, not origin. “The highest-grade asset a professional would approve” makes a far superior anchor than “whatever the camera captured.”
- Render photographs at full resolution. Otherwise, you systematically erase the fine detail photography does best.
- Isolate perceptual quality from prompt adherence. Keep them distinct, especially if your prompts are reverse-engineered from your references.
- Treat UI defaults as design data. Keep controls empty and log every user interaction.
- Publish agreement metrics alongside rankings. Rankings published without inter-rater agreement data project false certainty.
What this means for modern visual creators
This study does not prove that AI images have definitively beaten photography, nor does it establish that GPT Image 2 looks more “real” than a genuine camera capture. On photographic plausibility, the photograph received 51.7% of preferences, with the confidence interval still including 50%.
Instead, the practical implications are far more immediate. For commercial assets evaluated on a blended mix of brief compliance, plausibility, and aesthetic polish, photography is no longer an automatic ceiling. Creative teams should feel confident testing generative outputs right alongside scheduled reference shots rather than assuming the camera always wins.
Real photography remains preferable or necessary in several contexts:
- Scenes anchored in authentic human micro-expressions and real behavioral nuances
- Production work requiring uncompromising, full-resolution fine detail
- Strict legal documentation of real products, physical locations, or live events
- Aesthetic styles outside the commercial boundaries of standard benchmarks
The question we are left with
We set out to build a rigid ruler with a permanent, immovable top. Instead, the top moved. Part of that shift resulted from our own methodological mistakes—but another part hints at a deeper truth: “real” and “better” were never synonymous in the minds of visual gatekeepers.
So, once generative models cross the threshold of indistinguishable quality, what should we anchor our scales to?