Latest news
AI ModelsGPT ImageNano Banana 2FLUX

GPT Image 2.5 tops the Model League’s October test

GPT Image 2.5 topped the overall rankings in the October 2026 Model League with a score of 90/100. Four different models earned the highest score across the five architectural prompts.

Image of a two-story residence with a concrete and wood facade at blue hour, which earned GPT Image 2.5 the highest score in the Model League testAI image
Representative image, generated with AI.Image: 3dsınıfı / FCA AI

In brief

  1. GPT Image 2.5 ranked first overall with an average score of 90/100.
  2. FLUX.2 [pro] scored 93 on the prompt for a stone villa in Bodrum.
  3. Qwen-Image 3.0 Pro stood out on the kitchen detail prompt, while Nano Banana 2 led on the urban square prompt.
  4. GPT Image 2.5 had the lowest cost per image; Seedream 5.0 Pro had the longest average generation time.

What did we test, and how did we evaluate it?

In the October 2026 edition of the 3dsınıfı Model League, we compared six image-generation models using five architectural prompts. The prompts covered a residential exterior at twilight, a Scandinavian living room, a stone villa and pool in Bodrum, a kitchen material detail, and an aerial view of an urban square. Each model ran through the same API using its standard settings.

The images were evaluated by an AI judge that was not told which models had generated them. Four criteria were used: prompt adherence (30%), architectural accuracy (30%), photorealism (25%), and flawlessness (15%). Overall scores were calculated out of 100. The times below show average generation time, while costs are listed per image.

Rankings

RankModelScoreTimeCost
1GPT Image 2.590/10011.9 sec$0.011
2FLUX.2 [pro]85.8/10011.7 sec$0.045
3Nano Banana 285.6/10022 sec$0.103
4Nano Banana Pro85.3/10031.7 sec$0.138
5Qwen-Image 3.0 Pro84.9/10059.5 sec$0.04
6Seedream 5.0 Pro84.4/10077.7 sec$0.096

Standouts by prompt

Residential exterior at twilight: GPT Image 2.5 won with a score of 94.5. The two-story house, flat eaves, expansive glazing, blue-hour lighting, and gravel landscaping with olive trees were all judged to match the prompt. The landscaping obscured part of the front facade. Nano Banana Pro performed well on materials and warm interior lighting; Seedream’s sky was brighter, and its chimney broke the flat roofline.

Scandinavian living room: GPT Image 2.5 scored 90. Its geometry and daylight looked convincing, with two plants, a single pendant light, herringbone flooring, and a large window with sheer curtains. Qwen-Image 3.0 Pro also stood out for including the requested key elements and natural-looking materials. FLUX.2 [pro], however, showed three potted plants instead of the two specified in the prompt.

Stone villa and pool in Bodrum: FLUX.2 [pro] took first place with a score of 93. The stone villa, large windows, infinity pool overlooking the sea, and bougainvillea were clearly visible, while the geometry and lighting were rated highly. GPT Image 2.5 also delivered consistent building perspective. In Seedream’s image, some balcony and upper-floor connections looked complex, and the flowers appeared artificial in places.

Kitchen material detail: Qwen-Image 3.0 Pro scored 90. The brass faucet, ceramic bowl of lemons, and fluted cabinets stood out, and the window light looked natural. The image was clean despite a minor shape flaw at the faucet tip. In Nano Banana 2’s image, a human silhouette in the background was distracting.

Aerial view of an urban square: Nano Banana 2 won with a score of 90. The dense urban fabric, rows of trees, long water feature, and tram were clearly rendered; the architectural layout remained consistent despite the wide framing. GPT Image 2.5 also handled the tram, bicycles, and square geometry well. Nano Banana Pro was less consistent in its building perspective and square layout.

Which model is right for which task?

If you’re weighing overall score, speed, and cost together, GPT Image 2.5 was the most balanced option in this test: it earned the highest average score, took 11.9 seconds on average, and had the lowest listed cost per image. FLUX.2 [pro] stood out for local architecture and landscape scenes like the stone villa in Bodrum; its average time of 11.7 seconds also placed it among the faster models. Qwen-Image 3.0 Pro won the kitchen close-up, while Nano Banana 2 led on the urban square composition. These results are specific to the prompts tested and don’t mean the same ranking will hold for every task.

Limitations

Each prompt was generated once, so results may vary if the test is repeated. The evaluation was conducted by an AI judge, and scores for the criteria rely on visual judgment. Prompt phrasing, level of detail, and emphasis may also affect the outcome. The ranking should therefore be read as a comparison within this test set, not as a definitive quality guarantee for different project types.

See all the images side by side on the Model League page.

Sources

1 source
3T
3dsınıfı Test Laboratuvarı3dsinifi.com/model-ligi
Summary

Source texts are not republished; short quotes are marked, everything else is our own summary and commentary.

Frequently asked questions

Which model took first place overall in the October 2026 Model League?

GPT Image 2.5 ranked first with an average score of 90/100.

Which model had the lowest cost per image?

The test lists GPT Image 2.5 at $0.011 per image.

Are the results reproducible for every prompt?

Each prompt was generated only once, so results may vary. The wording of the prompts can also affect the evaluation.

Comments and the forum are in Turkish.Join the discussion
+

Related news