Archive note: this article captures GPT-5's launch and the author's assessment on 13 August 2025. Product limits, model rankings, reported behaviour and rollout announcements retain that historical context. Quotations translated from the Spanish body are distinct from the original English transcribed in screenshots. The article includes its own late clarification about context limits; both the initial account and that clarification are preserved.

“The floor of AI quality has risen significantly; the ceiling has not.” This comment by engineer and entrepreneur Andrzej Dąbrowski, seen on X just minutes after the GPT-5 presentation, is probably the best possible summary. Hundreds of millions of free users around the world are moving from a model such as GPT-4o to access to GPT-5, which can include reasoning. It is important to call on that reasoning when needed; we will return to that. This is an enormous leap. By contrast, more advanced users who regularly use frontier models, or simply ChatGPT Plus subscribers who routinely used o3, will notice only a small incremental improvement. That is entirely logical, given that o3 launched in April, just four months ago.

I was quite convinced that paying users tended to call on reasoning models for many tasks. But according to these figures from Altman himself, the share of users — even paying ones — using the best reasoning models was tiny, making the overall leap greater still. It is quite possible, and at the same time deeply troubling and disappointing — and makes newsletters such as this one more valuable than ever — that medical, legal and consulting professionals, as well as policymakers, were using a much “dumber” model than the one available to them for very important tasks, with all the consequences that could have for the results.

“The floor has risen” mainly means that almost everyone, even without advanced AI knowledge, will notice more useful and reliable answers, especially if they activate reasoning mode when the problem requires it. Alternatively, they can use the simple trick of asking the model to “think deeply” before answering, without spending uses of Thinking mode.

What GPT-5 brings — and what it does not

What it does bring: a unified system that routes between a “fast” model and a fairly good reasoning model, GPT-5 Thinking. The router “broke” on the first day, causing numerous problems and creating the impression — mine included — that this was a rather mediocre model. The problem was that it tried to solve complex questions with unsuitable models. There are fewer hallucinations, better instruction-following and visible improvements in coding and structured writing. It is also much cheaper and more efficient relative to its capabilities than its predecessors: the great advantage for OpenAI.

What it does not bring: for free and Plus ChatGPT users, the context window remains very modest compared with what is available through the API. In practice, it is tiny for free users, at 8,000 tokens, and 32,000 tokens on Plus — far from the million tokens Gemini can now offer.

What this means in practice for readers: in principle, you will notice greater reliability, a better tone and more developed plans, like the example we tested below. But you will not be able to put even 200 pages into ChatGPT Plus at once unless you have Pro or use the API, something uncommon among non-technical users. This is a tremendous limitation, which we raised with Sam Altman, alongside dozens of other users, during his Reddit AMA last Friday. They have promised to look into how much information the model can process as input and output in one interaction. But expanding that capacity carries large computing costs and, unlike Google with its own TPUs, OpenAI is currently very short of compute. I do not know whether we will see major improvements here soon, even though the model itself allows them: through the API, GPT-5 is served with a context window of up to 400,000 tokens.

IMPORTANT NOTE: it appears that the very small context window applies only to the “normal” GPT-5 model without Thinking, according to this clarification from an OpenAI employee just hours before this edition closed:

Yann Dubois clarifies historical Plus context limits; embedded post is truncated.

Original English post by Yann Dubois, dated 12 August 2025:

I saw a lot of people complaining about 32k context size in ChatGPT for plus users, which would be terrible for coding. But actually we are giving 196k context size for plus users when using GPT5 thinking and that’s the model you should use for coding use-cases!

32k is for the chat/non-reasoning model. If you have examples that require more than 32k for non-coding usecases please post them below.

Sorry, this wasn’t clear in the original release... :/

The embedded Mark Kretschmann post begins: “GPT-5 with a 32K context: Why it causes problems for @OpenAI users! A 32K token window for GPT-5 sounds generous until you try to do real work with it. In multi turn chats and coding sessions, that budget melts fast. Every message carries overhead you never see, including system”. The screenshot stops there; its hidden continuation is not reproduced.

The following chart, shared on X by AI communicator Jon Hernández, illustrates the leap and the importance of using Thinking mode for complex tasks, or at least the phrase “think deeply”. Several users report that this activates “thinking low”, which is not optimal but is already far better than the default model. This has not been officially confirmed: OpenAI has yet to make fully clear which specific model within the family we are using at any given moment.

Chart comparing model intelligence scores and routing modes.

Original chart: “ChatGPT 5 intelligence highly depends on the ‘router’”. Its subtitle says that underlying intelligence has improved, but routing to Thinking would be a huge upgrade for many regular GPT-4o users. Values shown on the Artificial Analysis Intelligence Index:

Scroll across the table to read all columns.

New model Score Earlier model Score
GPT-5 (high) 69 o3 67
GPT-5 (medium) 68 o4-mini (high) 65
GPT-5 mini 64 GPT-4.1 47
GPT-5 (low) 63 GPT-4.5 42
GPT-5 nano 54 GPT-4.1 mini 42
GPT-5 (minimal) 44 GPT-4o 40

Rows list the two series, not matched equivalents. The graphic marks GPT-5 minimal as the instant-answer default and medium as Thinking mode, a 24-point difference. It marks GPT-4o as the earlier ChatGPT default and a 27-point gain from selecting o3. Credit printed on the graphic: Artificial Analysis (https://artificialanalysis.ai/) and Peter Gostev (@PeterGostev; https://www.linkedin.com/in/peter-gostev/). These are historical chart labels, not verified descriptions of today's routing.

The model presentation was also rather flat and underwhelming, in contrast to the hype of recent weeks on social media, and heavily focused on selling the model's considerable coding strengths.

The initial rollout was troubled — “bumpy”, as Altman himself described it — with numerous complaints about a router that misdirected some queries, almost always choosing the “dumber, cheaper” model. There was also considerable criticism from highly advanced users expecting a much more spectacular leap from a much larger model. Altman acknowledged this in an open Reddit question-and-answer session, an AMA, last Friday and promised fixes.

Meanwhile, the company had to restore GPT-4o to the model picker following community pressure. Some users found 4o warmer and more creative — please see the original newsletter's Meme of the Week section. That is possible, since GPT-5 shines particularly in technical and coding tasks. Many users even experienced the change as the loss of a “loved one”, having formed a supportive emotional relationship with the previous model. This both fascinates and “terrifies” Sam Altman, who has commented on it on X.

Early this week, Altman posted an X thread on “changes to come”, from making manual selection of Thinking easier to adjusting the rollout's pace. He announced increased reasoning-model usage for Plus subscribers and reiterated that they would listen to intensive users, such as those of us active in Reddit's AI forums, as well as non-technical users.

I would like to pause over the question-and-answer session held by OpenAI's team, particularly Altman, on Reddit, to draw attention to one crucial absence.

Altman's AMA — and the missing angle

This AMA was a rare and valuable opportunity. The CEO of one of the most influential technology companies openly taking the community's questions, complaints and suggestions should set an example, including in politics. The discussion was dominated by technical questions — the router, earlier models and limits — pricing and power users' concerns. Altman acknowledged problems and promised corrections: better routing, an adjusted rollout pace, the return of options such as GPT-4o, and greater control over when to activate Thinking.

Yet we missed questions about uses in politics and government: government planning, emergency services, public hearings, procurement, urban planning. In these forums, the absence of advisers, consultants and public policymakers — even serving politicians — leaves the debate almost entirely to technical specialists and programmers. AI is a technology that will transform everything, not only software: how decisions are made, policies communicated and public services delivered.

This strikes me as so important, and so obvious in situations like this, that we will return to it at length in a future issue. The political world's abdication of responsibility over the development of this technology is serious. As we reported in the previous issue's news, the 2024 Nobel laureate in chemistry, Demis Hassabis, considers its impact likely to be ten times greater than the Industrial Revolution's in magnitude, and probably in speed too.

Precisely in relation to politics and government, GPT-5 brings — or at least its presentation and several benchmarks have shown us — some key improvements:

  • Hallucinations and safety: GPT-5 Thinking consistently reduces invented facts and adds clearer safeguards.
  • Greater honesty: when it lacks sufficient data or the question is impossible, it says so — “I don't know” or “information is missing” — instead of filling the gaps.
  • Fewer errors: it reduces the tendency to invent answers and deceptive behaviour in scenarios involving broken tools or incomplete requirements.
  • Safe answers: on sensitive subjects, such as biology and cybersecurity, it avoids dangerous instructions and offers high-level alternatives.
  • Oversight and testing: reasoning is monitored, red-teaming exercises are conducted, and layered controls aim to curb misuse.

OpenAI response-level error rates on de-identified ChatGPT traffic.

OpenAI chart, “Response-level error rate on de-identified ChatGPT traffic”: the share of responses with at least one error is 4.8% for GPT-5 with thinking, 11.6% for GPT-5 without thinking, 22.0% for OpenAI o3 and 20.6% for GPT-4o.

OpenAI claim-level hallucination rates on three evaluations.

OpenAI chart, “Hallucination rate on open-source prompts”. Claim-level hallucination rates shown:

Scroll across the table to read all columns.

Evaluation GPT-5 with thinking OpenAI o3
LongFact-Concepts 0.7% 4.5%
LongFact-Objects 0.8% 5.1%
FActScore 1.0% 5.7%

Why does this matter for politics and government? Fewer hallucinations mean fewer incorrect communications, fewer corrections and more reliable answers to citizens. When data is missing, the system asks for clarification rather than inventing an answer: that improves transparency, traceability and accountability.

Our practical test of GPT-5's improvement

At estrategIA, we tested a simple, accessible example, suggested among others by the model itself: a twelve-month plan for developing a public bike-sharing system in a medium-sized city. Using a two- or three-paragraph brief and a single prompt, we asked GPT-4o, ChatGPT's default model, and GPT-5 with reasoning to produce a plan. Below are links to their full responses, which we do not reproduce here to avoid making this article excessively long:

Public bike-sharing plan with GPT-4o: https://chatgpt.com/canvas/shared/68970fa1908c81919aad934a737871d6

Public bike-sharing plan with GPT-5 Thinking: https://chatgpt.com/share/68971020-8348-8011-b084-12c9862c7dd5

The differences in depth, development and structure are fairly obvious at first sight. But since we are not public-transport experts, we asked a rival AI, also one of the world's most advanced, to judge both responses without telling it which models had produced them. We supplied only the texts and asked for its verdict, which follows.

Comparative analysis and key strengths of Model 2

Gemini 2.5 Pro's assessment of the two responses, translated from the original Spanish output:

Of course. Here is a comparative analysis of the two plans, a summary and a breakdown of their strengths.

Summary and verdict

After analysing both plans, Model 2 is markedly superior. It offers a much more professional and detailed approach, aligned with international standards in the shared-mobility industry.

Whereas Model 1 presents a competent, well-structured project plan, Model 2 demonstrates a deep understanding of urban planning and the operation of public bike-sharing systems, making it a much more robust and reliable guide to implementation.

Gemini comparison of the two public bike-sharing plans.

English translation of the complete comparison table in Gemini's response. These are the evaluating model's judgements, not an independently validated transport assessment:

Scroll across the table to read all columns.

Comparison criterion Model 1: basic plan Model 2: advanced plan Verdict
Technical rigour Generic, without grounding in specific standards. Mentions NACTO in isolation. Based on international standards, NACTO and ITDP, defining network density and infrastructure. Model 2
Strategic approach Valid but general objectives, such as reducing traffic. Focused on specific mobility problems, such as the first/last kilometre. Model 2
KPIs and data Basic indicators, such as trips per month and satisfaction. Does not mention data standards. Operational and industry KPIs, such as trips per bike per day and TTR; requires open data through GBFS. Model 2
Cost detail Simple, fixed cost estimates per unit. Granular, realistic CapEx/OpEx breakdown by equipment type: fixed versus hub, mechanical versus electric. Model 2
Pilot methodology A simple trial in two areas without methodological rigour. A statistically valid A/B test to inform decisions about which technology to scale. Model 2
Accessibility and equity These concepts are not mentioned explicitly. Incorporates physical and digital accessibility, including WCAG, from the design stage, together with a territorial equity plan. Model 2
Final conclusion A competent, functional project draft. A detailed, professional implementation plan, ready to deploy with greater assurance of success. Model 2

To keep this article from becoming too long, Gemini 2.5 Pro's full response is available at this link.

In summary, Model 1 is a good draft, but Model 2 is a professional implementation plan ready to put into practice. It minimises risks and maximises the chances of success by drawing on global industry experience and data.

As the example shows, the leap for users who previously used ChatGPT in its default mode will be substantial if the router correctly directs complex questions like this to Thinking. The GPT-5 family is quite capable, with the Pro model already reaching very high levels of “intelligence”. The following chart shows how the duration of tasks AI can perform continues to increase, now beyond two hours, with a curve that remains clearly exponential. If that continues, within two or three years AI could perform tasks that would take humans months. This supports the conclusion that GPT-5, while no revolution, is a powerful and interesting model in its Thinking version, despite limitations such as router failures or its very small context window.

METR chart of task length at 50% success, with GPT-5 time horizon and uncertainty.

METR chart: “GPT-5 has a 50%-time-horizon of about 2 hr 15 min (95% CI: 1–4.5 hours)”. The vertical axis is task length at a 50% success rate, and the horizontal axis is model release date. The chart's fitted trend has a doubling time of 213 days, using data from 1 January 2019 to 1 March 2025, with R² = 0.98. Labels trace GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1, o3, Grok 4 and GPT-5. Example task levels are answering a question, counting words in a passage, finding a fact on the web, training a classifier and training an adversarially robust image model. The extension of the line beyond the observed models is a projection. Credit: METR, metr.org, CC-BY.

In its most powerful reasoning version, GPT-5 (high), the model is now clearly one of the most capable AIs:

Kol Tregaskes post and historical LMArena category rankings.

Original English post by Kol Tregaskes: “GPT-5 (high) now dominates the LMArena leaderboards, #1 in all the following categories: Hard Prompts; Coding; Math; Creative Writing; Instruction Following; Longer Query; Multi-Turn.” The visible ranking table is transcribed below. Ellipses preserve model identifiers truncated in the screenshot; the two separately listed Claude Opus 4 rows remain distinct by their displayed overall rank.

Scroll across the table to read all columns.

Model as displayed Overall Hard prompts Coding Maths Creative writing Instruction following Longer query Multi-turn
gpt-5-high 1 1 1 1 1 1 1 1
claude-opus-4-1-202… 2 1 1 3 1 1 1 1
gemini-2.5-pro 2 2 3 1 1 2 1 1
o3-2025-04-16 2 3 3 2 6 6 7 6
chatgpt-4o-latest-2… 4 4 3 11 2 4 2 1
gpt-4.5-preview-202… 4 5 4 6 2 3 2 1
grok-4-0709 5 6 8 1 2 4 5 6
qwen3-235b-a22b-ins… 5 2 1 1 4 2 1 5
claude-opus-4-20250… 7 4 2 5 2 2 1 6
kimi-k2-0711-preview 7 6 4 6 9 12 10 6
deepseek-r1-0528 8 6 6 6 8 12 10 11
glm-4.5 8 4 4 5 6 4 6 8
claude-opus-4-20250… 9 7 7 6 2 5 2 6
gemini-2.5-flash 11 14 20 5 4 9 7 13
grok-3-preview-02-24 11 8 13 18 6 8 6 11
gpt-4.1-2025-04-14 12 8 9 29 6 9 5 6
qwen3-235b-a22b-thi… 12 8 4 4 6 6 7 9
claude-sonnet-4-202… 15 8 3 6 8 5 3 6
o1-2024-12-17 16 14 17 6 13 11 12 20
o4-mini-2025-04-16 16 17 15 5 24 21 25 17
qwen3-235b-a22b-no-… 16 15 10 8 17 21 15 9
deepseek-r1 17 14 13 6 13 12 17 9
deepseek-v3-0324 17 18 15 29 8 21 17 8
claude-3-7-sonnet-2… 22 17 16 16 9 10 7 13
o1-preview 22 24 27 15 17 21 22 17

Conclusions from GPT-5's release

Humanity as a whole will use AI better after GPT-5: fewer glaring errors, greater competence by default, and a reasoning mode that makes a difference when used thoughtfully. There is no large leap in the ceiling — only four months have passed since o3 — but there is a clear rise in the floor, widening access to quality results. In the immediate term, OpenAI's challenge is not just technical. It concerns the product — interface, router and transparency — and communication: aligning expectations with real limits, such as context.

The corrections of these first days — restoring GPT-4o, admitting a “bumpy” launch, promising adjustments and substantially increasing usage allowances for Plus subscribers — are moving in the right direction. Today, 13 August, Altman announced a change to the model-selection system to address the usability problems caused by the malfunctioning router. The rest depends on listening to those who did not ask questions in the AMA: much of the wider public, the ordinary user who benefits most from this change, alongside OpenAI itself, which will use models that are more powerful but, above all, much cheaper. And for our own field — public administrations, policy consultants and communicators — that is where a “higher floor” becomes public value.

Fernando Nieto Lobato

Director of Digital Innovation at Institución Educativa ALEPH

This is a translation of the original Spanish essay published on 13 August 2025. Its claims, examples and forecasts retain that historical context. Read the original Spanish edition, including its accompanying illustrations.

Cite this essay

Fernando Nieto Lobato. “GPT-5: the floor rises sharply, the ceiling barely moves — and why that matters.” estrategIA, issue 098, 13 August 2025. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/098/

Back to the top ↑