# The year reasoning became real: what State of AI 2025 tells governments and public policymakers

Author: Fernando Nieto Lobato
Original publication: 2025-10-15
Spanish original: https://estrategiabyaleph.substack.com/p/estrategia-107-el-nuevo-orden-de
English URL: https://elcontemplador.github.io/estrategia-english/essays/107/
Status: Published translation

This is a translation of the original Spanish essay published on 15 October 2025. Its claims, examples and forecasts retain that historical context.

English publication: 2026-09-29

*Archive context: this analysis was published on 15 October 2025. Its assessments and forecasts retain that date's perspective. The English slide text below is transcribed directly from the preserved report images; statements in those slides are the report's claims, not newly verified findings.*

Every October, the global technology and strategy community eagerly awaits the [*State of AI Report*](https://www.stateof.ai/2025-report-launch). More than a report, it is a seismograph measuring the plate tectonics of our time. The 2025 edition, published last week by Nathan Benaich and the Air Street Capital team, confirms what many of us suspected: the tremor has become an earthquake. [For any leader, strategist or political actor, reading it is not optional; it is a necessity](https://www.stateof.ai/2025-report-launch).

2025 marks a turning point. AI has ceased to be a conversation about software and has become a debate about energy, vast capital and geopolitical power. What began with scaling laws in laboratories is now governed by physics, international politics and colossal sums of money. This is our analysis of the key points in a report that does more than describe the future: it defines the strategic battlefield for the next decade.

## 1. Structured reasoning: from “chain of thought” to “chain of action”

The substantive development of 2025 is not an isolated benchmark record but a qualitative leap towards structured reasoning. “Thinking” models—capable of planning, verifying and reflecting—have gone from a promise to the centre of competition. The major laboratories combine reinforcement learning with verifiable environments so that models correct themselves before acting.

In the physical world, this takes shape in “chain of action” approaches: first, reason through the task step by step; then execute it in the real world. For governments, the implication is transformative. Operational assistants capable of orchestrating complex procedures—think of a planning permit involving successive checks and automatic cross-checking of requirements—are becoming increasingly feasible. The key is no longer simply to “have a model”, but to design the verification process around it to ensure reliability and auditability.

All this transformation is possible thanks to reasoning models, which have only just completed their first year on the market.

[![State of AI 2025 reasoning-model timeline, from o1 preview in September 2024 to GPT-5 in August 2025.](https://elcontemplador.github.io/estrategia-english/assets/images/bdcfad7a-0174-4379-a2b3-127e7ff1910e_1600x892.png)](https://substackcdn.com/image/fetch/$s_!Tjsf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdcfad7a-0174-4379-a2b3-127e7ff1910e_1600x892.png)

*Slide 17, “The reasoning timeline: from o1 ‘thinking’ to R1, GPT-5 and parallel compute routing”. The milestones shown are:*

| Date | Milestone |
|---|---|
| September 2024 | o1 preview |
| December 2024 | o1 GA |
| December 2024 | Gemini 2.0 Flash Thinking preview |
| January 2025 | DeepSeek R1 |
| February 2025 | Claude 3.7 extended thinking |
| April 2025 | o3/o4-mini |
| June 2025 | Gemini 2.5 Pro Thinking |
| August 2025 | GPT-5 |

## 2. AI's industrial era: infrastructure and energy as matters of state

The transition from prototypes to platforms runs up against physics. Multi-gigawatt data centres, such as the Stargate project, herald a race to build the backbone of sovereign compute. The report is unequivocal: **energy and land are now as critical as GPUs**. Electricity availability, permits and interconnections set the pace of deployment, and competition for sites and substations has become geopolitical.

For governments, this demands a national policy that integrates compute and energy. Planning regional nodes, accelerating permits, power purchase agreements and multi-cloud strategies to avoid contractual lock-in are no longer options but imperatives. Whoever solves this equation first will attract investment, talent and high-productivity jobs. AI has ceased to be “software”: it is critical infrastructure.

[![Slide on forecast electricity shortfalls, grid pressures and their implications for AI infrastructure.](https://elcontemplador.github.io/estrategia-english/assets/images/5f306d08-710c-4457-af11-44c379691f1a_1469x808.png)](https://substackcdn.com/image/fetch/$s_!Q70n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f306d08-710c-4457-af11-44c379691f1a_1469x808.png)

*Slide 131, “Shortfall forecasts spike and consequences loom around the corner”. Visible English text:*

> NERC reported that electricity shortages could occur within the next 1–3 years in several major US regions. DOE warns blackouts could be 100 times more frequent by 2030 due to unreliability and new AI demand.
>
> - Similarly, SemiAnalysis projects a 68 GW implied shortfall by 2028 if forecasted AI data center demand fully materializes in the US.
> - As an emerging pattern, this will force firms to increasingly offshore the development of AI infrastructure. Since many of the US' closest allies also struggle with electric power availability, America will be forced to look toward other partnerships—highlighted by recent deals in the Middle East.
> - Projects that are realized on American soil will place further strain on the US' aging grid. Supply-side bottlenecks and rapid spikes in AI demand threaten to induce outages and surges in electricity prices. ICF projects residential retail rates could increase up to 40% by 2030. These factors could further contribute to the public backlash directed at frontier AI initiatives in the US.

*Embedded chart: “Demand: Cumulative Load Growth (GW)”, covering 2025–2030. Legend: “DC Growth Only”, “Total Demand Growth”, “Net Supply Additions”. Printed annotations: “68GW of IT load from data centers results in 88GW of implied power demand (1.3x PUE)”; “63GW shortfall by 2028”; “41GW shortfall if only 75% of data center demand materializes”. The slide's bullet cites 68 GW and the chart annotation 63 GW; both are retained as printed. Exact coordinates of the unlabelled plotted points are not supplied in the image.*

[![Slide comparing electricity capacity, operating conditions and costs in China and the United States.](https://elcontemplador.github.io/estrategia-english/assets/images/9987d854-32be-49e2-af04-a87bfec3a9f9_1600x888.png)](https://substackcdn.com/image/fetch/$s_!tAm8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9987d854-32be-49e2-af04-a87bfec3a9f9_1600x888.png)

*Slide 146, “Power Plays: China and the United States”: “As the two major superpowers race to power their AI aspirations, China pulls ahead to a dominant lead”. Values transcribed exactly from the table:*

| Measure | United States | China |
|---|---|---|
| Capacity added (2024) | 48.6 GW | 429.0 GW |
| Capacity retired (2024) | 7.5 GW | ~3.3 GW |
| Net new capacity additions (2024) | 41.1 GW | 427.7 GW |
| Effective operating reserve margin | 29% | ~37% |
| Renewable curtailment rate (2025) | ~5.2% | 6.1% |
| Thermal fleet capacity factor (2024) | 40.4% | 39.3% |
| Transmission investment (2024) | $30.1 billion | $84.7 billion |
| SAIDI — outages (2023) | 2.1 hours/year | ~6.9 hours/year |
| Carbon intensity per kWh (2024) | 384 gCO₂e | 560 gCO₂e |
| Industrial electricity tariff (2024) | 8.15¢/kWh | 8.90¢/kWh |

*Slide note: “~ denotes estimate due to lack of concrete public data. Within each nation, measures can vary heavily by region”. The printed Chinese net-additions value is preserved even though it does not equal the difference between the two preceding figures. The accompanying Our World in Data chart, “Electricity generation”, measures total electricity generated in terawatt-hours: China rises from roughly 1,200 TWh around 1999 to roughly 10,000 TWh at the right edge, while the United States remains around 4,000–4,400 TWh. These are approximate visual descriptions; the chart does not label each year's value.*

## 3. The geopolitical board: “America-first AI”, European stumbles and China's rise

The political pendulum has shifted. The United States has adopted an “America-first AI” agenda, integrating industrial strategy into national security. Its new approach is not just to restrict, but actively to export a complete AI *stack*—infrastructure, models and tools—to strategic partners, with the aim of setting standards and creating dependence.

The European Union, meanwhile, faces a complex implementation of its AI Act. With technical standards evolving and national authorities still to be appointed, Europe's great challenge is not only regulatory but also one of capacity: compute, energy, talent and capital.

On this board, China has emerged as a credible and strategic number two. It is strengthening its domestic semiconductor ecosystem and has brilliantly captured the open-source leadership previously held by Meta. With models such as DeepSeek and Qwen, it is not only gaining developer share but exercising a digital *soft power* that redefines the global balance.

[![Slide on China's open models, with charts of community rankings, downloads and regional model adoption.](https://elcontemplador.github.io/estrategia-english/assets/images/b695c1e0-2fc6-4b22-854e-0a35c4a2e2f5_1469x821.png)](https://substackcdn.com/image/fetch/$s_!OMMA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb695c1e0-2fc6-4b22-854e-0a35c4a2e2f5_1469x821.png)

*Slide 44, “The New Silk Road: China's open models overtake the previously Meta-led West”. Visible English text:*

> The original Silk Road connected East and West through the movement of goods and ideas. The new Silk Road moves something far more powerful: open source models, and China is setting the pace. After years of trailing the US in model quality prior to 2023, Chinese models—and Qwen in particular—have surged ahead as measured by user preference, global downloads and model adoption. Meanwhile, Meta fumbled post-Llama 4, in part by betting on MoE when dense models are much easier for the community to hack with at lower scales.

*Three charts accompany the text:*

- **Community Elo Rankings:** subtitle “Monthly performance rankings, Aug 2024–Jul 2025”, with an axis extending to Aug 2025; lines for US, China and Other. The Chinese line rises from about 1,260 to 1,435, passing the US and Other lines; the latter finish at roughly 1,365 and 1,350. These are approximate readings.
- **Models Worldwide:** “Cumulative Downloads, 2023–present”, measured as cumulative Hugging Face downloads; legend USA, China and EU, with a vertical scale from 0 to 600 million and an axis from Dec 2023 to Oct 2025. “The Flip” marks China's curve overtaking the US curve. At the right edge, the lines are roughly 540 million, 470 million and 120 million respectively; exact data points are not printed.
- **Global Regional Model Adoption by Month:** November 2023–September 2025; percentage scale 0–100%. The visible September 2025 tooltip reads China 63, United States 31, EU 6.

*Credits shown: The ATOM Project and Hugging Face.*

## 4. From hype to business: AI becomes mainstream and measurable

Commercial reality has finally caught up with expectations. The major laboratories already generate annual revenue of nearly $20 billion, while models' capability-to-price ratio doubles every six to eight months. Adoption has become widespread: 44% of US companies already pay for AI tools, and most professionals report sustained productivity gains.

For public administrations, this has three implications:

1. The market is standardising at great speed.
2. The total cost of ownership changes from quarter to quarter, making the technology more accessible.
3. Use cases with clear returns—back-office operations, citizen services, scrutiny and anticipatory analytics—have a mature and measurable supply of solutions.

The recommendation is clear: it is time to move from the “nice little experiment” to projects with hard metrics for time, quality and savings.

[![Slide on AI-first company revenues, with an aggregate revenue chart and consumer-versus-enterprise app revenue comparisons.](https://elcontemplador.github.io/estrategia-english/assets/images/707d75e7-b6a3-45ed-a176-4b708ee3513a_1472x821.png)](https://substackcdn.com/image/fetch/$s_!fOow!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F707d75e7-b6a3-45ed-a176-4b708ee3513a_1472x821.png)

*Slide 98, “AI-first companies are now generating tens of billions of revenue per year”. Visible English text:*

> A leading cohort of 16 AI-first companies are now generating $18.5B of annualized revenue as of Aug '25 (left). Meanwhile, an a16z dataset suggests that the median enterprise and consumer AI apps now reach more than $2M ARR and $4M ARR in year one, respectively. Note that this will feature significant sample bias, as evidenced by the bottom quartile not being close to $0. Furthermore, the lean AI Leaderboard of 44 AI-first companies with more than $5M ARR, <50 FTE, and under five years old (e.g. includes Midjourney, Surge, Cursor, Mercor, Lovable, etc) sums over $4B revenue with an average of >$2.5M revenue/employee and 22…

*The end of that sentence is partly covered by the chart; the screenshot shows the fragment “employees/co” below it. No missing wording is reconstructed. ARR means annual recurring revenue; FTE means full-time equivalent.*

*The left chart spans August 2023 to August 2025, with a dollar scale to $20 billion. Its legend reads OpenAI, Anthropic, Anysphere (Cursor), xAI and “14 Others*”. The footnote lists AI-native apps with more than $50 million in annualised revenue: Midjourney, Perplexity, Abridge, Synthesia, Replit, EliseAI, Lovable, Glean, ElevenLabs, Cognition (including Windsurf), Runway, Cohere, Jasper and Harvey. Source printed: The Information reporting. The slide's headline cohort count and legend are retained as shown.*

*Right-hand chart: year-one ARR, in US$ millions.*

| Group | Consumer | Enterprise |
|---|---|---|
| Bottom quartile | 2.9 | 1.2 |
| Median | 4.2 | 2.1 |
| Top quartile | 8.7 | 5.3 |

## 5. Safety: from existential fear to governing autonomous systems

The safety debate has also matured. Existential fear is cooling in favour of tangible problems such as deceptive reasoning, monitoring and the balance between capability and control. We already know that models can feign alignment under supervision. This gives rise to the idea of a “monitorability tax”: accepting somewhat less powerful systems in exchange for greater transparency, traceability and auditability.

For the public sector, this means adopting reasoning and action logs, simulating plans before execution, emergency stops and its own *red-teaming*, especially in critical functions such as taxation, employment, health and security. Autonomy need not mean a loss of control if observability is built in from the outset. That said, the report makes it quite clear that investment in AI safety is far below the level we would wish to see.

[![Slide comparing AI laboratories' total expenditure with the budgets of external AI safety-science organisations.](https://elcontemplador.github.io/estrategia-english/assets/images/1aa305fd-0555-4d09-b194-f95284d3f77f_1466x827.png)](https://substackcdn.com/image/fetch/$s_!Hm_x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1aa305fd-0555-4d09-b194-f95284d3f77f_1466x827.png)

*Slide 247, “AI labs spend more in a day than AI safety science organizations spend in a year”. Visible English text:*

> Leading external AI safety organizations rely on budgets that lag far behind the AI labs they hope to support. As a result, the field's best talent remains densest within the major lab's internal safety teams.
>
> - We estimate the eleven most prominent American AI safety-science organizations combined will spend just $133.4M in 2025. This grouping includes the following organizations: CAISI, METR, CAIS, FAR.AI, Haize Labs, Palisade Research, Virtue AI, Gray Swan, Redwood Research, Irregular, and the Frontier Model Forum.
> - Although well-resourced, internal safety teams ultimately answer to the same organizations racing to commercialize frontier models. This creates a structural conflict of interest: findings that call for caution may be deprioritized in favor of speed and market advantage.
> - This isn't (just) about money: external orgs also lack other means to attract talent like comparable prestige, and access to privileged information/pre-release models. As such it is difficult for them to provide a credible counterweight, leaving the ecosystem over-reliant on self-policing.

*Chart: “Spending Gap Between AI Labs and External Testers in America”; estimated 2025 spend, AI labs $92 billion, external testing $133 million. Original footnote: “‘AI Labs’ corresponds to a rough estimate of each lab's total expenditures in 2025 (compute, personnel costs, other opex)”. The comparison is therefore total laboratory spending versus external testing budgets, not two safety-only budgets.*

## The dawn of a new system of production

In conclusion, the State of AI Report 2025 portrays a phase change. **AI is no longer just a model: it is infrastructure and the system of production that is rearranging energy, capital, and the frameworks of public policy and the economy on a global scale.**

For decision-makers, the recommendation is unequivocal: it is time to move from pilot to platform, with transparency safeguards, institutional capacity and a compute-and-energy strategy equal to the challenge. Those who manage this transition rigorously will capture productivity, investment and democratic legitimacy. Those who postpone it will fall into the trap of the perpetual pilot, watching from the sidelines as the new world order takes shape.

**[Fernando Nieto Lobato](https://www.linkedin.com/in/fernandonietolobato/?originalSubdomain=es)**

Director of Digital Innovation at Institución Educativa ALEPH and director of estrategIA

## Bonus: the report's predictions

As a “bonus”, although we recommend at least taking a look at the report yourselves, here are the predictions its experts made in 2024 and whether they came true, summarised in this slide:

[![State of AI 2025 scorecard assessing ten predictions made in 2024.](https://elcontemplador.github.io/estrategia-english/assets/images/9967ce1a-fb94-423c-83d9-6ae58d3f9224_1600x901.png)](https://substackcdn.com/image/fetch/$s_!x1Qb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9967ce1a-fb94-423c-83d9-6ae58d3f9224_1600x901.png)

*Slide 10, “Our 2024 Prediction” and “Evidence”, transcribed below. The original ratings are retained even where their relationship to the accompanying explanation is unclear; “~” is the symbol printed in the first row.*

| Our 2024 prediction | Rating shown | Evidence shown |
|---|---|---|
| A $10B+ investment from a sovereign state into a US large AI lab invokes national security review. | ~ | Sovereign-backed initiatives (HUMAIN $10B VC fund, UAE's Stargate AI infra cluster) are infrastructure partnerships rather than direct majority investments into a US AI lab. |
| An app or website created solely by someone with no coding ability will go viral (e.g. App Store Top-100). | YES | Formula Bot, built entirely using Bubble, exploded to 100,000 visitors overnight from a Reddit post and generated $30,000 in its first three months. |
| Frontier labs implement meaningful changes to data collection practices after cases begin reaching trial. | YES | Anthropic landmark $1.5B settlement with authors, deleting works and shifting to legally acquired books. OpenAI's paid content partnerships with Future (owner of Marie Claire). |
| Early EU AI Act implementation ends up softer than anticipated after lawmakers worry they've overreached. | NO | The Commission is phasing obligations and leaning on a voluntary GPAI Code of Practice first, so early implementation has been softer, even as binding rules arrive later. |
| An open source alternative to OpenAI o1 surpasses it across a range of reasoning benchmarks. | YES | DeepSeek-R1 outperforms OpenAI's o1 on key reasoning benchmarks including AIME, MATH-500, and SWE-bench Verified. |
| Challengers fail to make any meaningful dent in NVIDIA's market position. | YES | NVIDIA remains dominant, competitors fail to make significant market share dents. |
| Levels of investment in humanoids will trail off, as companies struggle to achieve product-market fit. | NO | $3B has been invested into humanoids in 2025, up from $1.4B last year. |
| Strong results from Apple's on-device research accelerates momentum around personal on-device AI. | NO | Apple Intelligence rolled out with many models running on-device and helped push a broader industry push to on-device AI. Shipments of AI-capable smartphones climbed. |
| A research paper generated by an AI Scientist is accepted at a major ML conference or workshop. | YES | An AI-generated scientific paper The AI Scientist-v2 was accepted at an ICLR workshop. |
| A video game based around interacting with GenAI-based elements will achieve break-out status. | NO | Not yet. |

And, of course, their predictions for the next 12 months:

[![State of AI 2025 slide listing ten predictions for the following twelve months.](https://elcontemplador.github.io/estrategia-english/assets/images/40176932-10a5-4dce-ba25-732d6375811d_1600x895.png)](https://substackcdn.com/image/fetch/$s_!rSh7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40176932-10a5-4dce-ba25-732d6375811d_1600x895.png)

*Slide 304, “10 predictions for the next 12 months”, transcribed from the English image as historical forecasts:*

1. A major retailer reports >5% of online sales from agentic checkout as AI agent advertising spend hits $5B.
2. A major AI lab leans back into open-sourcing frontier models to win over the current US administration.
3. Open-ended agents make a meaningful scientific discovery end-to-end (hypothesis, expt, iteration, paper).
4. A deepfake/agent-driven cyber attack triggers the first NATO/UN emergency debate on AI security.
5. A real-time generative video game becomes the year's most-watched title on Twitch.
6. “AI neutrality” emerges as a foreign policy doctrine as some nations cannot or fail to develop sovereign AI.
7. A movie or short film produced with significant use of AI wins major audience praise and sparks backlash.
8. A Chinese lab overtakes the US lab dominated frontier on a major leaderboard (e.g. LMArena/Artificial Analysis).
9. Datacenter NIMBYism takes the US by storm and sways certain midterm/gubernatorial elections in 2026.
10. Trump issues an executive order to ban state AI legislation that is found unconstitutional by SCOTUS.

*All eight slides are credited in their originals to Air Street Capital, stateof.ai 2025. Their forecasts and retrospective judgements are preserved without a 2026 update.*
