English edition note: this article describes the September 2024 launch. Model availability, prices, capabilities and expectations are preserved in their historical context.

OpenAI has once again redefined the boundaries of artificial intelligence with last week’s launch of its latest model series, OpenAI o1. This revolutionary system represents a significant advance in machines’ reasoning and problem-solving capabilities.

The o1 series introduces an innovative approach to AI processing, designed to imitate human reasoning more closely than ever. Unlike previous models, which focused on rapid pattern recognition and generating responses, o1 has been designed to “think” more carefully before answering. This deliberate approach allows the model to tackle complex problems in areas such as science, programming and mathematics with unprecedented precision and depth.

Key features of o1

  • Improved reasoning: o1 employs a sophisticated “chain of thought” process, breaking complex problems down into manageable steps and iteratively refining its approach.
  • Better problem-solving: the model excels in fields requiring deep analytical thought, such as physics, chemistry, biology and advanced mathematics.
  • Code generation and debugging: o1 demonstrates superior abilities in generating and debugging complex code, making it a powerful tool for developers.
  • Safety-focused design: OpenAI has introduced new approaches to safety training that use o1’s reasoning capabilities to adhere more closely to ethical guidelines and safety rules. By reasoning more effectively about OpenAI’s safety policies, this model is less at risk of being used for malicious purposes despite its greater capabilities.

The o1-preview and o1-mini variants are already available to all paying GPT-4 customers, while the original model remains solely in OpenAI’s hands. It is rumoured that the company is using it to train the future “GPT-5”.

o1-preview: the full version, designed to tackle the most complex reasoning tasks across different domains.

o1-mini: a smaller, faster and more cost-effective version, specifically optimised for programming tasks. It offers 80% of the performance at 20% of the cost compared with o1-preview.

Benchmark performance

Benchmark chart comparing GPT-4o, o1-preview and o1 on mathematics, coding and PhD-level science questions.

Accessible transcription of the original chart. The numbers below are the values printed above the bars; no additional values are inferred from the shaded portions.

Scroll across the table to read all columns.

Test Measure GPT-4o o1-preview o1 Expert human
Competition maths: AIME 2024 Accuracy (%) 13.4 56.7 83.3 —
Competition coding: Codeforces Percentile 11.0 62.0 89.0 —
PhD-level science questions: GPQA Diamond Accuracy (%) 56.1 78.3 78.0 69.7

An em dash means that no bar is shown for that category.

The o1 model — the one currently available only internally at OpenAI — has demonstrated notable performance across a wide variety of tests, highlighting its advanced reasoning capabilities:

  • International Mathematical Olympiad qualifying examination: o1 achieved an impressive 83% accuracy, compared with GPT-4o’s 13%.
  • Codeforces competitions: o1 placed in the 89th percentile, demonstrating exceptional programming skills.
  • American Invitational Mathematics Examination: o1 ranked among the top 500 students in the US.
  • GPQA, a PhD-level physics, biology and chemistry examination: o1 exceeded the accuracy of humans with PhDs in these scientific domains.

Beyond the benchmark performance figures supplied by OpenAI itself — bearing in mind that many benchmarks are saturated and can no longer properly measure some models’ “almost superhuman” capabilities — the internet has filled with interesting examples from people using o1-preview. Remember that, in evaluations, the full o1 model appears far superior to this version, which is the one already available to us. Here is a small selection of links if you would like to see some of the new model’s incredible capabilities:

Applications and use cases

We should point out that for many of our “everyday” uses of AI — for example, working on this newsletter — this is not the appropriate model. As a very new model, it is still strongly focused on particular tasks where its reasoning ability sets it apart. We can expect these capabilities to be integrated before long with multimodality, larger context windows and even internet access for the model. On the other hand, o1’s enhanced reasoning capabilities open up a wide range of distinctive potential applications across industries, particularly scientific research, software development, financial analysis and healthcare.

As mentioned above, despite its impressive capabilities, o1 — and even more so the preview and mini versions available to us today — has limitations, including these two:

Functional gaps: as an early model, o1 lacks some features found in other AI systems, such as web browsing and image processing. At present, it is a text-only model.

Cost and speed: o1 is significantly more expensive to use today than previous models, with higher input and output costs. It can also be slower when processing complex queries.

For anyone wanting to learn more about the model’s capabilities and limitations, we recommend this excellent summary of an AMA — “Ask Me Anything” — with o1’s development team.

Conclusion: the future of AI reasoning

If I may offer a personal assessment, based on numerous comments from industry experts and even people working on the project itself, o1’s significance lies less in the model itself than in this new approach. According to OpenAI, it began to take shape in practice barely a year ago, and it is achieving very rapid, continuous improvements in models. It also gives us another dimension along which to improve beyond the sheer compute used in training: in this case, the inference time devoted to responses and the way the model “reasons”. This will make a difference, and we are probably at the beginning of another enormous leap in AI, whose results we will see over the coming months.

The introduction of o1 marks a significant milestone in the evolution of AI systems. By incorporating reasoning processes that more closely resemble human ones, o1 represents an important step towards artificial general intelligence.

Although o1 is still in its early stages and faces some limitations, its potential impact across industries and scientific disciplines could be immense. As OpenAI continues to refine and expand the o1 series, we are likely to see even more impressive advances in AI reasoning and problem-solving in the near future.

Important note: the new model seems to favour very different kinds of “instructions”. We address this further down in the same issue, in the prompts for GPT-4 section. We recommend reading that section and keeping this firmly in mind if you want to get the most out of o1-preview.

Archive note: the separate prompts section referred to above is available in the original Spanish newsletter; this edition translates the main article.

Fernando Nieto Lobato

Director of Digital Innovation at the ALEPH Educational Institution

This is a translation of the original Spanish essay published on 18 September 2024. Its claims, examples and forecasts retain that historical context. Read the original Spanish edition, including its accompanying illustrations.

Cite this essay

Fernando Nieto Lobato. “OpenAI o1: a new frontier in AI “thinking” and problem-solving.” estrategIA, issue 051, 18 September 2024. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/051/

Back to the top ↑