# February, the month of AI “deep research”: introducing and testing three major tools

Author: Fernando Nieto Lobato
Original publication: 2025-03-05
Spanish original: https://estrategiabyaleph.substack.com/p/estrategia-75-como-usar-la-ia-para
English URL: https://elcontemplador.github.io/estrategia-english/essays/075/
Status: Published translation

This is a translation of the original Spanish essay published on 5 March 2025. Its claims, examples and forecasts retain that historical context.

English publication: 2026-09-29

*Archive note: this comparison was published on 5 March 2025. Model capabilities, prices, access limits and interface screenshots are preserved as described at that time. The evaluations below are historical AI-generated assessments, not current product recommendations.*

February 2025 marked a milestone in artificial intelligence with the launch of three deep-research tools: **OpenAI Deep Research**, **Perplexity Deep Research** and **xAI DeepSearch**. During the month, the race to transform how we explore and analyse information reached new heights. In this article, we introduce them and offer an initial evaluation of capabilities that we believe could be tremendously useful in politics and government.

The pioneer in launching a deep-research tool was [Google, with Deep Research on 11 December 2024](https://blog.google/products/gemini/google-gemini-deep-research/). We have not included it in this comparison because, although it **deserves considerable credit for being first, it is still powered by Gemini Pro 1.5, a previous-generation AI model, and it would be unfair to compare it here with deep-research tools using far more powerful and advanced reasoning models**. Google is very likely to update the model soon, however, [just as it has developed an even more complex and powerful model for scientific research, Co-scientist, which we discussed last week](https://open.substack.com/pub/estrategiabyaleph/p/estrategia-74-la-revolucion-cientifica?r=gp8ei&utm_campaign=post&utm_medium=web&showWelcomeOnShare=false).

The term “deep research” refers to AI models' ability to carry out advanced searches, analyse multiple sources and synthesise information into understandable reports, often using deep-learning techniques and advanced reasoning. These models, with the differences we will discuss, are especially useful for professional tasks such as financial analysis, scientific research and product development. We believe they could also be useful in academic and professional work in politics and government.

## Introducing the models

First, let us recap the models introduced in February 2025 and their main features.

### 1. OpenAI Deep Research: the power of the most advanced reasoning

[Launched on 2 February](https://openai.com/index/introducing-deep-research/), this tool, powered by the o3 model—the most powerful AI publicly accessible anywhere in the world today—stands out for its focus on logical reasoning and its ability to generate long, detailed reports. Aimed at professionals, it was initially available only to Pro users, at $200 a month. Fortunately, since 26 February, Plus subscribers, paying $20—or €20—plus tax a month, have been able to carry out ten investigations a month.

[![ChatGPT's historical deep-research option, with a tooltip showing ten uses available until 30 March; visible text is translated below.](https://elcontemplador.github.io/estrategia-english/assets/images/db22168a-35b1-44db-a453-db959117bc4e_744x303.png)](https://substackcdn.com/image/fetch/$s_!G1P1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb22168a-35b1-44db-a453-db959117bc4e_744x303.png)

*Screenshot text: “How can I help you?”; “Ask anything”; “Search”; “Deep research”. The tooltip reads: “10 available until 30 March. Get detailed insights on any topic.” Other fully visible controls include “Analyse images” and “More”; some shortcut labels are obscured by the tooltip.*

### 2. Perplexity Deep Research: accessibility and efficiency

[On 14 February, Perplexity introduced its advanced-research offering](https://www.perplexity.ai/es-es/hub/blog/introducing-perplexity-deep-research), based on the Chinese DeepSeek R1 model, an open-source transformer optimised for reasoning. Its main advantage over competitors is that it can be used free, albeit with usage limits.

[![Perplexity's historical model menu, including Deep Research, R1 and o3-mini; all substantive visible options are translated below.](https://elcontemplador.github.io/estrategia-english/assets/images/d2a49dc1-38dc-428a-97c2-027632e5194a_797x649.png)](https://substackcdn.com/image/fetch/$s_!4MC9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2a49dc1-38dc-428a-97c2-027632e5194a_797x649.png)

*Screenshot text: “What do you want to know?”; “Ask something…” The model menu offers:*

- **Auto:** best for everyday searches.
- **Pro Search:** three times more sources and detailed answers.
- **Deep research:** in-depth reports on complex topics.
- **Reasoning with R1:** new DeepSeek model hosted in the United States.
- **Reasoning with o3-mini:** OpenAI's latest reasoning model.

*The fully visible background headline reads “Amazon launches quantum chip”; other background cards are partly hidden by the menu.*

### 3. xAI DeepSearch: next-generation search

Just a few days later, [xAI launched DeepSearch](https://x.com/xai/status/1892400134178164775) on 18 February as part of Grok 3. Integrated with real-time data from X and the web, it is available to Premium+ subscribers at $30 a month **and, temporarily, to free accounts too**. We encourage you to try it while that remains the case.

[![Grok 3's English input box, with DeepSearch selected and Think beside it.](https://elcontemplador.github.io/estrategia-english/assets/images/1a0bf84e-8404-4037-ad26-4eb0860879c1_1560x276.png)](https://substackcdn.com/image/fetch/$s_!6xGO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a0bf84e-8404-4037-ad26-4eb0860879c1_1560x276.png)

*The original English interface reads “How can Grok help?”, with “DeepSearch” selected beside “Think”, and “Grok 3” as the model.*

### Comparison at a glance

[![Historical comparison of the three research tools, launch dates, underlying models and access terms; a complete English table follows.](https://elcontemplador.github.io/estrategia-english/assets/images/2702feb0-b550-4eb3-826e-95250c26d682_1463x732.png)](https://substackcdn.com/image/fetch/$s_!sTTV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2702feb0-b550-4eb3-826e-95250c26d682_1463x732.png)

*English translation of the complete historical comparison table:*

| Model | Launch date | Underlying model | Access |
| --- | --- | --- | --- |
| OpenAI Deep Research | 2 February 2025 | o3 | Pro subscription, $200/month; from 26 February 2025, Plus includes ten investigations a month |
| Perplexity Deep Research | 14 February 2025 | DeepSeek R1 | Free with usage limits; coming soon to mobile apps |
| xAI DeepSearch | 18 February 2025 | Grok 3 | Premium+ subscription on X, $30/month; currently also available to free accounts |

## Preliminary tests of these models

To give you more information about these deep-research tools—all interesting, but with enormous differences in their results that make them more or less suitable for different uses—**we carried out two preliminary but in-depth tests. First, we asked for a more “theoretical and academic” investigation comparing electoral systems in different countries. Second, we requested a more applied, practical investigation closely tied to current affairs: an analysis of Spain's current electoral polling picture.**

As I had previously read from prominent experts and scientists such as Ethan Mollick and Derya Unutmaz, considerable subject expertise is needed to judge the research produced by OpenAI's model because of its power and quality. To make this preliminary analysis **more impartial**, I took advantage of our decision not to evaluate Google's advanced-research model, for the reasons explained at the beginning, and used **its most capable reasoning model to help make the question we put to the three models as effective as possible and, especially, to evaluate the content obtained after using all three systems for the same task**.

Below we share **the question**, **links to each model's response**, which you can and should examine for yourselves—we consider that the best way to get a sense of these models—and **a summary table of Gemini 2.0 Flash Thinking Experimental's evaluation of the answers**.

We also invite you to consult [this 119-page document containing the full questions—the prompts we submitted—the responses collected from the three models, and the complete evaluation by Google's most advanced reasoning model](https://docs.google.com/document/d/1WQ_Th0-vHBqe_lWBoE_eSSkmeSbZKrPMpcm5eVw6Dxc/edit?usp=sharing).

### First question

**A comprehensive comparative analysis of electoral systems in diverse democracies: an in-depth study of the United Kingdom, Germany and Brazil.**

Remember that [the full request—the prompt—is quite long and detailed and is available in the document](https://docs.google.com/document/d/1WQ_Th0-vHBqe_lWBoE_eSSkmeSbZKrPMpcm5eVw6Dxc/edit?usp=sharing). The wording above is only its title.

**The models' responses:** click each name to access its original response, which we recommend reading.

- **[OpenAI Deep Research](https://chatgpt.com/share/67c18c2d-1b14-8011-b314-c8d5deb5d393):** one key detail, which we will return to, is the very different amount of time these tools take. OpenAI's system spent more than eleven minutes producing that extensive answer. Although the others search for sources and “think”, they do not generally take more than a minute…
- **[Perplexity Deep Research](https://www.perplexity.ai/search/analisis-comparativo-exhaustiv-8EPhNrlBQeOsB9mTDcLfIw)**
- **[xAI DeepSearch](https://grok.com/share/bGVnYWN5_12ec1ce0-e16f-4c09-85e4-9d51c57f0698)**

Here is the summary table of **the evaluation that Google's most advanced reasoning model made of the content produced by the different investigations**:

[![Gemini's historical assessment of the electoral-systems research task across nine comparison rows; a complete English table follows.](https://elcontemplador.github.io/estrategia-english/assets/images/947f6bff-9b53-4750-b72c-0c782b935614_1471x603.jpeg)](https://substackcdn.com/image/fetch/$s_!uoTw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F947f6bff-9b53-4750-b72c-0c782b935614_1471x603.jpeg)

*Complete translation of the historical model-generated “Summary comparison table”:*

| Feature | OpenAI Deep Research | Perplexity | Grok Deep Search (detailed note) |
| --- | --- | --- | --- |
| Depth and technical detail | Very high | Medium–high | Medium |
| Academic rigour and evidence | High | Medium–high | Low |
| Structure and organisation | Excellent | Very good | Good |
| Explicit comparative analysis | Implicit, but present | Good | Less explicit |
| Integration of academic studies | Separate, referenced section | Briefly integrated, summarised | Very weak, informal |
| Comprehensiveness of sources | Very high | Medium | Low |
| Clarity and comprehensibility | Very good | Very good | Good |
| Academic format | Less formal, bracketed citations | More formal, numerical citations | Less formal, mentions |
| Overall ranking | 1st | 2nd | 3rd |

### Second question

**Electoral projection for Spain: an analysis of polling trends and political scenarios for March 2025.**

- **[OpenAI Deep Research](https://chatgpt.com/share/67c19875-37b4-8011-aa10-41360767512c)**
- **[Perplexity Deep Research](https://www.perplexity.ai/search/proyeccion-electoral-en-espana-iqOx4fmYSGqjlkiyrPEGhw)**
- **[xAI DeepSearch](https://grok.com/share/bGVnYWN5_c181bdb0-ef89-4677-a26f-400e566610ac)**

This was Gemini 2.0 Flash Thinking Experimental's summary evaluation of the responses to the second question:

[![Gemini's historical assessment of the Spanish electoral-projection task across seven comparison rows; a complete English table follows.](https://elcontemplador.github.io/estrategia-english/assets/images/052659ab-65e2-49ab-8f4b-a3943b170907_1459x496.jpeg)](https://substackcdn.com/image/fetch/$s_!xKJj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052659ab-65e2-49ab-8f4b-a3943b170907_1459x496.jpeg)

*Complete translation of the historical model-generated “Summary comparison table”:*

| Feature | OpenAI Deep Research | Perplexity | Grok Deep Search (detailed note) |
| --- | --- | --- | --- |
| Data collection and presentation | Extremely detailed and extensive | Less detailed, summarised | Minimal, superficial |
| Analysis of trends | Very deep and granular | Concise and effective | Superficial and generalised |
| Scenario projection | Very detailed and robust | Probabilistic and concise | Oversimplified and basic |
| Reasoning and justification | Very strong and transparent | Less detailed | Weak and vague |
| Academic rigour and detail | Excellent | Good | Low |
| Structure and format | Exemplary | Well structured | Reasonably structured |
| Overall ranking | 1st | 2nd | 3rd |

**This was Gemini's main conclusion after evaluating the answers in this preliminary test of deep-research models:**

> “OpenAI Deep Research, in this comparison, demonstrates an impressive ability to act as a powerful research and analysis tool, approaching the quality of professional political science analysis.”

*Translation note: this quotation and the model-generated evaluation tables are translated from the Spanish text preserved in the original article and screenshots.*

## Personal impressions

First, I should say that **I broadly agree with Gemini's conclusions**. We also informally shared the response with a couple of human experts in these fields so they could take a look, and their conclusions were similar to those offered by Google's reasoning model. Although they felt that some aspect was missing or that a source might, at first glance, have been better or more up to date, they were quite impressed by the power of OpenAI's model. Let us remember that, like all current AI, it is “the worst technology we will use for the rest of our lives”, and that it is improving very quickly.

Clearly, **OpenAI's Deep Research is a very powerful research tool. The problem for the other two is that they are a long way from that level, first because they do not use a reasoning model as powerful as o3, by far the world's most capable publicly accessible AI right now, and second because they are not designed in the same way. The fact that OpenAI's model spends more than ten minutes on each request, using all that computing capacity and o3's power, makes it very costly—hence the limit of just ten monthly requests for Plus users, and we are fortunate it has finally been opened to them at all. It cannot be compared with free use of Perplexity's model**. Even so, the latter works fairly well for its price, although personally I have not seen a very large improvement over Perplexity's normal search system, which I have been using as my regular AI search engine for many months. I have a paid account.

On the other hand—and here I should stress that these are two tests with an in-depth approach and methodology but very preliminary scope, since results could probably differ with another prompt or other kinds of sources—**I want to point out that, in my personal use on technology and computing topics, drawing on English-language sources such as specialist websites, Reddit and forums, xAI DeepSearch has been very useful and done an excellent job, despite the results of this small preliminary test. It is entirely possible that, although not geared towards academic research or research of the same depth, it could be very useful in everyday use, particularly if it remains free, drawing on other, less academic and multilingual sources**. But that is only an impression, and the reason I preferred an impartial evaluator such as an AI for these initial tests. I also know that just two tasks form a very small sample, offering only a glimpse that may be heavily influenced by the prompt, the languages or the geographical area in which we ask it to work.

[In his live broadcast last Friday](https://www.youtube.com/live/BYJhIFEsC_E?si=gO7loFl4CfZje47T), AI communicator **Patricio Fernández** spoke enthusiastically about OpenAI Deep Research and stressed a point that also emerges clearly from our preliminary analysis: **an attempt has been made, probably for commercial reasons, to give the “same name” to products that are not particularly alike in their capabilities or operation, and especially in their design and purpose**.

Before closing, I want to **briefly mention that there are other tools besides these three, including some open-source ones, with a similar approach**. They include **[Storm](https://storm.genie.stanford.edu/)**, [the tool developed by Stanford University that we discussed in issue 67](https://open.substack.com/pub/estrategiabyaleph/p/estrategia-67-predecir-para-influir?r=2tyybn&utm_campaign=post&utm_medium=web&showWelcomeOnShare=false), and **[Elicit](https://elicit.com/)**, which focuses entirely on academia and research papers. We have not yet been able to try Elicit, but our guest author in issue 58, [Fernando Domínguez Sardou](https://www.linkedin.com/in/fernando-dominguez-sardou/), has spoken very highly of it. And just hours before this issue of estrategIA was published, news emerged, as expected, that [Google is already working on a substantially improved version of its Deep Research](https://x.com/testingcatalog/status/1897010217004802418?s=67&t=zsdeJJXua1aBic7ko6En1A).

To sum up, despite the considerable limitations of this preliminary analysis, which also deals with three quite different products, I must convey **my considerable astonishment at the capabilities of OpenAI's model**. Even there, however, attention is needed: possible hallucinations should be reviewed, sources checked, and perhaps another model asked to improve the writing of certain passages. **My impression is that Perplexity's and xAI's models, while a long way behind OpenAI's, could be useful in many everyday situations**. Undoubtedly, **the best thing is for you to try these increasingly powerful and sophisticated “helpers” yourselves, and judge whether they could be useful to you already, or in the near future as their capabilities are refined further**.

Fernando Nieto Lobato

*Director of Digital Innovation at the ALEPH Educational Institution*
