Original editorial note: As an experiment, using this issue’s recommended “AI tool of the week”, Openvoice2, we have used a free, open-source artificial intelligence system — described in more detail in that section — to generate a narration of this issue’s main article and news.
English edition note: this translates the article published on 15 May 2024. Its references to availability, prices and forthcoming events describe that historical moment. The preserved source does not include a playable narration. Quoted wording is translated from Spanish unless already present in English in the images.
In barely a month, between mid-May and mid-June, we see the year’s greatest concentration of technology conferences. Yesterday, 14 May, was Google I/O, upstaged on the 13th by a brief but spectacular OpenAI conference and demonstration. Microsoft Build will take place from 21 to 23 May, and Apple’s Worldwide Developers Conference from 10 to 14 June. This year, artificial intelligence is the major — almost the only — theme at all of them. They will show us not just what is already here, but also what the immediate future might look like.
At estrategIA, we will keep you informed about everything coming from these technology giants which, alongside very few other companies — Meta, Tesla, Nvidia, Mistral… — set the pace for a technology as disruptive as AI. But let us begin with the first: OpenAI’s conference and demonstration last Monday. It has probably shown the most innovative capabilities and brought the biggest gift: GPT-4o FREE for all users.

Text in the announcement: “Introducing GPT-4o” (Spanish heading). “You can now try our newest model, GPT-4o. It’s faster than GPT-4, better at understanding images, and speaks more languages.” Button: “Try it now” (translated from Spanish).
At a 25-minute event held last Monday, 13 May, to upstage Google, OpenAI changed the artificial intelligence landscape with the launch of GPT-4o, in which the “o” stands for “Omni”. This model significantly surpasses its predecessor, GPT-4, demonstrating a 60-point Elo improvement on the LMSys benchmark and positioning itself as the world’s most powerful model, ahead of Gemini 1.5 Pro, Claude 3 and Llama 3-70B.

Accessible transcription of the original “Text Evaluation” chart. These are the historical figures shown in the source image; N/A is retained where a result is unavailable.
Scroll across the table to read all columns.
| Benchmark | GPT-4o | GPT-4T | GPT-4, initial release 23-03-14 | Claude 3 Opus | Gemini Pro 1.5 | Gemini Ultra 1.0 | Llama3 400b |
|---|---|---|---|---|---|---|---|
| MMLU (%) | 88.7 | 86.5 | 86.4 | 86.8 | 81.9 | 83.7 | 86.1 |
| GPQA (%) | 53.6 | 48.0 | 35.7 | 50.4 | N/A | N/A | 48.0 |
| MATH (%) | 76.6 | 72.6 | 42.5 | 60.1 | 58.5 | 53.2 | 57.8 |
| HumanEval (%) | 90.2 | 87.1 | 67.0 | 84.9 | 71.9 | 74.4 | 84.1 |
| MGSM (%) | 90.5 | 88.5 | 74.5 | 90.7 | 88.7 | 79.0 | N/A |
| DROP (F1) | 83.4 | 86.0 | 80.9 | 83.1 | 78.9 | 82.4 | 83.5 |
The most surprising and relevant aspect of GPT-4o is that it will be free for users, marking a radical shift in the sector’s business dynamics and finally putting a fairly advanced model within the general public’s reach. Previously, OpenAI had launched GPT-4 as a $20-a-month subscription, a decision that limited adoption. Now, thanks to efficiency improvements, OpenAI can offer GPT-4o at no cost, challenging competitors that charge for inferior models.
GPT-4o is not only superior at text processing. Perhaps its most significant feature, apart from being available for free, is that it is a multimodal model, able to handle text, audio, voice and images simultaneously. This brings it closer to the vision of an AI like the one in the film Her — something made very clear in the live demonstration — with advanced voice and video capabilities, displaying emotional responses and almost human behaviour in real time.
On OpenAI’s website, alongside a detailed account of all the new GPT-4o’s features, the company has added several very interesting videos and images showing capabilities different from those in the brief live technical demonstration. These are fascinating for their potential applications in politics too: image consultancy, generating text in images, brand placement…
We invite you to learn more about the model on OpenAI’s website and, above all, to try it, since it is already available today to free users too at https://chatgpt.com.
Google I/O¶
That happened on Monday and had a very considerable impact on the worldwide community of AI enthusiasts: both because the model is free, allowing anyone to get a better sense of AI’s current capabilities, and because of its multimodality, voice use and potential as a “work colleague”. Yesterday, Tuesday, Google held its traditional annual developer conference, Google I/O, and tried to respond with almost two hours of presentations focused entirely on integrating artificial intelligence into all its tools and developing new models.
One of the most notable announcements was Project Astra, a new Gemini-powered AI agent displaying remarkable spatial understanding and memory. In a live demonstration, Project Astra was able to remember where an employee had left their glasses in DeepMind’s London office, pointing out their location beside an apple on the desk. This offered a simple example of why it could be useful for multimodal models to have memory to help us with multiple tasks.
Google also presented Gemini 1.5 Pro, an advanced AI model offering a one-million-token context window, opening up entirely new possibilities for developers. Pichai announced that Gemini 1.5 Pro is being rolled out to all developers worldwide, and that the company is expanding it to an even longer context window of two million tokens.
Google also revealed updates to Google Photos, including Gemini-powered Ask Photos, which can summarise photographic memories and extract information from them. Users can ask questions such as “What is my car’s registration number?”, and Google Photos will recognise a frequently appearing car and provide that number.
On the Android front, Android 15 will be updated with “AI at its core”, according to Sameer Samat, president of Google’s Android ecosystem. The main improvements include AI-powered search at your fingertips, Gemini as Android’s new AI assistant, and on-device AI to unlock new experiences.
Google also showed Veo, its new AI video generator, which will compete with OpenAI’s Sora, and Imagen 3, the latest version of its AI image generator. In music, Google presented Lryia, its AI music generator.
Archive note: “Lryia” reproduces the spelling in the Spanish source.
Perhaps the most significant immediate transformation, although its initial launch is for the US only, is the change to the Google search engine we all routinely use. It will integrate AI into its answers by default, changing a model established for more than 25 years and very substantially affecting traffic to external websites, as well as search rankings and search advertising.
Overall, Google I/O 2024 underlined Google’s commitment to integrating AI into all its key products and services, to the point that the conference closed with a joke about how many times the words “artificial intelligence” appeared in the presentation script: 121.
Despite that whole display, and the presence on stage of Demis Hassabis, founder of DeepMind and leader of what is probably today’s most powerful scientific AI, Alpha Fold 3, revealed a few days ago, the AI enthusiast community has not been particularly convinced by Google’s offering. Many of the developments presented are still prototypes, and almost all the new features will not be rolled out immediately. The feeling remains that OpenAI is a year, or at least six months, ahead of Google in AI — and in artificial intelligence, that is an eternity.
Even so, the innovations seen over these two days in both presentations are beginning to show how AI will transform many aspects of our everyday experience.
Cite this essay
Gabriela Ortega. “Towards ubiquitous AI: the OpenAI and Google developments that will transform the future.” estrategIA, issue 033, 15 May 2024. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/033/