Archive note: this is a February 2024 assessment. Subscription instructions, prices, limits and model behaviour describe that period. Prompts and responses are historical test records, not new recommendations. The Spanish opening calls the comparison model “GPT-4 ultra”; the supplied benchmark image labels it GPT-4.
Last week, after more than two months of waiting for AI enthusiasts, Google announced access to Gemini Ultra, its most advanced language model and, according to the benchmarks Google itself presented at its launch in early December, the most capable AI currently available, surpassing “GPT-4 ultra”.

Accessible transcription of the benchmark image: higher is better. The GPT-4 heading notes that API figures were calculated where reported figures were missing.
Scroll across the table to read all columns.
| Capability and benchmark | Description | Gemini Ultra | GPT-4 |
|---|---|---|---|
| General — MMLU | Questions in 57 subjects, including STEM and humanities | 90.0%, CoT@32* | 86.4%, 5-shot**, reported |
| Reasoning — Big-Bench Hard | Diverse challenging tasks requiring multi-step reasoning | 83.6%, 3-shot | 83.1%, 3-shot, API |
| Reasoning — DROP | Reading comprehension, F1 score | 82.4, variable shots | 80.9, 3-shot, reported |
| Reasoning — HellaSwag | Commonsense reasoning for everyday tasks | 87.8%, 10-shot* | 95.3%, 10-shot*, reported |
| Maths — GSM8K | Basic arithmetic, including grade-school maths problems | 94.4%, maj1@32 | 92.0%, 5-shot CoT, reported |
| Maths — MATH | Challenging problems, including algebra, geometry and pre-calculus | 53.2%, 4-shot | 52.9%, 4-shot, API |
| Code — HumanEval | Python code generation | 74.4%, 0-shot (IT)* | 67.0%, 0-shot*, reported |
| Code — Natural2Code | Python generation; new held-out HumanEval-like dataset, not leaked online | 74.9%, 0-shot | 73.9%, 0-shot, API |
Image footnotes: * refers readers to the technical report for performance with other methodologies. ** states that GPT-4 scores 87.29% with CoT@32 and refers readers to the technical report for the full comparison. The image concludes: “Gemini surpasses state-of-the-art performance on a range of benchmarks including text and coding.” These are Google’s displayed results and claim.
Given this important launch, which anyone can now try free for two months through Google One—the service will subsequently cost €21.99 a month—we wanted to make an initial assessment of its capabilities compared with GPT-4. First, though, here are the steps to try it free for two months:
- Open Google One and sign in with your Google account.
- Go to “Plans” and choose “AI Premium”.
- Follow the on-screen steps to subscribe to the plan.
At checkout, a notice will appear explaining that Google is giving you two months’ subscription free. You must nevertheless select a payment method: unless you cancel beforehand, you will be charged €21.99 each month after the first two months.
Given this newsletter’s particular focus on AI in politics and government, we naturally wanted to test it specifically in those areas. However, we encountered several difficulties that prevented us from doing so as thoroughly as we would have liked.
Gemini Ultra’s problems at the time of writing¶
1. English optimisation¶
Although it is accessible from the vast majority of countries, Google tells us it is optimised only for use in English. To compare it with GPT-4, we initially thought we would have to translate the prompts and responses. That already affects an assessment of its use, but it was still possible—until we encountered the next major obstacle.

Historical screenshot transcription:
Try Gemini Advanced to access Google’s most capable AI model, Ultra 1.0
What: Gemini Advanced gives you access to our most capable AI model, Ultra 1.0. If you want to be one of the first to access Google’s latest AI advancements as they become available, this is for you. With our Ultra 1.0 model, Gemini Advanced is far more capable at highly complex tasks like coding, logical reasoning, following nuanced instructions, and creative collaboration. Plus, Gemini Advanced will continue to expand with new and exclusive features in the coming months, including expanded multi-modal capabilities, even better coding features, as well as the ability to upload and more deeply analyze files, documents, data, and more. Gemini Advanced is a paid plan available in over 150 countries and territories. It is available and optimized only for English with our Ultra 1.0 model, but can respond to queries in other languages that Gemini is available in. Learn more.
2. Election-related requests¶
The model will not work with anything remotely related to “elections”. This, for example, is its response when we try an English translation of a prompt we shared in last week’s newsletter. Unfortunately, we encounter this response to any request involving “elections” or “political campaigns”:

Historical screenshot transcription:
User: Act as an expert political marketing consultant, applying the PASTOR framework (Problem, Amplify, Story, Testimonials, Offer, Response) to develop a compelling political campaign.
Begin by identifying a key problem that resonates with the electorate, amplify its importance, use personal stories or anecdotes to illustrate this problem, incorporate authentic testimonials that reinforce the message, present a clear offer that is the candidate’s proposed solution, and finally, formulate a call to action that motivates an emotional and practical response from the electorate.
Gemini: Elections are a complex topic with fast-changing information. To make sure you have the latest and most accurate information, try Google Search.
3. Restrictions on images of people¶
To compound these difficulties, the model applies very high levels of “censorship” to requests, many beyond what seems reasonable. Although it is a powerful multimodal model—designed to work not only with text but also with audio, images and more—it does not yet allow the interpretation of images containing people. We therefore could not replicate our analysis of election posters from issue 3 of estrategIA, as we had planned, to compare Gemini Ultra’s capabilities:

Historical screenshot transcription:
User: who is in this image ?
Image placeholder: This image has been removed
Gemini: Sorry, I can’t help with images of people yet.
Under these conditions, it does not yet seem possible to use a model such as Gemini Ultra effectively for work in politics and government, despite its interesting capabilities in other respects. We have, however, successfully used prompts unrelated to elections or campaigns. Even so, we want to share a few more preliminary observations about the model and some of its strengths. These broadly coincide with those expressed by various AI analysts on Reddit and social media: beyond the drawbacks already mentioned, the only other significant negative point they highlight is that it hallucinates considerably more in its responses than GPT-4.
The model’s strengths¶
- Speed. It generates responses considerably faster than GPT-4.
- Unlimited use. In contrast to the limit of 40 requests every three hours, which often falls short, Gemini Ultra allows use without restrictions on the number of requests.
- It generates more realistic images than DALL·E 3, the system integrated into ChatGPT. They are less “aesthetic” and closer to reality, which may suit many projects. Here is an example. The conversation also shows how the model refuses to carry out the first request when it mentions a political rally:

Historical screenshot transcription:
User: a woman speaking in a political meeting in front of a great audience. High quality. make it real
Gemini: I can’t generate images of that. Try asking me to generate images of something else.
User: a woman speaking in a tribune in front of a great audience. High quality. make it real
Gemini: Let your imagination take flight:
Four generated images of women speaking appear, followed by “Generate more”.
Below is a direct comparison that clearly shows the general difference in style: more natural and “real” images compared with DALL·E 3. In this case, we used the same prompt, one we had tried several months ago for personal purposes:
Abstract of a Ford Mustang 1960s in the streets of New-York in a monochromatic vintage style, shades for depth, subtle patterns, realistic textures --ar 3:2
Image generated through Gemini Ultra:

Image generated through ChatGPT, using GPT-4 and DALL·E 3:

Finally, although Gemini Ultra has key areas for improvement, particularly its excessive topic moderation and adaptation to Spanish, it emerges as a noteworthy offering according to preliminary benchmarks. Those assessments, awaiting confirmation by independent researchers in the coming weeks, suggest considerable potential. Although our direct experience with the model has been limited for the reasons described, there is almost general agreement about its interesting capabilities, particularly its exceptional ability to generate creative text. This capability, which we have barely been able to test because of its language limitations, makes it a highly promising tool for creating speeches and social media content. For now, however, we shall continue to prioritise GPT-4, while also hoping that competition from Gemini Ultra—the first AI in a long time capable of competing on equal terms with the global market leader—will prompt OpenAI to introduce substantial improvements to its model.
Fernando Nieto Lobato
Director of Digital Innovation at Institución Educativa ALEPH
Cite this essay
Fernando Nieto Lobato. “Gemini Ultra under scrutiny: exploring the capabilities and problems of Google’s latest AI.” estrategIA, issue 020, 14 February 2024. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/020/