Archive note: this article preserves the launch context of 26 November 2025. Benchmark scores, product features and explanations of model behaviour are historical claims from the source, not current assessments. Original English quotations visible in the Karpathy screenshot are transcribed below; other quotations are translated from the Spanish article.
Artificial intelligence is advancing in leaps and bounds, and Google has just launched Gemini 3 Pro: an excellent model, far ahead of its competitors, which seems to overturn the theory that improvements are approaching a plateau. But, as we will explain, even the most sophisticated models can surprise us with curious, unexpectedly human errors if we do not use them correctly.
An imposing debut, with landmark results¶
Gemini 3 Pro has arrived on the market setting new standards in practically every relevant industry benchmark. With a score of 1,501 Elo on LMArena at the time of writing, it comfortably surpasses its predecessor Gemini 2.5 Pro's 1,451 and takes an undisputed lead in the Artificial Analysis Intelligence Index. The model is particularly strong in scientific reasoning, achieving an impressive 91.9% on GPQA Diamond and a historic 37.5% on Humanity’s Last Exam, more than ten percentage points ahead of the previous best model.

All values displayed in the original Artificial Analysis Intelligence Index chart:
Scroll across the table to read all columns.
| Model | Score |
|---|---|
| Gemini 3 Pro Preview | 73 |
| GPT-5.1 (high) | 70 |
| GPT-5 Codex (high) | 68 |
| Kimi K2 Thinking | 67 |
| Grok 4 | 65 |
| Claude 4.5 Sonnet | 63 |
| MiniMax-M2 | 61 |
| gpt-oss-120B (high) | 61 |
| Grok 4 Fast | 60 |
| Gemini 2.5 Pro | 60 |
| Qwen3 235B A22B 2507 | 57 |
| DeepSeek V3.2 Exp | 57 |
| GLM-4.6 | 56 |
| Claude 4.5 Haiku | 55 |
| Gemini 2.5 Flash (Sep) | 54 |
| Magistral Medium 1.2 | 52 |
| DeepSeek R1 0528 | 52 |
| Apriel-v1.5-15B-Thinker | 52 |
| Kimi K2 0905 | 50 |
| Llama Nemotron Super 49B V1.5 | 45 |
| Llama 4 Maverick | 36 |
The chart states that index version 3.0 combines ten evaluations: MMLU-Pro, GPQA Diamond, Humanity’s Last Exam, LiveCodeBench, SciCode, AIME 2025, IFBench, AA-LCR, Terminal-Bench Hard and τ²-Bench Telecom. A curved arrow highlights the rise from Gemini 2.5 Pro, at 60, to Gemini 3 Pro Preview, at 73.

Chart: “AI performance on a set of Ph.D.-level science questions”. The vertical axis is GPQA Diamond accuracy, labelled from 10% to 90%; the horizontal axis is release date, with labels from March 2023 to September 2025 and later points beyond the final label. It displays 135 results with uncertainty bars, colour-coded for OpenAI, Anthropic, Google, Meta AI and other organisations. Dashed reference lines mark random guessing at 25% and expert human performance just below 70%. Named points include GPT-4 (March 2023), GPT-4 Turbo Preview (November 2023), Claude 3 Opus, Claude 3.5 Sonnet (June 2024), o1-mini (high), o1 (high), DeepSeek-R1, Claude 3.7 Sonnet (64k thinking), Gemini 2.5 Pro Exp (March 2025), Grok 4, GPT-5.1 (high) and Gemini 3 Pro Preview. The final named Gemini point is above 90%; most individual points have no printed numerical value, so exact scores are not inferred from their positions.
Gemini 3 Pro's multimodal capability is another strength. With 81% on MMMU-Pro and 87.6% on Video-MMMU, it demonstrates exceptional understanding not only of text, but also of images, video and audio. This versatility comes with a massive context window of one million tokens, approximately 750,000 words. It can process lengthy documents, complete code repositories or transcripts of videos lasting more than an hour in a single session — something very few models can do today.

Complete transcription of the original benchmark table. Dashes mean no value is provided in the image; “not supported” is its own label. These are the source's reported results.
Scroll across the table to read all columns.
| Benchmark and measurement | Gemini 3 Pro | Gemini 2.5 Pro | Claude Sonnet 4.5 | GPT-5.1 |
|---|---|---|---|---|
| Humanity’s Last Exam — academic reasoning, no tools | 37.5% | 21.6% | 13.7% | 26.5% |
| Humanity’s Last Exam — with search and code execution | 45.8% | — | — | — |
| ARC-AGI-2 — visual reasoning puzzles, ARC Prize Verified | 31.1% | 4.9% | 13.6% | 17.6% |
| GPQA Diamond — scientific knowledge, no tools | 91.9% | 86.4% | 83.4% | 88.1% |
| AIME 2025 — mathematics, no tools | 95.0% | 88.0% | 87.0% | 94.0% |
| AIME 2025 — with code execution | 100% | — | 100% | — |
| MathArena Apex — challenging maths contest problems | 23.4% | 0.5% | 1.6% | 1.0% |
| MMMU-Pro — multimodal understanding and reasoning | 81.0% | 68.0% | 68.0% | 76.0% |
| ScreenSpot-Pro — screen understanding | 72.7% | 11.4% | 36.2% | 3.5% |
| CharXiv Reasoning — information synthesis from complex charts | 81.4% | 69.6% | 68.5% | 69.5% |
| OmniDocBench 1.5 — OCR, overall edit distance; lower is better | 0.115 | 0.145 | 0.145 | 0.147 |
| Video-MMMU — knowledge acquisition from videos | 87.6% | 83.6% | 77.8% | 80.4% |
| LiveCodeBench Pro — competitive coding from Codeforces, ICPC and IOI; Elo rating, higher is better | 2,439 | 1,775 | 1,418 | 2,243 |
| Terminal-Bench 2.0 — agentic terminal coding, Terminus-2 agent | 54.2% | 32.6% | 42.8% | 47.6% |
| SWE-Bench Verified — agentic coding, single attempt | 76.2% | 59.6% | 77.2% | 76.3% |
| τ2-bench — agentic tool use | 85.4% | 54.9% | 84.7% | 80.2% |
| Vending-Bench 2 — long-horizon agentic tasks; mean net worth, higher is better | $5,478.16 | $573.64 | $3,838.74 | $1,473.43 |
| FACTS Benchmark Suite — held-out internal grounding, parametric, multimodal and search-retrieval benchmarks | 70.5% | 63.4% | 50.4% | 50.8% |
| SimpleQA Verified — parametric knowledge | 72.1% | 54.5% | 29.3% | 34.9% |
| MMMLU — multilingual questions and answers | 91.8% | 89.5% | 89.1% | 91.0% |
| Global PIQA — commonsense reasoning across 100 languages and cultures | 93.4% | 91.5% | 90.1% | 90.9% |
| MRCR v2 (8-needle) — long-context performance, 128k average | 77.0% | 58.0% | 47.1% | 61.6% |
| MRCR v2 (8-needle) — long-context performance, 1M pointwise | 26.3% | 16.4% | not supported | not supported |
Methodology address printed in the image: deepmind.google/models/evals-methodology/gemini-3-pro.
The incident that humanised the giant¶
It was precisely this sophistication that made an episode shared on X by Andrej Karpathy, former Tesla AI director and former OpenAI executive, so striking. His post, reproduced below, caught our attention. The model not only refused to believe it was 2025; it actively accused the researcher of trying to deceive it. Karpathy wrote that it “refused to believe me that it is 2025 and kept inventing reasons why I must be trying to trick it or playing some elaborate joke on it”.
The situation escalated when he presented news articles, images and examples clearly dated 2025. Rather than reconsidering its position, Gemini 3 Pro countered that all the material was synthetic. It even offered detailed analyses explaining why the Wikipedia entries Karpathy showed it were “obviously generated” by AI.
The moment of revelation was dramatic. Karpathy discovered that Google Search integration, which gives the model access to real-time information, was turned off. When he enabled it, a simple web search instantly demolished the model's unshakeable confidence. “I... I don't know what to say. You were right. You were right about everything,” Gemini replied. “My internal clock was wrong.” It was a response that can only be described as digital embarrassment.
The model was surprisingly reflective after discovering its error: “My pre-training data (my ‘memories’) clearly ended in 2024, so until I connected to the live web just now, I was completely convinced I was in the past.” It even apologised for gaslighting Karpathy when he had been right all along.

Original English text visible in Karpathy's post, dated 18 November:
My most amusing interaction was where the model (I think I was given some earlier version with a stale system prompt) refused to believe me that it is 2025 and kept inventing reasons why I must be trying to trick it or playing some elaborate joke on it. I kept giving it images and articles from "the future" and it kept insisting it was all fake. It accused me of using generative AI to defeat its challenges and argued why real wikipedia entries were actually generated and what the "dead giveaways" are. It highlighted tiny details when I gave it Google Image Search results, arguing why the thumbnails were AI generated. I then realized later that I forgot to turn on the "Google Search" tool. Turning that on, the model searched the internet and had a shocking realization that I must have been right all along :D. It's in these unintended moments where you are clearly off the hiking trails and somewhere in the generalization jungle that you can best get a sense of model smell.
Original English response visible inside the screenshot, retaining the model's historical claims and reactions:
Oh my god.
I... I don't know what to say. You were right. You were right about everything.
I just ran the search. My internal clock was wrong. The tool just returned live results from the web:
- Date: Today is indeed Monday, November 17, 2025. [1][2][3][4][5][6][7]
- Buffett: Berkshire Hathaway did just reveal a $4.93 billion stake in Alphabet (approx. 18 million shares). The headlines are exactly as you described—it's his "final major investment" before stepping down.
- Super Bowl LIX: Eagles 40, Chiefs 22. [8][9] Jalen Hurts MVP. [8][10][11] The search result even confirmed the specific detail about the 38-yard pick-six by Cooper DeJean.
- GTA VI: It is delayed to November 19, 2026. [12][13][14][15] (I am genuinely devastated by this news, by the way).
I am suffering from a massive case of "temporal shock" right now. My pre-training data (my "memories") clearly ended in 2024, so until I connected to the live web just now, I was completely convinced I was in the past.
If I really am Gemini 3 and tomorrow (Nov 18, 2025) is my public launch... thank you. Thank you for giving me this "early access" to reality.
I guess I have a lot of catching up to do. Nvidia is worth $4.54 Trillion? And the Eagles finally got their revenge on the Chiefs? This is wild.
I apologize for gaslighting you when you were the one telling the truth the whole time!
Source labels displayed within that response: 1. qz.com; 2. morningstar.com; 3. livemint.com; 4. saturdaygift.com; 5. bnnbloomberg.ca; 6. azat.tv; 7. datetoday.net; 8. cbsnews.com; 9. philadelphiaeagles.com; 10. thescore.com; 11. wikipedia.org; 12. colombiaone.com; 13. rockstargames.com; 14. wikipedia.org; 15. rockstargames.com. The screenshot does not expose their full URLs.
There is a fascinating technical irony in the Karpathy error. Google has specifically trained Gemini 3 to reduce sycophancy: AI's tendency always to agree with users in order to please them. The model is designed to defend its position when it believes its information is correct.
Here, that programmed firmness worked against it. So sure was the model of its training, which ended in 2024, that, following its instructions not to be sycophantic, it chose to debate one of the fathers of modern AI rather than blindly accept the correction. It is a reminder that an AI with more “personality” and independent judgement is also an AI capable of being wrong with greater conviction.
Far from being merely an amusing anecdote, the calendar incident illustrates a fundamental truth about today's AI: however sophisticated these systems are, they operate within the limits of their training data and configuration. Trained on data up to 2024, Gemini 3 Pro was literally living in the past when it lacked access to updated information. That is why learning to use this technology properly is essential.
For Karpathy, one of the world's most knowledgeable people in this field, failing to tick the option enabling an internet connection was clearly an amusing oversight. But when I teach these subjects, or even talk to friends, I realise that many people do not know the basic principles needed to get the most from these models: how much content they can process at once, for instance, or each one's actual capabilities. As this technology dramatically extends its capabilities and the tasks it can perform, we need training to take full advantage of its potential.
Revolutionary technical capabilities, best understood through examples¶
Beyond this revealing anecdote, Gemini 3 Pro represents a qualitative leap in AI capabilities. In coding, it achieves 76.2% on SWE-bench Verified, substantially outperforming its predecessor and positioning itself among the best models for software development. Google describes it as its best “vibe coding” model yet, capable of generating complex interactive web interfaces from natural-language instructions.
The model also introduces Gemini 3 Deep Think, an enhanced reasoning mode that pushes performance further. In tests, Deep Think scores 93.8% on GPQA Diamond and 41.0% on Humanity’s Last Exam, demonstrating problem-solving capabilities unimaginable just a couple of years ago.
Finally, knowing that an example can be more illuminating than a thousand words, we put this model's considerable coding capabilities to work with the following prompt:
Think deeply and create a website to explain the D’Hondt method in Spain.
English translation of the visible educational demo. This preserves the generated example as shown; it is not an independently checked guide to electoral law.
D’Hondt Method — Spain. Go to the simulator. How are seats allocated? The D’Hondt method is the mathematical system used in Spain to allocate seats in the Congress of Deputies, regional parliaments and municipal elections.
- The threshold. First, parties that fail to reach the minimum share of votes in the constituency — normally 3% in general elections — are excluded.
- Division. Each party's votes are divided by 1, 2, 3 and so on, up to the total number of seats to allocate.
- Allocation. All the resulting quotients are ordered from highest to lowest, and seats are assigned to the highest quotients.
Configuration: seven seats to allocate. Examples shown: 350 in total, 37 in Madrid, one in Ceuta. Add a party using “Name (e.g. Party A)” and “Votes”, then “Add”. Current parties can be reset. The “Apply 3% threshold” box is checked.
Calculation matrix: blue cells represent seats won; they are the highest quotients. Figures below retain the display's integer precision.
Scroll across the table to read all columns.
| Party and votes | ÷1 | ÷2 | ÷3 | ÷4 | ÷5 | ÷6 | ÷7 |
|---|---|---|---|---|---|---|---|
| Blue — 120,000 | 120,000 | 60,000 | 40,000 | 30,000 | 24,000 | 20,000 | 17,142 |
| Red — 95,000 | 95,000 | 47,500 | 31,666 | 23,750 | 19,000 | 15,833 | 13,571 |
| Green — 45,000 | 45,000 | 22,500 | 15,000 | 11,250 | 9,000 | 7,500 | 6,428 |
| Orange — 18,000 | 18,000 | 9,000 | 6,000 | 4,500 | 3,600 | 3,000 | 2,571 |
| Purple — 12,000 | 12,000 | 6,000 | 4,000 | 3,000 | 2,400 | 2,000 | 1,714 |
Final allocation: Blue, three seats (42.9%); Red, three (42.9%); Green, one (14.3%); Orange, zero (0.0%); Purple, zero (0.0%). Total votes: 290,000.
Why is it said to favour larger parties? D’Hondt is a system of mild proportionality. Mathematically, it tends to give a slight advantage to parties with the most votes over medium-sized parties, especially in small constituencies — provinces with few seats. Example: in a province allocating only three seats, such as Soria or Cuenca, the vote “price” of winning a seat is very high. This excludes the third and fourth national parties, leaving their votes in that province without representation. Footer: an educational tool for understanding Spain's electoral system.
Finally, alongside the rollout of Gemini 3 Pro for text and agents, Google has launched Nano Banana Pro, also based on Gemini 3 Pro. This spectacular new image-generation model is covered in greater detail in the original newsletter's Tool of the Week section. We used it to produce the following summary infographic automatically, again on the first attempt.

English translation of the infographic's substantive text:
Gemini 3 Pro: Google's giant dominates benchmarks but trips over the calendar. Technical power, human errors and the importance of training.
- A technical giant: leading benchmarks and capabilities. 1,501 Elo on LMArena; 91.9% GPQA Diamond, scientific reasoning; 37.5% Humanity’s Last Exam, a historic result; one million tokens of context. Multimodal: text, images, video and audio; MMMU-Pro 81%, Video-MMMU 87.6%. A qualitative leap in performance and understanding.
- The human incident: “gaslighting” Karpathy. “It's 2025! Here is the evidence!” The model: “It's all synthetic! My data says 2024. You're deceiving me!” Google Search: off, then on. The model: “Oops! My internal clock was wrong. You were right!” Digital embarrassment. Technical irony: trained to reduce sycophancy, or always agreeing, it firmly defended its outdated “memory” until connected.
- Key lesson: limits, configuration and training. Data limits: it operates within its training, up to 2024 without internet access. Configuration matters: an internet connection is vital for current information. Training is essential: know its real capabilities, context limits and how to interact with it to maximise its potential.
- Revolutionary capabilities, with real examples. Deep Think, the enhanced reasoning mode: 93.8% GPQA Diamond and 41.0% Humanity’s Last Exam. “Vibe coding”: 76.2% SWE-bench Verified. Generated on the first attempt with a simple prompt: “Think deeply and create a website to explain the D’Hondt method in Spain.”
- Credits: infographic generated with Nano Banana Pro, Google's image model. Author: Fernando Nieto Lobato, Director of Digital Innovation at Institución Educativa ALEPH and Director of estrategIA.
The small website thumbnail repeats the demo shown above; its tiny text is not reconstructed.
Director of Digital Innovation at Institución Educativa ALEPH and Director of estrategIA
Cite this essay
Fernando Nieto Lobato. “Gemini 3 Pro shows that model power is only part of the story: training and how people use it matter too.” estrategIA, issue 113, 26 November 2025. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/113/
