Archive context: this article describes the author's tests and the services available on 3 September 2025. Product availability, access advice and performance comparisons belong to that date. The synthetic examples below are demonstrations, not photographs of real events or approved public projects. English text already visible in screenshots is transcribed as shown; Spanish prompts and captions are translated.

This week we bring you a fairly practical issue. For many weeks in August, a mysterious image AI model called “Nano banana” generated justified hype on social media and in expert forums (I briefly tried it on LMArena at the time and could confirm those expectations). “Nano banana” was finally launched at the end of the month—on 26 August, to be precise—and turned out to be the nickname of Google's latest artificial intelligence model for generating and editing images, officially known as Gemini 2.5 Flash Image. This system represents a qualitative leap in digital visual creation because it can both generate hyperrealistic images from text and, above all, modify existing photographs simply and precisely, even surpassing traditional applications such as Photoshop in speed and flexibility. It is not perfect yet, but it shows where AI's potential in this field is heading and the possibilities it will offer us in the very near future.

Following the example of issue 80, in which we explored some of the possibilities for politics and government offered by the image model OpenAI had then released, this time we bring you something similar, though a little shorter to avoid repeating what we have already created. We have been testing this Google model, which, as the benchmarks clearly show, is well ahead of everything that has existed in AI imagery to date, especially in photographic editing.

LMArena Image Edit Arena ranking, with Gemini 2.5 Flash Image Preview first; the displayed results are transcribed below.

Screenshot transcription: “Gemini-2.5-flash-image, aka ‘nano-banana’, debuts #1 on Image Edit Arena (with the largest score jump in history)”. The table's rank column is labelled “Rank (UB)”; confidence intervals are labelled “95% CI (±)”. The organisation name “Black Fores…” is truncated in the original.

Scroll across the table to read all columns.

Rank (UB) Model Score 95% CI (±) Votes Organisation Licence
1 gemini-2.5-flash-image-preview 1362 2 2,521,035 Google Proprietary
2 flux-1-kontext-max 1191 3 357,196 Black Fores… Proprietary
3 flux-1-kontext-pro 1174 2 2,015,530 Black Fores… Proprietary
3 gpt-image-1 1170 3 1,026,399 OpenAI Proprietary
5 flux-1-kontext-dev 1152 3 1,584,400 Black Fores… Proprietary
6 qwen-image-edit 1145 2 1,585,904 Alibaba Apache 2.0
6 seededit-3.0 1142 4 1,285,080 Bytedance Proprietary
8 gemini-2.0-flash-preview-image-generation 1093 3 1,700,785 Google Proprietary
9 bagel 1044 5 12,774 Bytedance Apache 2.0
10 step1x-edit 1017 4 138,399 StepFun Apache 2.0

Chart comparing Gemini 2.5 Flash Image's wins and losses against six image models.

Chart transcription: “Google's Gemini 2.5 Flash Image dominates arena. The new model won >85% of encounters on the LMArena, based on 2.5m votes.” The chart is headed “LMArena Image Edit Arena Win Rate”; green denotes Gemini wins and red denotes losses, with a 50% reference line. These are the displayed percentages, including the two rows below the headline's threshold.

Scroll across the table to read all columns.

Opponent Gemini wins Gemini losses
Flux 1 Context Max — Black Forest Labs 81% 19%
Flux 1 Context Pro — Black Forest Labs 82% 18%
GPT Image 1 — OpenAI 85% 15%
Qwen Image Edit — Alibaba Cloud (Qwen) 85% 15%
Flux 1 Kontext Dev — Black Forest Labs 86% 14%
Gemini 2.0 Flash Preview — Google DeepMind 93% 7%

The graphic credits LMArena, printing https://lmarena.ai/leaderboard/webdev, and Peter Gostev, https://www.linkedin.com/in/peter-gostev/.

First of all, you can try this model free of charge in Google's AI Studio. We have admittedly had some problems generating images when accessing it from Spain; in that case, we recommend using a VPN. It is also available in Gemini, although there we have found that the model refuses to perform many editing tasks, probably because of much stricter safety settings (which, in many completely innocuous cases we have tested, are rather hard to understand).

The first thing we did was replicate the great majority of the examples we created with ChatGPT's image model in issue 80, and the results were as good or, generally, better. There is one important caveat which seems to be, also in the judgement of several experts, one of the model's weak points: it is not especially good at generating text. Nor is it good at generating comics of the kind OpenAI's image model can produce. In the other areas, we found a very interesting model. These are some of the images it generated along the lines of what we asked ChatGPT's model to produce (to avoid making this newsletter incredibly long, we suggest consulting the results from that model there).

Source note: an orphaned editing marker, [[1]](#_msocom_1), splits the Spanish word “Tampoco” in the archived source. No corresponding comment survives; the word has been rejoined here.

Posters, public spaces and figurines

It is quite capable of creating images for election posters in different artistic styles and producing variations when asked in just a couple of words:

Generated Art Nouveau election poster with a woman, floral decoration and peacocks.

Poster text, translated: “Imagine a better future. Today, we are backing you.” “Vote — November.”

Variation on the same generated election poster, featuring Marge and Homer Simpson.

The variation repeats the same text: “Imagine a better future. Today, we are backing you.” “Vote — November.”

It is also perfectly capable of generating satellite images to show a possible redevelopment or the construction of new public infrastructure in an area, in a very similar way to ChatGPT:

Satellite photograph and generated version adding a sports complex in the southern part of the image.

Prompt, translated: “Generate an image of a sports complex in the southern part of this satellite image. The aim is to show what the area would look like.” The model's response adds football pitches, an athletics track, a swimming pool and courts.

For example, it takes the typical action-figure “toy” made from an image, which became fashionable a few months ago, to another level:

Generated boxed action figure labelled President Barack Obama, with a presidential office backdrop.

The fictional packaging reads “President Barack Obama”, “Official Presidential Action Figure” and “POTUS Collectibles”. These are generated labels, not an assertion of official endorsement.

Screenshot turning a painting of an armoured horseman into a collectible figurine on a computer desk.

Visible English prompt, ending where the screenshot cuts it off: “Create a 1/7 scale commercialized figurine of the characters in the picture, in a realistic style, in a real environment. The figurine is placed on a computer des...” The remaining text is not visible.

It is still not a good model for creating complex posters (and it makes more errors in the text). Designing with several elements is something artificial intelligence still clearly needs to improve. There is barely any progress here, and it handles the text worse:

Generated European election poster illustrating the model's imperfect handling of text and complex layouts.

Legible poster text, translated: “For a strong and serene Europe”; “Reason, resolve, future”; “Solid principles for a strong continent”. The party-name line is malformed in the generated Spanish image; it is retained visibly rather than silently repaired.

Editing photographs and combining elements

Where it really shines, though, is in its ability to transform images, as in this example we previously produced with ChatGPT from a photograph of a presidential debate on Televisión Española a few years ago. As you can see, it is much better with Google's new model.

Original image:

Original photograph of four debate participants standing in front of a Televisión Española backdrop.

ChatGPT (also shown in issue 80 of estrategIA):

ChatGPT's generated transformation of the debate participants into Hogwarts-style robes, with an altered background.

And the new Gemini 2.5 Flash Image, which preserves the background and faces much better and also adds wands to the candidates' hands:

Gemini's generated transformation, retaining more of the debate backdrop while adding robes and wands.

To avoid dwelling too long on a direct comparison with ChatGPT's model and the examples we generated at the time, our assessment, as already mentioned, is that it is probably a worse model for working with text and for the multi-panel comic format that is so interesting in OpenAI's model. But it can competently produce the other examples we previously asked ChatGPT for, and it shines very noticeably in image editing, as you can see in the previous example—generated on the first attempt—and in some of those we are about to show you.

This example is particularly fascinating because it shows its ability to understand instructions and introduce objects and people into any scene:

AI Studio prompt combining two portrait photographs and two football shirts into a new scene.

Prompt, translated: “Create an image of the first subject wearing the white shirt and the second subject wearing the blue and red shirt. Put them in the middle of a football pitch facing each other.” The pictured shirts carry the “Emirates Fly Better” and Spotify branding. The historical interface also displays: “Gemini 2.5 Flash Image does not currently support editing images of children.”

Generated result showing the two subjects facing each other on a football pitch in the specified white and blue-and-red shirts.

And I have “only”—although this is already astonishing—used four different elements: two subjects and two shirts. But the model reaches truly incredible limits, as this post on X shows, which will surely allow image specialists to work on highly imaginative ideas to a high standard:

Travis Davids post showing thirteen reference images combined into a fashion scene beside a pink car.

English text transcribed from Travis Davids's post (@MrDavids1):

New record? 13 images merged into a single image using Gemini 2.5 Flash Image (Nano Banana). This collage method is absolutely BANANAS! I'm actually amazed that it can do this however I feel like I'm reaching it's limit now but even at 13 elements it's still managing to obtain consistency, the detailed prompt however is very important once you start playing around with a crazy amount of elements like this. 🤯

Prompt: A model is posing and leaning against a pink bmw. She is wearing the following items, the scene is against a light grey background. The green alien is a keychain and it's attached to the pink handbag. The model also has a pink parrot on her shoulder. There is a pug sitting next to her wearing a pink collar and gold headphones.

The reference collage contains the model's portrait, a jacket, cap, car, trainers, sunglasses, pug, headphones, parrot, handbag, jeans, alien keyring and collar.

It also works very well at transforming a person's image to adapt it to different styles and historical periods. Here is an example of a prompt I took directly from a Google post on X, although I changed the reference subject and tried new historical styles:

Two generated historical transformations of the same portrait, using prompts about India in the 1920s and Spain in the 1420s.

The two English prompts shown are: “Make me look like I am in the 1920s in India.. Change my fashion, hairstyle and background” and “Make me look like I am in the 1420s in Spain. Change my fashion, hairstyle and background”. The images are the model's imaginative responses, not historical photographs.

There are many more uses generated by social media users in which we see potential for politics and government. Here are a few examples:

Illia's post transforming a night-time photograph of a building into an isolated daytime isometric view.

English text shown in the post by Illia (@Zieeett): “Nano Banana: Make Image Daytime and Isometric (Building Only)”. The screenshot is dated 26 August 2025, 9:15 p.m., and displays 762.4 thousand views.

Nano banana is also excellent for trying on clothes. Other models could already do this, but the consistency here is probably better:

Post demonstrating a generated T-shirt replacement while retaining a small clip-on microphone from the source photograph.

English text shown in AshutoshShrivastava's post (@ai_for_success):

RIP 379 startups..

nano-banana is 🔥

i can’t believe it replaced the entire t-shirt and still kept that tiny microphone intact from the original image.

The shirt reads “Even I use ChatGPT”. The screenshot is dated 25 August 2025, 6:35 p.m., and displays 83.5 thousand views. The post's claim is quoted as part of the original example.

There are many more uses, and this post could become almost endless. We particularly recommend taking a look at this X thread from Google itself, with some very interesting examples that are sure to give you new ideas for your projects. We also recommend looking at the “recommendation of the week” section if you want to see what the model can do in more “traditional” photographic editing.

Archive note: that recommendation is in the original newsletter's separate weekly section, outside the main article translated here.

Finally, among the examples we have created, we find this use especially useful—if it could handle text better, which it still struggles with. It can highlight monuments or buildings in images taken, for example, from Google Maps and add contextual information about them. Here is a simple example we made with Salamanca Cathedral:

Generated annotation highlighting Salamanca Cathedral and adding an information panel to the photograph.

Visible panel text: “Salamanca Cathedral”; “Construction began: 1513”; “Architects: Juan Gil de Hontañón, Rodrigo Gil de Hontañón, Juan de Álava”; “Style: Late Gothic, Renaissance”; “Designated. National Monument 1887”. These are the labels generated in the example, reproduced here without independent historical verification.

In any case, the model offers enormous possibilities and shows AI's immense potential in imagery, which also applies to its use in politics and public administration. We have not included these here, because they require a further step beyond the model itself, but combining Nano banana's capabilities with video generation can produce some very striking projects. Our invitation, especially given that the model can be tried easily and free of charge, is to use it, try to replicate some of the options we have shown you and, above all, explore whether its outstanding image-editing capabilities—where we see the clearest leap in quality—could be useful in your everyday work or projects. Ultimately, the tools will rapidly reach levels of technical excellence, and it is our imagination and knowledge that will make the difference in how we use their possibilities.

Fernando Nieto Lobato

Director of Digital Innovation at Institución Educativa ALEPH

This is a translation of the original Spanish essay published on 3 September 2025. Its claims, examples and forecasts retain that historical context. Read the original Spanish edition, including its accompanying illustrations.

Cite this essay

Fernando Nieto Lobato. “Gemini 2.5 Flash Image (Nano banana): a qualitative leap in image editing and generation.” estrategIA, issue 101, 3 September 2025. English edition, 29 September 2026. https://elcontemplador.github.io/estrategia-english/essays/101/

Back to the top ↑