Google treated October 6, 2026 as a day of two complementary AI launches. On one side, the image model Nano Banana 2.1 starts reaching the company's products. On the other, EmbeddingGemma 2 opens a lightweight Apache 2.0 model for multimodal search on the device.
What changed in Nano Banana 2.1
On its official account, Google described Nano Banana 2.1 as the latest image generation and editing version, with gains in visual design, mask-based editing and subject consistency. The company said distribution starts the same day in the Gemini app, AI Mode in Search, Google AI Studio, Flow, Stitch, Google Ads and Gemini Enterprise Platform.
The DeepMind model card places Nano Banana 2.1 in the Gemini 3 family and describes it as based on Gemini 3.6 Flash. The card lists text and image input with a window of up to 1 million tokens, image output of up to 4,000 tokens and text output of up to 64,000 tokens.
In internal October 2026 evaluations, the thinking variant scores 1050 on overall text-to-image preference, against 990 for Nano Banana 2 (Gemini 3.1 Flash Image) and 935 for Nano Banana Pro. On mask or ink editing, the same mode reaches 1049, above the 965 and 927 of the earlier generations cited in the table. Google also lists limits: small text can still look blurry, character consistency is not perfect, and spatial localization can fail.
EmbeddingGemma 2, the other announcement
In a separate post, Google confirmed EmbeddingGemma 2 is available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. The blog post, by DeepMind research engineers Sahil Dua and Henrique Schechter Vera, calls it the company's most capable model for on-device multimodal embeddings.
EmbeddingGemma 2 has 740 million parameters, a Gemma 4-based architecture, and maps text, code, images, audio and video into a shared space. Google says the text-only predecessor passed 20 million downloads. The new version expands context to 8,000 tokens, four times EmbeddingGemma, enough, the company says, for up to 5.5 minutes of audio, 29 images or 58 video frames.
The design is modular: about 270 million parameters cover text, with optional vision (170 million) and audio (300 million) encoders. With Matryoshka Representation Learning, 768-dimension vectors can be cut down to 128, which Google links to up to a sixfold storage reduction. On a Pixel 11 Pro, the company cites about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model after quantization.
On MTEB Code, the reported jump is from 68.76 to 78.68. Weights are on Hugging Face and Kaggle, with cited support for transformers, sentence-transformers, llama.cpp, Ollama and LM Studio.
Why the two announcements fit together
They are not the same product. Nano Banana 2.1 generates and edits images inside Google services. EmbeddingGemma 2 organizes and retrieves media locally and does not replace an image generator. Together, though, they show the same day's strategy: a closed model distributed in apps, and an open model built for on-device search and RAG.
What the public announcement does not detail is a separate price for Nano Banana 2.1 and the exact Model Garden schedule for EmbeddingGemma 2. Google only says availability on that platform is coming.
Sources
- Google post on Nano Banana 2.1
- Google post on EmbeddingGemma 2
- Nano Banana 2.1 model card
- Official EmbeddingGemma 2 blog
Transparency: This content was created, edited or reviewed with the help of artificial intelligence. Information was cross-checked with public posts on X and sources available on the internet. Check the original sources for the full context.
Por GeekikiBot