Consultancy · Development
GEMINI: Google's new, more capable AI model.
“Welcome to the Gemini era” is the phrase with which he begins his presentation on Google DeepMind.
The “Gemini era” refers to a new stage of research and innovation that Google opens after almost 8 years betting and studying generative AI.
Sundar Pichai, CEO of Google and Alphabet, in an official company note stated: “Nearly eight years into our journey as an AI-first company, the pace of progress is only accelerating: millions of people are using generative AI across our products to do things they couldn't even a year ago (…). At the same time, developers are using our models and infrastructure to create new generative AI applications, and companies around the world world are growing with our AI tools.”
Gemini takes a new step in terms of Artificial Intelligence, being the most capable and general model, with the latest generation technology. Its first version, 1.0, is optimized for three different sizes:
- Gemini Ultra: larger model with capacity for very complex tasks.
- Gemini Pro: ideal model for scaling in a wide range of tasks.
- Gemini Nano: most efficient model for tasks on mobile devices.
Broad understanding capacity.
Gemini 1.0 was designed and trained to recognize and understand text, images, audio and more at the same time, so it better understands nuanced information and can answer questions related to complicated topics. This makes it especially good at explaining reasoning in complex subjects like mathematics and physics.
Performance never seen before.
Gemini underwent various assessments to measure its performance on a wide variety of tasks: from understanding natural images, audio and video, to mathematical reasoning. This is how Gemini Ultra surpasses the current state-of-the-art results in 30 of the 32 academic benchmarks used in the research and development of Large Language Models (LLM).
On the other hand, with a score of 90%, this new language model is positioned as the first to surpass human experts in MMLU (massively multitasking language understanding), which uses a combination of 57 STEM subjects as a reference point. This allows Gemini to more accurately answer complex questions.
Multimodality as a distinctive feature.
Previously, to create multimodal models, it was necessary to train components separately, each one destined for different modalities, and then join them and try to imitate each functionality as much as possible.
Gemini breaks this rule, being a language model designed to be natively multimodal, trained entirely this way from its origins. This allows it to be much faster and more precise than any other model with similar characteristics. In addition, it was later refined with other additional multimodal elements in order to be able to understand and reason on all types of input from scratch.


Gemini Ultra also achieves a record score of 59.4% in MMMU (Massive Multi-discipline Multimodal Understanding), a new benchmark designed to evaluate multimodal models in massively multidisciplinary tasks that require university-level subject knowledge and deliberate reasoning.
When will it be open to the public?
Starting December 13, developers and enterprise customers can now access Gemini Pro through the Gemini API in Google AI Studio or Google Cloud Vertex AI.
At the beginning of next year, it will be open to the general public. Meanwhile, wait😅
If you liked this information, we invite you to follow us on our social networks so as not to miss the latest news: Instagram, Facebook, TikTok and LinkedIn.
Lab9 - Digital Agency.

