5Google DeepMind Blog·1mo ago

Gemini 2.5 Flash-Lite reaches general availability for production use

Google DeepMind has moved Gemini 2.5 Flash-Lite from preview to stable general availability. The model is positioned as a cost-efficient, small-footprint option within the 2.5 family, retaining key features including a 1 million-token context window and multimodal capabilities. It is now ready for scaled production deployment.

Long Context Evolution Frontier Model Releases Inference Economics Enterprise Deployment Patterns Gemini 2.5 Gemini-2.5-Flash-Lite Google DeepMind

Related guides (4)

Google DeepMind

Google DeepMind: Frontier AI Across Models, Robotics, and Scientific Discovery

Read asIn-depth

Frontier Model ReleasesTopic guide

Frontier Model Releases: The Race From Language to Action

Read asBeginner In-depth

Long Context EvolutionTopic guide

Long Context Evolution: From Bigger Windows to Smarter Memory

Read asBeginner In-depth

Enterprise Deployment PatternsTopic guide

Enterprise Deployment Patterns: From LLM Demo to Production Reality

Read asIn-depth

Related events (8)

7Google Deepmind Blog·1mo ago·source ↗

Gemini 2.0 Flash and Flash-Lite Reach General Availability

Google DeepMind has made Gemini 2.0 Flash-Lite generally available via the Gemini API, Google AI Studio, and Vertex AI for enterprise production use. This marks the transition of the Flash-Lite variant from preview to full GA status. The release expands developer and enterprise access to cost-efficient Gemini 2.0 inference capabilities.

Frontier Model Releases Inference Economics Google AI Studio Gemini-2.5-Flash-Lite Google DeepMind +3 more

8Google Deepmind Blog·1mo ago·source ↗

Gemini 2.5 Family Expansion: Flash and Pro GA, Flash-Lite Introduced

Google DeepMind has made Gemini 2.5 Flash and Gemini 2.5 Pro generally available, while simultaneously introducing Gemini 2.5 Flash-Lite, described as the most cost-efficient and fastest model in the 2.5 family. The announcement marks the full productization of the Gemini 2.5 generation. Flash-Lite targets latency- and cost-sensitive deployment scenarios.

Frontier Model Releases Inference Economics Gemini-2.5-Flash-Lite Google DeepMind Gemini-2.5-Pro +1 more

8Google Deepmind Blog·1mo ago·source ↗

Gemini 2.5: Updates to our family of thinking models

Google DeepMind has announced updates to the Gemini 2.5 model family, including Gemini 2.5 Pro reaching stable status, Gemini 2.5 Flash becoming generally available, and a new Gemini 2.5 Flash-Lite entering preview. These releases mark the maturation of DeepMind's 'thinking model' line with enhanced performance and accuracy. The updates span multiple tiers of the Gemini 2.5 family, from the flagship Pro to the lightweight Flash-Lite variant.

Long Context Evolution Frontier Model Releases Gemini-2.5-Flash-Lite Google DeepMind Gemini-2.5-Pro +1 more

6Google Deepmind Blog·1mo ago·source ↗

Gemini 3.1 Flash-Lite: Built for intelligence at scale

Google DeepMind has released Gemini 3.1 Flash-Lite, described as the fastest and most cost-efficient model in the Gemini 3 series. The announcement positions it as optimized for high-throughput, cost-sensitive deployments at scale. The body is sparse, offering no benchmark details or capability specifics beyond the efficiency framing.

Frontier Model Releases Inference Economics Google DeepMind Gemini 3.1 Flash Live Gemini +1 more

7Hacker News·1mo ago·source ↗

Gemini 3.5 Flash Released

Google has released Gemini 3.5 Flash, a new model in the Gemini family. The announcement appears on Google's official blog and has generated significant community discussion on Hacker News with 381 points and 304 comments. Gemini 3.5 Flash follows the Flash line of efficiency-focused models from Google DeepMind.

Frontier Model Releases Inference Economics Google Gemini 3.5 Flash Google DeepMind +3 more

8Google Deepmind Blog·1mo ago·source ↗

Gemini 3 Flash: frontier intelligence built for speed

Google DeepMind has announced Gemini 3 Flash, a new model positioned as a frontier-intelligence offering optimized for speed and cost efficiency. The announcement comes from the official DeepMind blog, indicating a formal product release. Specific capability details and benchmarks are not included in the available body text.

Frontier Model Releases Inference Economics Google DeepMind Gemini 3 Flash Gemini

8Google Deepmind Blog·1mo ago·source ↗

Introducing Gemini 2.5 Flash

Google DeepMind has released Gemini 2.5 Flash, described as their first fully hybrid reasoning model. The model allows developers to toggle 'thinking' (extended reasoning) on or off, combining standard and chain-of-thought inference modes in a single model. It is available to developers and represents a new architectural approach to balancing reasoning depth with inference cost.

Long Context Evolution Frontier Model Releases Gemini-2.5-Flash-Lite Google DeepMind Gemini-2.5-Pro +3 more

7Google Deepmind Blog·1mo ago·source ↗

Gemini 2.5 Pro Preview: Updated Version with Improved Coding Performance

Google DeepMind has released an updated preview version of Gemini 2.5 Pro ahead of its originally planned schedule, citing strong developer adoption and usage. The update focuses on improved coding performance. The early release reflects DeepMind's responsiveness to developer demand for the model.

Frontier Model Releases Agent and Tool Ecosystem Google DeepMind Gemini-2.5-Pro