Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter(opens in new tab) to selectively trade off cost for intelligence.
Modalities
Context
1M
Released
Sep 25, 2025
Knowledge Cutoff
Jan 2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e.
Gemini 2.5 Flash Lite Preview 09-2025 has a 1,048,576 token context window.
Gemini 2.5 Flash Lite Preview 09-2025 accepts text, images, files such as PDFs, audio and video as input and returns text.
Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash Lite and 26 more are other text models from Google.
Gemini 2.5 Flash Lite Preview 09-2025 was released on September 25, 2025. Its knowledge cutoff is January 31, 2025.
Token volume and request traffic to this model over time.