Skip to content

Gemini 3.5 Flash Lite

gemini-3.5-flash-lite
Google · chat · api
GA Alert me on changes
Intelligence
Context
1M
Max output
65.5K
Input $/1M
$0.30
Output $/1M
$2.50
Modalities
text
Released
21 Jul 2026
Intelligence Index via Artificial Analysis · 0–100, higher is better
Download image Share on X Share on LinkedIn
AI summary
● machine-written

Google releases Gemini 3.5 Flash Lite for agentic workflows

Gemini 3.5 Flash Lite is a low-latency, cost-efficient multimodal model designed for virtual assistant tasks and document parsing with high throughput. It supports text, image, video, audio, and PDF inputs and is optimized for multi-agent workflows and data extraction applications. The model offers a 1M token context window and 65.5K token maximum output.

What's new
  • Released July 21, 2026 on OpenRouter
  • Supports text, image, video, audio, and PDF inputs
  • Optimized for subagent execution in complex multi-agent workflows
  • Pricing at $0.30/$2.50 per 1M tokens (input/output)
Best for
Virtual assistant tasksHigh-throughput document parsingMulti-agent workflow subagentsCost-constrained applications where latency and API cost are primary constraints
Sources

Common questions about Gemini Flash →

Source: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite