Google has released Gemini 3.5 Flash, the latest in its family of efficiency-optimized AI models. The release delivers a 50% reduction in inference costs compared to Gemini 2.5 while maintaining comparable quality on standard benchmarks.
The Efficiency Focus
The model uses an optimized mixture-of-experts architecture, aggressive quantization techniques, and custom inference pipelines to reduce compute requirements without proportionally reducing capability.
Pricing Impact
At the announced pricing, Gemini 3.5 Flash costs approximately $0.075 per million input tokens and $0.30 per million output tokens — significantly below comparable models from OpenAI and Anthropic.
Market Positioning
The release targets the growing segment where cost efficiency matters more than maximum capability. Google's pricing strategy creates competitive pressure across the AI industry.