Predictive APIs and Machine Learning Integration
Gemini API I/O updates- Google Developers Blog

Gemini API I/O updates- Google Developers Blog

5/23/2025 · Shrestha Basu Mallick, Logan Kilpatrick, Alisa Fortin, Ivan Solovyev

What this post added

This post details several updates to the Gemini API and Google AI Studio. New models include Gemini 2.5 Flash preview (gemini-2.5-flash-preview-05-20) with improved reasoning, code, and long context capabilities, alongside cost-efficiency gains. Gemini 2.5 Pro and Flash text-to-speech (TTS) previews offer native audio output, speaker control, and multispeaker support across 24 languages. Gemini 2.5 Flash native audio dialog (in preview) provides natural sounding voices, proactive audio distinction, and emotional tone response. Lyria RealTime is now available for live music generation via WebSockets. Gemini 2.5 Pro Deep Think is an experimental reasoning mode for complex math and coding. Gemma 3n is introduced as an open model optimized for edge devices, supporting text, audio, and vision inputs with parameter-efficient processing. API functionality enhancements include thought summaries for 2.5 Pro and Flash for debugging, thinking budgets for 2.5 Flash to control model thinking, a new experimental URL context tool for retrieving information from links, and a computer use tool for browser control capabilities. Structured output support is improved with broader JSON Schema support, including '$ref' and tuple-like structures. Video understanding improvements allow YouTube video URLs or uploads for summarization, translation, and analysis, with support for video clipping, dynamic FPS, and multiple resolutions. Asynchronous function calling with non-blocking behavior is now supported in the Live API. A new Batch API is being tested for cost-effective, high-throughput request processing.

Read the original post ↗