
9/24/2024 · Logan Kilpatrick, Shrestha Basu Mallick
What this post added
This post announces the release of updated production-ready Gemini models (Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002), featuring significant price reductions (up to 64% for 1.5 Pro input tokens, 52% for output tokens), doubled rate limits for 1.5 Flash and tripled for 1.5 Pro, 2x faster output, and 3x lower latency. Model quality has improved by ~7% on MMLU-Pro, ~20% on MATH and HiddenMath benchmarks, and 2-7% on vision and code generation tasks. Default filter settings are now off by default. An experimental Gemini-1.5-Flash-8B-Exp-0924 model is also available.