Mixtral Sparse Mixture of Experts Model
Large Enough | Mistral AI

Large Enough | Mistral AI

7/24/2024

What this post added

Introduces Mistral Large 2, a 123B parameter model with a 128k context window. Highlights its performance on MMLU (84.0% accuracy), code generation benchmarks (on par with GPT-4o, Claude 3 Opus, Llama 3 405B), mathematical reasoning (GSM8K, MATH), and general alignment benchmarks (MT-Bench, Wild Bench, Arena Hard). Details improvements in instruction following, multilingual capabilities (especially on multilingual MMLU), and function calling/retrieval. Notes the model's single-node inference capability and cost efficiency. Mentions availability on la Plateforme (`mistral-large-2407`) and via cloud providers, and the extension of fine-tuning capabilities.

Read the original post ↗