Mixtral Sparse Mixture of Experts Model
Mistral 7B

Mistral 7B

9/27/2023

What this post added

Introduces Mistral 7B, a 7.3B parameter model, detailing its performance advantages over Llama 2 13B and Llama 1 34B. Highlights the use of Grouped-query attention (GQA) for faster inference and Sliding Window Attention (SWA) for efficient handling of longer sequences. Explains the technical details of SWA, including its linear compute cost and how stacked layers can attend beyond the window size. Discusses the optimization of attention cache using rotating buffers. Presents Mistral 7B Instruct, a fine-tuned version for chat, and its performance on MT-Bench.

Read the original post ↗