Mixtral Sparse Mixture of Experts Model
Mistral NeMo | Mistral AI

Mistral NeMo | Mistral AI

7/18/2024

What this post added

Introduces Mistral NeMo, a 12B parameter model with a 128k context window, highlighting its performance on reasoning, world knowledge, and coding benchmarks. Details the new Tekken tokenizer, which offers improved compression rates for various languages and source code. Presents instruction fine-tuning results demonstrating enhanced capabilities in instruction following, reasoning, and code generation. Mentions availability through HuggingFace, la Plateforme, and NVIDIA NIM.

Read the original post ↗