7/16/2024
What this post added
Introduces Codestral Mamba, a new model based on the Mamba architecture, highlighting its linear time inference and ability to model long sequences. Details its training for code and reasoning capabilities, benchmarked up to 256k tokens for in-context retrieval. Provides deployment instructions via mistral-inference SDK and TensorRT-LLM, and mentions availability of raw weights on HuggingFace. Contrasts its Apache 2.0 license with Codestral 22B's commercial/community licenses.