Torch Compile Caching
Fine-tuned models now boot in less than one second

Fine-tuned models now boot in less than one second

9/6/2023

What this post added

This post details a significant improvement in cold boot times for fine-tuned models on Replicate, reducing them to under one second. This is achieved by optimizing the model loading and initialization process, specifically targeting large language models like Llama 2 and image models like SDXL. The improvements are currently available for new fine-tuned models and are expected to be rolled out to all models in the future.

Read the original post ↗