
4/11/2025 · Michelle Chen, Jesse Kipp
What this post added
Introduced speculative decoding and prefix caching to significantly speed up inference times (2-4x) for models like Llama 3.3 70b. Launched an asynchronous batch API to handle large workloads more reliably, preventing immediate errors due to capacity and guaranteeing fulfillment. Expanded LoRA support to 8 models with larger ranks (up to 32) and file sizes (up to 300 MB). Added new models to the catalog and refreshed the dashboard for improved usability.