
8/11/2026
What this post added
Introduces Self-Serve Deployments as a new option within Crusoe Intelligence Foundry for deploying custom inference models. This feature provides dedicated inference endpoints with selectable optimization profiles (Throughput, Responsiveness, Balanced) to cater to different workload needs (high-volume concurrent requests vs. user-facing latency). It offers predictable pricing based on GPU per hour and integrates directly with models fine-tuned via Crusoe Serverless Fine-Tuning. The post also highlights the underlying MemoryAlloyTM technology for efficient inference.