Serverless Fine-Tuning
Self-Serve Deployments is now generally available for inference in Crusoe Intelligence Foundry

Self-Serve Deployments is now generally available for inference in Crusoe Intelligence Foundry

8/11/2026

What this post added

Introduces Self-Serve Deployments as a new option within Crusoe Intelligence Foundry for deploying custom inference models. This feature provides dedicated inference endpoints with selectable optimization profiles (Throughput, Responsiveness, Balanced) to cater to different workload needs (high-volume concurrent requests vs. user-facing latency). It offers predictable pricing based on GPU per hour and integrates directly with models fine-tuned via Crusoe Serverless Fine-Tuning. The post also highlights the underlying MemoryAlloyTM technology for efficient inference.

Read the original post ↗