BlogsCrusoeServerless Fine-Tuning

Serverless Fine-Tuning

Serverless Fine-Tuning

3
posts
2026

Crusoe has launched Serverless Fine-Tuning, a managed service within Crusoe Intelligence Foundry that allows users to customize open-source models using their own data. This service eliminates the need for users to provision GPU clusters or manage infrastructure, offering a pay-as-you-go model. It supports a variety of base models (Qwen, DeepSeek, Llama, Gemma, gpt-oss, etc.) and data formats (JSONL, Parquet). The pipeline includes data pre-processing (cleaning, tokenization, de-duplication), au. This post details the integration of NVIDIA's Nemotron 3 Ultra model with LangChain Deep Agents, running on Crusoe Cloud. It highlights how Crusoe's Managed Inference and the `langchain-crusoe` integration enable cost-effective deployment of frontier-class open agents, emphasizing harness engineering over model fine-tuning for performance gains. The post also introduces Crusoe's MemoryAlloy KV cache fabric for improved agentic workload performance and discusses tiered model deployment strategies using Nemotron 3 Ultra for orchestration and Nemotron 3 Nano Omni for execution.

2026

Nemotron 3 Ultra agents: 10x cheaper on Crusoe Cloud

8/11/2026

Introduces the integration of NVIDIA's Nemotron 3 Ultra model with LangChain Deep Agents, running on Crusoe Cloud. Details the use of Crusoe's Managed Inference and `langchain-crusoe` integration for deploying open agents. Explains harness engineering as a method for improving agent performance and cost-effectiveness. Highlights Crusoe's MemoryAlloy KV cache fabric for agent workloads and presents a blueprint for tiered model deployment using Nemotron models.

Self-Serve Deployments is now generally available for inference in Crusoe Intelligence Foundry

8/11/2026

Introduces Self-Serve Deployments as a new option within Crusoe Intelligence Foundry for deploying custom inference models. This feature provides dedicated inference endpoints with selectable optimization profiles (Throughput, Responsiveness, Balanced) to cater to different workload needs (high-volume concurrent requests vs. user-facing latency). It offers predictable pricing based on GPU per hour and integrates directly with models fine-tuned via Crusoe Serverless Fine-Tuning. The post also highlights the underlying MemoryAlloyTM technology for efficient inference.

Serverless Fine-Tuning: customize open models, no GPU ops

8/11/2026

This post introduces Crusoe's Serverless Fine-Tuning service, a new capability for customizing open-source AI models. It details the technical aspects of the service, including data ingestion and pre-processing, the use of LoRA for efficient fine-tuning, automated training and recovery mechanisms, and integrated evaluation strategies like early stopping and LLM-as-judge. It also highlights the seamless integration with Crusoe's Self-Serve Deployments for production inference, emphasizing the end-to-end model lifecycle management within the Crusoe Intelligence Foundry.