7/15/2026
What this post added
This post introduces the integration of Thinking Machines Lab's Inkling model, a 975B-parameter multimodal (text, image, audio) open-weight model with a mixture-of-experts architecture, onto the Baseten Platform. It details the technical challenges of serving such a large model, including its significant infrastructure footprint (2TB+ GPU memory for BF16, 600GB for NVFP4). The post highlights how the Baseten Inference Stack, specifically the Baseten Delivery Network for fast cold starts and Multi-cloud Capacity Management for global GPU pooling, enables day-0 support and reliable, scalable serving of Inkling. It also mentions optimizations in vLLM by Inferact for Inkling's deployment.