BlogsDecagonOpen Source LLM Productionization

Open Source LLM Productionization

Open Source LLM Productionization

3
posts
2026

Decagon is increasingly leveraging open-source Large Language Models (LLMs) for production workloads, particularly for customer service AI agents. This strategy is driven by the need for low-latency, highly specialized models that can be customized for specific tasks. The company's experience suggests a maturity curve for AI use cases, where early-stage applications benefit from general-purpose frontier models, while mature, high-volume use cases transition to fine-tuned open-source models for improved performance and cost-efficiency. Decagon Labs is developing a network of specialized models trained in-house, outperforming foundation models on their real-world use cases, focusing on precision, speed, and reliability for enterprise CX. This includes custom model architectures for functions like speech end detection, workflow execution, and hallucination detection, enabling low latency and high accuracy. The team is also developing frontier post-training techniques and agent architectures, with a current focus on voice agents and expanding proprietary model stack investments in evaluation, training infrastructure, and research.

2026

Everyone is wrong about open source AI in the enterprise | Decagon

7/9/2026

This post details Decagon's strategic shift towards using open-source LLMs for approximately 90% of its production workloads, driven by the critical requirement for low-latency models in customer service AI agents. It explains that the primary motivation is not cost, but the necessity of fine-tuning smaller models for specific tasks, which is more feasible with open-weight models. The post contrasts this with the broader enterprise trend of increasing reliance on closed-source frontier models for new, less mature use cases. It posits that as enterprise AI use cases mature, there will be a subsequent migration towards fine-tuned open-source models, though this process is expected to take considerable time due to the effort and expertise required for fine-tuning.

Introducing Decagon Labs | Decagon

3/24/2026

This post introduces Decagon Labs, the research and agent orchestration arm responsible for developing specialized, in-house trained models that outperform foundation models for enterprise customer experience. It highlights the architectural shift towards a network of specialized models for distinct functions (e.g., speech end detection, workflow execution, hallucination detection) to achieve low latency and high accuracy. The post emphasizes that research directly translates to production impact, with models powering millions of customer interactions. It also announces four accompanying deep dives into specific technical areas: off-policy training, prompt engineering optimization (GEPA), Bayesian VAD for turn detection, and reranker optimization for low-latency AI agents.

AI agents are never done: The new build-vs-buy calculus | Decagon

2/12/2026

This post elaborates on the continuous engineering effort required for AI agents beyond initial production deployment. It details the technical challenges and trade-offs involved in building and maintaining AI agents, including fine-tuning models, optimizing latency, ensuring reliability, evaluating providers, and implementing safety guardrails. It frames the build-vs-buy decision within the context of ongoing engineering investment, highlighting the costs associated with engineering time, performance gaps, operational risks, and infrastructure maintenance. The post emphasizes the importance of owning differentiating layers like workflows and domain logic while offloading infrastructure concerns. It critiques service-heavy vendor models and advocates for a product-first approach that enables direct user control and visibility into agent performance, referencing Decagon's specific product features that support continuous iteration and optimization.