H2O Danube Large Language Models
Open-weight H2O Danube3 Series

Open-weight H2O Danube3 Series

6/24/2026

What this post added

This post introduces the H2O Danube3 series of open-weight large language models. It details the training process involving three stages with varying data mixes and token counts (4.6T, 1.35T, and 0.05T tokens). The models are trained on approximately 6 trillion tokens using ~100 H100 GPUs and H2O LLM Studio. Performance highlights include Danube3-4B scoring over 80% on the 10-shot HellaSwag benchmark, surpassing AppleLLM OpenELM-3B-Instruct and competing with Microsoft Phi3 4B. Danube3-.5B outperforms Alibaba Qwen2-.5B and Apple OpenELM-.5B Instruct in 7 out of 12 academic benchmarks. The post also discusses applications such as cost-efficient on-device processing, enhanced privacy through local data processing, AI content detection, and guardrail LLMs for GenAI safety. The H2O AI Personal GPT mobile app is presented as an example application.

Read the original post ↗