BlogsH2O.ai Feature Trails

H2O.ai logo

H2O.ai Feature Trails

See how major capabilities shipped, upgraded, and evolved across H2O.ai's engineering blog.

Feature trails

11

Agentic AI Platform for Telecom

Active

H2O.ai's Agentic AI platform is designed to integrate Generative AI with Predictive AI for telecommunications. It offers specialized AI agents for various telecom workflows, including contact center automation, SQL query generation, churn prediction, and billing issue resolution. The platform supports sovereign AI deployments and aims to streamline operations, boost efficiency, and enhance customer experiences. Case studies highlight significant cost reductions and ROI improvements through fine-tuning. This post extends the platform's application to public sector agencies, emphasizing sovereign AI capabilities, FedRAMP compliance, and specific use cases like citizen-facing AI agents, document summarization, flood prediction, and anomaly detection in IT infrastructure, leveraging h2oGPTe and Enterprise LLM Studio for on-premise and air-gapped deployments.

2 posts

Timeline

AI Governance Framework

Active

This post introduces H2O.ai's AI Governance Framework, a set of rules and best practices for responsible AI adoption. It outlines four stages: Organizational Planning, Use Case Planning, AI Development, and AI Operationalization, with 11 key topics to ensure consistency, transparency, and risk mitigation in AI projects.

1 post

Timeline

Document AI

Active

H2O Document AI is a system designed to extract information from various document formats using a combination of AI technologies. It encompasses optical character recognition (OCR), intelligent character recognition (ICR), and natural language processing (NLP) to identify and extract key data points such as entity names, logos, and invoice numbers. The system supports automated data labeling, model training, and deployment via REST APIs, integrating with existing workflows and H2O MLOps for monitoring. This includes the development and release of multimodal foundation models like H2OVL Mississippi, which enhance OCR and document understanding capabilities by combining vision and language processing. These models, such as H2OVL Mississippi-0.8B and 2B, are optimized for efficiency and performance, outperforming larger models on specific benchmarks and offering advanced features like dynamic resolution and multi-scale adaptive cropping.

2 posts

Timeline

H2O Danube Large Language Models

Active

H2O LLM Studio provides a no-code framework for fine-tuning state-of-the-art Large Language Models (LLMs) and Small Language Models (SLMs). It leverages Deepspeed for distributed training on GPU clusters, enabling cost-effective and faster training compared to traditional LLMs. The studio supports various fine-tuning techniques including instruction/chat fine-tuning, causal classification and regression, and DPO/IPO/KTO optimization. It also facilitates distilling LLMs into SLMs using H2O LLM Data Studio, leading to lower TCO and potentially higher accuracy for specific use cases.

3 posts

Timeline

H2O Feature Store

Active

The H2O Feature Store is a system designed to extract, manage, and optimize the feature engineering and model building process for machine learning. It provides a central repository to connect information across disparate systems, enabling reuse of data and feature engineering efforts. Key capabilities include data versioning, lineage tracking, encryption for security and privacy, operational efficiency through API integration, proactive feature recommendations, and automated monitoring for data quality and drift. It supports collaboration by promoting feature sharing across teams and ensures accuracy and reliability through backtesting, time travel, and human feedback loops. The workflow involves data ingest, transformation (feature engineering), recommendation of existing features, monitoring for data quality and drift, feature serving (batch, stream, on-demand), and consumption by business users to understand feature influence on predictions.

1 post

Timeline

H2O Label Genie

Active

H2O Label Genie is a system designed to accelerate data annotation for machine learning by leveraging AI. It supports various data types (text, images, video, audio) and deep learning tasks (classification, regression, object detection, instance segmentation, entity recognition, summarization). It can be deployed as a managed or hybrid solution and is SOC2 compliant. It integrates with H2O Hydrogen Torch for model training and H2O MLOps for deployment.

1 post

Timeline

H2O MLOps

Active

H2O MLOps provides a comprehensive platform for managing, deploying, governing, and monitoring machine learning models in production. It streamlines the end-to-end model lifecycle, from experimentation to production deployment and ongoing monitoring. Key capabilities include model management and registry, 3rd party model ingestion, collaborative experiment repositories, model versioning, various deployment modes (real-time, batch, A/B testing), multi-environment support, automated scaling, and real-time scoring. This post details the challenges and solutions for bridging the gap between AI model development in an 'AI Factory' and their successful deployment and operation in production environments within the banking sector, emphasizing the importance of robust MLOps practices for achieving this. It highlights the need for continuous integration, deployment, and monitoring to ensure models deliver value reliably and at scale.

3 posts

Timeline

H2O Sparkling Water Integration

Active

H2O Sparkling Water enables users to combine H2O's machine learning algorithms with Apache Spark's data processing capabilities. It allows for seamless integration of Spark SQL queries, H2O model building and prediction, and subsequent use of results within Spark. The system supports driving computation from Scala, R, and Python, and offers easy deployment of H2O models (POJOs/MOJOs) for scoring within any environment. It is designed to run as a regular Spark application, initializing H2O services and accessing data from both Spark and H2O data structures. The product is available for various Spark versions and cloud platforms like Azure, AWS, and Google Cloud.

1 post

Timeline

H2O Wave AI App Development Framework

Active

H2O Wave is an open-source Python development framework for building real-time interactive AI applications. It provides a rich set of UI components and charts, enabling developers to create sophisticated visualizations and dashboards without requiring HTML, CSS, or JavaScript expertise. Wave's real-time server facilitates streaming dynamic information, and it integrates seamlessly with H2O's machine learning platform and AI Cloud for model building and deployment. The framework supports deployment across various operating systems and major cloud providers, and is extensible with popular ML libraries like TensorFlow, scikit-learn, PyTorch, and Pandas.

1 post

Timeline

H2O-3 Core Platform

Active

H2O Hydrogen Torch is an extension of the H2O-3 Core Platform, providing a no-code interface for state-of-the-art deep learning model training across various data modalities (text, image, 3D image, video, audio). It democratizes deep learning by enabling users to train, tune, and deploy models without coding, integrating with H2O MLOps for production deployment and H2O Wave for application development. Key capabilities include automated hyperparameter tuning, a library of problem types, and support for multi-GPU training.

3 posts

Timeline

Foundation Models for Tabular Data

Active

This post introduces tabH2O, a unified foundation model for tabular prediction that handles both classification and regression tasks in a single forward pass using in-context learning. Key technical advancements include a dual-head architecture for unified training, single-stage pretraining with stability improvements (bounded scalable softmax, inter-stage normalization, learnable residual scaling, logit soft-capping), and noise-aware pretraining to enhance robustness to irrelevant features. This model outperforms classical ML approaches and delivers accurate predictions in seconds with no training, tuning, or feature engineering required.

3 posts

Timeline