BlogsConfluentReal-time Enterprise RAG

Real-time Enterprise RAG

Real-time Enterprise RAG

63
posts
2019–2026

This release enhances Confluent Intelligence with generally available Real-Time Context Engine featuring low-latency querying capabilities (filters, ranges, compound queries, projections, ordering) directly on live tables, eliminating the need for external databases. It also introduces GA for Streaming Agents and the Agent Management Console, providing production-ready, event-driven AI agents with enterprise-grade operations and a centralized UI for management. New ML functions include Multiva. This release expands Confluent Intelligence with Agent2Agent (A2A) Integration for Streaming Agents, Multivariate Anomaly Detection for Built-in ML Functions, Vector Search for Azure Cosmos DB and Amazon S3 Vectors, AWS and Azure Private Link for Model Inference, External Tables, and Search, and Confluent Support for the open source MCP Server for Confluent Cloud.

2026

Customer Intelligence Hub for Real-Time GTM Insights

5/27/2026

Introduces the Customer Intelligence Hub (CIH), an internal application that centralizes customer signals from various systems (Salesforce, Zendesk, product telemetry, billing, Jira, etc.) into Kafka topics. It uses Apache Flink for stream processing and enrichment, correlating cross-system signals and generating embeddings. These are then stored in OpenSearch for vector similarity search. GenAI capabilities (AccountIQ) are used to provide contextual summaries grounded in CIH data, enabling GTM teams to get a prioritized view of customer accounts, detect risks, and identify expansion opportunities. Key capabilities include an Account Center for portfolio views, a Whitespace View for expansion potential, a Real-Time Activity Feed for prioritized events, AccountIQ for GenAI-powered context, and Executive Account Summaries.

RAG and GenAI for Customer 360 Architecture

5/26/2026

This post elaborates on the Real-time Enterprise RAG feature thread by detailing the specific architecture for an AI-powered Customer 360 system. It outlines a layered approach involving a real-time data engine (event streaming with Kafka, stateful stream processing with Flink, and a customer profile store), a knowledge and embedding pipeline, and a retrieval and AI layer (RAG). It also highlights the importance of observability and governance. The post details the data flow from customer events to AI responses, emphasizing continuous event ingestion, real-time profile updates, and context-aware retrieval for guardrailed generation.

Enterprise Knowledge Management with RAG

5/25/2026

Introduces a new capability: Real-time Enterprise RAG. Details the architecture including event streaming ingestion via CDC connectors, stream processing with Flink for chunking and PII redaction, real-time embedding generation, dynamic vector store upserts/deprecations, and a retrieval layer with RBAC. Highlights the data flow from content update to LLM output, emphasizing the role of Confluent as the event backbone for document ingestion, embedding updates, and context synchronization. Discusses core capabilities like continuous embedding updates and cross-system context synchronization with specific architectural enablers and real-life examples.

Agentic Fleet Management Architecture for Real-Time Operations

5/19/2026

Introduces agentic fleet management architecture, a real-time, event-driven system where autonomous agents continuously process streaming data for autonomous operational decisions and closed-loop feedback systems. It details a layered, event-driven architecture with components like edge ingestion, an event streaming backbone (Kafka), stream processing (Flink), agent orchestration, AI/ML inference, command and control feedback, and monitoring. Key capabilities enabled include dynamic route optimization, predictive maintenance, incident detection, autonomous dispatch, energy/fuel optimization, and multi-vehicle coordination. Design principles for production-grade systems emphasize decoupled event streams, stateful processing, exactly-once guarantees, resilient failover, governance, and multi-region support.

Confluent Cloud Q2 2026: dbt adapter, Confluent Intelligence updates

5/19/2026

Introduces a managed Model Context Protocol (MCP) server for Confluent Cloud, providing a scalable and highly available interface for AI agents to discover and debug Confluent resources. Also introduces Confluent Agent Skills, which encode Confluent best practices directly into AI agents to accelerate development and improve deployment quality. Enhances Flink accessibility with a dbt adapter for Confluent Cloud, enabling data engineers to define, test, and deploy streaming pipelines using dbt workflows. Introduces Materialized Tables for simplified Flink pipeline evolution with automated offset management and job orchestration. Adds Process Table Functions (PTFs) for custom, stateful stream processing logic in Java. Enables external connectivity for user-defined functions (UDFs) to interact with external APIs and services. Announces General Availability for Snapshot Queries, unifying batch and streaming workloads by enabling point-in-time SQL queries on historical and real-time data.

New in Confluent Intelligence: Real-Time Context Engine Upgrade, New Model Support, ML Functions, and More

5/18/2026

The Real-Time Context Engine is now generally available with enhanced querying capabilities including filters, ranges, compound queries, projections, and ordering, allowing agents to access rich context without external databases. Streaming Agents and the Agent Management Console are also GA, offering production-ready, event-driven agents with enterprise-grade operations and a centralized UI for creation and management. New ML functions include Multivariate Anomaly Detection (Open Preview), PII Detection (Early Access), and Sentiment Analysis (Early Access). Support for TimesFM (EA) for time-series forecasting and native support for Anthropic and Fireworks AI models have been added.

InfiniteWatch + Confluent: Turning Customer Interaction Data into Real-Time Intelligence

5/15/2026

InfiniteWatch utilizes Confluent as the event backbone for its AI-native customer interaction intelligence platform. The architecture ingests, buffers, processes, replays, enriches, and routes customer interaction data in real-time. Key capabilities enabled by Confluent include burst absorption, durable event preservation, service decoupling, replayable intelligence for AI systems, and multi-destination routing of events to various downstream systems like enrichment pipelines, AI analysis services, alerting systems, operational workflows, and analytical stores. The system processes data through Capture, Stream, Enrich, Understand, and Act stages.

How Agent Taskflow Scaled AI Agents to 1M with Confluent and AWS

4/8/2026

Details how Agent Taskflow scaled AI agents to 1M using Confluent Cloud and AWS, demonstrating a serverless, event-driven architecture for AI orchestration with low-latency responses and high reliability, even on minimal infrastructure. It highlights the use of Confluent Cloud as the streaming backbone for agent communications and Amazon Bedrock for LLM inference, emphasizing enterprise-grade security with IAM and Confluent Schema Registry for data quality.

New in Confluent Intelligence: A2A, Multivariate Anomaly Detection, and More

2/26/2026

Introduces Agent2Agent (A2A) Integration for Streaming Agents, enabling collaboration with external agents over Kafka. Adds Multivariate Anomaly Detection using the ML_DETECT_ANOMALIES_ROBUST function in Flink SQL for detecting anomalies across multiple correlated metrics. Expands Vector Search support to Azure Cosmos DB and Amazon S3 Vectors, allowing direct querying from Flink SQL. Enables AWS and Azure Private Link for secure, private connectivity from Flink to external models, databases, and vector stores. Provides vendor-backed support for the open source MCP Server for Confluent Cloud.

Focal Systems: Boosting Store Performance with an AI Retail Operating System and Real-Time Data

2/4/2026

Focal Systems details their implementation of an AI retail operating system using Confluent's managed Kafka and Flink. They describe how they ingest retail data (shelf images, product updates, inventory, sales), process it with computer vision and ML models on GCP, and use Flink SQL to compute real-time metrics and aggregates for shelf availability, out-of-stock events, and planogram compliance. The system outputs operational data to MongoDB and analytical data to Snowflake. The post highlights challenges with high-cardinality data, scalability, and the need for managed solutions, and explains how Confluent's platform addresses these by providing an event-driven streaming backbone and stream processing capabilities.

Thunai Automates Customer Support with AI Agents and Data Streaming

1/23/2026

Thunai details how they use Confluent's data streaming platform as their backbone for an agentic AI platform focused on automating customer support. They ingest live customer conversations, CRM updates, and user interactions as streaming context to continuously update their 'Thunai Brain' knowledge base. This enables real-time AI agent actions, L1 task deflection, and real-time agent assistance. The post highlights the technical challenges of batch processing for real-time AI and how Confluent's low latency, scalability, and reliability address these. A reference architecture is provided showing data ingestion from producers to Kafka topics, consumption by agents, normalization, enrichment, and storage as vector embeddings in MongoDB Vector DB, keeping the knowledge base updated without batch processing. They utilize multiple LLMs on top of this streaming context.

2025

Gartner Recognizes Confluent's DSP for Data Integration

12/23/2025

This post details how Confluent's Data Streaming Platform (DSP) enables AI to operate directly on data in motion, contrasting with traditional data integration tools. It highlights the Real-Time Context Engine and Streaming Agents for event-driven AI, and introduces Tableflow as a feature to materialize Kafka topics as open table formats (Iceberg, Delta Lake) for seamless integration with data lakes and warehouses. This allows AI systems to access fresh, enriched context in real-time, improving decision-making. The post also discusses the unification of batch and streaming processing with Confluent Cloud for Apache Flink® and the continuous reasoning about data.

Real-Time Context for Snowflake Intelligence with Confluent

12/8/2025

This post details the integration of Confluent Intelligence with Snowflake Cortex agents, enabling them to leverage real-time business context for more accurate and situationally aware insights and actions. It highlights the use of Tableflow to ingest Kafka data into Snowflake as Iceberg tables, Streaming Agents to build event-driven AI agents that interact with Cortex, and the Real-Time Context Engine to serve fresh business context to Cortex agents via MCP.

Detecting the Unexpected: Built-in Real-Time Anomaly Detection in Confluent Cloud for Apache Flink®

11/13/2025

Introduces built-in anomaly detection capabilities in Confluent Cloud for Apache Flink using the `ML_DETECT_ANOMALIES` function, which leverages ARIMA models. This function supports auto-ARIMA parameter determination, seasonal pattern detection via STL, adaptive training windows (`minTrainingSize`, `maxTrainingSize`, `updateInterval`), and confidence-based detection (`confidencePercentage`). The post details real-world use cases and provides an example SQL query for its implementation.

Confluent Cloud Q4 2025: Data in Motion for AI

10/29/2025

Phase 2 of Streaming Agents introduces Agent Definition (Open Preview) for simplified agent building and dynamic tool calling, Observability and Debugging (Open Preview) for immutable logging of agent interactions, and Real-Time Context Engine (Early Access) for a managed serving layer materializing streaming data into a low-latency cache. Confluent welcomes Airy to enhance Flink functions for AI agents.

Faster, Smarter, More Context-AwareStreaming Agents

10/29/2025

Introduces new capabilities for Streaming Agents including simplified agent definition with iterative tool calling, enhanced observability and debugging through structured immutable logs and Flink-powered recovery, and deeper integration with the Real-Time Context Engine for secure, reliable, and scalable delivery of live structured data to AI agents via MCP. Details new integrations with AWS, Google Cloud, Azure, various vector databases, and system integrators. Highlights use cases such as real-time fraud prevention, dynamic customer support, and intelligent supply chain optimization.

Introducing Real-Time Context Engine for AI

10/29/2025

Introduces the Real-Time Context Engine in Early Access, a new serving layer for streaming data that materializes enriched enterprise data sets into a fast, in-memory cache and serves them to AI systems via Model Context Protocol (MCP). This service is fully managed within Confluent Cloud and aims to solve the last-mile problem of serving real-time, contextualized data to AI applications. It unifies historical replay, continuous processing, and real-time serving on the Confluent Data Streaming Platform, abstracting Kafka and Flink complexity and providing built-in security, governance, and RBAC.

How Real-Time Streaming Prevents Fraud in Banking & Payments

9/30/2025

This post details how real-time streaming and stream processing can be applied to fraud detection in banking and payments, enabling immediate anomaly detection, transaction scoring, composite event identification, and automated fraud responses. It highlights the shift from batch processing to real-time analysis, the impact on business outcomes like financial savings and operational efficiency, and the importance of trust, governance, and scalability for financial-grade alerting.

How to Build Real-Time, Event-Driven Automation for Business Insights

9/30/2025

This post details how to build real-time, event-driven automation for business insights using event-driven workflows and AI agents. It outlines the architectural components, including event brokers (Apache Kafka) and processing layers (Apache Flink/Kafka Streams), and provides best practices for implementation such as using Schema Registry, testing under real load, and planning for anomalies. It also presents real-world use cases like fraud detection, customer churn prevention, and inventory optimization, emphasizing the benefits of immediate responsiveness, scalability, and personalization.

Why Enterprise AI Runs on Data Streaming

9/18/2025

This post argues that data streaming is the foundational technology for enterprise AI, enabling AI use cases by providing real-time, trustworthy, and accessible data. It contrasts traditional batch data processing with data streaming's advantages for AI, such as instant actionability, early data quality issue detection, and architectural simplification. The post highlights how data streaming shifts the data strategy upstream to operational estates and supports trends like becoming a software company and the blurring of operational and analytical data estates with technologies like Tableflow. It provides real-world examples of AI powered by streaming data in fraud detection, retail, and manufacturing, and introduces the Data Streaming Organization (DSO) framework to guide enterprise adoption of data streaming for AI.

Confluent Cloud Q3 '25 Unlocks Cost-Effective Data Streaming

8/19/2025

Introduces Streaming Agents in Open Preview, enabling the building, deployment, and orchestration of event-driven AI agents directly within stream processing pipelines on Apache Kafka and Apache Flink. Key features in preview include tool calling with MCP, Connections for secure and reusable integrations with models and databases, and external tables/search for data enrichment. GA features include AI model inference and embeddings within Flink SQL, and built-in ML functions for time-series forecasting and anomaly detection. Also introduces Custom Single Message Transforms (SMTs) for fully managed connectors, allowing users to upload and manage their own transformation logic. Expands Tableflow functionality with Delta Lake and Databricks Unity Catalog integration, including automated schema evolution and table maintenance, and upsert materialization for CDC. Cluster Linking is now available on Google Cloud.

Introducing Streaming Agents on Confluent Cloud

8/19/2025

Introduces Streaming Agents on Confluent Cloud, enabling event-driven AI agents natively on Apache Flink. This feature unifies stream processing and AI workflows, allowing agents to monitor and act on real-time business events. Key technical contributions include: event-driven architecture with Flink and Kafka for asynchronous processing and replayability; seamless integration with LLMs, tools, and data systems via Model Context Protocol (MCP) and native connectors; real-time context access for improved decisioning; built-in ML functions for forecasting and anomaly detection; secure and governed event flows with replayability for auditing and debugging; and enhanced external tables and vector search for data enrichment. The post details how Streaming Agents leverage Flink APIs for development and deployment, enabling engineers to build context-aware automation.

Modernize Legacy Data for Real-Time AI with Oracle & MongoDB

8/5/2025

Details a use case for modernizing legacy data (Oracle) for real-time AI applications by streaming changes to Kafka, enriching and embedding data with Flink and LLMs, and sinking embeddings to MongoDB for semantic search and personalization. Demonstrates building real-time AI solutions with event-driven architecture, even for non-cloud-native organizations. Includes Flink SQL examples for generating product descriptions using GPT-4 and generating vector embeddings using OpenAI models.

Transforming Confluent Operations With GenAI

7/18/2025

This post details the integration of NeuBird's Hawkeye, a GenAI-powered SRE assistant, with Confluent Cloud's observability stack. It describes how Hawkeye automates incident investigation and resolution by correlating Prometheus metrics, Confluent Cloud metrics, producer logs, and CloudWatch audit logs. The technical architecture includes Kubernetes deployment (Amazon EKS), Prometheus Alertmanager integration for incident triggering, and secure, read-only access for Hawkeye to telemetry sources. A specific scenario of authorization revocation is used to illustrate the traditional manual troubleshooting workflow versus the automated Hawkeye workflow, highlighting reduced MTTR.

Why Flink Agents Are the Future of Enterprise AI

7/3/2025

Introduces Flink Agents, a new Flink sub-project developed collaboratively with Alibaba, to bridge the gap in building enterprise AI agents. Key contributions include: 1. First-class agent semantics within Flink APIs for model inference, tool invocation, and contextual search. 2. Dynamic topology support for complex reasoning patterns like ReAct workflows, moving beyond traditional sequential data processing. 3. Enhanced observability for agent state, tool invocations, model inference calls, and decision traces. 4. Native Model Context Protocol (MCP) support for seamless integration with MCP-compatible tools and services. The post also highlights the developer experience with Flink's Table API for agent creation and application to data streams, emphasizing unified infrastructure, end-to-end consistency, fault tolerance, and replayability.

The Future of AI Agents Is Event-Driven | Confluent

5/13/2025

This post argues for Event-Driven Architecture (EDA) as the foundational infrastructure for scaling AI agents. It highlights the limitations of fixed workflows and LLMs, positioning agents as dynamic, context-driven systems. The core technical argument is that EDA provides the necessary loose coupling and asynchronous communication for agents to access data, use tools, and share information across systems without creating rigid dependencies, drawing parallels to the evolution from monoliths to microservices and the challenges of scaling early social networks. It emphasizes how EDA enables agents to function as independent microservices with informational dependencies, allowing for seamless integration into broader workflows and systems like CRMs and CDPs.

Building Real-Time Multi-Agent AI With Confluent

4/23/2025

This post details how Agent Taskflow leverages Confluent's data streaming platform to build a real-time multi-agent AI orchestration system. It emphasizes the event-first architecture where every agent action, thought, and decision is an event first, flowing through Kafka topics. Key technical aspects include using Kafka for multi-agent communication, enabling real-time interactivity and context sharing. Observability is achieved through replayable logs and per-event tracing. Fault tolerance and scalability are addressed by retrying failed steps and scaling agents independently. Identity and permissioning are managed by ensuring agents are aware of data access and action capabilities. The post also details the integration of Confluent connectors (PostgreSQL Sink, Iceberg Sink, custom webhook source) and Stream Governance features like Schema Registry for maintaining data quality and compatibility.

The AI Silo Problem: How Data Streaming Can Unify Enterprise AI Agents

4/3/2025

This post proposes a novel architecture for unifying enterprise AI agents by leveraging data streaming platforms as an event-driven backbone. It introduces the concepts of an agent registry for structured governance, dedicated Kafka topics for seamless event exchange between agents, and stream processing with LLM-based event mapping for intelligent orchestration. This approach aims to solve the 'AI silo' problem by enabling real-time communication and coordinated action among heterogeneous AI agents.

Real-Time Toxicity Detection in Games: Balancing Moderation and Player Experience

3/14/2025

Details a real-time AI/ML-based system for detecting and mitigating toxic messages in games, leveraging Confluent's data streaming platform and Databricks for advanced analysis. The system uses Apache Flink for initial triage and Databricks for deeper context-aware NLP, with Tableflow facilitating data integration. It emphasizes balancing moderation with player experience by processing the majority of messages with low latency and escalating complex cases for deeper analysis or human review, with continuous model refinement based on feedback.

Data-Driven Business Agility: Adapting to Market Changes in Real Time

3/10/2025

This post discusses how real-time data streaming platforms enable data-driven business agility, allowing organizations to respond rapidly to market changes and make informed decisions. It highlights the role of Apache Kafka and Apache Flink in capturing and analyzing real-time data, and presents Confluent Cloud as a fully managed DSP that supports data mesh principles, self-service infrastructure, and hybrid deployments for mission-critical use cases like real-time transaction processing and fraud detection. The post also emphasizes fostering a culture of agility, breaking down data silos, and ensuring data governance and security for successful real-time data implementation.

Your GenAI Project Needs a Data Streaming Platform

2/26/2025

This post addresses the 'data liberation problem' for Generative AI projects, arguing that traditional batch processing and centralized data architectures (ELT, reverse ETL) create data silos and delays that hinder real-time AI applications. It proposes data streaming platforms as a solution, enabling real-time data access, decoupled architectures, making unstructured data AI-ready, improving data quality, unifying ecosystems, and supporting scalable AI interactions. The post emphasizes a shift-left approach where data processing occurs closer to the source, ensuring data is fresh, accessible, and actionable for AI.

Building AI Agents and Copilots with Confluent, Airy, and Apache Flink

2/20/2025

This post details how Airy leverages Confluent Data Streaming Platform, Apache Flink, and AI Model Inference to build AI agents and copilots. It highlights the use of Confluent connectors for data ingestion, Flink for stream processing and data transformation, and AI Model Inference for creating embeddings and RAG. The post also mentions the use of Stream Governance for data contracts and schema management, and Tableflow for integrating with Iceberg. Key challenges addressed include data governance, accuracy of GenAI outputs, and scalable infrastructure.

Optimizing Supply Chains with Data Streaming and Generative AI

1/16/2025

This post details a specific implementation of a computer vision-aided AI stock monitoring solution for supply chain optimization. It describes how data streaming with Confluent, Azure AI Vision, Apache Flink, vector embeddings, Azure Cosmos DB, and OpenAI's LLM can be integrated to provide real-time guidance to shop assistants. The architecture involves streaming product photos to Azure AI Vision for classification, joining the results with real-time inventory data streamed from SQL Server, generating vector embeddings using Flink AI Model Inference, storing them in Azure Cosmos DB for RAG, and then using OpenAI to generate actionable advice for stock management.

2024

Predictive Analytics: How Generative AI and Data Streaming Work Together to Forecast the Future

12/20/2024

This post details the integration of generative AI and data streaming for predictive analytics. It explains how data streaming continuously feeds new information to generative AI models, which then refine predictive models and generate future scenarios. The post highlights Confluent's platform capabilities in enabling real-time data collection, processing, and AI model inference with Apache Flink for tasks like classification, text generation, and clustering, thereby enhancing predictive analytics and decision-making.

The Power of Predictive Analytics in Business: Using Generative AI and Confluent

12/20/2024

This post details how Confluent's real-time data streaming platform, combined with Generative AI, enhances predictive analytics capabilities. It explains how this integration allows businesses to move from reactive problem-solving to proactive opportunity seizing by analyzing historical data and real-time signals to forecast future trends and outcomes. The post highlights specific use cases in financial services, e-commerce, and air travel, emphasizing the role of data streaming in enabling real-time insights for AI-driven forecasting, scenario simulation, and operational optimization. It also touches upon business observability and dynamic supply chain optimization as key benefits.

The Power of Predictive Analytics in Healthcare: Using Generative AI and Confluent

12/20/2024

This post details how Confluent's data streaming platform can be leveraged for predictive analytics in healthcare, particularly when combined with generative AI. It highlights the importance of real-time data integration for AI-driven predictions, enabling faster clinical decisions, improved patient outcomes, reduced costs, and enhanced operational efficiency. Specific use cases such as fraud detection, streamlined claims processing, and vaccine distribution are discussed, emphasizing Confluent's role in creating responsive AI models through its Apache Kafka-based frameworks.

Let Flink Cook: Mastering Real-Time Retrieval-Augmented Generation (RAG) with Flink

9/4/2024

This post introduces the use of Apache Flink and Kafka in Confluent Cloud to build Retrieval-Augmented Generation (RAG) applications. It details how Flink SQL's new ML functions (`ml_predict()`, `federated_search()`) can be used for data preparation (chunking, vector embedding) and inference. It also discusses advanced RAG patterns like reasoning, workflows, and post-processing for reliability, emphasizing the benefits of an event-driven architecture for AI systems.

Using Kafka-Powered AI Models to Predict and Prevent Sepsis at City of Hope

7/16/2024

This post details City of Hope's implementation of a Kafka-powered AI system for sepsis prediction and prevention. It highlights the use of real-time data streams from Epic EHR, processed through Kafka topics (vitals, lab results, ADT, etc.) and microservices, to feed AI models. The system ensures timely predictions by bypassing traditional batch processing, with predictions fed back into Epic to trigger clinical workflows. Confluent Cloud on Azure is utilized for its scalability, reliability, and observability, enabling quick identification and resolution of data anomalies. The immutable log feature of Kafka is leveraged for uninterrupted processing during model crashes.

Build a Scalable and Up-to-Date Generative AI Chatbot with Amazon Bedrock and Confluent Cloud

6/24/2024

This post details the architecture and implementation of a Generative AI chatbot using Confluent Cloud and Amazon Bedrock. It showcases how Confluent's real-time event streaming capabilities and microservices architecture enable the chatbot to stay up-to-date with the latest AI advancements. Key technical components discussed include the KStream API for real-time processing, KCache for dynamic cache updates within the chatbot conversation processing microservice, and the use of Amazon Bedrock for AI model inference and prompt engineering to generate pre-approval forms. The post provides code snippets for the chatbot conversation processing and pre-approval orchestration, demonstrating the practical application of these technologies.

Next-Gen Customer Loyalty Programs with Data Streaming | Confluent

6/11/2024

This post details how Confluent's data streaming platform can be used to build next-generation customer loyalty programs by leveraging real-time data for personalized experiences and automated decision-making. It highlights the use of Flink for stream processing to join, enrich, and transform data in real-time, enabling instant reward calculations and personalized offers. The post also discusses the integration of AI/ML models for predictive analytics and real-time recommendations, leveraging Confluent's AI Model Inference feature in Flink.

Data Streaming in Healthcare: Achieving the Single Patient View

4/30/2024

This post details how Confluent's data streaming platform, including Apache Flink, can be used to build a Single Patient View (SPV) in healthcare, enabling real-time experimental treatment-matching services. It highlights the challenges of interoperability, regulatory compliance, scalability, observability, and disaster recovery in healthcare data integration and demonstrates how a managed, serverless data streaming platform addresses these issues, including a case study of a healthcare organization using Confluent Cloud for a treatment-matching service.

Making Predictive Customer Support a Reality for Telcos

3/26/2024

This post details how telcos can leverage data streaming for predictive customer support, ingesting and processing real-time data from customer behavior and network performance to train predictive models for preemptive issue detection and resolution. It highlights the use of Confluent's platform for connecting, processing, and governing data streams to build live data products for AI/ML models, enabling proactive network health monitoring and issue resolution. Specific technical details include using Confluent Platform and Confluent Cloud with connectors for IBM MQ, Splunk, PostgreSQL, and SQL Server, cluster linking for hybrid/multicloud deployments, Apache Flink for stream processing to enrich data streams (e.g., joining call drop rates with network performance), and Stream Governance features like Stream Lineage, field-level encryption, RBAC, and Schema Registry. An example Flink SQL query is provided for identifying unhealthy network situations based on call success rates and clearing codes.

Building Event-Driven GenAI Applications in 4 Steps

2/8/2024

This post details building GenAI applications using event-driven patterns, focusing on data augmentation (chunking, embeddings, vector stores), inference, workflows, and post-processing. It emphasizes how data streaming platforms integrate disparate data sources, manage LLM calls asynchronously, and decouple components for scalability and agility, directly supporting the RAG pattern for contextualized LLM prompts.

Real-Time Order Management: The Key to Streamlining Your Business Operations

1/16/2024

This post details how to leverage Confluent for real-time order management, focusing on data accuracy, order authenticity, and fraud detection using stream processing and connectors. It highlights the use of Flink for stream processing and provides a streaming architecture diagram and ksqlDB/Flink queries for enriching order data, validating against inventory, and sending notifications. The business impact includes reimagined customer excellence, ensured order integrity through Stream Governance, and real-time inventory replenishment.

2023

Modern data streaming platforms bring better customer travel experiences

12/19/2023

This post details how cruise lines are leveraging data streaming platforms, built on Apache Kafka, to provide real-time customer experiences. It highlights the challenges of legacy systems and limited bandwidth on ships, and describes an architecture involving CDC-based connectors, Kafka, Kafka Streams, and Confluent for Kubernetes on ships, and Confluent Cloud or self-managed clusters on shore. The solution enables real-time reservation updates and customer data synchronization across fleets and shore-based systems, transforming customer experiences in the travel industry.

Delivering Real-Time Manufacturing Predictive Maintenance

12/5/2023

This post details the application of Confluent Cloud and Apache Kafka for real-time predictive maintenance in engine manufacturing. It explains how manufacturing execution system (MES) data, including machining and testing results, can be streamed to Confluent Cloud. The post outlines a pattern for ingesting granular OT data, processing it with stream processing engines (Kafka Streams, ksqlDB, Flink) for both instant quality reactions and training ML models. It provides example ksqlDB queries for continuous learning and real-time scoring, emphasizing the benefits of reduced scrap, rework, and proactive maintenance. It also mentions the use of Confluent's connectors for integrating with downstream systems.

Data Portal for Confluent Stream Governance

12/4/2023

Introduces Data Portal, a self-service UI for discovering, exploring, and accessing Kafka topics on Confluent Cloud. It leverages Stream Catalog and Stream Lineage for data discovery and understanding. Features include searching topics by metadata, requesting access via an approval workflow, and querying data with Flink SQL. This enhances the user experience for data users interacting with Kafka data streams.

Real-Time Risk Analysis with Stream Processing

11/9/2023

This post introduces the capability of providing real-time risk analysis by leveraging data streaming. It details how to ingest data from various sources continuously, throughout the day, including portfolio positions, trading activity, market data, and news feeds, immediately as something happens. The core problem addressed is enabling rapid data access for risk managers to proactively assess uncertainties and take decisive action to effectively manage exposures, moving away from slow, batch-based risk analysis.

Generative AI: An Introduction to GPT, LLMs, and More

10/24/2023

This post introduces the application of Generative AI and LLMs within an enterprise context, leveraging the Kappa Architecture and Kora engine for context, memory, private data handling, and scaling. It details how stream processing, durable storage, and sink connectors to vector databases enable retrieval-augmented generation (RAG) for personalized and secure AI interactions.

Real-Time Demand Forecasting with Confluent

10/10/2023

This post introduces a demand forecasting capability leveraging data streaming to predict future customer demands accurately, enabling efficiency savings and market opportunity identification. It details how Confluent Cloud, along with Elasticsearch, Snowflake, and MongoDB, forms a data pipeline for ingesting, processing, and analyzing event data to provide real-time insights for various industries. The post highlights the migration from open-source Kafka to Confluent Cloud for scalability, reliability, and simplified operations, including eliminating patching and upgrades. The architecture involves ingesting event data into Confluent topics and then sinking it to Elasticsearch for real-time analysis, Snowflake for data warehousing, and MongoDB for semi-structured data storage. Snowpipe is used to automate data loading into Snowflake. The outcomes include data aggregation, enrichment, and standardized, scalable data delivery.

In-store Personalization with Confluent Data Streaming

10/6/2023

This post introduces strategies for enabling real-time personalization in brick-and-mortar retail environments by leveraging data streaming. It details the requirements for such a system, including high availability, scalability, and low latency, and outlines an implementation example that uses geofencing and app integration to deliver personalized offers and information to customers in-store. The solution emphasizes the use of data streaming platforms like Kafka for real-time data integration, stream processing, and message delivery to power these personalized experiences.

The Role of Real-Time Interoperability to Healthcare Payers

10/3/2023

This post details how real-time data streaming with Confluent can be used to address technical challenges in healthcare payer interoperability, specifically by overcoming data silos and legacy communication methods. It outlines an example architecture using Kafka Produce-Consume and CDC patterns to enable real-time data availability for downstream systems and microservices, ultimately improving care coordination, efficiency, and member experience.

AI Data Streaming: Building AI, ML, and LLM Apps Just Got Easier

9/26/2023

This post details Confluent's 'Data Streaming for AI' initiative, emphasizing the use of data streaming as a foundation for modern AI stacks. It highlights how Confluent enables real-time data integration for AI/ML and GenAI applications by connecting to diverse data sources, creating AI-ready streams with Apache Flink (preview of OpenAI API calls in Flink SQL), and sharing governed data with downstream AI applications. It specifically mentions partnerships with MongoDB, Pinecone, Rockset, Weaviate, and Zilliz for vector search capabilities and Retrieval Augmented Generation (RAG). The post also introduces the Confluent AI Assistant (private preview) for natural language interaction with the Confluent platform and mentions reference architectures with SI partners.

Building Event-Driven Architectures: Top 5 Best Practices

9/19/2023

This post introduces two patterns for integrating Confluent Cloud with AWS Lambda to build event-driven applications: the fully managed AWS Lambda Sink Connector and native event source mapping (ESM). It details the benefits and considerations of each pattern, including throughput, latency, scaling, and error handling. Additionally, it provides five best practices for building event-driven architectures with Confluent and Lambda: batching calls to Lambda, implementing robust error handling in Lambda code to avoid duplicates and stalled partitions, using Lambda functions idempotently, and keeping long-lived connections to Kafka outside the handler for better efficiency.

Real-time recommendations for mobile phones using data streaming

8/31/2023

This post details how a company built a real-time personalization system for mobile phones using Confluent Cloud, leveraging Kafka for data streaming, ksqlDB for real-time data enrichment and stream-table joins, and managed connectors for data ingestion and egress. It highlights the challenges of traditional systems in handling massive data volumes and delivering personalized content at scale, and how Confluent Cloud's capabilities like fanout, stream processing, and managed infrastructure addressed these issues. The post also touches upon Stream Governance features like Stream Lineage for debugging and understanding data flow.

Real-Time AI: Live Recommendations Using Confluent and Rockset

8/9/2023

This post details how Confluent and Rockset enable real-time AI applications by providing low-latency data streaming and vector search capabilities. It explains the architecture pattern of using Confluent for data streaming and Rockset for vector search and real-time analytics. It highlights the challenges of stale data in AI models and how this pattern addresses them. The post uses Whatnot's e-commerce platform as a case study, showcasing how they leverage Confluent Cloud and Rockset for real-time recommendations based on user activity and livestream data.

Maximizing the Power of AI with Data Streaming

7/14/2023

This post details how data streaming, powered by Apache Kafka, fuels AI by enabling continuous training on data streams, efficient data synchronization to ML platforms, real-time application of AI models, and high-volume data processing. It emphasizes the creation of a real-time data mesh to democratize data accessibility and reduce integration complexity. The post also highlights the importance of data governance and security, mentioning RBAC and ABAC for secure and scalable data access to AI systems.

2022

Build Streaming Data Pipelines Visually with Stream Designer

10/4/2022

Introduces Stream Designer, a visual interface for building, testing, and deploying streaming data pipelines natively on Kafka. It translates pipeline logic into ksqlDB code, integrates with Confluent connectors and Kafka topics, and allows switching between a graphical canvas and a SQL editor. It also supports importing existing ksqlDB pipelines and exporting them as SQL code.

Getting Started with Stream Processing: The Ultimate Guide

8/11/2022

This post introduces ksqlDB as a simplified stream processing solution, detailing its benefits for building real-time applications. It explains how ksqlDB simplifies architecture by handling event capture, transformations, aggregations, and materialized views within a single model, reducing infrastructure complexity. Technical use cases like streaming data pipelines, materialized caches, and event-driven microservices are explored. Key ksqlDB constructs such as Streams and Tables, Persistent Queries, and Push and Pull Queries are described, highlighting their role in processing data in motion and serving queries against it.

2021

Real-Time Data Streams in the Cloud, On-Prem, or Both with Confluent

6/22/2021

This post introduces a tutorial demonstrating an end-to-end data pipeline using Confluent Cloud and Confluent Platform. It covers setting up Confluent Cloud and MongoDB Atlas, using the Datagen connector to stream sample data into Kafka, joining data sets into a combined stream using ksqlDB, and streaming the results into MongoDB Atlas. It also shows the same pipeline in a self-managed deployment using Confluent Platform and MongoDB Server, and uses Confluent Replicator to bridge the two environments.

Real-Time Analytics with Apache Kafka and Pinot

3/9/2021

This post details how Apache Pinot integrates with Kafka for real-time analytics, covering data ingestion, partitioning, replication, indexing, and query processing. It explains Pinot's architecture for handling high-throughput analytical queries on streaming data, including mutable and immutable segment management, and strategies for optimizing query performance through data partitioning.

2020

Using Stream Processing to Prevent Fraud and Fight Account Takeovers

4/9/2020

This post details building a fraud prevention system using Kafka Streams to analyze website login attempts in real-time, detecting new devices, new locations, botnet agents, and brute-force attacks. It leverages Kafka Streams' DSL for stream processing, including KStream and KTable operations, joins, filtering, and windowing. The post also discusses integration with other systems using Kafka Connect and provides a deep dive into the implementation of device recognition, including topology design and event mapping.

2019

Building an Event-Driven Stock Platform at Euronext

10/8/2019

Euronext has built a brand new market infrastructure and event-driven trading platform called Optiq with Confluent Platform at the core. They replaced their market data gateway with one that handles billions of messages per day, sending market data to vendors and trading members. Confluent Platform also enables them to build applications that interface with clearinghouses, monitor market latency, perform replication for disaster recovery, and store records in a data warehouse in compliance with regulatory requirements. The platform provides a tenfold increase in capacity to ingest messages and an average performance latency as low as 15 microseconds.