7/13/2026 · Aanchal Goyal, Maroun Touma, Shahrokh Daijavad, Hima Patel, Sophie Jin
What this post added
Introduces the open-source Data Prep Kit (DPK) with an Apache 2.0 license. DPK offers 20+ modules (transforms) for data pre-processing for AI workloads, including ingestion, annotation, filtering, and redaction. It supports laptop-scale to datacenter-scale processing and integrates with Ray, Spark, and Kubeflow Pipelines for scalable execution. The post highlights DPK's use in preparing data for RAG, fine-tuning, and instruction-tuning, and its role in IBM's watsonx.data integration.