Data Prep Kit for LLM Workloads
Unleash the potential of LLMs through the Data Prep Kit

Unleash the potential of LLMs through the Data Prep Kit

7/13/2026 · Aanchal Goyal, Maroun Touma, Shahrokh Daijavad, Hima Patel, Sophie Jin

What this post added

Introduces the open-source Data Prep Kit (DPK) with an Apache 2.0 license. DPK offers 20+ modules (transforms) for data pre-processing for AI workloads, including ingestion, annotation, filtering, and redaction. It supports laptop-scale to datacenter-scale processing and integrates with Ray, Spark, and Kubeflow Pipelines for scalable execution. The post highlights DPK's use in preparing data for RAG, fine-tuning, and instruction-tuning, and its role in IBM's watsonx.data integration.

Read the original post ↗