Blogs›IBM›Data Prep Kit for LLM Workloads
Data Prep Kit for LLM Workloads
The Data Prep Kit (DPK) is an open-source toolkit with 20+ modules for pre-processing data for code and language. It supports ingestion, document annotation, filtering, and redaction of private information, enabling developers to build end-to-end data pipelines from ingestion to tokenization. DPK modules can scale from laptop to datacenter and integrate with frameworks like Ray and Spark, offering APIs for Python, Ray, Spark, and Kubeflow Pipelines. It has been used to produce pre-training data. Elyra extends Jupyter Notebooks with AI-centric extensions, including a visual editor for building Notebook-based AI pipelines, enabling the conversion of notebooks into batch jobs or workflows. It supports hybrid runtime environments via Jupyter Enterprise Gateway for distributed clusters like Spark and Kubernetes, and allows Python script execution. Git integration provides versioning and collaboration, and a shared configuration service simplifies workspace management. Elyra's pipeline visual editor was derived from IBM Watson Studio.