← Back to match
Data-Juicer
Data processing toolkit for cleaning, synthesizing and analyzing data for foundation models.
Desktopfree
Data-Juicer is an open-source data processing toolkit for preparing training data for foundation models, offering over 200 operators to clean, deduplicate and curate text, image, audio and video data. Pipelines are defined as reusable YAML recipes that can run from a laptop up to large Ray clusters, with a CLI, Python API and JupyterLab playground. It is aimed at ML engineers and data scientists building datasets for pre-training, fine-tuning or RAG.
Categories
Data PipelineAI InfrastructureData Processing

