Matchboxmatchbox
← Back to match

Data-Juicer

Data processing toolkit for cleaning, synthesizing and analyzing data for foundation models.

Desktopfree

Data-Juicer is an open-source data processing toolkit for preparing training data for foundation models, offering over 200 operators to clean, deduplicate and curate text, image, audio and video data. Pipelines are defined as reusable YAML recipes that can run from a laptop up to large Ray clusters, with a CLI, Python API and JupyterLab playground. It is aimed at ML engineers and data scientists building datasets for pre-training, fine-tuning or RAG.

Categories
Data PipelineAI InfrastructureData Processing

Full match profile

Behind the summary, Matchbox keeps a richer profile of Data-Juicer - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether Data-Juicer (or something else) actually fits.