Back to feed
arXiv cs.LG
arXiv cs.LG
7/24/2026
The original title is "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators"

The original title is "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators"

Original: DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Short summary

DataPrep-Bench is the first unified benchmark evaluating LLMs as training data preparators across two tracks: data construction (transforming raw sources into supervised data) and data quality evaluation (predicting downstream training utility). The authors release Data-Construction-Skill, a skill-guided agent that lifts the Dolly-only baseline by ~20 points on Llama-3.1-8B Finance, and DAS, a distribution-based evaluator achieving the strongest cross-model correlation in 4 of 6 domains. The benchmark covers six domains and multiple base models with downstream-grounded scoring.

  • First unified benchmark for LLM-driven training data preparation covering construction and quality evaluation
  • Data-Construction-Skill agent improves Dolly-only baseline by ~20 absolute points on Llama-3.1-8B Finance
  • DAS evaluator achieves r > 0.70 in Math, Science, and Medical, outperforming existing quality and diversity metrics

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more