A Bitter Lesson for Data Filtering

(arxiv.org)

5 points | by nujan_dev 7 hours ago ago

1 comments

  • nujan_dev 7 hours ago

    Large models don't just tolerate noisy or nominally "low-quality" web data but this paper suggests that they benefit from the distributional entropy in it. Over-curated datasets artificially compress the representation manifold, starving the model of the edge cases needed for robust generalization.