PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect…
PRISM indexes and ranks — it never republishes. The full piece lives with its author on arxiv.org.
Read on arxiv.org