CroissantMiner: a benchmark for end-to-end extraction of Croissant metadata from 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations
Read the original at arxiv.org→arXiv:2610.07132v1 Announce Type: new Abstract: Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of...
Original headline: "CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets"
Coverage timeline
- Oct 7, 04:00 UTC arXiv cs.CL lead source CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets