Hugging Face has launched the 'Data Is Better Together' initiative in collaboration with Argilla, marking a significant step towards improving dataset quality in the open-source machine learning (ML) community. This initiative has quickly attracted over 385 contributors, leading to the successful release of the DIBT/10k_prompts_ranked dataset, which is set to assist in building advanced ML models.
Key Details
- The primary aim is to create a dataset of 10,000 prompts, combining synthetic and human-generated inputs to refine prompt ranking tasks and improve synthetic data generation.
- The DIBT/10k_prompts_ranked dataset is already being utilized in developing new models such as SPIN.
- To address language disparities, the initiative has initiated the Multilingual Prompt Evaluation Project, focusing on translating 500 high-quality prompts into multiple languages, including Dutch, Russian, and Spanish.
- As part of the project, various community tools and guides have been created to empower contributors in building their domain-specific datasets.
- The overarching goal is to tackle existing inequalities in dataset representation across different languages, domains, and tasks in the open-source landscape.
By fostering a collaborative environment, Hugging Face aims to not only improve the quality of datasets but also ensure diverse representation, which is crucial for developing fair and reliable ML models. With the release of the initial dataset and ongoing community engagement, this initiative represents a significant stride toward an inclusive and effective open-source ML ecosystem.




