Hugging Face, a leader in open-source machine learning, has partnered with Argilla to launch the 'Data Is Better Together' (DIBT) initiative. This collaborative effort is designed to empower the ML community by fostering the creation of impactful datasets that can enhance model training and evaluation.
The initiative kicked off with a major contribution: a dataset of 10,000 prompts ranked by quality, which has rapidly attracted participation from over 385 community members within just a few days. Dubbed the DIBT/10k_prompts_ranked dataset, it is already being utilized to train new models, including the noteworthy SPIN model, which effectively leverages these ranked quality prompts.
Recognizing the gaps in language representation within the ML community, the initiative has also introduced the Multilingual Prompt Evaluation Project. This project aims to improve language representation by translating 500 high-quality prompts into multiple languages, including Dutch, Russian, and Spanish.
Key Points
- The DIBT/10k_prompts_ranked dataset has enabled the growth of innovative models aimed at addressing various challenges in machine learning.
- Community-driven efforts will extend toward building domain-specific datasets, ensuring that various niches are represented adequately.
- The initiative provides essential tools and documentation to facilitate participation, reinforcing community collaboration.
- This project highlights the importance of inclusivity in benchmarks, shedding light on the inadequacies in representation of diverse languages and domains within the open-source ML arena.
The 'Data Is Better Together' initiative not only aims to produce effective datasets but also facilitates a culture of collaboration in the machine learning ecosystem. The community is encouraged to partake actively in building datasets that address real-world needs and underrepresented languages.
For more details on the initiative, visit the official announcement at Hugging Face Blog.




