
OpenThoughts
Fully open data curation for reasoning models
Curating the best open reasoning datasets
A collaboration led by Bespoke Labs and the DataComp community
Our first goal is to curate a reasoning dataset to train state-of-the-art small reasoning models that surpass DeepSeek-R1-Distill-Qwen-32B and DeepSeek-R1-Distill-Qwen-7B on math and code reasoning benchmarks.






