New AI Datasets — Training & Eval Corpora
New training and evaluation datasets for AI — fresh corpora and benchmarks, with what's inside each one and how to use it.
10 releases tracked
- AlphaGenome Atlas — a prediction for every possible DNA letter change
AlphaGenome's predictions for every single-letter DNA change, precomputed so researchers can look one up instead of running the model.
- UltraData-RL-2609 — 86,000 checkable RL tasks behind MiniCPM5-2B
The open RL corpus OpenBMB used to post-train MiniCPM5-2B — 85,995 tasks that a machine can mark right or wrong.
- Agentic Task Ecosystem — Cohere Labs maps 696,291 AI tools to job tasks
Cohere Labs released ATE, an open dataset that links almost 700,000 MCP tools to the job tasks they can actually finish.
- 215,128 machine-made 'best software' pages — and Perplexity cites them
A citation audit of Perplexity's software recommendations, published with the full dataset and scripts.
- LAION-BVD — 10 million hours of open video for multimodal training
LAION released BVD, an open video dataset of 80 million videos and 10 million hours, built for multimodal pre-training.
- Ultra-FineWeb-L1 — a 1.3T-token open web corpus for pretraining
A 1.3T-token English web corpus from OpenBMB, cleaned from six 2025 Common Crawl snapshots and free under Apache 2.0.
- The Stack v3 — 5T tokens of code across 224M GitHub repos, open license
A fresh 5-trillion-token, 770-language open code corpus that resets the pretraining baseline for code LLMs.
- ClawBench — 153 Real-World Browser-Agent Tasks Across 144 Live Websites; Best Model at 33.3%
Agents score 70% in sandboxes; ClawBench shows the real-world number is 33%
- Nemotron-Personas-Korea — 7M Synthetic Personas from Official Korean Demographics
NVIDIA's Korean-language persona dataset: 7 million synthetic personas grounded in official government statistics, CC BY 4.0.
- OpenSpatial — 3M-sample spatial-intelligence data engine
A 3 million-sample, open-source data engine for 3D spatial reasoning — fine-tuned models gain around 19 percent relatively on spatial benchmarks.