In a significant move to support the growth of artificial intelligence research, Snorkel AI recently announced the inaugural group of projects to receive funding through its Open Benchmarks Grants initiative. With a substantial commitment of $3 million, this initiative aims to bolster open-source datasets, benchmarks, and evaluation research crucial for developing high-performance AI systems. Launched in February 2026, the Open Benchmarks Grants program has attracted hundreds of applications from a diverse range of researchers, labs, and engineers who are grappling with the pressing challenge of measuring AI systems' performance effectively.
Fred Sala, an assistant professor at the University of Wisconsin–Madison and member of the grants steering committee, expressed enthusiasm over the projects selected for funding: "These projects tackle some of the field's most difficult evaluation challenges, ranging from complex environments to expansive autonomy horizons and the generation of sophisticated outputs. I look forward to seeing the broader research community validate and build upon them."
The first funded projects encompass a variety of ambitious research efforts:
1.
Frontier-Bench: Developed in collaboration with the Laude Institute and the Harbor community, this project is a challenging successor to Terminal-Bench 2.1 that aims to assess AI in a more diverse range of domains. Its development follows an open, task-by-task methodology, subject to continuous adversarial review.
2.
Agents' Last Exam: In partnership with UC Berkeley RDI and the RDI Foundation, this project evaluates agents on long-horizon workflows that are deemed economically valuable across 55 sub-industries, including over 1,500 tasks that are under validation by industry experts.
3.
OSWorld 2.0: Created with the XLANG Lab, this benchmark evaluates computer-use agents across 108 long-horizon workflows within 31 self-hosted web environments and professional desktop applications.
4.
Continual Learning Bench: This initiative, developed with UC Berkeley SkyLab and the University of Wisconsin–Madison, aims to measure whether agents genuinely improve their performance across sequential tasks.
5.
SlopCode Bench: Involved with research teams at the University of Wisconsin–Madison, this project assesses how the quality of code degrades as coding agents modify and extend their solutions over time.
6.
Terminal-Bench 2.1: This project, a collaboration with Stanford University, the Laude Institute, and the Harbor community, analyzes agents' performance in terminal environments. Revised tasks have been validated continuously to enhance rigor in evaluation.
The development of
Terminal-Bench Science is also underway, extending existing frameworks to incorporate computational research workflows across various scientific domains, showcasing the far-reaching implications of the Open Benchmarks Grants initiative. Beyond these projects, Snorkel has made strides in its collaboration with Princeton University and the University of Wisconsin–Madison to create
Senior SWE-Bench, which evaluates coding agents on senior-level engineering challenges such as feature implementation from complex instructions and bug investigations.
Backed by prominent industry leaders including Hugging Face, Prime Intellect, and Factory, the Open Benchmarks Grants program remains open for applications that will be reviewed on a rolling basis. Interested applicants can find more information and submit their proposals at
benchmarks.snorkel.ai.
Founded out of the Stanford AI Lab in 2019, Snorkel AI focuses on building the data and environments crucial for driving next-generation AI systems. Their commitment to developing high-quality datasets, benchmarks, and evaluation frameworks positions them as a pioneer in enabling better outcomes in the AI landscape.