subinium GitHub avatar

Awesome-Scientific-LLM-Benchmarks

subinium

A curated list of benchmarks for evaluating large language models on scientific reasoning and discovery across mathematics, physics, chemistry, biology, materials science, and agentic science.

Stars

28

7-day growth

No data

Forks

1

Open issues

0

License

MIT

Last updated

2026-07-13

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

This repository aggregates over 100 scientific LLM benchmarks in one place, organized by discipline, making it an indispensable reference for researchers tracking the rapidly evolving landscape of scientific AI evaluation.

Who it is for

  • AI researchers
  • ML engineers
  • Science educators
  • Benchmark developers

Use cases

  • Select appropriate benchmarks for evaluating scientific reasoning in LLMs
  • Identify gaps in current evaluation suites
  • Stay informed about the latest benchmarks in various scientific domains
  • Compare the scope and difficulty of different benchmarks

Strengths

  • Comprehensive coverage: spans 6 scientific domains with 100+ benchmarks
  • Well-organized: each domain is listed alphabetically with links to papers and code
  • Includes diverse difficulty levels: from grade-school to expert/olympiad
  • Community-driven: uses Awesome list format with star counts and badges

Considerations

  • Only a reference list; does not provide model evaluation results or rankings
  • May require regular updates to stay current with new benchmarks
  • Selection may reflect the curator's focus; some domains might be underrepresented

README quick start

Awesome Scientific LLM Benchmarks

Benchmarks for evaluating large language models on scientific reasoning and discovery — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.

Listed alphabetically within each domain.

Contents

General / Multi-domain Science

Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.

BenchmarkOrgYearPaperCodeStarsDescription
AGIEvalMicrosoft2023papercodeHuman-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.
ARBDuckAI / Georgia Tech2023papercodeAdvanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.
ARC (AI2 Reasoning Challenge)Allen AI (AI2)2018papercode7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets.
C-EvalSJTU / HKUST2023papercode13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels.
EMMACUHK / Microsoft2025papercode2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models.
FrontierScienceOpenAI2026paper

Description

A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.

Related repositories

Similar projects matched by category, topics, and programming language.

makecindy
Featured
makecindy GitHub avatar

cindy

Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.

AI & Machine LearningLarge Language Models
958
hahhforest
Featured
hahhforest GitHub avatar

pi-textbook

A hands-on course that walks through building a Pi-style coding agent from scratch across 15 checkpoints, starting from an offline agent trajectory.

AI & Machine LearningLarge Language Models
625
0xwilliamortiz
Featured
0xwilliamortiz GitHub avatar

openclaude-improved

OpenClaude is an open-source CLI coding agent that runs on any platform and supports a wide range of LLM providers, offering the same tools and workflows as Claude Code.

AI & Machine LearningLarge Language Models
577