subinium GitHub avatar

Awesome-Scientific-LLM-Benchmarks

subinium

一个精选的基准测试集合,用于评估大语言模型在数学、物理、化学、生物、材料科学和智能体科学等领域的科学推理与发现能力。

Stars

28

7 天增长

暂无数据

Fork 数

1

开放 Issue

0

开源协议

MIT

最近更新

2026-07-13

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

该仓库汇集了超过100个科学LLM基准测试,按学科分区整理,是研究人员跟踪科学AI评估快速演进格局不可或缺的参考资源。

适合谁使用

  • AI研究人员
  • 机器学习工程师
  • 科学教育工作者
  • 基准测试开发者

典型使用场景

  • 选择合适的基准测试来评估LLM的科学推理能力
  • 识别当前评估套件中的空白
  • 及时了解各个科学领域的最新基准测试
  • 比较不同基准测试的范围和难度

项目优势

  • 覆盖全面:涵盖6个科学领域,100+基准测试
  • 组织清晰:每个领域按字母顺序排列,附有论文和代码链接
  • 难度多样:从基础年级到专家/奥赛级别
  • 社区驱动:采用Awesome列表格式,包含星标和徽章

使用前须知

  • 仅为参考列表,不提供模型评估结果或排名
  • 需要定期更新以保持时效性
  • 选择可能受策展人偏好影响,某些领域可能覆盖不足

README 快速开始

Awesome Scientific LLM Benchmarks

Benchmarks for evaluating large language models on scientific reasoning and discovery — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.

Listed alphabetically within each domain.

Contents

General / Multi-domain Science

Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.

BenchmarkOrgYearPaperCodeStarsDescription
AGIEvalMicrosoft2023papercodeHuman-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.
ARBDuckAI / Georgia Tech2023papercodeAdvanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.
ARC (AI2 Reasoning Challenge)Allen AI (AI2)2018papercode7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets.
C-EvalSJTU / HKUST2023papercode13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels.
EMMACUHK / Microsoft2025papercode2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models.
FrontierScienceOpenAI2026paper

项目描述

A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

makecindy
精选
makecindy GitHub avatar

cindy

Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.

AI 与机器学习大语言模型
958
hahhforest
精选
hahhforest GitHub avatar

pi-textbook

A hands-on course that walks through building a Pi-style coding agent from scratch across 15 checkpoints, starting from an offline agent trajectory.

AI 与机器学习大语言模型
625
0xwilliamortiz
精选
0xwilliamortiz GitHub avatar

openclaude-improved

OpenClaude is an open-source CLI coding agent that runs on any platform and supports a wide range of LLM providers, offering the same tools and workflows as Claude Code.

AI 与机器学习大语言模型
577