cindy
Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.
该仓库汇集了超过100个科学LLM基准测试,按学科分区整理,是研究人员跟踪科学AI评估快速演进格局不可或缺的参考资源。
Benchmarks for evaluating large language models on scientific reasoning and discovery — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.
Listed alphabetically within each domain.
Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| AGIEval | Microsoft | 2023 | paper | code | Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions. | |
| ARB | DuckAI / Georgia Tech | 2023 | paper | code | Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric. | |
| ARC (AI2 Reasoning Challenge) | Allen AI (AI2) | 2018 | paper | code | 7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets. | |
| C-Eval | SJTU / HKUST | 2023 | paper | code | 13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels. | |
| EMMA | CUHK / Microsoft | 2025 | paper | code | 2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models. | |
| FrontierScience | OpenAI | 2026 | paper | — |
A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.
根据分类、Topic 和编程语言匹配的相似项目。
Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.
A hands-on course that walks through building a Pi-style coding agent from scratch across 15 checkpoints, starting from an offline agent trajectory.
OpenClaude is an open-source CLI coding agent that runs on any platform and supports a wide range of LLM providers, offering the same tools and workflows as Claude Code.