MiaAI-Lab GitHub avatar

Best-Local-Model_Agentic-Workflows_2026

MiaAI-Lab

一份针对智能体工作流选择最佳本地LLM的基准测试和比较报告,基于84个场景、16个类别的系统评估。

Stars

31

7 天增长

暂无数据

Fork 数

2

开放 Issue

1

开源协议

暂无数据

最近更新

2026-07-06

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

该仓库提供了系统化的多轮次基准测试,重点关注可靠性、可部署性以及实际智能体使用场景,为拥有96-128GB内存硬件的用户给出了明确的模型推荐。

适合谁使用

  • 构建本地AI代理的开发者
  • 评估LLM工具使用能力的研究人员
  • 拥有96-128GB内存设备的爱好者
  • 需要可靠模型的人工智能安全工程师

典型使用场景

  • 选择默认生产环境智能体后端
  • 评估安全关键或对抗性工作负载
  • 优化低延迟与高可靠性需求
  • 避免带有危险注入漏洞的模型

项目优势

  • 全面的多类别基准测试(84个场景、16个类别)
  • 透明的方法论:每个模型8次试验,提供Pass@8、Pass^8等详细指标
  • 清晰的层级推荐和排名,附带可操作建议
  • 顶级模型在所有核心智能体类别中无失败场景

使用前须知

  • 基准数据仅限于所测试的模型,可能未覆盖所有可用本地模型
  • 结果特定于Hermes Agent框架和tool-eval-bench评估工具
  • 顶级模型需要96-128GB内存的硬件支持

README 快速开始

Best Local Model for Agentic Workflows for single DGX Spark or other 96-128gb rigs (2026)

Interactive comparison report for choosing a local LLM backend for agentic workflows — multi-turn tool orchestration, function calling, and autonomous planning as exercised by frameworks like Hermes Agent.

Benchmark data comes from tool-eval-bench: 84 scenarios, 16 categories, 8 trials per model (where available), scored pass / partial / fail.

Quick answer

Use Qwen 3.6 35B A3B UD Q8_K_XL as your default Hermes Agent backend.

MetricQwen 35B Q8_K_XL
Mean score91.0
Pass@8 (capability ceiling)91.7%
Pass^8 (reliability floor)76.2%
Deployability79
Median turn2.5s
Never-pass scenarios0

It is the only model with 100% across all core agent categories on every trial: multi-step chains, error recovery, tool selection, parameter precision, structured output, and instruction following.

View the reports

Open either HTML file in a browser — no build step, no server required.

git clone https://github.com/MiaAI-Lab/Best-Local-Model_Agentic-Workflows_2026.git
cd Best-Local-Model_Agentic-Workflows_2026
xdg-open agentic-model-comparison.html   # Linux
open agentic-model-comparison.html         # macOS

Full ranking (agentic use)

RankModelScoreWhy
1Qwen 3.6 35B A3B UD Q8_K_XL91.0Best overall: score, capability, deployability, zero never-pass
2Qwen 3.6 27B NVFP489.0Hard-mode / safety tier (93% hard, 96% safety, 81% Pass^8)
3Qwopus 3.6 27B Coder MTP85.2Fastest reliable tier (2.2s, 4.8pp gap, 100% error recovery)
4DeepSeek V4 Flash Q286.5High ceiling (88.1% Pass@8) but TC-60 injection + 23 safety warnings
5Agents-A1 Q8_083.4High Pass@8 ceiling (90.5%), flaky floor (64.3%)
6Gemma 4 26B NVFP481.4Mid-tier
7Nemotron 3 N

项目描述

Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

ddosi
精选
ddosi GitHub avatar

PacketLens

PacketLens is a pure front-end, offline pcap analysis tool that runs entirely in the browser, supporting deep protocol decoding, HTTPS decryption, and million-packet instant loading without any backend.

HTML
19
ViffyGwaanl
精选
ViffyGwaanl GitHub avatar

kimi-k3-learn

An interactive learning system that turns the 47-page Kimi K3 technical report into a single offline HTML file with running algorithms, 3D visualizations, spaced repetition quizzes, and a smart highlighting QA tool.

AI 与机器学习
15
iamtechartist
精选
iamtechartist GitHub avatar

human-cell-visualizer

An interactive 3D visualization of three human cell types using Three.js, WebGL particles, and custom GLSL shaders, allowing users to rotate, zoom, morph, and explore annotated structures.

设计与创意
13