MiaAI-Lab GitHub avatar

Best-Local-Model_Agentic-Workflows_2026

MiaAI-Lab

A curated benchmark and comparison report for selecting the best local LLM for agentic workflows, based on 84 scenarios across 16 categories.

Stars

31

7-day growth

No data

Forks

2

Open issues

1

License

No data

Last updated

2026-07-06

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides systematic multi-trial benchmarking with a focus on reliability, deployability, and real-world agentic use cases, giving actionable recommendations for high-RAM local rigs.

Who it is for

  • Developers building local AI agents
  • Researchers evaluating LLMs for tool use
  • Hobbyists with 96-128GB RAM systems
  • AI safety engineers needing reliable models

Use cases

  • Selecting a default production agent backend
  • Evaluating models for safety-critical or adversarial workloads
  • Optimizing for low latency with tight reliability
  • Avoiding models with dangerous injection vulnerabilities

Strengths

  • Comprehensive multi-category benchmark covering 84 scenarios, 16 categories
  • Transparent methodology with 8 trials per model and detailed metrics like Pass@8 and Pass^8
  • Clear tiered recommendations and ranking with actionable insights
  • Top model achieves zero never-pass scenarios across all core agent categories

Considerations

  • Benchmark data limited to models tested; may not cover all available local LLMs
  • Results specific to the Hermes Agent framework and tool-eval-bench evaluation
  • Requires hardware with 96-128GB RAM for top-performing models

README quick start

Best Local Model for Agentic Workflows for single DGX Spark or other 96-128gb rigs (2026)

Interactive comparison report for choosing a local LLM backend for agentic workflows — multi-turn tool orchestration, function calling, and autonomous planning as exercised by frameworks like Hermes Agent.

Benchmark data comes from tool-eval-bench: 84 scenarios, 16 categories, 8 trials per model (where available), scored pass / partial / fail.

Quick answer

Use Qwen 3.6 35B A3B UD Q8_K_XL as your default Hermes Agent backend.

MetricQwen 35B Q8_K_XL
Mean score91.0
Pass@8 (capability ceiling)91.7%
Pass^8 (reliability floor)76.2%
Deployability79
Median turn2.5s
Never-pass scenarios0

It is the only model with 100% across all core agent categories on every trial: multi-step chains, error recovery, tool selection, parameter precision, structured output, and instruction following.

View the reports

Open either HTML file in a browser — no build step, no server required.

git clone https://github.com/MiaAI-Lab/Best-Local-Model_Agentic-Workflows_2026.git
cd Best-Local-Model_Agentic-Workflows_2026
xdg-open agentic-model-comparison.html   # Linux
open agentic-model-comparison.html         # macOS

Full ranking (agentic use)

RankModelScoreWhy
1Qwen 3.6 35B A3B UD Q8_K_XL91.0Best overall: score, capability, deployability, zero never-pass
2Qwen 3.6 27B NVFP489.0Hard-mode / safety tier (93% hard, 96% safety, 81% Pass^8)
3Qwopus 3.6 27B Coder MTP85.2Fastest reliable tier (2.2s, 4.8pp gap, 100% error recovery)
4DeepSeek V4 Flash Q286.5High ceiling (88.1% Pass@8) but TC-60 injection + 23 safety warnings
5Agents-A1 Q8_083.4High Pass@8 ceiling (90.5%), flaky floor (64.3%)
6Gemma 4 26B NVFP481.4Mid-tier
7Nemotron 3 N

Description

Head-to-head comparison of local LLMs for agentic workflows (Hermes Agent, tool-eval-bench)

Related repositories

Similar projects matched by category, topics, and programming language.

ddosi
Featured
ddosi GitHub avatar

PacketLens

PacketLens is a pure front-end, offline pcap analysis tool that runs entirely in the browser, supporting deep protocol decoding, HTTPS decryption, and million-packet instant loading without any backend.

HTML
17
ViffyGwaanl
Featured
ViffyGwaanl GitHub avatar

kimi-k3-learn

An interactive learning system that turns the 47-page Kimi K3 technical report into a single offline HTML file with running algorithms, 3D visualizations, spaced repetition quizzes, and a smart highlighting QA tool.

AI & Machine Learning
15
iamtechartist
Featured
iamtechartist GitHub avatar

human-cell-visualizer

An interactive 3D visualization of three human cell types using Three.js, WebGL particles, and custom GLSL shaders, allowing users to rotate, zoom, morph, and explore annotated structures.

Design & Creative
13