OSAC-Bench is a public-preview benchmark for evaluating operating-system agents on continuous state diagnosis, bounded tool use, and accountable collaboration across two tracks of five tasks each.

Stars

6

7-day growth

No data

Forks

0

Open issues

0

License

NOASSERTION

Last updated

2026-07-29

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It introduces a structured collaboration model (intent owner, evidence worker, independent verifier) and deterministic checking for OS agent memory reuse and stale-fact rejection, filling a gap in reproducible evaluation of OS agents.

Who it is for

  • OS agent researchers and developers
  • Academic researchers studying tool-use and memory benchmarks
  • Open-source communities building Linux-based automation agents
  • Engineers designing agent evaluation frameworks

Use cases

  • Evaluating an OS agent's ability to diagnose system state changes across package, service, or runtime configurations
  • Testing agent memory reuse policies with scoped fingerprint validation and stale-fact rejection
  • Benchmarking multi-agent collaboration patterns (e.g., owner-worker-verifier) under deterministic checkers
  • Reproducible integration testing of agent runners against a frozen public-development task suite

Strengths

  • Domain-level tracks (system and runtime state) that are distribution-agnostic and use synthetic frozen fixtures
  • Clear role-based collaboration model with observable responsibilities and deterministic checker contracts
  • Public-development answer key and reference validator for integration testing without dependencies
  • Methodological alignment with established benchmarks while contributing original OS-agent task content

Considerations

  • Public preview – not an official OS certification suite or leaderboard; results require exact version tracking
  • CC BY-NC 4.0 license restricts commercial use without separate written permission
  • Requires Node.js 18+ and no held-out golden answers in the public repository (evaluator must provide private bundle)

README quick start

Quick start

Description

Research benchmark for evidence-grounded OS-agent collaboration, continuous state diagnosis, scoped memory reuse, and stale-state rejection.

Related repositories

Similar projects matched by category, topics, and programming language.

mshumer
Featured
mshumer GitHub avatar

Claude-of-Duty

A browser-based first-person shooter built entirely with procedural generation and orchestrated AI agents, featuring 55k lines of Three.js/WebGL2 code and no art assets.

AI & Machine LearningAI Agents
1,234
gnipbao
Featured
gnipbao GitHub avatar

story-to-handdrawn-video

A Remotion-based tool that converts Chinese story text or ordered hand-drawn images into vertical hand-drawn diary-comic animation with handwritten captions, left-to-right reveals, optional page-curl transitions, and silent H.264 output for post-production dubbing.

AI & Machine LearningAI Agents
683
0xwilliamortiz
Featured
0xwilliamortiz GitHub avatar

ponytail-improved

Ponytail is a plugin for AI coding agents that enforces a disciplined ladder of reuse before writing code, reducing code volume by roughly 54% while preserving safety.

AI & Machine LearningAI Agents
545