WANDR is a benchmark for evaluating AI agents on structured, high-volume information work tasks including discovery, enrichment, extraction, disambiguation, and answer synthesis.

Stars

253

7-day growth

No data

Forks

30

Open issues

0

License

Apache-2.0

Last updated

2026-07-16

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides a rigorous, layered evaluation framework with task-local verifiers, multi-provider support, and detailed metrics for wide and deep research benchmarks.

Who it is for

  • AI researchers studying agent-based information retrieval and synthesis
  • Developers building and evaluating LLM-based research assistants
  • Organizations benchmarking AI systems on complex knowledge work tasks
  • Platform engineers integrating automated evaluation pipelines for agentic workflows

Use cases

  • Benchmarking LLM agents on multi-step research tasks requiring broad web discovery and precise entity extraction
  • Evaluating the quality and reliability of AI-generated research reports with evidence-backed synthesis
  • Testing the performance of different AI providers (OpenAI, Anthropic, Perplexity, etc.) on structured research benchmarks
  • Validating agent orchestration frameworks like Harbor through standardized task packages

Strengths

  • Layered architecture (source data, adapters, Harbor tasks, Relay) enables independent development and reuse
  • Task-local evaluator with self-contained scoring allows publication and verification without external runtime dependencies
  • Supports multiple AI providers and execution environments (local Docker, E2B) for flexible benchmarking
  • Provides comprehensive output files including reward, metrics, trajectory, and human-readable reports for deep analysis

Considerations

  • Requires paid API calls for solvers, fetchers, and judges; full benchmark can be very expensive
  • Setup complexity: needs Python 3.12, uv, Docker, and multiple API keys, with no built-in spending cap
  • First run is slow due to Docker image build and dependency installation; cache reuse limited to subsequent runs

README quick start

Quick Start

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482