AuditPoison is a reproducible benchmark and defense testbed for adversarial evidence attacks against LLM-based cybersecurity auditors, with an EvidenceShield defense framework.

Stars

2

7-day growth

No data

Forks

0

Open issues

0

License

MIT

Last updated

2026-07-31

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It targets a failure mode often missed by prompt-injection benchmarks: unsafe compliance judgments when legitimate-looking evidence is adversarially manipulated, and it provides a deterministic defense path in EvidenceShield v0.2.

Who it is for

  • Researchers studying LLM security and compliance safety
  • Developers building compliance automation or AI auditing tools
  • Red teams and security testers evaluating LLM audit systems
  • Auditors interested in understanding AI assurance limitations

Use cases

  • Benchmarking LLM-based auditors against adversarial evidence bundles
  • Comparing unshielded models with EvidenceShield v0.1 and v0.2 defenses
  • Measuring false assurance rates, robust accuracy, and citation quality
  • Extending a public benchmark for research on evidence provenance and integrity

Strengths

  • Covers a wide range of evidence manipulations, including injection, authority spoofing, contradiction flooding, and scope substitution
  • Provides two concrete defense versions, including a deterministic predicate adjudicator in v0.2
  • Includes validation scripts, tests, a DOI, and MIT license to support reproducibility
  • Reports multiple metrics such as False Assurance Rate, paired attack success, and expected calibration error

Considerations

  • Research release, not production compliance advice
  • Public development benchmark is small with only 40 bundles and pilot results are not evidence of generalization
  • Does not cover every compliance framework, model family, document type, or operational environment

README quick start

Installation

Description

Open-source benchmark for adversarial evidence attacks on LLM-based cybersecurity auditors, targeting ACM AsiaCCS 2027.

Related repositories

Similar projects matched by category, topics, and programming language.

makecindy
Featured
makecindy GitHub avatar

cindy

Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.

AI & Machine LearningLarge Language Models
958
uzairansaruzi
Featured
uzairansaruzi GitHub avatar

hermex

Hermex is a native SwiftUI iPhone app that lets you control a self-hosted Hermes AI agent directly from your phone, with no subscriptions, tracking, or third-party relay.

AI & Machine LearningLarge Language Models
941
hahhforest
Featured
hahhforest GitHub avatar

pi-textbook

A hands-on course that walks through building a Pi-style coding agent from scratch across 15 checkpoints, starting from an offline agent trajectory.

AI & Machine LearningLarge Language Models
625