AI evaluation benchmark
SkillAudit
Skill-centered assessment for agent skills across utility, efficiency and cost, and safety, backed by sandboxed execution evidence.
- Year
- 2026
- Status
- active
- Focus
- Python · Docker · Browser Extension
What it does
SkillAudit accepts an arbitrary agent skill package, derives capability-aligned evaluation tasks, runs paired experiments in isolated sandboxes, and produces auditable reports covering utility, efficiency, cost, and safety. A Chromium extension surfaces the results when developers are deciding whether to install a skill.
Unlike a fixed benchmark suite, the evaluation is generated around the capabilities claimed by each submitted skill. The paired setup compares agent behavior with and without the skill, while preserving execution traces and safety evidence for review.
Evidence and scope
This page is a concise guide to the work. Use the primary paper for the experimental setup, quantitative results, limitations, and formal claims; use the repository for the implementation and current reproduction instructions.