AI evaluation benchmark

SkillAudit

Skill-centered assessment for agent skills across utility, efficiency and cost, and safety, backed by sandboxed execution evidence.

Year
2026
Status
active
Focus
Python · Docker · Browser Extension

What it does

SkillAudit accepts an arbitrary agent skill package, derives capability-aligned evaluation tasks, runs paired experiments in isolated sandboxes, and produces auditable reports covering utility, efficiency, cost, and safety. A Chromium extension surfaces the results when developers are deciding whether to install a skill.

Unlike a fixed benchmark suite, the evaluation is generated around the capabilities claimed by each submitted skill. The paired setup compares agent behavior with and without the skill, while preserving execution traces and safety evidence for review.

Evidence and scope

This page is a concise guide to the work. Use the primary paper for the experimental setup, quantitative results, limitations, and formal claims; use the repository for the implementation and current reproduction instructions.