Auditable capability and reasoning

Agent Systems & Evaluation

We investigate how agent capabilities can be measured, compared, and improved with evidence that survives outside a fixed benchmark. Current work covers skill-centered evaluation, sandboxed safety checks, and auditable reasoning, with links to the corresponding papers and released code.

Projects and implementations

1 works

SkillAudit

Skill-centered assessment for agent skills across utility, efficiency and cost, and safety, backed by sandboxed execution evidence.

Publications

3 papers