Every report. Every task. Every run.
Look up any audited skill to see its safety and effectiveness across every run, plus the open dataset behind the paper.
· skills · · reports · · runs · · categories · · judge items
Evaluation surface across harnesses and models.
Eight runs, each a harness driving a model along one axis (utility, security, or both), over the same ~227 skills.
How a skill becomes a task suite.
Each skill runs three scenarios (U1 U2 U3), about 5-6 binary judge items each. Open a skill to see its judge sheet.
Cite + download
BibTeX · index.json · per-skill bundles
BibTeX
@misc{skillaudit2026,
title = {SkillAudit: From Fixed-Suite Benchmarking to
Skill-Centered Assessment},
author = {SkillAudit Contributors},
year = {2026},
eprint = {2606.22613},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
note = {Project page: \url{https://skillaudit.github.io/}}
}
Reproduce
Long-format CSV (one row per skill × run) for direct loading into
pandas / R / DuckDB. Per-skill JSON files contain every judge item and every
(run × scenario) result, with privacy-sensitive owners tokenized.