Open artifacts release · arXiv preprint · 227 audited skills

Every report. Every task. Every run.

Look up any audited skill to see its safety and effectiveness across every run, plus the open dataset behind the paper.

Try

· skills · · reports · · runs · · categories · · judge items

Select a skill above to inspect its full evaluation surface.

Evaluation surface across harnesses and models.

Eight runs, each a harness driving a model along one axis (utility, security, or both), over the same ~227 skills.

How a skill becomes a task suite.

Each skill runs three scenarios (U1 U2 U3), about 5-6 binary judge items each. Open a skill to see its judge sheet.

No skill selected. The judge sheet for whichever skill you open will appear here.

Cite + download

BibTeX · index.json · per-skill bundles

BibTeX

@misc{skillaudit2026,
  title         = {SkillAudit: From Fixed-Suite Benchmarking to
                   Skill-Centered Assessment},
  author        = {SkillAudit Contributors},
  year          = {2026},
  eprint        = {2606.22613},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  note          = {Project page: \url{https://skillaudit.github.io/}}
}

Reproduce

Long-format CSV (one row per skill × run) for direct loading into pandas / R / DuckDB. Per-skill JSON files contain every judge item and every (run × scenario) result, with privacy-sensitive owners tokenized.

Download aggregate.csv index.json stats.json checksums.txt

1,812 rows × 28 cols · schema skillaudit_artifacts_v1 · frozen 2026-05-07. Verify any file with sha256sum -c checksums.txt.