I stopped maintaining this website in September 2025. Work on what AI companies should do in terms of safety and what they are doing is now being done by METR, Guidelight, Midas, and others.

Anthropic

Risk assessment 44%

50%
Evals: domains, quality, elicitation
15%
Evals: accountability
50%
Adversarial evaluation for alignment
75%
Model organisms

Evals: domains, quality, elicitation

50%
Click to show details/rubric

Evals: accountability

15%
Click to show details/rubric

Adversarial evaluation for alignment

50%
Click to show details/rubric

Model organisms

75%
Click to show details/rubric