I stopped maintaining this website in September 2025. Work on what AI companies should do in terms of safety and what they are doing is now being done by METR, Guidelight, Midas, and others.

Risk assessment

Evals: domains, quality, elicitation

Weighted 55% of category
Anthropic
50%
more
DeepMind
50%
more
OpenAI
50%
more
Meta
10%
more
xAI
10%
more
Microsoft
2%
more
DeepSeek
0%

Evals: accountability

Weighted 25% of category
Anthropic
15%
more
DeepMind
1%
more
OpenAI
20%
more
Meta
0%
xAI
0%
Microsoft
0%
DeepSeek
0%

Adversarial evaluation for alignment

Weighted 10% of category
Anthropic
50%
more
DeepMind
0%
more
OpenAI
10%
more
Meta
0%
xAI
0%
Microsoft
0%
DeepSeek
0%

Model organisms

Weighted 10% of category
Anthropic
75%
more
DeepMind
10%
more
OpenAI
0%
Meta
0%
xAI
0%
Microsoft
0%
DeepSeek
0%