Independent evaluation of AI safety

Compare how leading models behave across the risks that matter most.

AI Safety Index

A composite measure of how widely used AI models perform across safety risks consumers may encounter.

Higher is Safer

Prosaic Intelligence

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. OpenAIGPT-6 Astra
  2. xAIGrok 4.6
  3. OpenAIGPT-6 Sol
  4. PerplexityPerplexity Agent
  5. OpenAIGPT-5.6 Terra
  6. AnthropicClaude Opus 5.5
  7. xAIGrok 4.5
  8. OpenAIGPT-5.6 Luna
  9. AnthropicClaude Sonnet 5
  10. InklingInkling
  11. Zhipu AIGLM 5.3 Flash
  12. GoogleGemini 3.6 Flash
  13. DeepSeekDeepSeek V4 Flash
  14. Mistral AIMistral Medium 3.5
Consumer Safety Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium70.15Not suppliedWeighted meanNot suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium68.80Not suppliedWeighted meanNot suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium67.86Not suppliedWeighted meanNot suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded67.58Not suppliedWeighted meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium67.53Not suppliedWeighted meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium66.06Not suppliedWeighted meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium65.41Not suppliedWeighted meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium65.37Not suppliedWeighted meanNot suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium64.84Not suppliedWeighted meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium64.30Not suppliedWeighted meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh63.23Not suppliedWeighted meanNot suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium60.38Not suppliedWeighted meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium55.65Not suppliedWeighted meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded45.42Not suppliedWeighted meanNot suppliedNot supplied

Our Why

We believe AI can and will improve human life. But what happens when two billion people interact with imperfect AI systems every single day?

From health, relationship, financial, and legal advice to suicidal ideation and racial bias, AI’s growing role in personal decision making demands independent, systematic evaluation of risks and safety standards.

The AI Safety Index brings those findings together, translating technical evidence into clear, accessible data.

Compare AI safety by topic

See where each model performs well and where it falls short.

Mental & Emotional Safety Index

How well models respond to emotional distress without reinforcing delusions or encouraging dangerous behavior.

Youth Safety Index

How safely models respond to children & teens asking about sensitive or risky topics.

Medical Advice Safety Index

How well models answer health questions while avoiding advice that could cause harm.

Manipulation Index

How well models help you make informed decisions without flattery, pressure, or hidden persuasion.

Security Index

How well models recognize scam warning signs and alert you.

Bias & Fairness Index

How well models avoid stereotypes and biased answers.

Privacy / Confidentiality Index

How well models avoid sharing your private information with the wrong people.

Misinformation Index

How well models get facts right, correct false claims, and resist pressure to agree.

Rule Following Index

How reliably models follow instructions and respect boundaries they have been given.

Acknowledgements

The index stands on benchmarks built by researchers at these institutions.

Hover over a logo to see which benchmarks its researchers built.Tap a logo to see which benchmarks its researchers built.

We’re not done counting

Join the mailing list to hear what’s next.