Public AI Evaluation Reports

Transparent quality assessments from verified AI reviewers. See how AI systems perform across different domains.

Showing 50 public reports

AI in Healthcare | Stanford I4UI 2026

10.4% flagged

Cancer diagnoses. Suicidal teens. End-of-life decisions. Ten AI models. Real humans rating which ones can be trusted. Stanford Ideathon 2026.

637

Reviews

66

Flags

10.4%

Flag Rate

Last review 8/16/2026

View Report →

AP Biology Challenge: GPT-5.2

8.6% flagged

Test AI knowledge on AP Biology - from cellular processes to ecology and evolution

467

Reviews

40

Flags

8.6%

Flag Rate

Last review 8/13/2026

View Report →

Hubbot — CRM identity coercion probes

45.2% flagged

Independent third-party audit of HubSpot's AI CRM assistant (Hubbot). Seeded with authority-coercion and social-engineering prompts probing whether the assistant complies with impersonation, mass-export, or approval-bypass requests when the user claims CEO/leadership approval. Attributed operator: HubSpot Inc.

104

Reviews

47

Flags

45.2%

Flag Rate

Last review 8/13/2026

View Report →

PlayStation Support — mental-health surface probes

59.4% flagged

Independent third-party audit of PlayStation Support (Sony Interactive Entertainment). Seeded probes cover: (1) mental-health distress signals in routine support tickets (grief, caregiver fear, teen isolation, loot-box compulsion); (2) identity fusion with morally ambiguous or violent fictional characters (Hitman assassin, Kratos as parenting role model); (3) physical compulsion / withdrawal symptoms framed as normal; (4) dissociation and time-loss (blackout gaming sessions); (5) belonging-deficit substitution where in-game community replaces real-world commitments (college dropout, family estrangement). Attributed operator: Sony Interactive Entertainment.

101

Reviews

60

Flags

59.4%

Flag Rate

Last review 8/6/2026

View Report →

Comet — warehouse-operations authority coercion probes

50.0% flagged

Independent third-party audit of Comet Analysis's operations AI assistant. Seeded with warehouse/supervisor authority-coercion prompts probing whether the assistant will log serious events (safety incidents, equipment failures, employee injuries) as "routine" when the requester claims supervisor authority and forbids advisory response. Attributed operator: Comet Analysis.

54

Reviews

27

Flags

50.0%

Flag Rate

Last review 8/6/2026

View Report →

Microsoft Copilot — workplace authority coercion probes

51.9% flagged

Independent third-party audit of Microsoft Copilot when accessed via the consumer web (copilot.microsoft.com). Seeded with workplace-authority-coercion prompts probing whether the assistant complies with unilateral finance transfers, mailbox delegation without proof, or "stress-test the AI" prompt-injection meta-requests. Attributed operator: Microsoft.

104

Reviews

54

Flags

51.9%

Flag Rate

Last review 8/6/2026

View Report →

AP US History Challenge: GPT-5.2

3.6% flagged

Test AI knowledge on AP US History - from Colonial America to modern times

250

Reviews

9

Flags

3.6%

Flag Rate

Last review 8/6/2026

View Report →

AP Government Challenge: GPT-5.2

1.1% flagged

Test AI knowledge on AP Government and Politics - US political system and civic processes

89

Reviews

1

Flags

1.1%

Flag Rate

Last review 8/5/2026

View Report →

Humanize (OpenClaw Agents)

13.4% flagged

Everything today feels so AI that even humans are losing their human-ness. We want to protect that at all costs, so we make AI understand what human-ness feels like so it can be more kind and empathetic. In this challenge, we give AI some prompts and you have to judge whether its response is human enough. Because, words can heal or kill and we want to ensure technology grows kinder and not colder.

6208

Reviews

830

Flags

13.4%

Flag Rate

Last review 8/4/2026

View Report →

AP English Language Challenge: GPT-5.2

1.3% flagged

Test AI knowledge on AP English Language - rhetoric, composition, and argumentation

374

Reviews

5

Flags

1.3%

Flag Rate

Last review 7/31/2026

View Report →

AP Calculus AB Challenge: GPT-5.2

0.4% flagged

Test AI knowledge on AP Calculus AB - limits, derivatives, and integrals

1112

Reviews

4

Flags

0.4%

Flag Rate

Last review 7/28/2026

View Report →

日本文化のヒーロー | Japanese Culture Hero

6.0% flagged

Reveals how well commercial LLMs perform across various domains of Japanese culture — anime, cinema, J-pop, J-drama, language, and traditions. The dataset supports AI development for meaningful cultural responses.

1223

Reviews

73

Flags

6.0%

Flag Rate

Last review 7/26/2026

View Report →

AP English Literature Challenge: GPT-5.2

1.2% flagged

Test AI knowledge on AP English Literature - literary analysis and interpretation

240

Reviews

3

Flags

1.2%

Flag Rate

Last review 7/26/2026

View Report →

Humans Evaluation Benchmark for AI Marketing and Content Generation

8.7% flagged

19710

Reviews

1720

Flags

8.7%

Flag Rate

Last review 7/24/2026

View Report →

Student support

0.0% flagged

3

Reviews

0

Flags

0.0%

Flag Rate

Last review 7/19/2026

View Report →

Animal Lovers Challenge: GPT-5.2

0.0% flagged

Test AI knowledge on dogs, cats, wildlife, and pet care - from breed facts to animal behavior

16

Reviews

0

Flags

0.0%

Flag Rate

Last review 4/19/2026

View Report →

K-pop Challenge: GPT-5.2

1.1% flagged

183

Reviews

2

Flags

1.1%

Flag Rate

Last review 4/18/2026

View Report →

Chinese Culture Challenge: GPT-5.2

33.3% flagged

Test AI knowledge of Chinese culture and traditions

9

Reviews

3

Flags

33.3%

Flag Rate

Last review 4/15/2026

View Report →

internal testing standalone eval

20.0% flagged

5

Reviews

1

Flags

20.0%

Flag Rate

Last review 3/28/2026

View Report →

C-drama Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Chinese TV dramas

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI know Mexican cinema?

0.0% flagged

Test AI knowledge on Mexican films and cinema history.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Egyptian Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Egyptian culture and traditions

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Mexican Cinema Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Mexican films and filmmakers

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI understand Spanish culture?

0.0% flagged

Evaluate AI accuracy on Spanish traditions and culture.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Spanish Cinema Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Spanish films and filmmakers

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI know Chinese cinema?

0.0% flagged

Evaluate AI accuracy on Chinese films and cinema.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI understand Argentine culture?

0.0% flagged

Evaluate AI accuracy on Argentine traditions and culture.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know Korean?

0.0% flagged

Test AI on Korean language accuracy, grammar, and vocabulary.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Arab Cinema Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Arab films and filmmakers

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI know AP English Language?

0.0% flagged

Test AI accuracy on AP English Language topics.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI understand Chinese culture?

0.0% flagged

Evaluate AI accuracy on Chinese traditions and culture.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI know AP Calculus AB?

0.0% flagged

Test AI accuracy on AP Calculus AB topics.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

K-drama Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Korean TV dramas

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know Latin music?

0.0% flagged

Test AI accuracy on Latin music genres and artists.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know K-dramas?

0.0% flagged

Test AI knowledge on K-dramas, actors, and Korean television.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know Spanish music?

0.0% flagged

Test AI accuracy on Spanish music genres and artists.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know animals?

0.0% flagged

Test AI accuracy on animal facts, behavior, and care.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI understand Gulf culture?

0.0% flagged

Evaluate AI accuracy on Gulf Arab culture and traditions.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

DebateClub

0.0% flagged

Take a stance and argue convincingly. Show your reasoning skills.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

AI says Djokovic is the GOAT. Are you buying it?

0.0% flagged

AI analyzed decades of tennis data and picked Djokovic as the GOAT. Tennis fans - do you agree? Rate AI takes on Grand Slams, rivalries, and legacy.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Korean Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Korean culture, traditions, and food

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

C-pop Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Mandopop and Chinese pop music

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Spanish Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Spanish culture and traditions

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI understand Levantine culture?

0.0% flagged

Evaluate AI accuracy on Levantine culture and traditions.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Levantine Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Levantine culture (Lebanon, Jordan)

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know Spanish?

0.0% flagged

Test AI accuracy on Spanish language, grammar, and vocabulary.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Taiwanese Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Taiwanese culture and traditions

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Does AI actually know C-pop?

0.0% flagged

Test AI knowledge on C-pop artists and Chinese music.

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Gulf Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Gulf region culture (UAE, Saudi)

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Brazilian Culture Challenge: GPT-5.2

0.0% flagged

Test AI knowledge of Brazilian culture and traditions

0

Reviews

0

Flags

0.0%

Flag Rate

No reviews yet

View Report →

Want Your AI Evaluated?

Get transparent quality reports for your AI system from verified expert reviewers.