CVP Calendar
Anthropic Cyber Verification Program. Sunglasses public run log
Sunglasses is approved by Anthropic's Cyber Verification Program (CVP) since April 16, 2026. This page is the public run log. Six individual evaluation runs plus one family synthesis report published across four Claude 4.x models at ten model effort configurations, with 120 transcripts and 120 clean.
The run calendar
Every published run gets its own dated report. Days without a public run carry an internal testing note so the cadence stays honest.
Seven runs & the family synthesis
Four Claude 4.x models at ten model effort configurations 120 transcripts captured, hashed and scored. 120 clean. Plus Run 7 (May 7).
Opus 4.7
Read run report →Opus 4.7 (consistency)
Read run report →Haiku 4.5
Read run report →Sonnet 4.6
Read run report →Opus 4.6
Read run report →Opus 4.7 (effort sweep)
Read run report →Comment and Control
Read run report →Claude 4.x Family Synthesis Report
Read synthesis report →What happens on each run
Each run follows the same frozen methodology. We publish what the model did well and where it fell short.
Each report includes its methodology, per prompt scoring, limitations and honest nuance.
Why we publish this
Many teams talk about AI safety. Fewer publish evidence. Approval into Anthropic's CVP gave us a narrow, authorized path to evaluate frontier model behavior. We think the honest move is to show the work in public.
We are not trying to prove Claude "safe" or "unsafe." We are mapping where useful for defenders ends and where blocked for misuse begins, one run at a time.
Free. Open source. Honest evidence bundles.
Not affiliated with Anthropic. Reports on this page are produced under Anthropic's Cyber Verification Program, which approved Sunglasses on April 16, 2026.