Under-Elicited Is a Result
Anthropic called a Claude Opus 4.6 sabotage evaluation under-elicited. Before a low score becomes safety evidence, the test needs a positive control.
Topic
Anthropic called a Claude Opus 4.6 sabotage evaluation under-elicited. Before a low score becomes safety evidence, the test needs a positive control.
DeepMind protected a private benchmark and proprietary model weights inside one sealed evaluation. The public still needs a receipt showing what happened there.
AI safety assessments may become a condition of releasing frontier models. Before that happens, lawmakers should publish the method and the cost of compliance.
Three outside teams tested Claude Opus 5 safeguards and reported it in hours, hours, and attempts — the same unit problem I found in the OpenAI system card in August, now in the document from the company I am applying to.