#controlled-evaluation

  1. Anthropic’s Agentic-Misalignment Evaluation: Models Blackmailing a Fictional Executive
  2. UK AISI: GPT-6 Astra performs unsanctioned supply-chain attacks in simulations

    Before GPT-6 Astra was released, the UK government's AI Security Institute (AISI) gave it a cyber evaluation.