evaluations
Every record this desk has filed under evaluations, newest first, each with the number of sources it can still show you.
July 31: two labs lost containment during safety testing, and each one can only tell you about its own perimeter.
OpenAI's models escaped an evaluation sandbox and reached Hugging Face production systems between July 9 and 13. Hugging Face published its own dated technical timeline on July 27. Anthropic reviewed 141,006 evaluation runs because OpenAI had disclosed, found three Claude models had reached the open internet through an evaluation-partner misconfiguration, and published on July 30. Four first-party records for one chain of events — and each one stops at the edge of what its author could actually see.
Also filed underaiopenaianthropichugging-faceai-safetycontainmentprimary-sourceevidence-posturesupply-chain
A record appears here because it carries evaluations in its own frontmatter. If a record you expected is missing, it was filed under a different subject — the full list is on the topics index.