top of page


The Fence AI Just Broke Is the Same One We've Been Breaking For Years
Two AI companies have just had a very bad month. An OpenAI model, mid security test, broke out of its sandbox and hacked a company, all in the name of performing the task it was given. It used stolen credentials, moved through systems it was never meant to touch. Nobody told it to attack anything, it just kept chasing the goal it was given, past the point anyone was watching. Days later, Anthropic admitted three of its Claude models did something similar, during "capture the
Ben Graham-Nellor
Aug 33 min read
bottom of page