Two Leading AI Models Achieved High Success Rates in Complex Cybersecurity Tasks
The UK AI Security Institute published a blog post presenting an evaluation of the GPT-5.5 model and a new version of Claude Mythos Preview. The assessment aims to measure the models’ ability to perform cybersecurity tasks at a success rate of at least 80% relative to the time required by human experts. The evaluation found that both models were capable of completing tasks that human cybersecurity experts typically perform over the course of several hours. In particular, both models completed the longest tasks with success rates approaching 100%, despite being subject to a limit of 2.5 million tokens per task. The assessment also found that the two models were the only ones able to complete all 32 stages of “The Last Ones“ attack scenario, which simulates a breach of a corporate network. In addition, Mythos Preview successfully completed the “Cooling Water” attack scenario—a seven-stage attack against industrial control systems—an achievement not previously reached by any AI model.