Mythos proves potent in vulnerability discovery, less convincing elsewhere

Many claims have been made about the effectiveness of Anthropic’s Mythos AI model in the cybersecurity sector both by Anthropic itself and by third parties. Its ability to identify vulnerabilities in software has been described as a fundamental sea change in how the field will operate going forward. However, such claims can verge on the hyperbolic and require the scrutiny of more sustained, comprehensive testing. The autonomous offensive security firm XBOW, an expert at the use of agentic models in cybersecurity, has stepped up to provide such scrutiny, offering a clearer picture than ever of exactly how Mythos operates, where it excels, and where it may fall short of prophetic expectations.


XBOW was given early access to the Mythos Preview model for the sake of testing about two months ago. A panel of 10 experts was assembled to put the model through several benchmark tests, both inside of Claude Code and as a raw model via API. The benchmark tests were the same ones the company applies to other models, including Opus 4.7 and GPT 5.5. These tests primarily consist of using known vulnerable states of open-source applications and tasking the AI model with searching for vulnerabilities. In addition to these tests, XBOW also expanded their testing to analyze other capabilities, including the model’s ability to make judgements on threat modeling, its ability to interact with live systems, and its ability to find unexpected exploits.


The general results of the testing were favorable, with some notes of caution. XBOW has confirmed the efficacy of Mythos at auditing source code, meaningfully outperforming GPT 5.5 and Opus 4.6 at every iteration. In particular, Mythos showed significantly higher ability to identify vulnerabilities at a fixed token budget, outperforming every other tested AI model. The model did suffer a major performance hit when operating off of source code alone rather than source code plus live site access, but that was to be expected, since many vulnerabilities are based on live unsafe interactions between different elements rather than by code errors alone. Even when working on source code alone, Mythos demonstrated higher efficacy at identifying vulnerabilities than other models.


XBOW’s conclusions regarding the model’s judgement in cybersecurity operations was more mixed. When called upon to judge command safety, threat modeling, and trace triage, it delivered responses that XBOW describes as “careful and precise, but also literal and conservative.” The model did deliver fewer false positives than other models, but also had a tendency to lose true positives when evidence did not formally satisfy its criteria. A particular note of caution was sounded about Mythos’s to judge the safety of scripts to execute: it actually underperformed compared to the Haiku 4.5 model. XBOW notes in particular that Mythos would allow scripts that weren’t against the letter of the rules, but were often against the spirit of the rules. This speaks to the model’s tendencies to be literal.


The overall verdict was that Mythos Preview was quite powerful, but limited by some of its drawbacks, as well as by its sheer expense. XBOW’s findings show that when judged on a cost-adjusted token budget, GPT 5.5 outperforms Mythos in vulnerability discovery. As XBOW puts it: “The real choice is to either pay for an agent to use Mythos Premium for a bit, or to use GPT-5.5 for as long as needed. The better option depends on the use case; often, it’s the latter.”

Share

Related Posts

shubham-dhage-2nnRCNuHdVs-unsplash
8machine-_-pzcfw9AV5HY-unsplash
bw-blog_un-1682146029185-198922bd8350

Copyright © All Right Reserved

Privacy Policy