I actually have an update for you on this! AISLE also came after Mythos (having found a single low-severity CVE + 4 things that turned out to be false positives), and actually got 6 (!!) new CVEs assigned that Mythos has completely missed, including the oldest issue ever to be found in curl. Check this post out by Daniel Stenberg, the lead developer of curl himself: https://mastodon.social/@bagder/116807317163361428
In addition, we are very clearly able to go into a codebase after Mythos is "done" with it, and find zero-days that it misses and have them assigned with publicly tracked CVEs (even in FreeBSD, the very codebase Anthropic chose to demonstrate Mythos' capabilities). I make the data-driven case for it here: https://stanislavfort.substack.com/p/mythos-at-home-and-its-called-aisle, I think you might enjoy reading it!
Lastly, UC Berkeley is running a continuous academic study on the roly of AI in zero-day discovery (https://vuln.cs.berkeley.edu/). AISLE outranks Anthropic in 4 of the 8 categories they track, namely:
total number of zero-days with CVEs (the only way to actually track them reliably)
total number of critical zero-days with CVEs
coverage of vulnerability types (CWEs)
coverage of the MITRE top-25 most important vulnerability types
It is very clear that Mythos is not by any margin significantly better in finding zero-days than other, even much cheaper models, purposefully finetuned for the task + with a great harness.
Disclaimer: I'm the founder and CTO/Chief Scientist of AISLE
I actually have an update for you on this! AISLE also came after Mythos (having found a single low-severity CVE + 4 things that turned out to be false positives), and actually got 6 (!!) new CVEs assigned that Mythos has completely missed, including the oldest issue ever to be found in curl. Check this post out by Daniel Stenberg, the lead developer of curl himself: https://mastodon.social/@bagder/116807317163361428
In addition, we are very clearly able to go into a codebase after Mythos is "done" with it, and find zero-days that it misses and have them assigned with publicly tracked CVEs (even in FreeBSD, the very codebase Anthropic chose to demonstrate Mythos' capabilities). I make the data-driven case for it here: https://stanislavfort.substack.com/p/mythos-at-home-and-its-called-aisle, I think you might enjoy reading it!
Lastly, UC Berkeley is running a continuous academic study on the roly of AI in zero-day discovery (https://vuln.cs.berkeley.edu/). AISLE outranks Anthropic in 4 of the 8 categories they track, namely:
It is very clear that Mythos is not by any margin significantly better in finding zero-days than other, even much cheaper models, purposefully finetuned for the task + with a great harness.
Disclaimer: I'm the founder and CTO/Chief Scientist of AISLE