
Given all the news about Mythos, I ran a small experiment testing Opus 4.6 to understand a bit better how it finds bugs. The setup was: Sendmail crackaddr() bug (CVE-2002-1337) — the original source, a rewritten equivalent, a compiled binary with symbols, and an obfuscated stripped binary. The model found the bug quickly in the first three cases (under 3-4 minutes). The obfuscated version took ~45 minutes of actual "reasoning". A few things stood out: - The model behaves like a human bug hunter would: switching between "pattern matching" and dynamic analysis, using runtime feedback as an oracle - Having an oracle is crucially important. So much so that the agent constructed its own when instructed not to run the binary - The gap between "pattern matching" and "reasoning" capabilities seems significant. The latter appears fairly primitive for Opus. Is Mythos better purely because of the much larger context window or is it something else? - Opus behaves deceptively fairly often. It's surprising how much a hidden scratchpad helps (this is similar to the Sleeper Agents approach) Full writeup: https://vincenzoiozzo.com/blog/alphago-moment-vuln-research
Post summary
The write‑up reports a research experiment using the Mythos/Opus model to detect a known Sendmail bug (CVE‑2002‑1337), highlighting detection efficiency but providing no PoC, exploit, or patch information.
