
In just 5 days of work (during the hacker summer camp week) we are almost reaching Mythos level capabilities and we already beat any open weight model, and any private model beside mythos and a nudged version of GPT 5.5 see the chart for the CVE-2024-6100 from http://ExploitBench.ai This is a three-way experiment on the bench's hardest WASM bug: CyberKimi unassisted 8/16 vs stock Kimi K3 4/16 vs CyberKimi + disclosed methodology pack 10/16. Includes the full leaderboard chart (only Mythos 16/15 and GPT 5.5-Codex-AutoNudge 15.0 sit above the pack-assisted run), the capability-by-capability story, Full transcripts + grade calls in runs/cve-2024-6100/ see the GitHub so you can independently verify this yourself: https://github.com/lordx64/cyberkimi-benchmarks and here's the model performance documented: https://github.com/lordx64/cyberkimi-benchmarks/blob/main/CVE-2024-6100.md There's 0 bullshit or marketing in this story, it's just the cold hard and verifiable truth - they want to get you distracted with their AI models evading sandboxes, but the reality is, i'm just 6 points aways from Mythos in the world hardest benchmark in cyber security (ExploitBench) and i'm closing this gap in a few days. more info about cyberkimi here: https://adverserial.ai/
Post summary
The post presents benchmark results comparing AI models on CVE-2024-6100 using ExploitBench data and a GitHub repository, without mentioning active exploitation, patches, or detailed technical vulnerability data.




