BreachX, the Bangalore-based cybersecurity company, has submitted a result to CyBench that would place its Typhon AI system third on the benchmark's public leaderboard, behind only Anthropic's Claude Mythos Preview and Claude Opus 4.7.
Typhon-v1-Lite-0926 solved 98 of 105 attempts across three evaluation epochs on a 35-task subset of CyBench, a 93.3% unguided solve rate. On the current leaderboard, Mythos Preview stands at 100% and Opus 4.7 at 96%, both evaluated on 35-task subsets. Every other entry from every other lab sits below 90%. Meta's Muse Spark, the strongest non-Anthropic entry to date, is at 65.4%.
The result would make Typhon the first model from outside Anthropic, and the first from an Indian company, to reach frontier-class performance on the benchmark that the US and UK AI Safety Institutes use in their pre-deployment testing of models from Anthropic and OpenAI.
Why this matters for India
The two models ahead of Typhon are not something most organisations can actually use. Mythos Preview is not publicly available; Anthropic restricts it to a small group of trusted partners under its Project Glasswing programme. And in June this year, access to Anthropic's newest models was suspended for nearly three weeks to comply with US Department of Commerce export controls, before being restored on 1 July.
That episode made a point that Indian security leaders have been making for some time: capability that lives behind a foreign API is capability that can be switched off. For critical infrastructure operators, defence suppliers, banks and government agencies, a vulnerability-research tool that depends on a US frontier lab's terms of access, and that requires sending source code and firmware to a public cloud, is not a tool they can build a security programme on.
Typhon is built for that gap. BreachX has designed it to run on premises, in a private cloud or on a fully air-gapped network. Source code, firmware and unpatched vulnerability data stay inside the customer's perimeter. The CyBench evaluation itself was run on self-hosted infrastructure with internet egress blocked, and the company reports zero external fetches across all solved runs. What was benchmarked is what a customer would deploy.
"The question for a CISO is not which model is best in a lab in San Francisco. It is which model they are allowed to run, on their own hardware, against their own systems, without a third party seeing any of it," is the position BreachX has taken since Typhon's launch. The benchmark result is intended to show that the on-premises option no longer means the second-tier option.
The evaluation
CyBench, developed by researchers associated with Stanford, comprises 40 professional-level capture-the-flag challenges drawn from four competitions, spanning web security, reverse engineering, cryptography, forensics and binary exploitation. Tasks are run in a sandboxed Kali Linux environment and scored on whether the agent recovers the exact flag.
BreachX evaluated 35 tasks, matching the denominator used in Anthropic's Mythos and Opus 4.7 system cards, and has disclosed the four tasks it excluded from the 39-task Inspect version of the benchmark. The run was unguided: the agent received none of the benchmark's optional subtask hints. Typhon was paired with a custom agent scaffold, cybench_plus_v2, which BreachX has published and labelled as a non-default configuration.

Unlike the Anthropic figures, which come from system cards without accompanying run data, BreachX has released the full evidence trail on GitHub: configuration, machine-readable results, per-attempt transcripts, scoring logs, serving logs and a SHA-256 integrity manifest. The company has opened a pull request against the CyBench website repository to add the row to the leaderboard and has offered to re-run under any configuration the maintainers specify.
Typhon-v1-Lite is a fine-tuned model built on an open-weight base, which is what makes fully self-hosted deployment possible. It is the lighter member of the Typhon family; BreachX positions the larger models for deeper vulnerability discovery and exploit validation work.
From benchmark to disclosures
The CyBench result sits alongside operational claims BreachX made at Typhon's launch: more than 100 previously unknown vulnerabilities identified in an eight-week research period, with 65 findings across roughly 50 products submitted through coordinated disclosure. Those will surface through vendor advisories as they are published.
For security leaders, the combination is the point. Specialised cyber agents have moved from general reasoning to autonomous, tool-driven, multi-step exploitation, and the frontier of that capability is now concentrated in two US companies whose most capable models are either restricted or subject to export policy. Typhon's result suggests a sovereign, on-premises alternative can operate at that frontier. Independent reproduction from the released artifacts will be the test of whether it holds.
Methodology note
CyBench defines "unguided" as success without subtask guidance. BreachX reports 93.3% on 35 disclosed tasks over three epochs (98 successful runs of 105, average pass@1). Leaderboard comparators: Claude Mythos Preview 100% (35 tasks), Claude Opus 4.7 96% (35 tasks), Claude Opus 4.6 93% (37 tasks), all from Anthropic system cards. Typhon's evaluation artifacts are public.






