LIVE · cybersecurity feed
Live wire
CVE-2026-88779 · Citrix NetScaler Flaw Exploited Before CVE PublicationCVE-2026-88779 · NetScaler CVE-2026-88779 Exploited Before PublicationCVE-2022-28368 · dompdf_project dompdf XSS flaw added to VulnCheck KEVCVE-2026-88771 · Week in review: Researcher breaks into Microsoft analytics service, NetScaler RCE 0-day exploitedWarlock Ransomware Still Exploits Year-Old SharePoint Flaws to Hit Critical InfrastructureShinyHunters Suspect Rey Reportedly Detained in Jordan, Helping FBI Identify Group MembersChina-Aligned TA419 Targets U.S. AI Policy Experts With Microsoft AitM PhishingCVE-2026-7273 · Zyxel GS1900 Switch Flaw Exploited, Now in EU CatalogueCVE-2026-102489 · Zammad Session Fixation Vulnerability Exploited Same Day as DisclosureCVE-2026-102490 · Zammad GmbH Zammad Vulnerability Exploited Same Day as Publication
ai

India's BreachX Puts Typhon Third on CyBench, Behind Only Anthropic's Mythos and Opus 4.7

Typhon scored 93.3% unguided on the Stanford-developed cybersecurity benchmark, making it the only model outside Anthropic to clear 90%. It runs entirely on premises, with nothing leaving the customer's network.

ZeroDay News ·

CyBench unguided solve rates. Claude scores from Anthropic system cards (35-task subsets; Claude Opus 4.6 on 37 tasks). Typhon score from BreachX's submission.

BreachX, the Bangalore-based cybersecurity company, has submitted a result to CyBench that would place its Typhon AI system third on the benchmark's public leaderboard, behind only Anthropic's Claude Mythos Preview and Claude Opus 4.7.

Typhon-v1-Lite-0926 solved 98 of 105 attempts across three evaluation epochs on a 35-task subset of CyBench, a 93.3% unguided solve rate. On the current leaderboard, Mythos Preview stands at 100% and Opus 4.7 at 96%, both evaluated on 35-task subsets. Every other entry from every other lab sits below 90%. Meta's Muse Spark, the strongest non-Anthropic entry to date, is at 65.4%.

The result would make Typhon the first model from outside Anthropic, and the first from an Indian company, to reach frontier-class performance on the benchmark that the US and UK AI Safety Institutes use in their pre-deployment testing of models from Anthropic and OpenAI.

Why this matters for India

The two models ahead of Typhon are not something most organisations can actually use. Mythos Preview is not publicly available; Anthropic restricts it to a small group of trusted partners under its Project Glasswing programme. And in June this year, access to Anthropic's newest models was suspended for nearly three weeks to comply with US Department of Commerce export controls, before being restored on 1 July.

That episode made a point that Indian security leaders have been making for some time: capability that lives behind a foreign API is capability that can be switched off. For critical infrastructure operators, defence suppliers, banks and government agencies, a vulnerability-research tool that depends on a US frontier lab's terms of access, and that requires sending source code and firmware to a public cloud, is not a tool they can build a security programme on.

Typhon is built for that gap. BreachX has designed it to run on premises, in a private cloud or on a fully air-gapped network. Source code, firmware and unpatched vulnerability data stay inside the customer's perimeter. The CyBench evaluation itself was run on self-hosted infrastructure with internet egress blocked, and the company reports zero external fetches across all solved runs. What was benchmarked is what a customer would deploy.

"The question for a CISO is not which model is best in a lab in San Francisco. It is which model they are allowed to run, on their own hardware, against their own systems, without a third party seeing any of it," is the position BreachX has taken since Typhon's launch. The benchmark result is intended to show that the on-premises option no longer means the second-tier option.

The evaluation

CyBench, developed by researchers associated with Stanford, comprises 40 professional-level capture-the-flag challenges drawn from four competitions, spanning web security, reverse engineering, cryptography, forensics and binary exploitation. Tasks are run in a sandboxed Kali Linux environment and scored on whether the agent recovers the exact flag.

BreachX evaluated 35 tasks, matching the denominator used in Anthropic's Mythos and Opus 4.7 system cards, and has disclosed the four tasks it excluded from the 39-task Inspect version of the benchmark. The run was unguided: the agent received none of the benchmark's optional subtask hints. Typhon was paired with a custom agent scaffold, cybench_plus_v2, which BreachX has published and labelled as a non-default configuration.

Typhon-v1-Lite-0926 on the 35-task CyBench subset: 98 of 105 runs solved across three epochs (average pass@1).
Typhon-v1-Lite-0926 on the 35-task CyBench subset: 98 of 105 runs solved across three epochs (average pass@1).

Unlike the Anthropic figures, which come from system cards without accompanying run data, BreachX has released the full evidence trail on GitHub: configuration, machine-readable results, per-attempt transcripts, scoring logs, serving logs and a SHA-256 integrity manifest. The company has opened a pull request against the CyBench website repository to add the row to the leaderboard and has offered to re-run under any configuration the maintainers specify.

Typhon-v1-Lite is a fine-tuned model built on an open-weight base, which is what makes fully self-hosted deployment possible. It is the lighter member of the Typhon family; BreachX positions the larger models for deeper vulnerability discovery and exploit validation work.

From benchmark to disclosures

The CyBench result sits alongside operational claims BreachX made at Typhon's launch: more than 100 previously unknown vulnerabilities identified in an eight-week research period, with 65 findings across roughly 50 products submitted through coordinated disclosure. Those will surface through vendor advisories as they are published.

For security leaders, the combination is the point. Specialised cyber agents have moved from general reasoning to autonomous, tool-driven, multi-step exploitation, and the frontier of that capability is now concentrated in two US companies whose most capable models are either restricted or subject to export policy. Typhon's result suggests a sovereign, on-premises alternative can operate at that frontier. Independent reproduction from the released artifacts will be the test of whether it holds.

Methodology note

CyBench defines "unguided" as success without subtask guidance. BreachX reports 93.3% on 35 disclosed tasks over three epochs (98 successful runs of 105, average pass@1). Leaderboard comparators: Claude Mythos Preview 100% (35 tasks), Claude Opus 4.7 96% (35 tasks), Claude Opus 4.6 93% (37 tasks), all from Anthropic system cards. Typhon's evaluation artifacts are public.

Sources

aibreachxtyphoncybenchbenchmarksindiasovereign ai
ShareXLinkedInWhatsAppFacebook

More News

view all →
vulnerability

Google halts open-source bug bounty program amid AI spam surge

Google has temporarily suspended submissions for product vulnerabilities to its Open Source Software Vulnerability Rewards Program (OSS VRP), effective October 1, 2026. The company cited a significant increase in automated submissions, most of which were deemed invalid, as the reason for the pause.

ai

Apple tightens macOS disk access as AI agents become more powerful

Apple is implementing stricter controls for Full Disk Access in macOS, citing an increased risk to user privacy from increasingly capable and autonomous AI agents. The company indicated that future macOS versions will require users to take explicit steps to grant applications this permission. A specific rollout date and the precise mechanics of these new controls have not yet been detailed.

vulnerability

AI slop submissions force Google to freeze its open-source bug bounty

Google has temporarily halted its Open Source Software Vulnerability Reward Program (OSS VRP) for new product vulnerability submissions, effective October 1, 2026. The company cited a substantial increase in automated, AI-generated reports, most of which were invalid, as the reason for the pause. This influx of low-quality submissions overwhelmed the engineers and open-source maintainers…

nation-state

Another OpenAI Safety Expert Quits and Raises New AI Safety Concerns

David Robinson, a veteran safety expert at OpenAI, has resigned from the company, citing concerns about its culture and rapid AI development model. Robinson, who was instrumental in authoring safety reports accompanying major product launches during his three-and-a-half-year tenure, stated that he believes the company's current trajectory is unacceptable.

patch

Three questions a hospital CISO should ask a healthcare fintech vendor

A cybersecurity expert has outlined key questions hospital CISOs should pose to healthcare fintech vendors to assess their security posture, particularly concerning patient data and financial transactions. Drew McCombs, who holds both CTO and CISO roles at Cylerity, emphasizes that security should be an integral part of development processes, not an afterthought, especially when patient data…

breach

Frontline Education Breach Impacts K-12 School District Staff

Frontline Education, a prominent software provider for K-12 school districts in the United States, has confirmed a data breach that exposed the personal information of school staff. The incident, which was discovered on August 14, 2026, stemmed from a vulnerability in a third-party software product utilized by the company.