LIVE · cybersecurity feed
Live wire
vulnerability

Using LLMs to Find and Prioritize Vulnerabilities Is No Easy Task

The latest large language models have high false-positive rates and fail to take into account the context of scans, leading to more work for AppSec professionals.

zeroday.news · 11d ago

Recent analysis indicates that the application of large language models (LLMs) in the domain of vulnerability discovery and prioritization presents significant challenges for application security (AppSec) professionals. The primary issues identified are a high rate of false positives and a failure by these models to adequately consider the contextual nuances of security scans. This suggests that while LLMs offer potential, their current implementation in this area is not yet mature enough to reduce, and may even increase, the workload for human analysts.

The reported high false-positive rates mean that a substantial number of potential vulnerabilities flagged by LLMs are not, in fact, exploitable flaws. This necessitates manual review and verification by AppSec teams, diverting resources that would otherwise be spent on addressing genuine threats. This class of issue is common in automated security tools that rely on pattern matching or heuristic analysis, where a lack of deep understanding of code logic or system architecture can lead to misinterpretations.

Furthermore, the inability of LLMs to account for the context of security scans is a critical limitation. Context in AppSec can include factors such as the specific application environment, the intended functionality of a piece of code, the presence of compensating controls, or the overall threat model of a system. Without this contextual awareness, an LLM might flag a benign code pattern as a vulnerability or prioritize a low-risk issue over a more critical one, simply because it lacks the broader understanding of how the component fits into the larger system.

This limitation means that the output from LLM-driven vulnerability scanners often requires extensive human interpretation and refinement. AppSec professionals must manually sift through the reported findings, applying their expert knowledge of the application, its architecture, and its operational environment to distinguish between true vulnerabilities and false alarms. They also need to re-prioritize issues based on actual risk, rather than the raw output of the model.

For organizations considering or currently employing LLMs in their vulnerability management processes, typical mitigation guidance for this class of issue would involve robust post-processing and human oversight. This includes implementing a multi-stage review process where initial LLM findings are filtered and validated by other automated tools or, more critically, by experienced security analysts. Continuous feedback loops, where human corrections are used to retrain or fine-tune the LLM, could also help improve accuracy over time.

The findings highlight a broader trend in the cybersecurity industry regarding the integration of artificial intelligence and machine learning. While these technologies hold immense promise for automating and enhancing security operations, their deployment often uncovers practical limitations related to accuracy, contextual understanding, and the need for human expertise. The current state suggests that LLMs are best viewed as assistive tools that augment, rather than replace, the critical judgment and experience of AppSec professionals in the complex task of identifying and prioritizing software vulnerabilities.

vulnerabilityai
ShareXLinkedInWhatsAppFacebook

More News

view all →
vulnerability

Microsoft blames massive Microsoft 365 outage on maintenance bug

Microsoft says a bug in its automated network maintenance request system caused Thursday's massive outage by mistakenly removing IP routes from more devices than intended, disrupting Azure and Microsoft 365 services. [...]

breach

Hermes AI agent used to automate attack on Thai Finance Ministry

A threat actor used the open-source Hermes AI agent in unattended "YOLO" mode to automate post-exploitation activity during an alleged breach of Thailand's Ministry of Finance. [...]

security

Hackers hijack hotel Wi-Fi DNS to steal Microsoft 365 accounts

Hackers are changing the DNS settings on Wi-Fi devices at hotels and conference centers to redirect users to fake Microsoft 365 login pages. [...]

security

BGP ORIGIN attribute manipulation and its impact on the Internet

By doing in-depth testing, we found nearly 70% of BGP paths experience ORIGIN attribute rewrites by transit providers seeking traffic advantages. We examine the global impact of this practice and argue for deprecating ORIGIN in route selection.

security

Andy Burnham signals continuity on UK cyber policy, reappoints minister despite scrapping ministry

The new British prime minister is retaining Liz Lloyd in a cyber policy role, making her one of the few Keir Starmer allies remaining in government.

security

'Wrench' attacks against crypto holders appear to be on the rise

There are more reports than ever before of strong-arm tactics like home invasions and kidnappings against cryptocurrency holders, researchers say.