VEX-Bench: a benchmark for evaluating LLM agents on software supply chain vulnerability triage

Sep 29, 2026 | Featured News, News

by Yuan Tang, Senior Principal Software Engineer, Red Hat

Researchers and engineers from Red Hat and Purdue University have introduced VEX-Bench, the first benchmark for evaluating LLM agents’ ability to determine whether a known vulnerability in a third-party dependency is actually exploitable in a downstream software project. The work has been accepted for presentation at the Conference on Empirical Methods in Natural Language Processing (EMNLP), held October 24-29 in Budapest, Hungary. EMNLP is widely recognized as one of the premier, highly influential international tier-one venues for publishing peer-reviewed research in the fields of natural language processing and artificial intelligence.

Figure 1: VEX-BENCH sources task instances from real-world projects by pairing codebases with third-party dependency vulnerabilities (CVEs) to determine real-world exploitability. Provided with the project’s source code and CVE identifier, LLM agents search external information and analyze code to produce a vulnerability status (Affected / Not Affected), a justification label (e.g., code_not_reachable), and reasoning grounded in the codebase.

Software composition analysis (SCA) tools such as GitHub Dependabot and OSV-Scanner help surface potential exposures by matching package names and versions against vulnerability databases. This approach is deliberately broad: it flags any project that carries an affected dependency version, regardless of whether the vulnerable code is ever reached. The result is a high volume of false alerts (in the VEX-Bench dataset, 70.7% of scanner-flagged cases are not exploitable in practice), and security analysts spend substantial time triaging them case by case. Reachability analysis partially addresses this, but determining whether a call path is feasible, and whether project configuration or runtime environment prevents exploitation, remains beyond current automated tools. LLM agents, with their ability to read code across repositories and reason about advisory text, are plausible candidates for this task, yet until now no benchmark existed to measure how well they perform it.

Table 1: This table reports the composition and codebase size of VEX-BENCH, covering 75 cases across 67 CVEs and 35 real-world projects in Go, Python, and Java, where LOC denotes lines of code and Files denotes the number of source files in the target codebase.

Figure 2: Label distribution of VEX-BENCH across 75 cases, where the inner ring shows the binary vulnerability status (affected vs. not affected) and the outer ring breaks down each status into its fine-grained justification category.

VEX-Bench fills that gap with 75 real-world cases drawn from the top-starred open source repositories on GitHub, covering Go, Python, and Java. Each case pairs a target codebase with a CVE identifier for a known vulnerability in one of its third-party dependencies. The agent operates directly on the full project (a median of 273K lines of code across more than 2,600 files) with no curated file selection, no pre-computed call graph, and no advisory text beyond what it retrieves externally. It must produce a binary exploitability status and, when the verdict is Not Affected, a justification drawn from four categories adapted from the CISA Vulnerability Exploitability eXchange (VEX) vocabulary: “code_not_reachable”, “code_not_present”, “requires_configuration”, and “requires_environment”. Ground-truth labels were assigned by five security experts through a calibration-then-scaling annotation protocol, with an inter-annotator Fleiss’ Kappa of 0.667 at calibration and an 88.3% label agreement rate during the review phase. All experiments run inside isolated Docker containers with no access to benchmark labels, eliminating the possibility that agents circumvent the task by reading evaluation artifacts.

Table 2: Performance of agents on VEX-BENCH, reported as mean±standard deviation over three runs. Vulnerability status metrics treat the task as binary classification; justification metrics evaluate fine-grained reasoning. Token and cost columns report per-case averages. Best and second-best mean values are bolded and underlined, respectively; lower is better for tokens/cost.

The authors evaluated nine models, spanning both closed-source and open weight systems, across three agent harnesses: Claude Code, Codex CLI, and OpenCode. On the binary status task, Claude Opus 4.6 achieves the highest F1 of 81.6% and GPT-5.5 the highest precision at 88.7%; for fine-grained justification, only GPT-5.5 exceeds 70% macro-F1. By comparison, traditional SCA tools (OSV-Scanner, Trivy) achieve F1 scores of 40.5% and 36.1% in the same cases. This is the classic high recall but low precision that reflects precisely the false-alert burden practitioners experience.

Performance consistently drops from the binary to the justification task across all nine configurations, with gaps ranging from roughly 10 to 15 percentage points, indicating that current agents can identify affected projects with reasonable reliability but fall short of the depth of analysis needed to explain why a vulnerability is not exploitable. Among open weight models, DeepSeek-V4-Pro and GLM-5.1 offer accuracy within a few points of the top closed-source configurations at roughly one-tenth the inference cost. The harness framework was found to have a substantially larger effect on per-case token usage than on classification accuracy.

The paper is available on arXiv and will be available throughout the EMNLP 2026 proceedings. Code and data are available at https://github.com/steven1518/vex-bench.

Related Stories

DevConf.cz in your future? Here are some research-related highlights

DevConf.cz in your future? Here are some research-related highlights

DevConf.cz is back in Brno June 18-19. This premier community-led event features a packed agenda of AI, security, and open research talks. It's a strong showcase for how collaborative innovation tackles tough tech challenges. Here we’ve highlighted a few presentations...

 Find AI and far edge research and emerging tech at Red Hat Summit

 Find AI and far edge research and emerging tech at Red Hat Summit

If you’re headed to Red Hat Summit, May 19-22, join engineers from Red Hat Research and Emerging Technologies for labs and breakout sessions that let you get your hands on what they’ve been doing and dive deep into far edge and AI/LLM topics. Click on a presentation...