VEX-Bench: a benchmark for evaluating LLM agents on software supply chain vulnerability triage

by Yuan Tang and Edward Tsien Researchers and engineers from Red Hat and Purdue University have introduced VEX-Bench, the first benchmark for evaluating LLM agents’ ability to determine whether a known vulnerability in a third-party dependency is actually exploitable in a downstream software project. The work has been accepted for presentation at the Conference on … Continue reading VEX-Bench: a benchmark for evaluating LLM agents on software supply chain vulnerability triage