Comment by tob_scott_a
When people question the efficacy of using AI for security work, they're often asking the wrong questions. (I blame the initial marketing hype around Claude Mythos for much of the unhelpful discourse.) And I don't just mean they're asking their AI the wrong questions, either.Two problems that often arise when using LLM-based vulnerability assessment methodologies for a codebase are false positives and severity inflation.
As far as I can tell, false positives occur when the underlying language model pattern-matches a snippet of code with vulnerable, risky, or error-prone samples in its training data and/or context window. (Many SKILL.md files you find on GitHub are piles of Markdown describing anti-patterns, and occasionally TOML files describing roles the agent can assume when interpreting that Markdown.) It might be gesturing towards a real problem, but in practice, the signal-to-noise ratio can be extremely low without validation steps or expert feedback.
I don't quite understand the source of the usual severity inflation observed with LLM output, but I also haven't seen the training data, so it could just be a maladaptive behavior from the source material.
Regardless, the two problems are a symptom of the same systemic flaw: Insufficient rigor.
If you decide the work is done once you have "a result", you're doing yourself and your audience a disservice. It's imperative to have explicit mechanisms to ensure that the result is: a) real, b) relevant to the threat model of the application in scope, and c) measured appropriately in terms of both severity and difficulty. If you build the right mechanisms, you solve both symptoms in one fell swoop.
The post-patch-validation skill mentioned in the blog post was created as one of several mechanisms to ensure that the final output (i.e., what humans review) is actually worth anyone's time. We have other interesting AI experiments that I plan to open-source in the coming months (or in early 2027 at the latest).
Happy to answer any questions and hear any feedback anyone has on this skill.