Can Agentic AI Find the Security Vulnerabilities Other Tests Miss?
- Erik Schön
- 8th Oct 2026
- 8 min of reading time
Security teams are good at finding problems. The difficult part is finding the right problems across a growing codebase, a changing application and the connections between them. A scanner can flag a known pattern quickly. A skilled tester can follow a complicated line of enquiry. Neither has unlimited time, and there will always be a question nobody thought to ask.
Could AI agents take on more of that investigative work? I think the answer is yes, provided we are precise about what “find” means. Agents can produce leads, test assumptions and uncover issues that established checks have missed. A security vulnerability, though, has to survive examination in the context of the real system.
That distinction matters more as agents become available to both defenders and attackers. A familiar weakness can become more relevant when agents have the time to find it, test it and follow where it leads. Defenders have a good reason to look again at parts of their systems that were too laborious to examine thoroughly.
Think about the code paths that work perfectly well in normal use but behave differently with a surprising input or an unusual amount of data. Fuzzing (testing software with unexpected or malformed inputs) helps explore these cases; agents can then read the surrounding code, try a more specific hypothesis and keep working through leads that would consume too much of an engineer’s day if checked one by one.
Some issues only become apparent when you understand how the application is used. A finding may look harmless in isolation until you establish who controls the input, which permissions are involved or what the application is meant to prevent. Agents can explore a sequence of actions and check how the application responds at each step, which may reveal a problem that would be hard to spot from an isolated code warning.
There is also the possibility that one weakness opens the way to another. A low-priority issue could expose useful information; a forgotten endpoint might give access to a service assumed to be isolated. Agents can follow a discovery into the next step, although its account of that path still needs to be checked. The full route through a system may matter more than the severity of any one weakness in it.
I would be careful about grouping all of this under one claim that “AI finds what traditional testing misses”. Code analysis tools, fuzzing tools, formal verification and agents conducting a permitted penetration test are doing different jobs. Their findings have different levels of evidence. Formal verification is also starting to gain traction, with code agents proving useful for generating the specifications it requires. The useful question is where agents give the team another line of enquiry and how that enquiry will be validated.
We have seen one version of this in our own work. Lukas Backström, a senior developer at Erlang Solutions and a member of the Ericsson Erlang/OTP core development team, built a custom workflow using Claude Code to examine the Erlang/OTP repository for security vulnerabilities, bugs and performance issues.
This codebase already receives human review, regular automated functional and performance testing, and vulnerability scanning. Even so, agents surfaced cases where expensive encoding or decoding could be performed on untrusted data that was very large or unbounded, creating a possible denial-of-service risk. The operation itself was not mysterious. Finding the relevant combination of input and behaviour across a large repository was the challenge.
Our published case study reports approximately 20 previously unknown, potentially severe security vulnerabilities, alongside around 400 critical or high findings and roughly 4,000 findings overall, most of them low-severity bugs. These are reported findings from the investigation, not a claim that thousands of confirmed security vulnerabilities were waiting to be fixed. And, this is one of the world’s most thoroughly tested code bases.
The project gives us evidence for a particular point: agents helped investigate parts of an extensively tested codebase that would have been difficult to cover manually. It does not prove that agents will outperform every scanner or security team. The details of the task and the review process made the result useful.
Lukas gave agents a verification mode, including access to a Docker container to check findings and rank their severity and confidence. He also found that a fresh context and a focused second look at an issue helped expose false positives. That separate work took substantial compute, and the agents still sometimes returned a high-confidence finding that was wrong. It could also mistake a bug for a security issue because it misunderstood the threat model.
An engineer has to ask whether an attacker can reach the relevant code, what input they control, what limits exist elsewhere and what happens in a real deployment. A persuasive explanation from an agent is a reason to investigate, not evidence that every assumption is true.
There is a second challenge once agents start finding more faults. Most teams do not need a larger pile of unverified tickets. They need to know which findings are real and which ones deserve attention first. Every result takes time to assess, and asking agents to assess it again can add to the compute cost. Finding more vulnerabilities only helps if the team has enough context to act on what matters.
The investigative agents must also be secure in its own right. It may read code, retrieve documents and run tools, so its permissions and execution environment deserve scrutiny. Untrusted material could contain instructions intended to redirect its behaviour. The team needs to decide what the agents can access and do before giving it sensitive systems to examine.
I would treat an agentic investigation as an experiment with a clear question, rather than asking agents to “find everything” and judging success by the length of its report:
So, can agentic AI find security vulnerabilities other tests miss? Yes, it can uncover leads that existing checks have not surfaced, particularly when an issue takes sustained investigation to find. Our Erlang/OTP work showed that in a codebase with established reviews, automated tests and vulnerability scanning. It was not a test-by-test comparison, so we cannot say which individual check missed each issue.
Finding a lead is only part of the answer. An experienced engineer still needs to establish whether it is a genuine security vulnerability and what should be fixed. The opportunity is to investigate more of the questions a team has never had the time to pursue, then give those findings the scrutiny they deserve.
Dmytro Lytovchenko builds a live chiptune synthesizer in Erlang using equations and generated audio.
Brian Underwood explains how to build a Phoenix PubSub adapter backed by EventStore.
Learn how to use Lua for flexible configurations in Erlang and Elixir with Lee Sigauke.