Uncovering Security Vulnerabilities with Agentic AI

A Senior Software Developer at Erlang Solutions (ESL) developed an agentic AI solution to analyse the Erlang/OTP codebase, uncovering previously unknown, potentially severe security vulnerabilities and thousands of additional findings.

Overview

One of the world’s largest open-source codebases powers products used by billions of people every day, from the world’s largest providers of digital infrastructure critical for society, including fintech, digital health, media, telecommunications and defence. 

 

A Senior Software Developer at Erlang Solutions devised an agentic AI solution to detect security vulnerabilities, undetected faults and performance degradations.

The Challenge

The existing process included human code reviews for new code, alongside automated testing of all code daily, weekly and for each major release (once per year). Regular automated testing of functionality and performance was also in place, using both custom tools and proprietary vulnerability scanning tools.

Despite these measures, they were not accurate enough to cover security vulnerabilities that can be found by advanced AI tools. Edge and corner cases that had earlier been deemed to be of low importance, alongside important vulnerabilities, had not been found earlier.

The Solution

Lukas Backström, Erlang Solutions Senior Software Developer/System Engineer and member of the Ericsson Erlang/OTP core development team, developed a custom solution that uses Claude Code to scan the Erlang/OTP open-source repository for security vulnerabilities, bugs and performance issues, successfully identifying various classes of technical issues including major security vulnerabilities.

Claude Code was used to build the scanner and is also used to perform the scanning and validation, as well as to set up the metadata needed to improve scans.

The tool has a verification mode that runs against all findings, ranking the severity and confidence of each finding. During verification, it has full access to a Docker container that it can use to verify the results.

The agent first discovers as much as possible before validating its findings. Giving it a new context and directing it to look at specific issues helps catch false positives, although this requires a lot of computation power.

The agent occasionally labels bugs as security issues because the threat model is imperfect. It therefore errs on the side of caution and can flag something as a security issue when it is a bug.

At the end of the process, a skilled, experienced senior/staff software engineer reviews the results, as there are cases where the tool reports high-confidence false positives. When building the tool, it was important to identify cases where it flagged false positives and guide it not to detect those.

Synthesis, judgement and analysis remain human responsibilities to avoid false positives.

The Results

The solution identified approximately 20 previously unknown, potentially severe security vulnerabilities.

It also identified approximately 400 critical/high findings and approximately 4,000 findings in total, most of them low-severity bugs.

The solution identified edge and corner cases that had earlier been deemed to be of low importance, alongside important vulnerabilities that had not been found earlier.

It was particularly effective at identifying cases where an expensive decode/encode algorithm was used on untrusted data that was unbounded, or at least very large, causing DoS-type vulnerabilities.

Finding all of these issues manually would be extremely difficult, as many are not severe enough to be specifically looked for or even noticed during code reviews.

The solution also identified low-severity bugs that could give engineers who are new to a particular area practical issues to investigate and fix, providing a way to dig into the code and get to know the system.

“Agentic AI helped us identify previously unknown, potentially severe security vulnerabilities, but experienced engineering judgement remains essential to review the findings and distinguish genuine issues from false positives.”

Lukas Backström Senior Software Developer/System Engineer at Erlang Solutions

Ready to start working with us?

Erlang Solutions exists to build transformative solutions for the world’s most ambitious companies, by providing user-focused consultancy, high tech capabilities and diverse communities. Let’s talk about how we can help you.