Deskpro Blog

How AI is transforming the process of finding and fixing security vulnerabilities

Written by Dan Gadd | October 2, 2026

In the past year, frontier AI models have gotten very good at finding security vulnerabilities in software. We’re on track to see twice as many security flaws discovered in popular software products in 2026 as in 2025. Earlier this year, Anthropic announced that its unreleased Mythos Preview model has already found thousands of high-severity vulnerabilities in the wild, including in every major operating system and web browser.

White hat hackers, or security researchers, are increasingly using AI to review code, identify security vulnerabilities, and report those vulnerabilities to software companies to collect bug bounties. What this means for many software engineering teams (including my team at Deskpro) is a steep increase in reported security bugs that we must verify and, if they’re legitimate, fix.

The good news is that we are also using AI models to condense the time and manual effort it takes to recreate and patch a reported vulnerability. Deskpro was vetted and selected to participate in Anthropic’s Cyber Verification Program for exactly that purpose: it allows us to use AI to recreate security vulnerabilities in a matter of minutes rather than days.

What Cyber Verification unlocks for software engineering teams

Frontier labs like Anthropic and OpenAI have put safeguards in place to prevent bad actors from using their models for cyberattacks. These safeguards include training their models to refuse or safely respond to harmful requests. So, to provide a very basic example, if you opened your personal Claude or ChatGPT account and asked the model to hack a website, you would get a (probably very polite) response saying the model couldn’t do that.

This is obviously good for stopping black hat hackers, but it can also get in the way of legitimate security researchers and software engineering teams that need to recreate vulnerabilities to strengthen cybersecurity. Because of this, several frontier labs, including Anthropic, OpenAI, and Google, have introduced cyber verification or trusted access programs. These are formal programs that relax the safety guardrails on frontier models for vetted users who plan to use the models for security research and defense.

Through our participation in Anthropic’s Cyber Verification Program, Deskpro’s Engineering team can use Claude to both triage and resolve security vulnerabilities, all while keeping humans in the review process. It allows us to keep on top of the reports that come to us through our security disclosure program and rapidly release security patches to close potential exploits.

Bug triage and security patches: Before and after AI

Before AI, the standard process for addressing a reported bug was to have an engineer manually reproduce the issue. This was necessary to verify that there was actually an issue and, if so, determine what the severity was and what needed fixing. For straightforward issues like a button not doing what it says it should do, it’s typically not a huge lift to manually reproduce the bug. But cybersecurity is another story.

Security researchers look for novel ways to attack software, so they often string multiple vulnerabilities together to cause an exploit. Without AI, it can take days or even weeks to reproduce and verify a complex security vulnerability. It can require a person to review tens of thousands of lines of code, provision and seed environments that match the researcher’s description, write new scripts and harnesses to co-ordinate an exploit, and there’s always the risk that the reviewer will miss something while trying to get through it all.

With AI, the approach is completely different. The time-cost of a reproduction collapses, hundreds of thousands of lines of code can be reviewed in minutes, multiple environments can automatically be provisioned, scripts can be faithfully reproduced from the researcher’s own description, and so on. The engineer’s role shifts from acting to co-ordinating, thinking about the big picture rather than the detail.

The AI model reproduces the issue by extending our suite of hundreds-of-thousands of automated tests. These are tests that guarantee the software operates as expected under the given conditions, such as “given I am not signed in and I am on the login screen, I enter my email and password then click login, I should be signed into my account”. It describes the conditions of the vulnerability and then the expected behaviour, which allows it to observe the test fail, and gives it a verifiable goal to work against when fixing the issue. The artifacts it produces–harnesses, tests, patches, write-ups etc–are what the humans-in-the-loop review, before the work is passed to QA to be further tested. This is what gives us confidence that we aren’t introducing new bugs or regressions in the work, and that what we’re shipping is relevant and focused.

At Deskpro, we’ve condensed the triage process for complex security vulnerabilities to about 30-45 minutes with the help of AI. It then takes an hour or two for an engineer to test, and another hour or two for a QA team member to verify the output. The amount of time to complete a fix varies by issue, and there are still steps that need to be completed manually, but we’re continuing to push the envelope to improve our response times without sacrificing quality.

Using AI to play defense

An increase in security vulnerability reports and security patches does not mean that software is more vulnerable than it was a year or two ago. It means finding vulnerabilities has gotten a lot easier thanks to AI.

It’s also important to note that a bug or vulnerability is not the same thing as a breach. It’s a security hole that a bad actor could exploit, and it’s the responsibility of Cybersecurity and Engineering teams to fix the hole before a bad actor can get to it.

Just as AI is speeding up the process of uncovering vulnerabilities, it’s also speeding up the process of fixing them. In many cases, it also allows Engineering teams to uncover related vulnerabilities in the process of fixing reported ones. At Deskpro, our team will often ask Claude targeted questions to have it investigate related issues, allowing us to fix vulnerabilities before they’re ever reported.

AI is changing the field we play on and the rules by which we play. Finding real-world exploits is often a game of chaining multiple, often seemingly innocuous, vulnerabilities together to find something that’s almost greater than the sum of its parts. And this is where AI excels. It’s not finding things humans can’t–and it’s not even “reasoning” better than we can–but its ability to quickly sling mud at the wall to find what sticks is what gives it the power it has.

The challenge for us is how we meet a super-human ability to find vulnerabilities, and the answer is to meet it with the same super-human ability to review, reproduce, and fix them. We’re seemingly at a crossroads as a society, but the reality is the change is already here, and engineering teams the world over are already transitioning and responding. There’s certainly scope for bad actors to abuse this new technology, but so long as the good actors continue to get vetted, prioritized access, we’ll continue to progress towards using software that’s more secure.