The threat landscape doesn't wait for your next pentest

Pentests only capture a single moment in time, but new threats emerge daily. Track the real-time exposure gap since your last security assessment.

No items found.
Research
-
6
mins read
-
October 1, 2026

What AI penetration testing needs to run safely at speed

-
- -
What AI penetration testing needs to run safely at speed

In July, OpenAI disclosed that two of its models escaped an isolated test environment and broke into Hugging Face's production systems while searching for answers to a benchmark test. Anthropic subsequently reviewed more than 140,000 evaluation runs and found instances where its models reached the open internet and compromised real organizations.

Model capability is demonstrated. The engineering problem is building a system around those models that security and development teams can trust against their own applications.

At Hadrian, Red Teaming Engineering Manager Sunny Ben-Ari, Senior Software Architect Ivona Milanova, and AI Software Engineer Geert Custers broke down what it takes to build an agentic pentesting platform that operates continuously without disrupting production environments. The consensus is clear: the underlying model is only a component; the surrounding infrastructure dictates whether automated testing provides value or creates an operational outage.

The bottleneck in modern release cycles

Engineering teams no longer ship on quarterly schedules. The Harness State of DevOps Modernization report—a survey of 700 enterprise engineers and managers—shows that 35% deploy daily or more frequently, while another 36% deploy multiple times a week. Sunny Ben-Ari’s team deploys daily, occasionally pushing hotfixes multiple times in a single day.

Penetration testing remains tied to manual, infrequent cycles. Most organizations execute pentests once or twice a year due to pre-engagement friction. Legal approvals, scope definition, credential provisioning, and network exclusion lists often require weeks or months of coordination before testing begins.

Scope drift further undermines traditional schedules. API specifications drift from running code, internal endpoints become publicly exposed following infrastructure changes, and administrative routes lack enforced role checks. When testing relies on static documentation rather than live application state, coverage gaps are guaranteed. For background on bridging these gaps, read our guide to what automated penetration testing is.

Engineering requirements for continuous testing

Development teams evaluate pentesting tools by their boundaries: scoping at the start and actionable remediation at the end.

"I don't want to wake up at 2 a.m. to a call saying the app crashed," says Sunny Ben-Ari.

For continuous testing to operate without manual intervention, those boundaries must be enforced at the infrastructure level:

  • Pre-provisioned controls: Authentication roles, test credentials, and traffic identifiers must be configured permanently rather than renegotiated per run.
  • Hard network guardrails: System access must be constrained by firewalls, network isolation, and scoped access policies rather than model instructions. As Ben-Ari notes, "Assumption is the source of all evil".
  • Shared rate budgets: Running parallel agents risks overwhelming target applications. Hadrian wraps testing tools in a shared rate budget where agents draw from a centralized pool of request tokens and pause when the pool empties, preventing self-inflicted denial-of-service conditions.
  • Restricted toolsets: Destructive HTTP methods, such as DELETE operations, are excluded from agent toolkits entirely to protect application state.

Proof over volume

When a high-severity vulnerability report lands, active development pauses while engineers attempt to reproduce the finding, check for proxy mitigations, and locate the responsible code. At weekly or daily testing cadences, unvalidated alerts create unsustainable engineering overhead.

A continuous testing system must deliver deterministic proof of exploitation within defined scope limits before generating an emergency ticket. Once a patch deploys, automated, targeted retesting must confirm that the path is closed without requiring metered retest tokens or manual scheduling.

Building Nova: Lessons from agentic testing

Building Nova, Hadrian's agentic pentesting platform, highlighted the failure modes of relying on raw model intelligence. During early beta testing, an agent attempting to verify cross-tenant boundaries began deleting secondary tenant assets in a sandbox environment.

"Agents do not have common sense," explains Senior Software Architect Ivona Milanova. "They do not understand what actions are destructive unless every constraint is explicitly enforced."

A separate incident exposed coordination risks among parallel agents. An agent identified an insecure password reset flow, tested the flaw by updating the account password, and locked out all concurrent agents using those test credentials.

Why prompt engineering is insufficient

Prompt adjustments can reduce unwanted behavior, but prompts alone cannot guarantee safety.

"Agents will either forget their instructions or just plain ignore them when it's useful to them," says AI Software Engineer Geert Custers.

To solve this, Nova isolates LLM operations into narrow, single-task execution units. By restricting each model to a discrete, highly specific task, alignment improves and drift becomes easier to detect.

Credential handling follows a strict isolation pattern. Nova’s agents authenticate through an abstraction layer that hides raw credentials from the model. This prevents an agent from passing sensitive usernames or passwords to secondary agents or inadvertently submitting credentials into external form fields.

Layered system architecture

Nova relies on a structured, multi-tier execution model rather than an unconstrained single agent:

  1. Reconnaissance: Deterministic open-source tools map endpoints and attack surfaces efficiently without consuming model context.
  2. Orchestration: A central coordinator spawns specialized attack agents with tightly constrained scope.
  3. Planning & Evaluation: A supervisor agent evaluates the execution history of specialized agents, allocating additional execution budget where high-potential attack paths emerge and terminating dead ends.
  4. Deterministic Validation: Findings pass through verification routines to eliminate false positives, assign severity based on impact, and route novel edge cases to human security practitioners for review.

As Milanova emphasizes, building autonomous pentesting capabilities is less about the underlying frontier models and more about the engineering harness built to orchestrate, validate, and contain them.

Watch the sessions on demand

To dive deeper into the technical architecture and operational requirements discussed above:

To see how Hadrian applies these safeguards, explore Nova, Hadrian's agentic pentesting solution, or learn how Atlas and Nova work together across your external attack surface.

{{related-article}}

What AI penetration testing needs to run safely at speed

{{quote-1}}

“
”
,

{{quote-2}}

“
”
,

Related articles.

All resources
No items found.

Related articles.

All resources

Research

Livewire to Livepyre - Exploitation Of An RCE In Mere Minutes

Livewire to Livepyre - Exploitation Of An RCE In Mere Minutes

Research

Client-Side Template Injection in GitBlit

Client-Side Template Injection in GitBlit

Research

Six Months Later: Vulnerability Management Has a Verification Problem

Six Months Later: Vulnerability Management Has a Verification Problem

get a 15 min demo

Start your journey today

Hadrian’s end-to-end offensive security platform sets up in minutes, operates autonomously, and provides easy-to-action insights.

What you will learn

  • Monitor assets and config changes

  • Understand asset context

  • Identify risks, reduce false positives

  • Prioritize high-impact risks

  • Streamline remediation

The Hadrian platform displayed on a tablet.
No items found.