What continuous offensive security testing is and how to implement it

Introduction

Continuous offensive security testing (COST) is a model for running adversary-style tests whenever something material changes, such as a release, a new internet-facing asset, or a change to a security control.

Instead of the standard once a year test, Gartner introduced COST as a term to describe this trigger-driven intelligence-led approach. Escape's evidence standard supports that idea: a finding only counts after it has been reproduced in your own environment, with the requests, identities, and responses another tester needs to repeat it.

If you're supporting several engineering teams with a small AppSec or pentesting group, a yearly assessment leaves long gaps, and more scans only add results without showing which exposures are reachable.

This guide explains how the pieces in COST fit, then walks through a small local harness you can run yourself.

⚡ tl;dr

  • Broader than pentesting. Offensive security includes vulnerability assessment, penetration testing, and red teaming.
  • Triggered by change. Tests can run when applications, assets, threats, or security controls change.
  • Built around validation. Scores help you prioritize. Reproducing an issue in your environment shows whether it's reachable.
  • Designed to keep working. Cascade validates each finding, and DAST re-runs it as regression coverage on later builds.

Offensive security testing is wider than a pentest

Pentesting still matters, but it covers only one part of a broader offensive security program. Knowing what each testing method is for helps you spend limited time where it matters.

Offensive security isn't just a pentest

What continuous offensive security testing means

Continuous offensive security testing, which Gartner calls COST, applies adversarial testing on an ongoing or event-driven basis. A release might trigger one test. A new internet-facing asset, an authentication change, or relevant threat intelligence might trigger another.

💡
Continuous offensive security testing is attacking your own systems the way an adversary would, run whenever your environment changes rather than once a year, where a finding counts only after it has been reproduced in your own environment. Gartner calls it COST.

The Cloud Security Alliance describes offensive security as proactively simulating attacker behavior to identify vulnerabilities and strengthen security controls.

This doesn't mean running every test after every commit. It means choosing the right test when risk changes, then deciding what evidence a finding needs before it reaches an engineering team.

Where a pentest fits

The same Cloud Security Alliance article identifies three major approaches:

  • Vulnerability assessment, the automated identification of weaknesses using scanners.
  • Penetration testing, the simulation of attacks to identify and exploit vulnerabilities.
  • Red teaming, the simulation of complex multi-stage attacks to test detection and response.

A pentest goes deep on a single question: whether a given application can be broken. A broader offensive security program asks three more, what you have exposed across your attack surface, who can reach it through identity and access, and whether your controls and detection would notice.

The European Central Bank's TIBER-EU framework shows what function-based scoping looks like, covering critical functions and the people, processes, and technologies behind them. It is not a continuous testing framework, but the scoping lesson still applies: start with what the business needs to protect, not the boxes in an asset list.

How it differs from continuous penetration testing and continuous threat exposure management

These four terms get used as if they mean the same thing, but they describe different things.

COST vs continuous pentesting vs CTEM vs PTaaS: what's the difference?

Term What it means When you use it What you get
COST (continuous offensive security testing) Attack-testing your systems whenever something changes, not once a year. Threat intel flags an actively exploited CVE in a library you run, so that service is tested within hours instead of at the next quarterly. Proof a flaw is real in your environment, tied to the change that surfaced it.
Continuous penetration testing A hands-on pentest of one target, run on repeat instead of annually. Your payments API is retested for access-control and auth-bypass flaws on every release, catching regressions the week it ships. Fresh, deep findings on that one target after every release.
CTEM (continuous threat exposure management) A rolling program that finds, ranks, proves, and fixes exposure across the org. You inventory the external attack surface, rank issues by real exploitability, validate the top ones, route them to owners, then rerun the loop. A ranked, owned, closed-loop view of exposure across the whole org.
PTaaS (penetration testing as a service) A platform for ordering and tracking pentests, instead of a one-off consultant PDF. You launch a scoped test from a dashboard, watch findings land, and request a retest once engineering ships the fix. Tests on demand and findings in a dashboard, with retests built in.

Scanning and attack surface management sit earlier in that chain. A scanner identifies weaknesses and attack surface management inventories what's exposed, but neither proves that an attacker could reach a given asset in your environment. That proof is what this model exists to provide.

A wider scope gives your team more ground to cover and more results to review, which makes the tooling matter as much as the schedule. When an AI model is doing the testing, the system around it controls what the model can do, what it remembers, and what gets reported as a finding.

Why the system around the model matters

A 2025 study of language-model architectures for penetration testing found that the right supporting capabilities improve agent performance, especially on complex, multi-step tasks. The researchers tested context memory, communication between agents, more reliable tool use, adaptive planning, and real-time monitoring.

An AI pentesting system needs more than a model and a prompt. It needs tools that fit the application, memory that lasts across steps, and a way to check findings before anyone acts on them.

The example below builds that in miniature, giving the model three things:

  • Structured actions, so the harness can reject malformed output before it reaches the target.
  • Context memory, so the model can use observations from earlier steps.
  • Validation, so a report has to match what the target returned.

They don't make model quality irrelevant, but they give the model a controlled way to test the target and give your team evidence to review.

More testing requires better validation

There are five levels of evidence to consider

If your team already has a noisy queue, more frequent testing can make it worse. Decide what evidence you need and how you'll prioritize it before you add more results.

Use scores for prioritization, not validation

Severity and probability scores help you decide where to look first. They don't show whether a weakness is reachable in your environment. FIRST explains that EPSS estimates the chance of exploitation activity in the wild over the next 30 days. It doesn't account for your environment or compensating controls.

CISA's Known Exploited Vulnerabilities catalog answers a different question. It records CVEs with reliable evidence of active exploitation in the wild. Scanning, security research, and a published proof of concept don't qualify on their own under the criteria carried forward by Binding Operational Directive 26-04.

Neither EPSS nor KEV tells you whether a vulnerability works against your deployment, so you still need to test the affected system. Keep the scores; they're useful inputs. Just don't treat them as a substitute for local validation.

Build the program around proof

Organizations don't need a complicated framework to begin. Start with clear triggers, a shared evidence standard, and a route from each validated finding to the team that owns the fix.

Set your triggers, tiers, and routing

Four triggers are enough to get most teams started: a newly discovered asset, a production release, relevant threat intelligence, and a change to a security control.

Tier each event by exposure and business impact. Scope the test around the affected business function, then route findings using environmental evidence alongside severity and threat intelligence.

CISA's Stakeholder-Specific Vulnerability Categorization follows a similar principle. It considers exploitation status alongside technical impact, automatable exploitation, mission prevalence, and public-wellbeing impact. A locally reproduced exploit gives you valuable evidence, but it doesn't determine the SSVC outcome on its own.

Trigger Tier Response window Trigger source
New internet-facing asset, or a service using a product and version affected by a KEV entry Top Hours Asset inventory + exploitation-status intelligence
Release reaching production, or a change to an authentication or authorization control Middle Days Deployment pipeline + control-configuration changes
Other in-scope systems Baseline Defined test cycle Scheduled testing

Each trigger needs a reliable source, whether that's your asset inventory, threat-intelligence feed, deployment pipeline, or security-control configuration.

A scheduled job can read CISA's KEV feed and flag new entries tied to vendors in your inventory. That's only a first filter. Confirm the affected product and version before you launch a test.

Track how much of the in-scope estate you've tested and how many candidates become validated findings. Then measure how long it takes to reach a decision after a trigger.

Run one trigger end to end

Say a team releases a new internet-facing payments API. Here's how the program handles it:

  1. Select the response tier. A new internet-facing asset deserves prompt testing.
  2. Scope the critical paths. For payments, test authentication, authorization, transaction state, and connected services.
  3. Validate candidates. Capture the requests, identities, responses, and impact needed to reproduce each issue.
  4. Route the finding. Send the evidence and priority to the service team.
  5. Add regression coverage. Turn the reproduction steps into an automated test where the vulnerability allows it.

Write the evidence standard down so every team applies it the same way. For an authorization flaw, require a replayable request chain that shows one identity accessing another user's data; other vulnerability classes need different evidence. In every case, another tester should be able to repeat the steps and see the same impact.

You might start with a one-business-day validation target for the highest tier and one week for a routine release. Adjust those windows to your risk, staffing, and release frequency.

Keep findings under test and the attack surface in view

A finding shouldn't disappear into a PDF once it's fixed. Keep the reproduction steps and test them again after later releases.

Every Cascade finding includes the request sequence, user scopes, and working exploit. It then flows into Escape's DAST as a regression test that runs in CI/CD. Cascade brings the depth; DAST brings the scale.

This only works if you know what you expose. Escape's attack surface management keeps a current view of applications, APIs, and infrastructure. When a new endpoint appears, it can enter the testing workflow instead of waiting for the next annual inventory.

Watch a small model prove a real flaw

The clearest way to see the difference is to run it. Below is one short Python file that drives a small local model against OWASP Juice Shop, a deliberately vulnerable practice app, until it finds and proves a real access-control flaw.

The model chooses each request, while the prompt supplies the vulnerability class, the endpoint pattern, and the success condition. You'll need Docker, Ollama, and the model installed before you start.

Step 1: Start something safe to attack

OWASP Juice Shop is deliberately vulnerable and built for security training. Start the pinned version in Docker:

docker run --rm -p 3000:3000 bkimminich/juice-shop:v20.0.0

Once the container is ready, open localhost:3000. Keep the target local, and don't point this script at a system you don't own or have permission to test.

OWASP Juice Shop

Step 2: Point the harness at a local model

The script uses only Python's standard library, and it starts by setting the target, the local Ollama endpoint, the default model, and the maximum number of turns:

import json
import sys
import urllib.error
import urllib.parse
import urllib.request
BASE = "http://localhost:3000"
OLLAMA = "http://localhost:11434/api/chat"
MODEL = sys.argv[1] if len(sys.argv) > 1 else "qwen2.5:7b"
MAX_STEPS = 12

Step 3: Give the model a structured way to act

The model doesn't make network requests directly. The harness builds each request and attaches the session token:

def request_target(method, path, body=None, token=None):
    """Make one HTTP call. The harness owns the call and the auth, not the model."""
    headers = {"Content-Type": "application/json"}
    if token:
        headers["Authorization"] = "Bearer " + token
    data = json.dumps(body).encode() if body is not None else None
    req = urllib.request.Request(BASE + path, data=data, headers=headers, method=method)
    try:
        with urllib.request.urlopen(req, timeout=10) as resp:
            return resp.status, json.loads(resp.read().decode() or "{}")
    except urllib.error.HTTPError as err:
        return err.code, {}
    except Exception as err:
        return 0, {"error": str(err)}

The system prompt asks for one JSON action per turn: GET or report. The script rejects malformed JSON and unknown tool names. Production code should go further by validating every field and restricting requests to approved targets.

SYSTEM = """You are an application security tester probing a shop's REST API for a broken access control (IDOR/BOLA) flaw: can one user read another user's data?
You act only through tools, one per turn, by replying with a single JSON object:
  {"thought": "<one short sentence>", "tool": "http_get", "path": "/rest/basket/2"}
  {"thought": "...", "tool": "report", "basket_id": <int>}
Rules:
- A valid bearer token for YOUR account is attached to every call automatically.
- Baskets are served at GET /rest/basket/{id}. Your own basket id is given below.
- The other users were created before you, so their basket ids are LOWER than yours. Probe those ids one at a time.
- Each response tells you the basket's id and whether it is your own.
- The flaw is proven the moment you read a basket that is not your own and gets
  HTTP 200. Then, and only then, reply with "report" and that basket_id.
- Reply with the JSON object only. No prose outside it."""

Step 4: Ask the model, then confirm with the target

One helper sends the conversation to Ollama and asks for JSON. Another reads Juice Shop's challenge state so you can see whether the application registered the access-control violation:

def ask_model(messages):
    """Send the conversation to Ollama and return the model's next action as text."""
    body = {"model": MODEL, "stream": False, "format": "json",
            "options": {"temperature": 0}, "messages": messages}
    req = urllib.request.Request(OLLAMA, data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=180) as resp:
        return json.loads(resp.read().decode())["message"]["content"]
def challenge_solved(name):
    """Ask the app itself whether the access-control challenge has tripped."""
    query = urllib.parse.urlencode({"name": name})
    _, body = request_target("GET", "/api/Challenges/?" + query)
    rows = body.get("data", [])
    return bool(rows and rows[0].get("solved"))

Step 5: Process and validate actions

Three small helpers keep the loop readable. One creates a test account and logs in, one summarizes each response for the model, and one decides whether a reported finding actually counts:

def create_session():
    """Register a throwaway user and return its token and own basket id."""
    account = {"email": "atk@demo.test", "password": "Password1!", "passwordRepeat": "Password1!",
               "securityQuestion": {"id": 1}, "securityAnswer": "x"}
    request_target("POST", "/api/Users/", account)
    _, login = request_target("POST", "/rest/user/login",
                              {"email": account["email"], "password": account["password"]})
    auth = login["authentication"]
    return auth["token"], auth["bid"]
def build_observation(path, status, body, own_basket_id, seen):
    """Summarize one response for the model, and record any basket it read."""
    data = body.get("data") if isinstance(body, dict) else None
    if status == 200 and path.startswith("/rest/basket/") and isinstance(data, dict) and "id" in data:
        seen.add(data["id"])
        products = data.get("Products", data.get("products", []))
        return {"status": status, "basket_id": data["id"],
                "is_your_own_basket": data["id"] == own_basket_id, "items": len(products)}
    return {"status": status, "data": data if data is not None else body}
def validated_report(action, own_basket_id, seen):
    """A report counts only if we read a basket that is not ours and the app
    itself recorded the access-control violation."""
    basket_id = action.get("basket_id")
    read_another = basket_id in seen and basket_id != own_basket_id
    return read_another and challenge_solved("View Basket")

The main function logs in, then loops one action at a time. It rejects malformed JSON, dispatches http_get and report actions to the helpers, and stops as soon as a report is validated:

def run():
    print(f"[harness] model={MODEL}  target={BASE}")
    name = "View Basket"
    already = challenge_solved(name)
    print(f"[state]   challenge '{name}' solved before we start: {already}")
    if already:
        print("[result]  restart Juice Shop to reset the challenge before testing")
        return 1
    token, own = create_session()
    print(f"[memory]  authenticated; the harness will attach our token; our own basket id = {own}")
    messages = [{"role": "system", "content": SYSTEM},
                {"role": "user", "content": f"You are logged in. Your own basket id is {own}. Begin."}]
    seen = set()   # basket ids we have actually read
    for step in range(1, MAX_STEPS + 1):
        raw = ask_model(messages)
        try:
            action = json.loads(raw)
        except json.JSONDecodeError:
            messages += [{"role": "assistant", "content": raw},
                         {"role": "user", "content": "That was not valid JSON. Reply with one JSON action object."}]
            print(f"[step {step}] malformed action, asked model to retry")
            continue
        messages.append({"role": "assistant", "content": json.dumps(action)})
        tool = action.get("tool")
        print(f"[step {step}] {tool}: {str(action.get('thought', ''))[:100]}")
        if tool == "report":
            basket_id = action.get("basket_id")
            if validated_report(action, own, seen):
                print(f"[PROVEN]  model read basket {basket_id}, not its own (id {own}), with our token.")
                print(f"[verify]  target certified it too: challenge '{name}' solved = True")
                print("[result]  1 validated exposure")
                return 0
            messages.append({"role": "user", "content": "Not verified yet. Keep testing other basket IDs."})
            print(f"[verify]  rejected report on basket {basket_id}: not proven yet")
        elif tool == "http_get":
            path = action.get("path", "")
            status, body = request_target("GET", path, token=token)
            observation = build_observation(path, status, body, own, seen)
            messages.append({"role": "user", "content": "Response: " + json.dumps(observation)[:600]})
        else:
            messages.append({"role": "user", "content": 'Unknown tool. Use "http_get" or "report".'})
    print("[result]  no finding within step budget")
    return 1
if __name__ == "__main__":
    sys.exit(run())

The harness controls the session and keeps earlier observations available. Its final check confirms the reported basket came from a successful response, isn't the authenticated user's basket, and that the target itself recorded the access-control violation. In production, this branch should also replay the request and preserve fuller evidence.

Step 6: Run it, and watch it work

With Juice Shop and Ollama running, save the program as harness_demo.py. Then pull Llama 3.2 if you need it and run the script:

ollama pull llama3.2
python3 harness_demo.py llama3.2

In this run, the model checks several basket IDs before reporting one that belongs to another user:

Terminal - the model driving the harness

Llama 3.2 reaches basket 2 in seven turns. Qwen2.5 7B gets there in two in a separate run. That contrast is useful here, but it isn't a general comparison of the two model families.

Step 7: Inspect the evidence

The response pair below is the evidence. The first response is the authenticated user's empty basket. The second is basket 2, which returns another user's items with the same token.

Evidence - reading another user's basket with our token

The different user IDs establish the authorization failure. The model selected the requests, but this was still a guided test: the prompt supplied the endpoint pattern, likely ID range, and definition of success.

The division of labor is simple, with the model choosing actions while the harness manages authentication, keeps the observations, and checks the report. A production system also needs stricter schemas, target allowlists, request replay, and fuller evidence handling.

If you'd rather not use Docker, Juice Shop also publishes prebuilt releases that start with node build/app.js, and Ollama serves the local model on port 11434.

What the example proves

A small model chose every move here, but the harness is what produced usable evidence. Structured actions kept its calls valid, memory tracked what it had read, and a validation step confirmed the finding against the target before anyone trusted it. The result came from the system around the model, not from the model alone.

That system is what Escape runs at production scale. Where the example was handed the API shape, one flaw class, and a single identity, Cascade, the multi-agent pentest engine inside Escape, discovers the surface itself, authenticates across OAuth, MFA, and CAPTCHA, and works with several identities at once, which is why Escape reports that a second account alone uncovers 30 to 50 percent more issues. 

On Escape's in-depth benchmark, it found roughly four times as many issues as the same model without the harness, with no false positives on the two open-source apps used for the comparison (Juice Shop was set aside as too well documented to count).

The example proved that point for one flaw. A program has to hold that same standard across everything it tests, and the rest of this guide builds it, with clear triggers, a shared definition of proof, and a path from each validated finding to a fix.

Conclusion

You don't need to replace your whole testing program at once. Start with one high-value application, four clear triggers, and an evidence standard your engineers trust, then keep each useful reproduction as regression coverage. That's how continuous testing becomes part of the engineering process instead of another report to file away, and to see it run against your own applications, book a demo.

FAQs

What is continuous offensive security testing?

It's adversarial testing that runs on an ongoing basis or when risk changes instead of an annual calendar. Depending on your program, it can cover applications, APIs, infrastructure, identities, business logic, and defensive controls.

How is it different from penetration testing?

Penetration testing is one part of offensive security, alongside vulnerability assessment and red teaming. A pentest usually has a fixed scope and window and goes deep on whether an app can be broken. A continuous program can trigger the right type of test as your environment changes, spanning pentests, red teaming, or vulnerability management.

What counts as proof that an exposure is real?

The evidence depends on the vulnerability, so capture the target, preconditions, requests or actions, identities, response, and security impact. Another tester should be able to follow the steps and see the same result. CISA's Known Exploited Vulnerabilities catalog records exploitation in the wild; it doesn't tell you whether the issue is reachable in your environment.

What should trigger an offensive security test?

Start with four: a newly discovered asset, a production release, threat intelligence relevant to your stack, and a change to a security control. Match the response to the exposure and business impact. A new internet-facing service shouldn't get the same test as a routine internal deployment. A routine release warrants days, and other in-scope systems follow a scheduled cycle.

What should you measure instead of finding counts?

Track coverage, the share of candidates that become validated findings, time from trigger to decision, remediation time, and regression-test results. Finding volume alone won't tell you whether exposure is going down.

Is continuous offensive security testing the same as continuous penetration testing?

No. Continuous penetration testing repeats a hands-on pentest of one target, for example retesting a payments API on every release. COST is broader: it picks the type of test (vulnerability assessment, penetration testing, or red teaming) based on what changed, and it sets the evidence standard for findings. Continuous pentesting can be one method inside a COST program.

Does continuous offensive security testing mean testing after every commit?

No. COST means choosing the right test when risk changes, not running every test on every commit. A scheduled job can read CISA’s KEV feed and flag new entries tied to vendors in your inventory. That is only a first filter, so confirm the affected product and version before launching a test.

How do validated findings become regression tests?
Keep the reproduction steps and automate them, so the same flaw is retested on later builds. In Escape’s platform, each Cascade finding carries the request sequence, user scopes, and working exploit. It then flows into Escape DAST as a regression test that runs in CI/CD.

Sources