A penetration test is an authorised, time-boxed attempt by skilled testers to find and demonstrate exploitable weaknesses in a defined set of systems. Organisations buy them for compliance, before major launches, and after significant change. Many of those tests produce a long PDF that is read once and a set of findings that are still open a year later. The testing was competent, but the system around the testing was not.

This article treats a pentest as an architecture with inputs, a bounded process and tracked outputs. It covers how to scope a test and write rules of engagement, the phases recognised methodologies share, how to keep testing inside its authorisation with code, how to structure findings so they can be scored, assigned and retested, and how to run the remediation loop to closure. It is written for the engineers and security leads who commission, host and act on tests, and for testers who want their work to lead to fixes. It stays at the engagement and programme level and contains no exploitation techniques.

Advertisement

What a pentest is, and what it is not

ActivityQuestion it answersTypical shape
Vulnerability scanWhich known issues are detectable automatically?Automated, frequent, broad, shallow
Penetration testCan a skilled attacker exploit this scope, and how far could they get?Human-led, scoped, days to weeks, deep
Red team exerciseWould we detect and respond to a realistic adversary pursuing an objective?Objective-based, covert, tests people and detection too
Bug bountyWhat do many independent researchers find over time?Continuous, pay per valid finding, uneven coverage

The distinction matters when commissioning. A scan report relabelled as a pentest is a common failure, and a pentest is not a test of your detection and response; that is what red teaming and purple-team exercises are for. Pentests are strongest at depth: chaining small weaknesses into a real impact, testing business logic that scanners cannot understand, and validating that a high-scoring scanner result is actually exploitable in your environment. They complement, rather than replace, automated testing in the pipeline described in SAST and DAST.

The engagement as a system

A penetration test as a system: authorised inputs, bounded testing, tracked outputsScopeassets, exclusionsRules of engagementwindows, contacts, stopsThreat modelwhat matters mostTesting phases1. Planning2. Discovery3. Vulnerability analysis4. Controlled validation5. Reportingscope guard on every toolFindings storeevidence, score, statusReportsummary + reproducible findingsRemediation ticketsowner, SLA due dateRetestverify fix, close or reopenlessons feed the next threat modelCritical finding or production impactstop condition: pause, notify the named contact, resume only on written go-aheadThe report is not the output; closed, retested findings are.
Scope, rules of engagement and the threat model are inputs. Every tool runs through a scope guard. Findings flow into a store, then a report, then tickets with due dates, and close only after retest.

The inputs are signed documents, not conversations. The testing phases operate only within them, and a stop condition can pause everything. The outputs are not the report but individual findings with owners and due dates, and the loop closes only when a retest confirms each fix. Lessons from the test, such as which component produced most findings and which assumption was wrong, feed back into the next threat model.

Advertisement

Scoping

Scope defines exactly what may be tested. A good scope is a list of concrete assets, IP ranges, hostnames, applications, API base URLs and mobile app builds, plus explicit exclusions. It states the test perspective: black box with no information, grey box with user credentials and documentation, or white box with source code and architecture.

Start scoping from the threat model, not from the asset inventory. If threat modeling says the highest risk is a tenant reading another tenant's data, the scope must include multi-tenant authorisation paths with at least two test tenants, even if that is a small part of the estate. Scope also lists environments. Testing production gives the most realistic results and the most risk; a production-like staging environment is often the better compromise, provided it truly matches configuration and data shape.

Third parties need explicit attention. You can only authorise testing of what you own or control. Cloud providers, SaaS vendors and hosting partners publish their own testing policies, which may permit, restrict or forbid certain activities, so read the current policy for each provider involved and record the result in the scope.

Rules of engagement

Rules of engagement turn scope into operating constraints. At minimum they cover:

  • Testing windows, in UTC, and blackout dates such as release days or peak trading.
  • Source addresses the testers will use, so the operations team can tell test traffic from a real attack.
  • Prohibited activities, typically denial of service, destructive changes to data, social engineering unless agreed, and accessing real customer data beyond the minimum needed to prove a finding.
  • Credentials and accounts provided, and how they are revoked at the end.
  • Data handling: where evidence is stored, encryption, retention and deletion.
  • Stop conditions and escalation: a named contact on each side, reachable during testing, and an immediate-notification rule for critical findings or signs of a prior compromise.

If testers find evidence that someone else got there first, the engagement becomes an incident, and the rules should say who decides what happens next.

Methodology phases

Several public methodologies describe the work, and they agree more than they differ. The Penetration Testing Execution Standard (PTES) has seven sections: pre-engagement interactions, intelligence gathering, threat modeling, vulnerability analysis, exploitation, post-exploitation and reporting. NIST SP 800-115 describes four phases: planning, discovery, attack and reporting. For web applications, the OWASP Web Security Testing Guide, whose latest stable release is 4.2 with version 5.0 in development, provides a detailed catalogue of test cases by area such as authentication, session management, authorisation, input validation and business logic.

In practice an engagement moves through planning (scope, rules, access), discovery (mapping the attack surface: hosts, services, endpoints, roles), vulnerability analysis (hypotheses about weaknesses from configuration, behaviour and known issues), controlled validation (demonstrating impact with the least intrusive method that proves it), and reporting. Validation is deliberately bounded. The goal is evidence that a weakness is real and what it would allow, not maximum damage. Using a recognised methodology matters for compliance as well: PCI DSS v4.0 requirement 11.4 asks for a documented methodology based on industry-accepted approaches, and gives NIST SP 800-115 as an example.

Keeping testing inside its authorisation, in code

Most out-of-scope incidents are mistakes: a typo in an address, a wildcard hostname that resolves to a vendor, or a test run after the window closes. Policy documents do not stop typos, but a wrapper does. Every tool in the engagement goes through a scope guard that reads the signed rules of engagement, refuses anything outside them, and logs each run.

# scope_guard.py: every tool in the engagement runs through this wrapper.
# It refuses targets or times outside the signed rules of engagement and logs every run.
import ipaddress, json, shlex, subprocess, sys, datetime as dt

ROE = json.load(open("roe.json"))   # signed scope, exported from the engagement record
IN_SCOPE = [ipaddress.ip_network(n) for n in ROE["in_scope_cidrs"]]
EXCLUDED = [ipaddress.ip_network(n) for n in ROE["excluded_cidrs"]]
HOSTS = set(ROE["in_scope_hosts"])            # exact hostnames, no wildcards
WINDOW = (dt.time.fromisoformat(ROE["window_start_utc"]),
          dt.time.fromisoformat(ROE["window_end_utc"]))

def target_allowed(target: str) -> bool:
    try:
        ip = ipaddress.ip_address(target)
    except ValueError:
        return target in HOSTS                # hostnames must be listed exactly
    if any(ip in net for net in EXCLUDED):
        return False                          # exclusions win over inclusions
    return any(ip in net for net in IN_SCOPE)

def in_window(now=None) -> bool:
    t = (now or dt.datetime.now(dt.timezone.utc)).time()
    return WINDOW[0] <= t <= WINDOW[1]

def run(tool_cmd: str, target: str):
    if ROE.get("paused"):
        sys.exit("engagement paused by stop condition")
    if not target_allowed(target):
        sys.exit(f"REFUSED: {target} is not in signed scope")
    if not in_window():
        sys.exit("REFUSED: outside the agreed testing window")
    started = dt.datetime.now(dt.timezone.utc).isoformat()
    result = subprocess.run(shlex.split(tool_cmd) + [target], capture_output=True, text=True)
    with open("activity_log.jsonl", "a") as log:  # the client can reconcile this with their SIEM
        log.write(json.dumps({"ts": started, "tester": ROE["tester_id"], "target": target,
                              "cmd": tool_cmd, "exit": result.returncode}) + "\n")
    return result

The activity log has a second use. The client can reconcile it against their SIEM afterwards and ask which test activity was detected and which was not. That turns a pentest into a modest detection exercise at no extra cost, and the gaps are often as valuable as the findings.

Findings as data, and how to score them

A finding is the unit of work for everything after testing, so it should be structured data from the start, not prose assembled at the end.

from dataclasses import dataclass, field
from datetime import date, timedelta

SLA_DAYS = {"critical": 7, "high": 30, "medium": 90, "low": 180}   # example policy only

@dataclass
class Finding:
    id: str                      # stable across report versions and retests
    title: str
    asset: str                   # must match an in-scope asset
    cvss_vector: str             # base vector, recorded verbatim
    severity: str                # agreed rating after environmental context
    evidence: list = field(default_factory=list)   # redacted requests, screenshots, timestamps
    reproduction: str = ""       # steps a defender can follow to confirm the issue
    remediation: str = ""
    owner: str = ""              # team, not person
    reported: date = date.today()
    status: str = "open"         # open -> fixed_pending_retest -> closed | reopened | risk_accepted

    def due(self) -> date:
        return self.reported + timedelta(days=SLA_DAYS[self.severity])

    def overdue(self, today: date) -> bool:
        return self.status in ("open", "reopened") and today > self.due()

Scoring usually starts with CVSS. FIRST released CVSS v4.0 on 1 November 2023; whichever version you use, record the full vector, not just the number, so others can see the assumptions. The base score describes the vulnerability in general; your environment changes the real risk. An injection flaw on an internal admin tool behind single sign-on is not the same risk as the identical flaw on a public endpoint holding payment data. Agree a final severity with the client that takes exposure, data sensitivity and compensating controls into account, and record why it differs from the base score when it does.

Chains deserve special treatment. Three medium findings that together let an anonymous user reach another tenant's data are a critical issue, and the report should present the chain as its own finding with references to the parts, so that fixing any one link is recognised as breaking it.

Reporting

A good report has two audiences. The executive summary states, in a page, what was tested, the most serious outcomes in business terms, and the overall trend compared with the last test. The technical body gives each finding a title, affected assets, severity with its reasoning, evidence with sensitive data redacted, reproduction steps that a defender can follow to confirm the issue, and specific remediation guidance. An attack narrative describing how findings were chained helps engineers understand why a 'medium' matters. Deliver findings into the client's tracker, not only as a PDF, so that the identifiers in the report and in the tickets are the same.

Remediation and retest

After delivery, each finding becomes a ticket with an owning team and a due date derived from its severity. An example policy is 7 days for critical, 30 for high, 90 for medium and 180 for low; the right numbers are a business decision, but having them written down is not optional. When the owner marks a fix done, the finding moves to fixed-pending-retest, not closed. A retest, ideally by the original testers, confirms the fix and checks for obvious bypasses. Only then does it close. A finding that cannot be fixed in time is either risk-accepted by a named person with an expiry date, or escalated.

Many root causes are not local. If several findings come from the same library or base image, fix them through supply chain controls rather than one service at a time. PCI DSS v4.0 requirement 11.4 makes this loop explicit for in-scope environments: internal and external testing at least every twelve months and after significant change, correction of exploitable vulnerabilities followed by retesting, and segmentation testing every twelve months, or every six for service providers.

Worked example: an annual test of a SaaS application

A B2B SaaS company commissions a two-week grey-box test of its web application and public API, with two tenant accounts at each of three roles, staging as the primary target and read-only checks against production. The threat model puts tenant isolation first. The testers report eleven findings: one critical, a chain in which an export endpoint accepted a tenant identifier from the request body and a predictable job identifier let one tenant download another tenant's export; two high; five medium; three low. The critical is notified within an hour of confirmation under the rules of engagement, fixed in four days by deriving the tenant from the session, and retested the next day. At the 30-day review, both highs are closed, one medium is risk-accepted for 90 days with a compensating rate limit, and the activity-log reconciliation shows that the SIEM alerted on none of the export requests, which becomes its own detection ticket.

Failure modes

  • Scope drift. Testers follow a link into a vendor's system. A scope guard and exact host lists prevent most of it.
  • Compliance theatre. The same narrow scope every year, chosen to pass, never the risky parts. Rotate depth according to the threat model.
  • Report rot. Findings live only in a PDF. Put them in the tracker with due dates on day one.
  • Skipped retests. Fixes are assumed to work. Partial fixes are common, so retest before closing.
  • Evidence leakage. Screenshots containing customer data sit in shared drives. Redact, encrypt and delete on schedule.
  • Testing a staging environment that is nothing like production. Findings and, worse, non-findings do not transfer.

What to do next

  1. Start from your threat model and write down the three outcomes you most need the next test to examine.
  2. Write a scope with concrete assets, exclusions, perspective and environments, and check each third party's testing policy.
  3. Agree rules of engagement with windows, source addresses, prohibited activities, stop conditions and named contacts.
  4. Require testers to run tools through a scope guard and to share their activity log for SIEM reconciliation.
  5. Ask for findings as structured data with full CVSS vectors and agreed contextual severity.
  6. Create tickets with owners and SLA dates on delivery, and track overdue findings weekly.
  7. Retest every fix before closing it, and feed recurring root causes into the next threat model and your pipeline controls.
Key takeaway: A penetration test is only as useful as the system around it. Signed scope and rules of engagement define what is authorised, a recognised methodology structures the work, a scope guard keeps it inside its limits, and structured findings with contextual scores feed a remediation loop that ends with a retest. Run that way, a pentest becomes a measured reduction in risk rather than an annual document.