Skip to content

Penetration Testing

Penetration testing is an authorized, simulated attack against a system to find and exploit vulnerabilities the way a real adversary would. Where automated scanners (SAST, DAST) find known patterns, skilled pentesters find business-logic flaws, chained exploits, and creative attack paths that tools cannot. Pentesting validates that defenses work against a thinking attacker, not just against a known signature library.

Types and scope

Type Knowledge When to use
Black-box None — external attacker simulation Assessing the external perimeter as an unauthenticated attacker would see it
Grey-box Partial (credentials, architecture docs) Most common; cost-effective balance between depth and realism
White-box Full (source, infra, architecture) Deepest technical review; maximum coverage; often used for critical systems pre-launch

By target:

  • Web application — OWASP Top 10, authentication/authorization, business logic, API surfaces.
  • Mobile — OWASP MAS (MASVS/MASTG), client-side storage, traffic interception, reverse engineering.
  • API — OWASP API Security Top 10, BOLA/IDOR, rate limiting, schema validation.
  • Cloud and infrastructure — IAM misconfigurations, privilege escalation paths, network segmentation.
  • AI/LLM applications — prompt injection, insecure output handling, data leakage, and excessive agent permissions; see the OWASP Top 10 for LLM Applications and the OWASP GenAI Red Teaming Guide.
  • Social engineering — phishing, vishing, physical (if scoped); tests the human layer.

Define scope, rules of engagement, and authorization clearly and in writing before any testing begins. A "get out of jail free" letter signed by the authorizing executive protects both the testing team and the organization.

Methodologies

Structured methodologies make tests repeatable, defensible, and comprehensive:

Methodology Best for
OWASP WSTG Web application testing
OWASP MASTG Mobile application testing
OWASP API Security REST/GraphQL API testing
NIST SP 800-115 Technical guide to security testing and assessment; common baseline for government and regulated work
MITRE ATT&CK Framing red team and adversary-emulation scenarios against real-world TTPs
PTES Full-scope penetration tests with pre- and post-engagement phases
OSSTMM Rigorous, metrics-based security testing across channels
TIBER-EU Threat intelligence-led red team exercises for financial institutions

Pentest vs red teaming vs purple teaming

Exercise Objective Stealth Blue team aware?
Penetration test Find as many vulnerabilities as possible in scope No — speed over stealth Usually (announced)
Red team Test detection and response against realistic adversary objectives Yes — stealth is part of the test No (unannounced)
Purple team Improve detection and response collaboratively No — works with the blue team Yes (collaborative)

Red teaming goes beyond vulnerability finding — it tests whether the organization would detect and respond to a real attacker pursuing a specific objective (e.g., "exfiltrate customer data"). See Breach and Attack Simulation for continuous coverage between exercises.

Running an effective pentest engagement

Scoping

  • Define target systems, IP ranges, domains, and explicit exclusions.
  • Specify what types of attacks are permitted (exploitation, DoS, social engineering, physical).
  • Agree on data handling: what happens to credentials or data discovered during testing.
  • Agree on communication protocols: how critical findings are reported immediately vs. in the final report.

Pre-engagement information gathering

Before active testing:

  • Architecture diagrams, data flow diagrams, and threat models if available (white/grey-box).
  • Previous pentest reports and remediation status.
  • Known high-risk areas or recent changes the team wants prioritized.

During the engagement

  • Regular check-ins to confirm scope and raise critical findings immediately (do not wait for the final report).
  • Document all steps taken — reproduction steps are as important as the findings themselves.
  • Follow rules of engagement strictly; escalate ambiguous situations before acting.

Reporting

A good pentest report contains:

  • Executive summary — risk posture, critical findings, business impact in non-technical language.
  • Technical findings — for each issue: description, evidence/screenshot/PoC, CVSS score, business impact, remediation steps.
  • Prioritized remediation roadmap — ordered by risk, with suggested owner and timeline.
  • Positive findings — controls that worked; what the testers could not bypass.

Continuous and modern approaches

Annual pentests no longer match continuous delivery — a feature shipped the day after a pentest has been untested for almost a year. Complement manual pentests with:

  • PTaaS (Penetration Testing as a Service) — platforms like Cobalt, Synack, HackerOne, and Bugcrowd connect organizations with vetted pentesters on demand, with continuous engagement models and integrated finding management.
  • Automated continuous testing — tools like Pentera or Horizon3.ai NodeZero run automated exploitation tests continuously between manual engagements to catch regression.
  • A tight remediation loop — findings flow into Vulnerability Management with named owners and SLAs; retesting by the original pentesters (or an independent reviewer) confirms fixes. A finding closed without retest is a finding with unknown status.

Common pitfalls and anti-patterns

  • Pentesting without remediating — a pentest that produces a report that sits unactioned is a compliance exercise, not a security improvement. Commit to remediation SLAs before the engagement.
  • Scope too narrow — testing only the front-end while skipping APIs, admin panels, and cloud infrastructure misses the majority of modern attack surface.
  • No retesting — remediations are often incomplete or introduce new bugs. Always include a retest phase in the engagement scope.
  • Annual only — a once-a-year pentest against a system that ships 50 releases per year is not sufficient coverage. Supplement with PTaaS and automated testing.
  • Findings not prioritized by exploitability — a list of 40 findings sorted by CVSS score alone does not help defenders decide what to fix first. Add business impact, exploitability, and reachability context.

Maturity progression

Starter — Annual grey-box web application pentest by a qualified firm. Findings tracked in the vulnerability management system. Remediation of critical/high findings within 30 days.

Intermediate — Annual pentest for all tier-1 applications. PTaaS for continuous coverage of critical apps. Remediation loop with retesting included in scope. Pentest findings feed the SAST/DAST rule tuning process.

Advanced — Annual red team exercise testing detection and response. PTaaS with a rotating researcher pool for fresh perspective. Automated continuous exploitation testing (Pentera/NodeZero). Purple team collaboration to improve SIEM detection rules based on pentest findings. Pentest coverage metrics by application tier.

Metrics and KPIs

  • Critical/high findings per pentest — track trend; declining findings indicate improving security posture.
  • Mean time to remediate pentest findings — by severity tier; compare against SLA.
  • Retest pass rate — percentage of fixed findings confirmed closed by the tester; low rate indicates incomplete fixes.
  • Coverage rate — percentage of tier-1 applications that have been pentested in the last 12 months.
  • Findings that were missed by automated tools — quantifies the value of manual testing vs. scanner-only approach.

Tools1

Open-source

  • BloodHound — Maps attack paths in Active Directory and Entra ID (Azure) identity graphs; essential for assessing privilege escalation and lateral movement paths.
  • Burp Suite Community — Web security testing proxy and toolkit; intercepts, modifies, and replays HTTP/S traffic; industry standard for web application testing.
  • ffuf — Fast web fuzzer for directory/file discovery, virtual host enumeration, and parameter fuzzing.
  • Metasploit Framework — Exploitation framework with a large module library; used for post-exploitation, lateral movement, and payload generation.
  • Nmap — Network discovery, port scanning, and service enumeration; the starting point for almost every infrastructure engagement.
  • SQLmap — Automated SQL injection detection and exploitation; used to confirm and exploit injection findings.
  • ZAP — Formerly OWASP ZAP; web application scanning and manual testing proxy; open-source alternative to Burp for intercepting and scanning web traffic.

Commercial

  • Burp Suite Professional — Advanced web application penetration testing; adds active scanner, Collaborator (out-of-band detection), and advanced automation to the Community edition.
  • Cobalt — Pentest-as-a-Service platform with a curated researcher network, integrated finding management, and API-driven workflow.
  • Horizon3.ai NodeZero — Autonomous penetration testing platform that safely chains real exploits across network, identity, and cloud; suited to continuous validation between manual tests.
  • Pentera — Automated, continuous penetration testing for network and cloud environments; runs continuously rather than on engagement cycles.
  • Synack — PTaaS with a vetted researcher community; strong on financial services and government compliance requirements.


  1. Listed in alphabetical order. ↩