Penetration Testing¶
Penetration testing is an authorized, simulated attack against a system to find and exploit vulnerabilities the way a real adversary would. Where automated scanners (SAST, DAST) find known patterns, skilled pentesters find business-logic flaws, chained exploits, and creative attack paths that tools cannot. Pentesting validates that defenses work against a thinking attacker, not just against a known signature library.
Types and scope¶
| Type | Knowledge | When to use |
|---|---|---|
| Black-box | None — external attacker simulation | Assessing the external perimeter as an unauthenticated attacker would see it |
| Grey-box | Partial (credentials, architecture docs) | Most common; cost-effective balance between depth and realism |
| White-box | Full (source, infra, architecture) | Deepest technical review; maximum coverage; often used for critical systems pre-launch |
By target:
- Web application — OWASP Top 10, authentication/authorization, business logic, API surfaces.
- Mobile — OWASP MAS (MASVS/MASTG), client-side storage, traffic interception, reverse engineering.
- API — OWASP API Security Top 10, BOLA/IDOR, rate limiting, schema validation.
- Cloud and infrastructure — IAM misconfigurations, privilege escalation paths, network segmentation.
- AI/LLM applications — prompt injection, insecure output handling, data leakage, and excessive agent permissions; see the OWASP Top 10 for LLM Applications and the OWASP GenAI Red Teaming Guide.
- Social engineering — phishing, vishing, physical (if scoped); tests the human layer.
Define scope, rules of engagement, and authorization clearly and in writing before any testing begins. A "get out of jail free" letter signed by the authorizing executive protects both the testing team and the organization.
Methodologies¶
Structured methodologies make tests repeatable, defensible, and comprehensive:
| Methodology | Best for |
|---|---|
| OWASP WSTG | Web application testing |
| OWASP MASTG | Mobile application testing |
| OWASP API Security | REST/GraphQL API testing |
| NIST SP 800-115 | Technical guide to security testing and assessment; common baseline for government and regulated work |
| MITRE ATT&CK | Framing red team and adversary-emulation scenarios against real-world TTPs |
| PTES | Full-scope penetration tests with pre- and post-engagement phases |
| OSSTMM | Rigorous, metrics-based security testing across channels |
| TIBER-EU | Threat intelligence-led red team exercises for financial institutions |
Pentest vs red teaming vs purple teaming¶
| Exercise | Objective | Stealth | Blue team aware? |
|---|---|---|---|
| Penetration test | Find as many vulnerabilities as possible in scope | No — speed over stealth | Usually (announced) |
| Red team | Test detection and response against realistic adversary objectives | Yes — stealth is part of the test | No (unannounced) |
| Purple team | Improve detection and response collaboratively | No — works with the blue team | Yes (collaborative) |
Red teaming goes beyond vulnerability finding — it tests whether the organization would detect and respond to a real attacker pursuing a specific objective (e.g., "exfiltrate customer data"). See Breach and Attack Simulation for continuous coverage between exercises.
Running an effective pentest engagement¶
Scoping¶
- Define target systems, IP ranges, domains, and explicit exclusions.
- Specify what types of attacks are permitted (exploitation, DoS, social engineering, physical).
- Agree on data handling: what happens to credentials or data discovered during testing.
- Agree on communication protocols: how critical findings are reported immediately vs. in the final report.
Pre-engagement information gathering¶
Before active testing:
- Architecture diagrams, data flow diagrams, and threat models if available (white/grey-box).
- Previous pentest reports and remediation status.
- Known high-risk areas or recent changes the team wants prioritized.
During the engagement¶
- Regular check-ins to confirm scope and raise critical findings immediately (do not wait for the final report).
- Document all steps taken — reproduction steps are as important as the findings themselves.
- Follow rules of engagement strictly; escalate ambiguous situations before acting.
Reporting¶
A good pentest report contains:
- Executive summary — risk posture, critical findings, business impact in non-technical language.
- Technical findings — for each issue: description, evidence/screenshot/PoC, CVSS score, business impact, remediation steps.
- Prioritized remediation roadmap — ordered by risk, with suggested owner and timeline.
- Positive findings — controls that worked; what the testers could not bypass.
Continuous and modern approaches¶
Annual pentests no longer match continuous delivery — a feature shipped the day after a pentest has been untested for almost a year. Complement manual pentests with:
- PTaaS (Penetration Testing as a Service) — platforms like Cobalt, Synack, HackerOne, and Bugcrowd connect organizations with vetted pentesters on demand, with continuous engagement models and integrated finding management.
- Automated continuous testing — tools like Pentera or Horizon3.ai NodeZero run automated exploitation tests continuously between manual engagements to catch regression.
- A tight remediation loop — findings flow into Vulnerability Management with named owners and SLAs; retesting by the original pentesters (or an independent reviewer) confirms fixes. A finding closed without retest is a finding with unknown status.
Common pitfalls and anti-patterns¶
- Pentesting without remediating — a pentest that produces a report that sits unactioned is a compliance exercise, not a security improvement. Commit to remediation SLAs before the engagement.
- Scope too narrow — testing only the front-end while skipping APIs, admin panels, and cloud infrastructure misses the majority of modern attack surface.
- No retesting — remediations are often incomplete or introduce new bugs. Always include a retest phase in the engagement scope.
- Annual only — a once-a-year pentest against a system that ships 50 releases per year is not sufficient coverage. Supplement with PTaaS and automated testing.
- Findings not prioritized by exploitability — a list of 40 findings sorted by CVSS score alone does not help defenders decide what to fix first. Add business impact, exploitability, and reachability context.
Maturity progression¶
Starter — Annual grey-box web application pentest by a qualified firm. Findings tracked in the vulnerability management system. Remediation of critical/high findings within 30 days.
Intermediate — Annual pentest for all tier-1 applications. PTaaS for continuous coverage of critical apps. Remediation loop with retesting included in scope. Pentest findings feed the SAST/DAST rule tuning process.
Advanced — Annual red team exercise testing detection and response. PTaaS with a rotating researcher pool for fresh perspective. Automated continuous exploitation testing (Pentera/NodeZero). Purple team collaboration to improve SIEM detection rules based on pentest findings. Pentest coverage metrics by application tier.
Metrics and KPIs¶
- Critical/high findings per pentest — track trend; declining findings indicate improving security posture.
- Mean time to remediate pentest findings — by severity tier; compare against SLA.
- Retest pass rate — percentage of fixed findings confirmed closed by the tester; low rate indicates incomplete fixes.
- Coverage rate — percentage of tier-1 applications that have been pentested in the last 12 months.
- Findings that were missed by automated tools — quantifies the value of manual testing vs. scanner-only approach.
Tools1¶
Open-source¶
- BloodHound — Maps attack paths in Active Directory and Entra ID (Azure) identity graphs; essential for assessing privilege escalation and lateral movement paths.
- Burp Suite Community — Web security testing proxy and toolkit; intercepts, modifies, and replays HTTP/S traffic; industry standard for web application testing.
- ffuf — Fast web fuzzer for directory/file discovery, virtual host enumeration, and parameter fuzzing.
- Metasploit Framework — Exploitation framework with a large module library; used for post-exploitation, lateral movement, and payload generation.
- Nmap — Network discovery, port scanning, and service enumeration; the starting point for almost every infrastructure engagement.
- SQLmap — Automated SQL injection detection and exploitation; used to confirm and exploit injection findings.
- ZAP — Formerly OWASP ZAP; web application scanning and manual testing proxy; open-source alternative to Burp for intercepting and scanning web traffic.
Commercial¶
- Burp Suite Professional — Advanced web application penetration testing; adds active scanner, Collaborator (out-of-band detection), and advanced automation to the Community edition.
- Cobalt — Pentest-as-a-Service platform with a curated researcher network, integrated finding management, and API-driven workflow.
- Horizon3.ai NodeZero — Autonomous penetration testing platform that safely chains real exploits across network, identity, and cloud; suited to continuous validation between manual tests.
- Pentera — Automated, continuous penetration testing for network and cloud environments; runs continuously rather than on engagement cycles.
- Synack — PTaaS with a vetted researcher community; strong on financial services and government compliance requirements.
Links¶
-
Listed in alphabetical order. ↩