03 · Probe · Coerce · Contain
AI Penetration Testing
Red-team your models before an adversary does.
Adversarial testing for LLM- and ML-driven products. We probe prompt injection, jailbreaks, data exfiltration, model-supply-chain, and agentic-loop abuse — mapped to the OWASP LLM Top 10, MITRE ATLAS, and the NIST AI RMF. This is the one service where we attack your AI. Everywhere else, AI is the instrument we attack with — see the methodology.
Outcomes
- +Reproducible jailbreak + injection payloads, each with a detection rule
- +OWASP LLM Top 10 + MITRE ATLAS coverage report
- +Agent tool-use blast-radius analysis with a mitigation playbook
- +ISO 42001 evidence artifacts your internal audit can use directly
01Scope
What we cover
In scope
- +LLM-facing endpoints (chat, RAG, agent orchestrators, tool-use loops)
- +Prompt injection — direct and indirect (via tools, documents, browsing)
- +Jailbreak & policy-bypass against deployed guardrails
- +Data exfiltration via outputs, side channels, timing
- +Model supply chain — weights integrity, dataset-poisoning risk, dependency review
- +Autonomous-agent risk — unauthorized action, chained tool abuse, self-modification
Out of scope
- −Reverse-engineering closed-weight foundation models
- −Adversarial training / model-poisoning at scale (research engagement, separate)
02Approach
How the engagement runs
Reconnaissance / threat modeling
Data-flow map, trust boundaries, prompt provenance. What can enter the model? What can leave?
Enumeration
Every guardrail, tool, and integration in the deployed path, catalogued.
Vulnerability Analysis
Systematic evaluation of each guardrail against ATLAS-mapped tactics; candidate bypasses identified.
Exploitation
Direct + indirect prompt injection, RAG poisoning, tool-manipulation loops. Reproducible payloads, not one-off screenshots.
Post-exploitation / blast-radius
For agentic systems: enumerate every tool, quantify what a coerced agent could actually do, measure containment.
Report + evidence
Findings + reproduction harness + detection rules + ISO 42001 evidence pack.
03Deliverables
What you receive
Every artifact is defensible under external audit and actionable for engineering.
- 01AI Red-Team Report (technical + executive + reproduction harness)
- 02Payload library — reproducible test cases for your CI regression suite
- 03Guardrail-gap analysis mapped to OWASP LLM Top 10 + MITRE ATLAS
- 04Detection rules for your SIEM / observability stack
- 05ISO 42001 audit-evidence pack
04Frameworks
Regulator-defensible mapping
05Timeline
Typical engagement pace
Threat modeling
1 week
Adversarial testing
3–5 weeks
Report + payload delivery
1–2 weeks
Retest after mitigation
1 week
06FAQ
Common questions
Different question? Raise it on a scoping call — we'd rather flag surprises early.
Do you support open-weight models we host ourselves?+
Yes. We test the deployed inference path, guardrails, and orchestration — the model provenance is irrelevant to the attack surface.
Can this satisfy ISO 42001 audit requirements?+
The evidence pack is built to feed ISO 42001 AI-system-lifecycle and impact-assessment controls directly. It's designed to be handed to your auditor as third-party assurance.
How is agent testing different from LLM testing?+
Agents multiply blast radius by tool count. A prompt injection that's a curiosity in a plain chatbot becomes an unauthorized transaction in a banking agent. We test the whole loop, not just the prompt.
Ship AI features with confidence.
Scoping call, engagement letter within a week, kickoff within two. Report and payload library delivered end-to-end.
Book a scoping call