Prompt Intruder monogram

Private offensive security · AI systems

We enter where
the prompt does.

A boutique atelier for AI penetration testing. We assess large language models, agents, retrieval pipelines, and the infrastructure that holds them — with the same seriousness leading firms bring to classic red teaming, applied to the attack surface that actually matters now.

FocusAI & LLM systems
PostureSenior operators only
OutputBoard-ready evidence

Traditional pentests stop at the application. The model is now the application.

The house

Quiet operators. Full-stack intrusion.

Prompt Intruder exists for organisations shipping copilots, agents, and RAG systems who need more than a jailbreak screenshot. We test the model, the tools it can call, the documents it can read, the cloud it lives in, and the people who deploy it.

Engagements are scoped like a private intelligence brief: limited seats, named operators, no junior farm, no unvalidated scanner noise. Equivalent in coverage to the AI practices of Bishop Fox, Trail of Bits, and specialist LLM red teams — delivered as a house, not a platform.


Practice

What we test

All services
01

AI & LLM assessments

Prompt injection, guardrail bypass, system-prompt extraction, data leakage, model extraction, and denial-of-wallet. Mapped to OWASP LLM Top 10 and MITRE ATLAS.

02

Agentic red team

Tool-use abuse, MCP servers, multi-agent trust, goal hijacking, and excessive agency. We treat every function call as a privilege boundary.

03

RAG & data planes

Indirect injection via documents, retrieval poisoning, cross-tenant leakage, and embedding-store authorisation. Retrieved context is untrusted input.

04

Hybrid application testing

The AI feature and the classic app around it: APIs, authn/z, business logic, and the places where model output becomes a side effect.

05

AI-focused red team

Multistep adversary emulation against the AI pipeline: OSINT, operator phishing, cloud pivots to model artefacts, exfiltration and extortion scenarios.

06

Lifecycle & supply chain

Fine-tune poisoning, CI/CD gate tampering, model provenance, secrets in training and inference, isolation and trust-boundary review.

Attack surface

What leading programmes cover. What we cover.

Language layer

Direct and indirect prompt injection, multi-turn jailbreaks, encoded and multilingual payloads, system prompt leakage, policy bypass, and sensitive data regurgitation.

Action layer

Unsafe tool invocation, privilege escalation through agents, arbitrary side effects, plugin and MCP compromise, and chained actions that look benign in isolation.

Data layer

RAG corpus poisoning, connector over-permission, tenant isolation failure, training-data extraction, and context-window smuggling.

Platform layer

Inference endpoints, vector databases, GPU estates, secrets management, resource exhaustion, and supply-chain compromise of models and fine-tunes.

Human layer

Operators, MLOps, and vendors. The shortest path to a model artefact is still often a person with a deploy key.

Why a house

Boutique on purpose.

Enterprise firms scale with platforms and benches. We scale with discretion. Every finding is human-validated. No unconfirmed scanner output reaches a client. Reports are written for counsel, CISO, and the engineer who has to fix it.

Engagements

A private briefing, not a demo.

Tell us the system, the risk you cannot sleep on, and the date you ship. We will tell you whether we are the right house.

Request a briefing