AI Agent Assessment
We attack the AI feature you shipped, as deployed.
- Configuration and credentials first: debug surfaces, exposed internal endpoints, service-account scope, secrets, admin interfaces
- Tool and retrieval boundaries: authorisation on every callable action, tenant and document isolation, cross-user leakage, and the tool descriptions themselves, MCP servers included
- What the tools run inside: whether the runtime reaches a cloud metadata service, an internal registry or a package mirror, what egress it has, which secrets are readable from the agent process, and whether one agent can read what another wrote
- Identity, memory and transcripts: SSO and directory integration, token handling, chat logs, and whether one session can write something that steers the next
- Prompt-injection chains and jailbreaks, pursued until they reach a real action or a real disclosure
- Availability and cost: token exhaustion, reasoning loops that do not terminate, and deliberate amplification of your inference bill