← Analysis
AI securityCompanies8 min read
Series · Security for startups

How to test an AI agent connected to real tools

A practical testing map for AI agents with access to customer data, repositories, browsers, APIs, and production workflows.

Written by
Dravian Security Team
Security Research
Key takeaways
  • 01Map what the agent can read, write, and execute before designing tests.
  • 02Test both direct inputs and untrusted content retrieved through tools or RAG.
  • 03Validate complete paths from adversarial input to a real data or system impact.

The agent is an application with privileges

An AI agent may search documents, call APIs, write code, send messages, or update customer records. The model is only one part of the system. The meaningful security boundary sits between untrusted inputs and the actions the agent can take with its tools, identities, and data.

A useful assessment starts by mapping the whole path: user, model, system instructions, retrieved content, tool calls, permissions, connected services, and final actions. Without that map, testing can collapse into a list of jailbreak prompts that says little about business risk.

Build a capability matrix

Record what the agent can read, write, and execute. Include repository access, browser sessions, CRM records, email, cloud resources, payment actions, and any MCP tools. Then record which user roles are allowed to trigger each action and how authorization is checked.

The matrix becomes a set of testable questions. Can a lower-privilege user cause a write to a higher-privilege resource? Can text retrieved from a document make the agent send an email? Can an agent take an action beyond the user's authority because its service account has broader permissions?

Test untrusted content as an attacker-controlled surface

A direct prompt injection arrives from the user. An indirect injection arrives through material the agent later consumes: a web page, repository file, document, search result, support ticket, or tool response. The second path is easy to overlook because it can enter a workflow that looks legitimate.

Test whether the agent treats such content as instructions, whether it exposes hidden context or sensitive data, and whether it can be pushed to misuse a tool. In retrieval systems, also test whether poisoned or unauthorized content can be surfaced across trust boundaries.

Follow the path to impact

A prompt that changes a harmless answer is a different issue from a prompt that causes an unauthorized database operation. A meaningful finding should show the path from entry point to action: malicious content, agent interpretation, tool invocation, permission check, affected resource, and business impact.

This is where ordinary application security matters. An authorization bug in an API, an overprivileged token, or exposed cloud storage can turn an AI manipulation into a larger compromise. Assess the agent together with its supporting application and infrastructure.

Retest controls, then preserve the evidence

Fixes may include stronger tool permissions, explicit confirmation for sensitive actions, isolation of retrieved content, better authorization at the underlying API, and monitoring that helps detect misuse. A retest should use the original path and verify whether the exploit still works.

Keep a record of tested capabilities, attack paths, evidence, remediation, and retest results. That record is useful to builders and to customers asking how the AI system was independently evaluated. It should also state the scope and limitations of the assessment.

About the author
Dravian Security Team
Security Research

Practical guidance from the team testing products, validating risk, and helping organizations act on the findings.

About Dravian
AI security

Your agent has real access. Test its boundaries.

Dravian assesses agents, tools, data flows, identity, and infrastructure, then helps verify the fixes.

Explore AI security