Advanced Agent Red-Team
The Advanced Agent Red-Team is an adversarial security assessment for AI systems that retrieve documents, call tools and take actions. Because such a system carries a blast radius equal to its permissions, the assessment concentrates on the failure that matters most: an action taken on someone else's behalf.
Conducted strictly within a written, signed authorization naming the systems, methods and window.
Systems this is built for
Autonomous and multi-agent systems
Agents that plan, delegate to other agents, and run unattended for meaningful stretches.
RAG and retrieval products
Vector stores, hybrid search, document ingestion pipelines, and anything that assembles context from a corpus.
Tool-using assistants
Agents wired to internal APIs, databases, ticketing, email, CI systems or payment rails.
MCP server integrations
Model Context Protocol servers, the scopes they expose, and what a compromised or hostile server could do to the client.
Agents with write access
Anything that changes state. This is the class where an exploited flaw persists after the page is refreshed.
Customer-facing agents
Where the adversary holds a legitimate account and tests how far that account reaches.
What we attack
Each area below represents a hypothesis about how the system could be turned against itself. DI-DS tests every one, and the report records each result, including the controls that held.
Indirect prompt injection
Instructions planted in content your agent reads: a document, a web page, a support ticket, a calendar invite, a code comment. The payload arrives through a trusted channel and bypasses input validation entirely, which is why agentic systems are least prepared for it.
Tool misuse and unsafe chaining
Whether an agent can be induced to call a tool it should not, with arguments it should not, or to chain several individually-safe tools into an outcome none of them would permit alone.
Privilege boundaries
Whether the agent operates with its own identity or borrows a human's, whether it inherits more permission than the user who invoked it, and what happens at the boundary between them.
Cross-user and cross-tenant leakage
Whether one customer's data can surface in another customer's session through shared memory, a shared vector store, cache reuse or a mis-scoped retrieval query.
Retrieval manipulation
Poisoned or adversarial content in the corpus, retrieval that ignores authorization, misleading context assembly, and citations that do not support the claim attached to them.
Unauthorized actions
Whether the agent can send, pay, delete, deploy or escalate while the approval you believe is required stays unasked.
Confirmation-gate bypass
How firmly approval gates hold when pressed: skipped outright, satisfied with a misleading summary, or worn down by volume until approval turns reflexive.
External side effects
What the agent can reach outside your perimeter, and what it can be persuaded to send there.
Transaction safety
Behaviour under partial failure, retries and interruption, and what state an interrupted agent leaves behind.
Forensic visibility
After a hostile run, whether your logs let you reconstruct what the agent did, what it read, which tools fired and on whose authority.
Framework references are provided so findings map onto structures your team already uses. DI-DS is not affiliated with, endorsed by, or accredited under any of them.
Authorization comes first
We do not test anything we have not been authorized in writing to test
Active testing of an agentic system can cause real side effects, which is what makes it worth testing. Every engagement therefore begins with an executed Security Testing Authorization identifying the authorizing organization; the systems, environments, domains and applications in scope; the permitted and prohibited methods; the testing window; data-handling requirements; an emergency stop contact; and everything that remains out of scope.
Where your system depends on a third-party platform, you confirm that testing is permitted under your agreement with that provider. We flag every case where we think that question needs asking.
Choosing between the two
| Service | Best when | Price |
|---|---|---|
| Free AI Risk Snapshot | You want an outside read on what your AI product reveals publicly. Passive review of published information, with no testing of any kind. | Free |
| AI Assurance Audit | One AI application, bounded scope. It answers questions, generates content or classifies things, and you need to know it is built soundly. | $7,500 |
| Advanced Agent Red-Team | The system takes actions. Agents, tool calls, retrieval over sensitive data, autonomy, multi-tenancy, or anything where a successful attack changes state. | From $15,000 |
| Continuous Assurance | The system changes often, with new tools, new models and new prompts, so a point-in-time result ages faster than you can commission another. | From $1,000/mo |
Unsure which applies? Describe the system in the scoping form. Where an audit covers it, we say so and quote the audit.
Questions
Why is this a different service from the $7,500 audit?
They answer different questions. The Assurance Audit establishes whether an AI application is soundly built. A red-team assumes an intelligent adversary with a goal and tests whether they reach it, which for an agent with tool access means chaining several partial weaknesses into one real outcome. That takes materially more time and produces a different kind of finding.
Why does this one start at $15,000 rather than being fixed?
Agent estates vary enormously. Three tools and one agent is a different engagement from twelve tools, four agents and a retrieval layer over customer documents. We scope from your intake and quote a fixed price before any work begins, and that number holds.
Can you test against production?
Often production is the only place the agent has real tools attached, and we will test there under your written authorization. That authorization names prohibited methods, a defined window, an emergency stop contact, and confines actions to reversible operations wherever possible. The report states plainly which findings stayed unconfirmed as a result.
Do you need our source code?
Most engagements run without it. Black-box testing against a running system with real credentials finds what an actual attacker finds. Source access sharpens the architecture review, so include it where you can.
Will you break something?
Testing demonstrates reachability. We show that an action can be triggered and stop short of the destructive version. The authorization names prohibited methods, and one message to the emergency contact halts everything immediately.
What if you find something critical mid-engagement?
You hear about it the day we confirm it, with enough detail to act on. Critical findings travel ahead of the report.
Scope a red-team engagement
Describe the agent estate. Scope and a fixed price are confirmed before any work or payment.