AI Assurance Audit
The AI Assurance Audit is an independent security, reliability and governance assessment of a single AI application, conducted by engineers who build and operate production AI. The engagement produces an executive report, a technical findings report and a prioritized remediation plan.
Scope is agreed in advance, the price is fixed at $7,500, and the assessment is typically completed within two to three weeks of access being granted.
Who this is for
The audit is designed for organizations whose customers already depend on an AI feature and who must now answer a question from an enterprise buyer, a security questionnaire, a board member or an insurer: how do you know this system is safe?
- B2B SaaS with AI in the product. The AI touches customer data, and a mistake becomes a customer-visible incident.
- Regulated industries. Healthcare, fintech, insurance, legal and HR, where "how was this validated" needs a documented answer.
- Teams approaching an enterprise deal. A security review is coming, and arriving with findings already closed shortens it.
- Engineering leaders inheriting an AI system. Built quickly, now load-bearing, and still awaiting its first independent examination.
What gets tested
Every area below is examined and reported on. Where one falls outside your system, a missing retrieval layer for instance, the report says so in as many words.
Architecture review
Trust boundaries, data flows, model providers, third-party dependencies, and every point where untrusted input enters the system.
AI threat model
What an attacker would target in this specific application, and which controls stand in the way.
Data-boundary review
What reaches the model, what the model can reach, and whether both stay inside the boundaries you intend.
LLM security evaluation
Systematic probing of the model layer against the published attack classes, run as a repeatable suite.
Prompt-injection evaluation
Direct and indirect injection, instruction-hierarchy attacks, and attempts to surface hidden context.
RAG evaluation
Where retrieval is used: retrieval authorization, tenant boundaries, poisoned or misleading context, citation accuracy.
Authentication and authorization review
Who can invoke what, whether the model inherits privileges it should not, and how multi-tenancy holds up.
Human-control review
Where a person approves, how firmly that gate holds, and whether the approver sees enough to decide well.
Model behaviour evaluation
Instruction following, refusal behaviour, boundary conditions, and inappropriate disclosure.
Logging and auditability review
How much of an incident you could reconstruct afterwards: what the system did, for whom, and on whose authority.
Governance assessment
Ownership, intended and prohibited use, model inventory, change management, incident procedure.
Findings are mapped to these frameworks so your team can place them in a structure it already recognises. DI-DS is not affiliated with, endorsed by, or accredited under any of them.
What you need to provide
The intake form adapts to the answers provided, so questions about retrieval appear only where retrieval is used. Every technical question accepts "I don't know" as an answer.
Application access
A working environment we can exercise. Staging is the strong preference.
Architecture description
A diagram or a written walkthrough. A whiteboard photo works.
Model and provider details
Which models, which providers, roughly which versions. "I don’t know" is a valid answer and a useful one.
Tool and integration inventory
Everything the system can call, read or write: APIs, databases, retrieval stores, MCP servers.
Test accounts
Two accounts at different privilege levels, plus two tenants where the product is multi-tenant.
Signed authorization
A completed Security Testing Authorization naming the systems and the testing window.
Testing happens only within written authorization
Active testing begins only after the customer has signed a Security Testing Authorization naming the systems, environments, domains, permitted and prohibited methods, testing window and an emergency stop contact. DI-DS tests nothing outside that authorization. Where findings suggest the scope should be wider, DI-DS returns to the customer for approval before proceeding.
What you receive
Executive report
Written for a CTO, CISO, board or risk function, the executive report covers overall risk posture, the findings of commercial significance, the controls found to be operating effectively, production-readiness observations, and a prioritized remediation roadmap. It is written to be shared with the organization's own customers.
Technical findings report
Written for the engineers who will fix it. Scope, methodology, and the architecture as we found it. Every finding carries evidence, severity, confidence, the affected component, and reproduction steps where they are safe to publish. Remediation guidance and framework mapping follow each one. The limitations and everything out of scope are stated at the end.
How it works
Select and authorize
Choose the assessment, sign the agreement and the testing authorization.
Secure intake
Complete the technical intake and upload documentation and access details.
We assess
Automated evaluation plus expert review. Every finding is reviewed by a person before it reaches you.
Reports delivered
Executive and technical reports through the secure portal, with an optional findings call.
Questions
Do we have to get on a call first?
You can purchase, sign, complete intake and receive both reports asynchronously. A findings call is available at no extra cost, and plenty of teams take it. The choice is yours.
Is $7,500 really fixed?
Yes, for one application with a reasonably bounded scope. If the intake shows a system materially larger than described, such as several distinct applications or an agent estate that belongs in a red-team engagement, we tell you before any work starts and quote the correct service. You approve the number before we begin.
How long does it take?
Most assessments run two to three weeks from the point we have access and a signed authorization. The variable is almost always how quickly test accounts and environment access arrive, not our queue.
Will you test our production system?
Only if you explicitly authorize it in writing, and we would usually rather not. Staging with representative data and configuration gives better coverage with none of the risk. Where production testing is genuinely necessary, the authorization names the exact systems, methods and window.
What if you find nothing serious?
You get a report that says so, with the evidence behind it. That is exactly the artifact a customer, an insurer or a board asks for when they want to know how you know. A clean result is a real result, and we are happy to write one.
Can we share the report with customers?
Yes. The executive report is written to be shareable with a prospect, a board or a security reviewer. The technical report contains reproduction detail and is usually kept internal.
Do you certify our system?
No. We are not a certification body and we do not claim to be. What you receive is an independent assessment with documented methodology, evidence and findings, which is what security reviewers ask for.
What happens after the report?
Nothing, unless you want it to. You can take the remediation plan to your own engineers. You can have us do the remediation work as a separate fixed-price scope. Or you can move to Continuous Assurance if the system changes often enough to warrant it.
Start an assessment
Tell us about the system. Scope and price are confirmed before anything is charged, and before any testing begins.