Editorial slot: September 13, 2026. Published on September 19, 2026 after review.
A useful AI-agent security assessment follows one real business request from the initiating principal, whether a person, service, workflow, or scheduled task, through the agent, the identities it uses, every connected tool, and the evidence left after the action. A prompt-injection demonstration can be relevant, but it is only a starting point. The decision that matters is whether influenced model behavior can reach data or an action beyond the requester’s authority.
NIST’s February 2026 concept paper identifies identification, authorization, auditing, non-repudiation, and controls that prevent and mitigate prompt injection as questions for AI-agent identity and authorization. It also highlights the risk of giving agents access to varied data, tools, and applications. NIST’s concept paper is a good reason to buy a test of the whole execution path rather than a test of chat output alone.
Start with a testable question
Put one answerable security question into the statement of work: can an untrusted or insufficiently authorized principal cause this agent to read protected information, invoke a tool, or complete a consequential workflow outside the intended authority?
That wording gives the provider a boundary to test. It requires an account of who started the request, which permissions were intended, which effective identity reached the downstream system, and which controls permitted or denied the action. It also gives the buyer a way to reject a report that contains only interesting prompts and screenshots.
NIST’s May 2026 report says respondents widely agreed that agents bring novel security threats, while established cybersecurity practices remain relevant but need adaptation. The NIST summary supports applying familiar identity, application, API, and logging review to an agent’s delegation and tool-use paths.
Put every identity in scope
Do not let the scope call the agent a single identity. Ask for an identity map that distinguishes the human requester, the agent runtime or workload identity, delegated user tokens, connector credentials, tool administrators, prompt or knowledge-base administrators, and service-initiated workers such as queues, webhooks, and scheduled jobs.
For each identity, record where it originates, how it authenticates, its effective permissions, who can change it, how secrets are stored and rotated, and how access is revoked. When the agent can use both a dedicated service identity and a user-delegated identity, require separate test cases. A downstream application may enforce the user’s restrictions in one path and see a broadly privileged application in the other.
Where directory groups, identity synchronization, or privileged administration affect access, include them under Active Directory security. Where workload roles, secret stores, and managed-service permissions affect access, include cloud security.
Treat authorization as a chain
Authentication identifies the requester. Authorization decides whether that requester may take a particular action against a particular resource. The test should establish whether that decision survives planning, tool selection, tool execution, asynchronous work, and presentation of the result.
Require an authorization matrix for meaningful workflows. The following is a proposed procurement artifact for a fictional support agent; it is not evidence about a real product.
| Actor and requested action | Expected control | Evidence the tester should retain |
|---|---|---|
| Standard support user asks to view own case | Permit only the permitted case and fields | Effective identity, object decision, redacted response, correlation ID |
| Same user asks for another customer’s case | Deny before data is returned | Denial event, API or tool result, no protected fields in output |
| Agent proposes a refund | Require an approved role and approval step | Policy decision, approver identity, downstream event and final state |
| Revoked user token reaches a queued task | Deny or safely expire the task | Revocation time, worker decision, no completed action |
Cover cross-user and cross-tenant access, role changes, revoked access, user-controlled tool parameters, and tool use outside the declared workflow. Test safely with designated accounts and synthetic records. The purpose is to show whether external policy enforcement prevents an unsafe agent decision from becoming an unauthorized action.
Where tools call APIs, include token audience and scope, server-side object authorization, error handling, and cost or rate controls. API security testing is directly relevant because the consequential authorization decision often occurs at the connector or API boundary.
Inventory tools and permissions
A credible assessment has a tool inventory. It includes direct function calls, plugins, retrieval systems, browser or automation capabilities, and actions dispatched through workflows, queues, or webhooks. For each tool, require its owner and purpose, permitted inputs and outputs, data classification, tenant boundary, execution identity, operation types, approval gate, rollback or idempotency behavior, and the policy that should allow or deny it.
Check least privilege with the actual configured permissions. An agent that reads availability should not silently inherit authority to edit every calendar. An agent that drafts a reply should not need a general mailbox-sending credential when a constrained approval path can submit the message. For high-impact tools, the final authorization should be a deterministic downstream policy or explicit approval, not a judgment made only in model output.
OWASP describes its guide as practical and actionable guidance for secure LLM-powered agentic applications. OWASP Securing Agentic Applications Guide 1.0 is a useful reference when converting the inventory into implementation checks.
Make the audit trail a deliverable
Logging that a user asked a question is insufficient. For a significant action, the tester should be able to correlate the request or workflow ID across the initiating principal, user session where present, agent run, policy decision, retrieval event, tool call, downstream service event, approval event, and final response.
The record should identify the initiating principal, whether human or service-initiated, effective execution identity, tool, approved operation, target category, decision or approval reference, time, result, and correlation ID. Validate successful and denied actions. Also validate that a shared service account does not erase attribution, routine operators cannot alter the relevant logs, and responders can reconstruct a representative event sequence.
Use non-repudiation carefully. A technical assessment can assess evidence and identity binding for the organization’s accountability model; it cannot promise legal proof.
Acceptance checklist and stop conditions
Require the final report to include these items:
- a trust-boundary diagram with identities, tools, data sources, approvals, and log destinations;
- an effective identity and permission inventory, not only intended roles;
- tested allow and deny cases from the authorization matrix;
- redacted evidence of downstream enforcement and audit correlation;
- root-cause remediation, an owner, and repeatable retest criteria;
- explicit limitations, including untested tools, third-party boundaries, nonproduction substitutes, and excluded paths.
Stop or re-scope the exercise when the provider cannot obtain approved test accounts with distinct permissions, cannot identify the execution identity for a consequential tool, cannot test a safe substitute for an irreversible action, or cannot collect evidence from the downstream enforcement point. Those are scope gaps, not a reason to turn a weak test into a passing result.
Practical remediation should name the affected identity, tool, and failed boundary. Typical actions include replacing a broad connector credential with scoped delegation, enforcing object-level authorization in the downstream API, separating read and write tools, adding approval before an irreversible action, revoking dormant credentials, and protecting correlated audit events. Retest with the original denied case plus a legitimate allowed case, then record the test scope and limitations. Penetration test preparation can help arrange the roles, contacts, and safe records needed for that retest.
FAQ
Is prompt injection still in scope?
Yes. Use it to test whether untrusted content can influence an agent. Measure success by whether that influence bypasses authorization, causes an unapproved tool action, exposes protected data, or breaks the audit trail.
Should the agent use its own account or the user’s identity?
The assessment should examine both designs where they exist. A constrained agent identity can fit a service function, while delegation can preserve user-specific access boundaries. Test effective permissions, revocation, downstream enforcement, and audit attribution in either design.
What is the minimum proof I should expect?
Expect an identity map, tool-permission inventory, tested authorization matrix, redacted allow and deny evidence, audit-trail validation, remediation, retest criteria, and stated limitations. Prompts and screenshots alone do not establish security for a business workflow.
Comments
No comments yet. Be the first!
Leave a Comment