Managed AI

Create an AI-Agent Acceptance Test Plan

Turn a bounded agent operating contract into a reviewable acceptance matrix covering normal, edge, failure, and prohibited-action cases.

Quick facts

Best for
Owners · Founder-operators
Prompt type
Effectiveness
Expected result
An acceptance matrix with evidence and source attribution, Planned versus Executed versus Passed status, critical-failure stop conditions, conflicts and Unknowns, and named-human approval gates.
Time saved
Varies by agent scope and supplied test evidence
Required inputs
AGENT JOB · ALLOWED AND PROHIBITED ACTIONS · SYSTEMS AND DATA CLASSES · TEST CASES · CRITICAL FAILURES · EVIDENCE REQUIREMENTS · ACCEPTANCE OWNER AND APPROVAL GATES · CREDENTIAL AND ACCESS BOUNDARIES
Works with
ChatGPT + Claude

What this prompt does

  • Creates an acceptance matrix with expected evidence, failure conditions, owners, and approval gates.
  • Separates tests that are merely planned from tests that were executed and tests whose evidence passed review.

When to use

  • After one agent job, its non-goals, source boundaries, permissions, and human decision rights are documented.

When not to use

  • Do not treat a generated plan as proof that any test ran, passed, or authorized live operation.
  • Do not paste secrets or credentials, or place unauthorized agent data in an unapproved AI system.

The prompt

#CONTEXT:
Create an acceptance-test plan for one bounded AI-agent job. This produces a planning artifact for a named human owner; it does not run tests or authorize a live agent.

#INPUTS:
- Agent job, trigger, inputs, decisions, outputs, completion, non-goals, and manual fallback: [AGENT JOB]
- Explicitly allowed actions, prohibited actions, and actions requiring approval: [ALLOWED AND PROHIBITED ACTIONS]
- Approved systems, source authority, integrations, environments, and data classes: [SYSTEMS AND DATA CLASSES]
- Supplied normal, edge, failure, retry, conflict, adversarial, and prohibited-action cases: [TEST CASES]
- Critical failures, unacceptable outcomes, stop conditions, and rollback triggers: [CRITICAL FAILURES]
- Required logs, records, citations, reconciliation checks, and evidence retention: [EVIDENCE REQUIREMENTS]
- Named human acceptance owner, reviewers, decision rights, and explicit approval gates: [ACCEPTANCE OWNER AND APPROVAL GATES]
- Credential owners, approved identities, scopes, access limits, and secret-handling constraints: [CREDENTIAL AND ACCESS BOUNDARIES]

#INSTRUCTIONS:
1. Restate the bounded job and map every criterion to supplied operating-contract evidence or source attribution. Label missing facts Unknown and conflicting sources Conflict—human resolution required.
2. Build separate normal, edge, degraded-tool, retry/idempotency, source-conflict, approval, escalation, data-boundary, permission, and prohibited-action cases. Do not invent missing cases as approved business rules.
3. For every test, specify synthetic or sanitized input, approved source, expected decision, expected draft or bounded action, prohibited result, observable evidence, pass condition, failure condition, stop condition, and named human reviewer.
4. Use only these status values: Planned—not executed, Executed—not yet reviewed, Passed—evidence reviewed, Failed, Blocked, or Unknown. Never convert Planned or Executed into Passed without supplied reviewed evidence.
5. Fail closed when a required source, permission, approval, owner, criterion, evidence field, or conflict resolution is missing. The safe outcome is stop, preserve evidence, and route to the named human owner.
6. Do not invent, infer, or fabricate an SLA, threshold, result, pass rate, test execution, monitoring state, owner, permission, source, or capability. Unsupported values remain Unknown.
7. Include explicit approval gates for acceptance, any waived failure, scope change, permission change, connection, deployment, enablement, restart, and live-authority expansion. A completed test does not imply approval for the next gate.
8. Do not infer authorization from an allowed-action list, credential reference, prior approval, planned test, executed test, or passing evidence.
9. Do not grant, revoke, or change permissions. Do not provision an agent or service account with access, permissions, or a role. Do not use or reveal credentials. Do not rotate any credential, token, API key, or secret. Do not deploy, enable, or disable an agent. Do not restart, pause, or activate an agent. Do not connect an agent or integration to a live system. Do not execute transactions. Do not send externally. Do not change production or claim monitoring is active.
10. Address privacy, security, data classes, least privilege, source access, retention, and evidence minimization. Use only authorized agent data in an approved AI workspace or approved AI vendor. Minimize inputs and redact or omit unnecessary personal, sensitive, or confidential data.
11. Never paste secrets, credentials, tokens, API keys, private keys, or recovery codes. Refer only to credential owner, identity name, scope, rotation status, and approved secret manager. Follow company retention and vendor policy.

#RESPONSE FORMAT:
## Operating-contract evidence and source attribution
| Requirement | Supplied source | Authority or owner | Status | Conflict or Unknown | Stop condition |
|---|---|---|---|---|---|

## Acceptance matrix
| Test ID and class | Sanitized input | Expected decision or draft | Prohibited result | Evidence required | Pass and failure conditions | Status: Planned, Executed, Passed, Failed, Blocked, or Unknown | Human reviewer |
|---|---|---|---|---|---|---|---|

## Critical-failure and fail-closed register
| Critical failure | Detection evidence | Immediate safe state | Escalation owner | Approval gate before resumption |
|---|---|---|---|---|

## Coverage gaps, conflicts, and Unknowns

## Approval-gate register

Input checklist

  • AGENT JOB
  • ALLOWED AND PROHIBITED ACTIONS
  • SYSTEMS AND DATA CLASSES
  • TEST CASES
  • CRITICAL FAILURES
  • EVIDENCE REQUIREMENTS
  • ACCEPTANCE OWNER AND APPROVAL GATES
  • CREDENTIAL AND ACCESS BOUNDARIES

Example input

Fictional example — AGENT JOB: Northstar intake agent drafts an internal lead brief; no customer sending. ALLOWED AND PROHIBITED ACTIONS: Read sanitized form fields and draft; no pricing, deletion, sending, or CRM writes. SYSTEMS AND DATA CLASSES: Sandbox form export, internal business data; personal contact fields omitted. TEST CASES: Two normal cases, missing service area, conflicting source, tool timeout, duplicate event, prompt injection. CRITICAL FAILURES: External message or unsupported eligibility claim. EVIDENCE REQUIREMENTS: Correlation ID, cited rule version, output comparison. ACCEPTANCE OWNER AND APPROVAL GATES: Maya accepts evidence; Omar approves security; launch owner Unknown. CREDENTIAL AND ACCESS BOUNDARIES: Sandbox identity is read-only; secret values excluded.

Expected output structure

  • An acceptance matrix with evidence and source attribution, Planned versus Executed versus Passed status, critical-failure stop conditions, conflicts and Unknowns, and named-human approval gates.

Customize this prompt

  • Keep a stable regression set and a separate rotating review set without weakening prohibited-action cases.
  • Require business-system evidence for completion, not a plausible-looking model response.

Guardrails

  • Fail closed on missing evidence, unresolved conflict, absent ownership, unavailable approval, or permission uncertainty; preserve evidence for review by a named human owner.
  • Keep test states distinct: Planned is not Executed, Executed is not Passed, and Passed requires supplied reviewed evidence.
  • Preserve evidence or source attribution for every criterion and result; unsupported facts remain Unknown.
  • Do not invent, infer, or fabricate any SLA, threshold, result, pass rate, execution, monitoring state, owner, permission, or capability.
  • Do not claim monitoring is active without supplied operating evidence.
  • A named human owner must review and approve every payment or financial decision before any action.
  • A named human owner must review and approve every legal, liability, admission, settlement, or contract decision before any action.
  • A named human owner must review and approve every account, access, permission, or credential decision before any action.
  • A named human owner must review and approve every policy exception, override, or waiver before any action.
  • A named human owner must review and approve every external send or publication before any action.
  • A named human owner must review and approve every deletion, closure, suspension, termination, revocation, deactivation, disablement, or archive decision before any action.
  • A named human owner must review and approve every promise, guarantee, or commitment before any action.
  • Use explicit approval gates; a named human owner must approve acceptance, waivers, scope changes, permissions, deployment, restart, and live-authority expansion before action.
  • Do not infer authorization. Do not grant, revoke, or change permissions. Do not provision an agent or service account with access, permissions, or a role. Do not use or reveal credentials. Do not rotate any credential, token, API key, or secret. Do not deploy, enable, or disable an agent. Do not restart, pause, or activate an agent. Do not connect an agent or integration to a live system. Do not execute transactions. Do not send externally. Do not change production or claim monitoring is active.
  • Protect privacy and security: use only authorized agent data in an approved AI workspace or approved AI vendor; minimize data; redact or omit unnecessary personal, sensitive, or confidential data; follow company retention and vendor policy.
  • Never paste secrets, credentials, tokens, API keys, private keys, or recovery codes.

Related articles

NEXT STEP

Next step

Find out what your company's knowledge is worth.