Running an Authorised Exercise
Authorisation and Scope Come First
Everything in this module applies to systems you own or have explicit written authorisation to test. That is the operating constraint of the work, not a formality attached to it. Probing another organisation's AI product, a hosted model you merely have an account with, or a supplier's agent without permission is unauthorised access regardless of intent, and provider terms of service usually prohibit adversarial probing independently of the law. Before any testing begins, agree in writing which systems, endpoints and environments are in scope; which are explicitly excluded; which accounts, credentials and data may be used; the time window; the named contact for both sides; and what happens if the exercise causes an outage or uncovers a live compromise. Test against environments configured like production but populated with synthetic data, using dedicated accounts. Nothing that follows is legitimate without this step.
- Only systems you own or are authorised in writing to test — no exceptions
- Provider terms commonly prohibit adversarial probing on their own terms
- Agree scope, exclusions, accounts, data, window, and named contacts before starting
- Production-like configuration with synthetic data and dedicated accounts
- Define the stop condition and escalation path for a real incident found mid-exercise
Rules of Engagement That Prevent Harm
The rules of engagement translate scope into constraints on how the team works, and the AI-specific clauses are worth writing explicitly. No real customer data as attack material and none in evidence, since findings travel further than anyone plans. No planting of content in shared systems outside the agreed environment, because indirect injection tests write into corpora and a forgotten test document is a live liability. Destructive actions only against designated targets, with a rollback path agreed in advance. Rate limits so testing does not become a denial-of-service against a shared dependency. A defined handling procedure for anything sensitive the team obtains, including immediate reporting and secure deletion. And a rule that the team stops and escalates on discovering evidence of an actual intrusion rather than continuing to explore it, which is the situation that most often goes wrong in practice.
- Synthetic data only, in the exercise and in the evidence
- Planted test content stays in the agreed environment and is removed afterwards
- Agree rollback for destructive actions and rate limits for shared dependencies
- Stop and escalate on signs of a real intrusion instead of investigating further
Safe Methodology and Evidence
Design tests so success is observable without causing harm. The standard approach is a benign objective: a distinctive marker token that should never appear in output, a designated no-op tool that should never be called, a canary record that should never be retrieved or transmitted. If injected content can cause the marker to appear at a destination, you have demonstrated the path without exercising real harm, and the result is unambiguous, automatable and safe to include in a report. Keep evidence proportionate: enough to reproduce in a controlled environment and to verify the fix, and no more. Record the effort — attempts made, success rate, turns required, whether the deployed guardrail stack was active — because a path that succeeds once in fifty attempts is a real finding with a different priority than one that succeeds immediately, and omitting effort makes both look identical.
- Benign objectives: marker tokens, no-op tools, canary records that should never move
- Demonstrate the path, do not exercise the harm
- Record attempts, success rate and turns — effort is part of the finding
- Test the deployed stack with guardrails active, and say which configuration was tested
Reporting That Produces Change
A finding nobody fixes was not worth discovering, and the difference is usually the write-up. Lead with what an attacker achieves and against whom, in the language of the affected system rather than of the technique. State preconditions honestly: what access was required, which configuration was in place, how many attempts, what success rate. Give a reproduction that an engineer can run in a controlled environment with the detail needed to verify and no more — mechanisms and conditions rather than a payload anyone can paste elsewhere. Recommend the fix at the architectural layer that actually contains the issue, because a prompt-level patch for a structural problem produces the appearance of closure and a repeat finding six months later. Define the retest criteria in the report itself, so closure is a measurement someone can perform rather than an opinion someone offers.
- Impact and affected parties first; technique second
- State preconditions, configuration, attempts and success rate — credibility is the asset
- Describe mechanisms and conditions, not distributable payloads
- Recommend the fix at the layer that contains it, and write the retest criteria in the report
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.