1. Have you tested adversarial or ambiguous prompts on purpose?
2. Can the model expose private context, system instructions, or another tenant's data?
3. What happens when the model times out, refuses, or returns junk?
4. Do write / send / delete actions require an approval gate?
5. After the last prompt or model change, did you recheck core journeys?
6. Who owns go / no-go for this AI release?

Want this checked on your real product?

Bring this output to a triage call — we'll turn it into a scoped plan.