This post is the boundary document.

What it is for

A structured self-check can:

That is enough value for a free tool. It does not need to pretend to be more.

What it is not

Not a security certification.

Nobody should wave a self-score at a customer as third-party assurance.

Not a substitute for a checklist with owners.

A label without open items and names does not ship a product safely.

Not proof you tested everything.

It measures what you claim about coverage. Claims still need probes.

Not a number to optimize.

Re-answering until the result looks calm only trains self-deception.

Not legal, compliance, or insurance advice.

If you need those, talk to qualified people. We will not invent that authority.

Not a quote.

We will not turn a self-check into a fabricated price.

How misuse shows up

If any of those are happening, put the tool down and write the open Critical questions on a whiteboard.

How to use it well

  1. Take it with someone who did not write all the prompts.
  2. Pick the weaker option when unsure.
  3. Screenshot the result into the launch channel.
  4. Immediately run the gap finder or checklist.
  5. End the week with a written ship / wait / block call.

The score starts the conversation. The gate ends it.

Trust is specific

Readers trust teams that admit limits. A credible internal note sounds like:

“Risk self-check landed medium/high because write gates and fallbacks are weaker than we want. We are keeping the beta flag on, closing those two gaps this week, and re-running go / no-go on Thursday.”

That is more trustworthy than “our risk score is green.”

If you need a public narrative, talk about process, not a badge:

No invented customers required.

A note on confidence theater

Dashboards, green badges, and self-scores feel like progress. They become harmful when they replace probes.

Ask one filter question before you share any score externally or in a launch email:

“Can a skeptical engineer retest the claim behind this number this week?”

If the answer is no, keep the score internal and go gather evidence.

Read the answers, not only the badge

Two teams can land in similar urgency for different reasons. One might lack fallbacks. Another might have ungated writes. The next actions differ.

If teammates get different scores, do not average them. Resolve the factual disagreement: was tenant isolation probed or not?

The full guide

How to interpret urgency bands and turn them into a one-week plan:

Score your AI launch risk (and what to do next)

Where we stand

We test. We do not build. A free score is homework, not theater, and not a certificate with our logo on your launch.

If honest answers say urgency is high and the date is real, book a free triage. Bring the answers. Skip the polished deck.