- A launch risk score is a conversation starter, not a compliance certificate.
- Answer the prompts honestly. Optimistic clicks waste your own time.
- Higher urgency usually means missing probes on safety, fallbacks, or write actions.
- Next steps: gap finder → checklist → go / no-go → scoped testing if needed.
- Do not put the score in a customer deck as “proof we tested.”
- Re-run after major prompt, model, or tool changes, not until you like the number.
Why a score helps at all
Teams often know something feels thin but cannot prioritize. Calendar pressure fills the gap. A structured self-check forces plain answers in front of other people:
- Have you tested adversarial or messy prompts on purpose?
- Do you know what happens on model failure?
- Can the system expose data it should not?
- Are write / send / delete actions gated?
That is enough to change a launch meeting from vibes to a short list of unknowns.
Our interactive version takes about a minute:
Launch risk score (free, no signup)
What the score is
It is:
- A structured set of honest questions about coverage gaps that matter for AI products
- A way to start a launch thread with less hand-waving
- A pointer to what to do next
- Optional homework before a triage call
It is directional. Two teams can get similar urgency for different reasons. Always read the answers, not only the label.
What the score is not
It is not:
- A security certification
- A substitute for a checklist with owners
- Proof for a customer that “we tested everything”
- A statistical model of your company
- A number you re-roll until it looks calm
- Legal, compliance, or insurance advice
If you need those things, you need different processes and qualified advisors. This tool will not pretend otherwise.
How to take it without kidding yourself
- Sit with someone who did not write all the prompts.
- If you are unsure, pick the weaker option.
- If the product has write actions or private data, do not mentally skip those rows.
- Save or screenshot the result for your launch thread so the discussion has a baseline.
- Do not “practice” until you get a calmer result. That only trains self-deception.
A useful facilitation script: “Answer as if a skeptical teammate will retest tomorrow.”
How to read the result
Exact labels can change as we tune tool copy. The useful reading is directional.
Lower urgency
Your answers suggest more of the basics are covered. That is not “done.”
Still do:
- A checklist with owners
- A regression plan for the next prompt / model change
- A written ship / wait / block call
Medium urgency
You have partial coverage. Prioritize open gaps before widening rollout.
Typical moves:
- Close safety / fallback / write-action unknowns first
- Keep the audience narrow (flag, internal, single workspace)
- Put names and dates on open High items
Higher urgency
Pause feature expansion theater. Probe safety, permissions, and fallbacks before you celebrate the demo.
Typical moves:
- Block or Wait on broad launch
- Run the gap finder and checklist the same day
- Consider an independent pass if the date is real and Critical questions remain
If the score surprises you, that is useful. If it does not surprise you and you still planned to ship tomorrow with open safety questions, the score did its job.
What to do next (practical sequence)
- Gap finder — tick what you already cover
- Checklist — 24-point launch checklist for owners and tracking
- Go / no-go — ship / wait / block helper
- Independent pass — if date risk is real, book a free triage
That order keeps you from jumping straight from anxiety to a random test idea.
Full method: How to test an AI app before launch.
Turning a medium/high score into a one-week plan
Example plan (adjust to your product):
- Day 1: Write the risk half-page (actions, data, blast radius). Run gap finder.
- Day 2: Force fallback and timeout paths in staging. Fix or flag UX.
- Day 3: Document and run a small adversarial / ambiguous prompt set.
- Day 4: Probe tenant / role boundaries and write-action gates.
- Day 5: Severity list, owners, go / no-go, and a note to leadership with open Critical/High only.
You will not finish every nice-to-have. You can finish the expensive unknowns.
Facilitating the score in a launch meeting
Try this 20-minute agenda:
- Project the tool. One person drives. Everyone else can challenge answers.
- For each question, ask “what would we show a skeptical teammate tomorrow?”
- Screenshot the result into the channel before debate starts.
- Spend the remaining time only on Gap or Weaker answers.
- Leave with owners and dates, not with a vibes rematch.
If the meeting turns into arguing about whether the badge is “fair,” you have lost the plot. The unanswered probe is the plot.
Comparing score language to release language
Use the score for urgency. Use go / no-go for the decision. They are related, not identical.
- Higher urgency often leads to Wait or Block for broad audiences
- Lower urgency does not auto-approve Ship without checklist evidence
- Medium urgency is usually “narrow the blast radius while you close gaps”
Keep both artifacts. One without the other recreates optimism with better fonts.
How GenCodeQA uses scores in triage
On a triage call we care more about your answers than the number:
- Product type and model setup
- Write actions and data sensitivity
- What you have already tested
- Launch timing
The score is optional homework. It helps you arrive with sharper questions. It is not an intake exam and it is not a quote. We will not invent a dollar price from a self-check.
After you ship: keep the score honest
A launch is not the end of risk. When you change prompts, models, tools, or retrieval sources, coverage can quietly decay.
Simple rule:
- Re-run the score when the AI surface changes in a user-facing way
- Diff the answers against last time
- Any answer that moved from stronger to weaker is a release note item
This is cheaper than discovering decay through support tickets.
Limits of self-assessment
Self-checks inherit your blind spots. Teams that built the prompts often under-rate adversarial gaps. Teams under deadline under-rate fallbacks.
Mitigations:
- Always include one non-author in the scoring session
- Prefer weaker answers when unsure
- Verify Gap answers with an actual probe the same week
- Bring disputed answers to an independent reviewer if the date is hard
A score is a mirror. Mirrors do not replace inspections.
Common false confidence
- Running the score alone and declaring the product safe
- Re-running until you get a calmer result by changing answers
- Confusing model vendor marketing with product risk
- Ignoring medium scores because the UI looks polished
- Pasting the score into a sales deck as third-party validation
- Treating lower urgency as permission to skip regressions
Free tool
Companion post: What a risk score is not.
When to get a second pair of eyes
If the score (or your gut after honest answers) says urgency is high, and you have a real launch date, bring the output to a triage. We will turn it into scope, not theater.
FAQ
Is this score scientific?
It is a structured self-check based on failure modes we see in AI product QA. It is not a peer-reviewed statistical model of your company.
Can we use it for every release?
Yes as a quick gate before wider rollout. Pair it with checklist items that changed since last time.
Does a low score mean we do not need QA?
No. It means your answers suggest lower immediate urgency. Journeys, regressions, and eval hygiene still matter.
Should we share scores with customers?
Usually no. Share what you tested, what remains open, and how you gate releases. A self-score is an internal tool.
What if two teammates get different results?
That disagreement is the finding. Align on facts (was tenant isolation probed or not?) before you argue about the label.