We use three outcomes. They are boring on purpose. Boring language survives calendar pressure better than vibes.

The three outcomes

Ship — critical probes done; open issues ranked with owners; rollout audience matches actual confidence.

Wait — coverage partial; keep blast radius small while you close gaps; named owner and date required.

Block — data exposure, ungated writes, or missing fallbacks are still unknown on a user-facing flow.

“We feel ready” is not on the list. Neither is “sales needs it Friday” without a named person accepting residual risk.

Why three beats two

A binary go/no-go sounds clean and fails in practice. Teams need a middle path that is still disciplined: narrow the audience, flag the feature, remove write actions, fix fallbacks, then reconsider.

Wait is not a soft forever. If there is no owner and no date, you do not have Wait. You have avoidance.

What usually forces a Block

In AI products, Block often comes from unknowns, not from taste debates:

You can love the product and still Block external rollout. Those can both be true.

What Ship still allows

Ship does not mean perfect. It means:

Medium findings with owners are normal. Silent Critical findings are not.

Who decides

Someone with authority to delay the launch. Testing partners and QA can recommend. They should not be the only signature on residual risk.

Write the call down:

Put it in the launch channel. Private optimism does not count.

A 15-minute facilitation script

  1. Answer the four coverage questions (adversarial, leakage, fallbacks, writes) out loud.
  2. Add one product-specific fifth.
  3. List open Critical/High only.
  4. Pick Ship, Wait, or Block.
  5. If Wait: set audience limit + owner + date.

If step 4 takes half an hour of debate with no facts, you are missing probes, not vocabulary.

Examples of clean calls

Ship (narrow): Core probes done for a read-only assistant. Two Medium copy issues owned. Audience is internal flag. Retest pack exists for prompt changes.

Wait: Write path exists; confirmation is soft; fallback on timeout is blank. Keep flag on, fix this week, do not expand to external users.

Block: Product can see multi-tenant data; isolation not probed; model can update records with little friction. External launch waits.

Notice none of these require fake metrics. They require facts you can point to.

When people push back

“We will look bad if we Wait.”

You will look worse if users find Critical issues you never checked.

“The model vendor is safe.”

Their model card is not your product test.

“It is only a soft launch.”

Soft launch still needs gates proportional to data and write power.

Pushback is normal. Facts still decide.

Building a culture that can say Block

Teams that never Block are not braver. They are usually later to learn.

Practical habits:

Culture is just repeated release notes. Write better ones.

Tools and the longer guide

Interactive helper: go / no-go

Full criteria, examples, and ownership rules:

Go / no-go criteria for shipping AI features

Pair with the launch checklist so Wait items do not evaporate.

Where GenCodeQA fits

We do not replace your release owner. We provide evidence for their call: severity-ranked findings, locations, and retests when scoped.

If your team is stuck arguing Ship vs Wait, an independent pass often ends the loop faster than another internal demo. Free triage. Bring the disagreement, not a performance review of the model vendor.