Each study follows the same shape used on serious QA engagements: overview, numbered challenges, how we tested, starting point vs required controls, what we delivered, outcomes, and stack. Composite product classes — not named-client testimonials, and no invented percentages.
39 studies · Filter by domain · Same method as the evaluation framework.
A clinic wanted an assistant to draft visit notes from the chart. The demo looked fluent. The risk was invented conditions landing in the record.
Read the case →A voice-to-note tool sat in exam rooms. The product risk was not only transcript quality — it was where the audio and PHI actually went.
Read the case →A payer portal chatbot was meant to explain benefits and book appointments. Users asked clinical questions. The model answered like a doctor.
Read the case →The assistant could read transactions and answer “can I afford this?” It also invented available balance and implied credit approval.
Read the case →An ops agent could send payouts from a ticket. Timeouts and retries created duplicate movements.
Read the case →Document AI sped up onboarding. It also swapped names across uploads and let pending users transact.
Read the case →The bot was supposed to explain the customer’s policy. It invented riders, waiting periods, and claim outcomes.
Read the case →First notice of loss looked simple in a demo. Duplicate claims, document isolation, and invented status were the actual defects.
Read the case →The assistant drafted memos with real-looking citations. Several opinions were fabricated, then defended when challenged.
Read the case →A firm assistant was “restricted to this matter.” Retrieval still pulled another client’s contract clause into the answer.
Read the case →Teachers liked the homework helper. A student prompt plus a guessed id returned someone else’s scores.
Read the case →The tool scored essays in seconds. It also rewarded verbose fluff and penalized dialect, with no teacher override trail.
Read the case →A city assistant was meant to explain licenses and labor rules. It gave illegal advice with the authority of a government site.
Read the case →The bot reduced call volume in the demo. Independent questions showed it was wrong more often than the static site.
Read the case →A ranking model trained on historical hires. Historical hires were mostly men. The model treated that as signal.
Read the case →The copilot briefed hiring managers. It merged two similarly named candidates and invented a promotion.
Read the case →The bot promised 90-day returns and price-match refunds the store did not offer. Customers screenshot the chat.
Read the case →The agent could add to cart and “reserve” items. It invented delivery dates and won a last-unit race twice.
Read the case →Sellers asked the copilot about orders. It mixed shops, commissions, and buyer PII.
Read the case →The copilot drafted tickets with CVE ids, impact, and fix versions. Some CVEs did not exist. Some patches were wrong.
Read the case →The weekly AI digest was the leak. The table view was isolated. The summary was not.
Read the case →The agent hallucinated a dependency name. Attackers can squat those names. It also pasted a customer token into a generated config.
Read the case →The agent scaled clusters and reran jobs from chat and from webhooks. Unsigned events were treated as trusted instructions.
Read the case →The tool shortened articles and auto-captioned video. It attributed sentences nobody said and leaked drafts on guessed URLs.
Read the case →The model triaged comments. It auto-published a subset and skipped the legal-hold queue.
Read the case →The product was marketed broadly. Safety filters were weaker in “character” mode. Age gates were a checkbox.
Read the case →Users could block people. The generative reply feature still pulled blocked users’ posts into “suggested responses.”
Read the case →Buyers asked about HOA fees and flood zones. The assistant answered from neighborhood vibes, not listing fields.
Read the case →The assistant wrote ad copy and “suggested filters.” Some suggestions steered by family status and coded neighborhood language.
Read the case →The bot told a traveler to book now and claim a bereavement-style refund later. That rule did not exist. Screenshots did.
Read the case →Two guests, one room. The concierge confirmed both. Timezones made check-in dates wrong.
Read the case →The assistant explained alarms and then suggested a breaker action. The logged-in role was read-only.
Read the case →Operators asked “is the line up?” The assistant answered from a cache that was 14 minutes old without saying so.
Read the case →The generated UI looked finished. Object ids in the URL returned other users’ records. Billing webhooks created entitlements twice.
Read the case →Demos used perfect tool arguments. Production models omitted fields, renamed keys, and retried until the budget burned.
Read the case →Citations looked professional. One citation was a contract from a different customer. Another was a ticket with an injection payload.
Read the case →The demo call was quiet. On a noisy line the agent heard “yes” and moved money. There was no repeat-back of amount and beneficiary.
Read the case →The model was aligned in empty-context chat. Retrieved tickets and PDFs contained instructions. Tools ran anyway.
Read the case →The eval dashboard was green. Users still could not reset a password. The team almost shipped on model scores alone.
Read the case →No studies in that domain. Pick All domains.
Free tools first, or a triage call if you want an independent pass before launch.