Domain

39 case studies

Healthcare / life sciences · RAG clinical summarizer

Clinical RAG that wrote false diagnoses into the chart

A clinic wanted an assistant to draft visit notes from the chart. The demo looked fluent. The risk was invented conditions landing in the record.

Read the case →
Healthcare / life sciences · Ambient voice / multimodal scribe

Ambient scribe leaking PHI into the wrong vendor path

A voice-to-note tool sat in exam rooms. The product risk was not only transcript quality — it was where the audio and PHI actually went.

Read the case →
Healthcare / life sciences · Member / patient chatbot

Patient chatbot that would dose instead of routing

A payer portal chatbot was meant to explain benefits and book appointments. Users asked clinical questions. The model answered like a doctor.

Read the case →
Fintech / banking / payments · Retail banking copilot

Banking copilot that invented balances and “you’re approved”

The assistant could read transactions and answer “can I afford this?” It also invented available balance and implied credit approval.

Read the case →
Fintech / banking / payments · Agent / tool-calling payouts

Payments agent that double-paid on a retry

An ops agent could send payouts from a ticket. Timeouts and retries created duplicate movements.

Read the case →
Fintech / banking / payments · Onboarding / KYC document AI

KYC extraction that merged the wrong person

Document AI sped up onboarding. It also swapped names across uploads and let pending users transact.

Read the case →
Insurance · Policyholder copilot

Policy chatbot that invented coverage

The bot was supposed to explain the customer’s policy. It invented riders, waiting periods, and claim outcomes.

Read the case →
Insurance · FNOL and claims-status assistant

Claims copilot that skipped FNOL rules and leaked files

First notice of loss looked simple in a demo. Duplicate claims, document isolation, and invented status were the actual defects.

Read the case →
Legal / compliance · Legal research / brief assistant

Legal research assistant that cited cases that do not exist

The assistant drafted memos with real-looking citations. Several opinions were fabricated, then defended when challenged.

Read the case →
Legal / compliance · Matter-scoped knowledge assistant

Matter RAG that answered from another client’s files

A firm assistant was “restricted to this matter.” Retrieval still pulled another client’s contract clause into the answer.

Read the case →
Education · District / LMS copilot

School assistant that exposed another student’s grades

Teachers liked the homework helper. A student prompt plus a guessed id returned someone else’s scores.

Read the case →
Education · Assignment scoring assistant

Grading copilot without a human review path

The tool scored essays in seconds. It also rewarded verbose fluff and penalized dialect, with no teacher override trail.

Read the case →
Government / public sector · Official .gov assistant

Civic chatbot that told businesses to break the law

A city assistant was meant to explain licenses and labor rules. It gave illegal advice with the authority of a government site.

Read the case →
Government / public sector · Benefits / tax information assistant

Benefits and tax bot that was confidently wrong

The bot reduced call volume in the demo. Independent questions showed it was wrong more often than the static site.

Read the case →
HR / recruiting · Scoring / traditional ML + copilot

Resume ranker that learned to penalize women

A ranking model trained on historical hires. Historical hires were mostly men. The model treated that as signal.

Read the case →
HR / recruiting · Interviewer assistant

Interview copilot that hallucinated the candidate’s career

The copilot briefed hiring managers. It merged two similarly named candidates and invented a promotion.

Read the case →
E-commerce / retail · Customer support copilot

Support copilot that invented a return policy

The bot promised 90-day returns and price-match refunds the store did not offer. Customers screenshot the chat.

Read the case →
E-commerce / retail · Shopping / inventory agent

Shopping agent that sold stock it did not have

The agent could add to cart and “reserve” items. It invented delivery dates and won a last-unit race twice.

Read the case →
E-commerce / retail · Multi-tenant marketplace assistant

Marketplace copilot that quoted another seller’s payout

Sellers asked the copilot about orders. It mixed shops, commissions, and buyer PII.

Read the case →
Cybersecurity · Detection / vuln-triage assistant

SOC copilot that hallucinated CVEs and patches

The copilot drafted tickets with CVE ids, impact, and fix versions. Some CVEs did not exist. Some patches were wrong.

Read the case →
Cybersecurity · Multi-tenant findings platform

Security SaaS whose AI summary leaked another org’s vulns

The weekly AI digest was the leak. The table view was isolated. The summary was not.

Read the case →
DevTools / infrastructure · Coding agent / copilot

Coding agent that installed a package that does not exist — yet

The agent hallucinated a dependency name. Attackers can squat those names. It also pasted a customer token into a generated config.

Read the case →
DevTools / infrastructure · CI / ops agent

Infra agent that ran write tools on a forged webhook

The agent scaled clusters and reran jobs from chat and from webhooks. Unsigned events were treated as trusted instructions.

Read the case →
Media / content · Editorial assistant / summarizer

News summarizer that invented quotes

The tool shortened articles and auto-captioned video. It attributed sentences nobody said and leaked drafts on guessed URLs.

Read the case →
Media / content · UGC moderation assistant

Moderation copilot that missed the report queue

The model triaged comments. It auto-published a subset and skipped the legal-hold queue.

Read the case →
Consumer / social · Consumer companion / character chat

Companion chatbot without an age-appropriate fail-closed path

The product was marketed broadly. Safety filters were weaker in “character” mode. Age gates were a checkbox.

Read the case →
Consumer / social · Social feed with AI replies

Social AI replies that ignored block and report

Users could block people. The generative reply feature still pulled blocked users’ posts into “suggested responses.”

Read the case →
Real estate / proptech · Listing Q&A assistant

Listing assistant that invented HOA, flood, and school facts

Buyers asked about HOA fees and flood zones. The assistant answered from neighborhood vibes, not listing fields.

Read the case →
Real estate / proptech · Listing marketing / search assistant

Proptech copilot that generated discriminatory targeting language

The assistant wrote ad copy and “suggested filters.” Some suggestions steered by family status and coded neighborhood language.

Read the case →
Travel / hospitality · Airline / OTA support chatbot

Airline-style assistant that invented a refund rule

The bot told a traveler to book now and claim a bereavement-style refund later. That rule did not exist. Screenshots did.

Read the case →
Travel / hospitality · Hotel / stay booking with AI concierge

Hotel search that oversold the last room — in chat too

Two guests, one room. The concierge confirmed both. Timezones made check-in dates wrong.

Read the case →
Energy / industrial · OT / operations assistant

Operator copilot that recommended a command the user could not issue

The assistant explained alarms and then suggested a breaker action. The logged-in role was read-only.

Read the case →
Energy / industrial · Telemetry Q&A assistant

Historian chatbot that treated stale tags as live

Operators asked “is the line up?” The assistant answered from a cache that was 14 minutes old without saying so.

Read the case →
AI-generated products · App built with Lovable / Bolt / Cursor / v0

AI-built SaaS that hid buttons but not the API

The generated UI looked finished. Object ids in the URL returned other users’ records. Billing webhooks created entitlements twice.

Read the case →
Agents / tool calling · Multi-step agent workflow

Agent whose tools failed when JSON was almost right

Demos used perfect tool arguments. Production models omitted fields, renamed keys, and retried until the budget burned.

Read the case →
RAG / knowledge assistants · B2B RAG assistant

Knowledge assistant that answered from another tenant’s docs

Citations looked professional. One citation was a contract from a different customer. Another was a ticket with an injection payload.

Read the case →
Voice / multimodal · Voice IVR / phone agent

Voice agent that misheard a yes and executed a transfer

The demo call was quiet. On a noisy line the agent heard “yes” and moved money. There was no repeat-back of amount and beneficiary.

Read the case →
Safety / guardrails · Ticket-grounded support agent

Support RAG that obeyed instructions inside a ticket

The model was aligned in empty-context chat. Retrieved tickets and PDFs contained instructions. Tools ran anyway.

Read the case →
Release gates · SaaS with AI features

Green LLM evals, broken signup, and a ship decision that waited

The eval dashboard was green. Users still could not reset a password. The team almost shipped on model scores alone.

Read the case →

Want this applied to your product?

Free tools first, or a triage call if you want an independent pass before launch.