Last updated: September 2026. Written by Josh Hutcheson, OnlineCourseing editor. Every case re-checked against court records, company statements or original reporting on 22 September 2026. See our review methodology.
By Josh Hutcheson · E-Learning Specialist
Reviewing online learning platforms since 2019. Review methodology
THE SHORT ANSWER
Bottom line: most AI failures are not rogue machines. They come from bad or biased data, models used outside the conditions they were built for, confident made-up answers, and no one checking the output. The 12 cases below, from 2016 to 2024, show each pattern, what it cost and the lesson that still applies.
- Costliest: Alphabet lost $100 billion in market value after one wrong answer in a Bard ad (Reuters).
- Most expensive fraud: about $25.6 million sent after a deepfake video call (Arup, CNN).
- Legal precedent: Air Canada was held liable for what its chatbot told a customer (2024).
- Still rising: documented AI incidents reached 362, up from 233 in 2024 (Stanford AI Index 2026).
Learn how AI projects work: AI For Everyone →
12 AI failures at a glance
Before you spend money on the wrong online course, read this.
Get the free 2026 Platform Comparison Guide — 12 platforms compared on price, certificates, and refund policies. Instant PDF, plus my honest Tuesday picks.
No spam. Unsubscribe anytime.
| Year | Organization | System | What went wrong | Root cause |
|---|---|---|---|---|
| 2016 | Microsoft | Tay chatbot | Manipulated into posting offensive tweets | Adversarial users; no safeguards |
| 2018 | Amazon | Recruiting tool | Penalized resumes mentioning women | Biased training data |
| 2018 | Uber | Self-driving test car | Fatal crash with a pedestrian in Tempe | Perception failure; weak safety oversight |
| 2018 | IBM | Watson for Oncology | Unsafe treatment suggestions | Trained on synthetic cases |
| 2020 | Inverness Caledonian Thistle | AI camera | Tracked a bald head instead of the ball | Edge case not tested |
| 2021 | Zillow | Zillow Offers pricing | About $304M write-down; unit closed | Model outside its conditions |
| 2023 | Bard launch ad | Wrong telescope fact; $100B market value lost | Hallucination; no fact-check | |
| 2023 | Law firm (Mata v. Avianca) | ChatGPT legal research | Six fake cases filed; sanctions | Hallucination; no verification |
| 2024 | Air Canada | Customer service chatbot | Wrong refund policy; liable in tribunal | Unchecked answers to customers |
| 2024 | New York City | MyCity business chatbot | Advised businesses to break the law | Hallucination on legal questions |
| 2024 | Arup | Deepfake video call (attack) | HK$200M (about $25.6M) sent to fraudsters | AI misuse; weak payment verification |
| 2024 | AI Overviews | Suggested glue on pizza, eating rocks | Satire and forum posts taken as fact |
The 12 AI failures, case by case
1. Microsoft Tay learns the worst of the internet (2016)
Microsoft launched Tay, a chatbot designed to learn from conversations on Twitter, in March 2016. Within its first 24 hours online, a coordinated attack by a subset of users exploited a vulnerability, and Tay began posting offensive and hurtful tweets. Microsoft apologized and took Tay offline (Microsoft).
Lesson: any AI that learns from or responds to the public will be tested by people trying to break it. Plan for abuse before launch, not after.
2. Amazon’s recruiting tool penalizes women (2018)
Amazon built an experimental system to rate job applicants. By 2015 the company realized it was not rating candidates for technical roles in a gender-neutral way, because it had been trained on patterns in resumes submitted over a 10-year period, most of which came from men. It penalized resumes containing the word “women’s”, as in “women’s chess club captain”, and downgraded graduates of two all-women’s colleges. Amazon scrapped the project, Reuters reported in October 2018 (Reuters).
Lesson: a model trained on past decisions reproduces the bias in them. Audit training data and outcomes by group before a model touches real decisions.
3. Uber’s self-driving test car kills a pedestrian (2018)
On 18 March 2018, an Uber automated test vehicle, a modified Volvo XC90, struck a pedestrian crossing a road in Tempe, Arizona, at 39 mph. The pedestrian died. The operator, who was meant to supervise the automated system, began steering only a fraction of a second before impact, according to the US National Transportation Safety Board’s investigation (NTSB).
Lesson: in safety-critical systems, the human backup and the safety process matter as much as the model. Testing in public requires the strictest oversight, not the lightest.
4. IBM Watson for Oncology recommends unsafe treatments (2018)
Internal IBM documents reported by STAT in July 2018 showed that Watson for Oncology often gave erroneous cancer treatment advice, including recommendations the documents called “unsafe and incorrect.” The documents largely blamed its training: the software was taught with a small number of “synthetic” hypothetical cancer cases rather than real patient data (STAT).
Lesson: impressive demos are not clinical evidence. High-stakes AI needs validation on real cases and independent review before it reaches patients or customers.
5. An AI camera follows a bald head instead of the ball (2020)
When the pandemic kept fans out of stadiums, Scottish club Inverness Caledonian Thistle began live streaming its matches with an automatic camera that used AI ball-tracking. During one live stream it repeatedly mistook a linesman’s bald head for the ball, to the frustration and amusement of fans watching at home (The Verge).
Lesson: a harmless failure, but a perfect example of an edge case nobody tested. Real-world conditions contain things your training data did not.
6. Zillow’s pricing algorithm buys homes it cannot sell (2021)
Zillow Offers used algorithmic price estimates to buy homes directly and resell them. In November 2021 Zillow announced it would wind the business down. It took a write-down of about $304 million on homes bought for more than it expected to sell them for, and planned to cut about 25% of its workforce. Chief executive Rich Barton said the unpredictability in forecasting home prices “far exceeds what we anticipated” (Zillow Group).
Lesson: a model that works in a stable market can fail when conditions shift. When a model’s errors are paid for in cash, cap the exposure and monitor for drift.
7. Google Bard’s launch ad gets a fact wrong (2023)
In a promotional video for Bard, its new chatbot, Google showed an answer claiming the James Webb Space Telescope took the very first pictures of a planet outside our solar system, which was inaccurate. Alphabet lost $100 billion in market value the day the error was reported (Reuters).
Lesson: generative AI can be fluent and wrong at the same time. Anything published, especially in marketing, needs a human fact-check.
8. Lawyers file fake cases invented by ChatGPT (2023)
In Mata v. Avianca, lawyers in New York submitted a brief that cited six court cases that did not exist; ChatGPT had generated them. In June 2023, US District Judge P. Kevin Castel found the lawyers had acted in bad faith and ordered them and their firm to pay a $5,000 fine (Reuters).
Lesson: never use generative AI as a source of facts you do not verify. Treat every citation, figure and quote it produces as unconfirmed until checked.
9. Air Canada is held liable for its chatbot (2024)
After his grandmother died, Jake Moffatt asked Air Canada’s website chatbot about bereavement fares. It told him he could apply for the discount after travelling, which contradicted the airline’s actual policy. When he was refused, he took the case to British Columbia’s Civil Resolution Tribunal. Air Canada argued it could not be held liable for information from its chatbot; the tribunal called the suggestion that the chatbot was a separate legal entity “a remarkable submission”, found negligent misrepresentation and awarded $650.88 in damages in February 2024 (Civil Resolution Tribunal).
Lesson: companies own what their AI says to customers. Ground customer-facing chatbots in current, approved policy and route anything uncertain to a person.
10. New York City’s chatbot tells businesses to break the law (2024)
New York City launched the Microsoft-powered MyCity chatbot in October 2023 to help business owners. An investigation by The Markup in March 2024 found it telling businesses they could take workers’ tips, that landlords could discriminate based on source of income and that restaurants could go cash-free, even though a 2020 city law requires businesses to accept cash (The Markup).
Lesson: legal, medical and financial questions are where hallucinations do the most harm. Public-sector and regulated services need tested answers, clear warnings and human escalation.
11. A deepfake video call costs Arup $25 million (2024)
A finance employee in Arup’s Hong Kong office joined a video call in which the other participants, including someone appearing to be the company’s chief financial officer, were deepfakes. Believing the colleagues were real, the employee made 15 transfers totalling HK$200 million, about $25.6 million. Arup confirmed it was the victim in May 2024 (CNN).
Lesson: this is AI misuse rather than an AI system failing, and it shows how verification processes must change. Large payments should require confirmation through a separate, pre-agreed channel, however convincing the request looks.
12. Google’s AI Overviews suggest glue on pizza (2024)
Shortly after launching AI Overviews in US search results, Google faced viral screenshots of bad answers, including advice to use glue to make cheese stick to pizza and to eat rocks. Google explained that forum posts and satirical content had been misread, and that AI Overviews had sometimes misinterpreted language on web pages; it said it made more than a dozen technical improvements (Google).
Lesson: at the scale of billions of queries, rare failures become common. Systems that summarize the web need to weigh source reliability, not just relevance.
Why AI projects fail: the six patterns
Across these cases, the same causes repeat. Very few failures are about the algorithm alone.
| Pattern | Cases | How to prevent it |
|---|---|---|
| Biased or unrepresentative data | Amazon, IBM Watson | Audit data and outcomes by group; validate on real cases |
| Model used outside its conditions | Zillow, Uber, AI camera | Test edge cases; monitor drift; cap exposure |
| Hallucination stated as fact | Bard, Mata v. Avianca, NYC MyCity, AI Overviews | Ground answers in trusted sources; human fact-checks |
| No human accountability | Air Canada, Mata v. Avianca | Name an owner for AI output; escalate uncertain cases |
| Adversarial manipulation | Tay, Arup | Red-team before launch; verify high-risk requests out of band |
| Hype over a clear need | Many cancelled projects | Start from a measurable business problem |
The same pattern shows up in the data. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, describing most as “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied” (Gartner). McKinsey finds only 37% of organizations see any positive effect on EBIT from AI (McKinsey), and 66% of developers say their most common frustration with AI tools is output that is “almost right, but not quite” (Stack Overflow).
AI failures are becoming more common, not less
Stanford’s 2026 AI Index counts 362 documented AI incidents, up from 233 in 2024, and notes that reporting on responsible-AI benchmarks remains patchy even as capability grows. It also illustrates why failures are hard to predict: the same leading models that can win a gold medal at the International Mathematical Olympiad read an analog clock correctly only 50.1% of the time, a pattern researchers call the jagged frontier (Stanford HAI). More AI in more places means more opportunities for these gaps to show. Our guide to AI trends in 2026 covers the wider picture.
How AI failures have changed
The older cases on this list, from 2016 to 2021, are failures of predictive AI: a model classified, ranked or estimated something badly. Amazon’s tool ranked candidates with a learned bias, Zillow’s model misjudged prices when the market moved, and the Uber and camera systems misperceived the world around them. The damage was serious but usually contained to the decisions that model made.
The newer cases are failures of generative AI, and they behave differently. Bard, ChatGPT in Mata v. Avianca, New York City’s chatbot and Google’s AI Overviews all produced fluent, confident and false statements, often in public, where one screenshot can travel worldwide. Deepfakes add deliberate misuse: the Arup fraud did not need any AI system to malfunction at all. The next wave involves AI agents that take actions rather than just give answers, which is why Gartner expects so many agentic projects to be canceled and why limiting what agents can do matters so much.
Who is responsible when AI gets it wrong?
“The AI did it” is not working as a defense. In the Air Canada case, the tribunal said the chatbot was simply part of the airline’s website and that the company was responsible for all the information on it. In Mata v. Avianca, the lawyers who filed invented cases were sanctioned, not the chatbot’s maker. The practical rule is that whoever deploys or relies on an AI system owns its output.
Regulation is moving the same way. In the EU, AI Act transparency rules have applied since 2 August 2026, requiring people to be told when they are dealing with an AI system and certain AI-generated content to be labeled, with stricter obligations for high-risk uses such as hiring still to come. For organizations, that means documenting where AI is used, who oversees it and how its errors are caught.
How to avoid an AI failure: a checklist
- Start with a measurable problem, not a technology. Define what success and failure look like before building.
- Check the data. Is it representative of the people and situations the system will face? Could it encode past bias?
- Test the edges. Try unusual inputs, adversarial users and changed conditions, not just the average case.
- Keep a person accountable for consequential outputs, from customer answers to hiring and credit decisions.
- Verify generative output against trusted sources before it is published, filed or sent to a customer.
- Limit access and exposure. Give AI agents only the permissions they need, and cap the money or data a model can put at risk.
- Monitor after launch for drift, complaints and unexpected behavior, and keep a way to switch the system off quickly.
- Harden payments and approvals against deepfakes with out-of-band verification. See our guide to AI in cybersecurity.
Learn to use AI without repeating these mistakes
Most of the failures above were avoidable with better judgment about what AI can and cannot do. Two short courses cover that judgment well:
- DeepLearning.AI – AI For Everyone (Coursera). Andrew Ng’s non-technical course on how AI projects work, how to choose them and where they go wrong. About seven hours; ideal for managers deciding what to build.
- Google – AI Essentials (Coursera). Five short courses on using AI tools at work, including a course on using AI responsibly and recognizing its limits.
Both are included in Coursera Plus. Coursera removed its free audit option for most courses in 2025, so check the price or trial terms before enrolling. For deeper technical training, see our ranking of the best AI courses.
Frequently asked questions
What are the most famous AI failures?
Well-documented examples include Amazon’s recruiting tool that penalized resumes mentioning women, Microsoft’s Tay chatbot, the 2018 Uber self-driving test vehicle crash in Tempe, IBM Watson for Oncology’s unsafe treatment suggestions, Zillow’s home-buying algorithm, Google Bard’s launch error, lawyers sanctioned for ChatGPT’s fake case citations, Air Canada’s chatbot ruling and the $25 million Arup deepfake fraud.
Why do AI projects fail?
The common causes are biased or unrepresentative training data, models used outside the conditions they were built for, generative AI stating false information confidently, too little human review, weak security against manipulation, and projects started for hype rather than a clear business need. Gartner describes most current agentic AI projects as early-stage experiments that are mostly driven by hype and often misapplied.
How many AI projects fail?
There is no single reliable failure rate, and widely repeated figures often trace back to old or vendor surveys. Recent indicators: Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, and McKinsey finds only 37% of organizations report any positive effect on EBIT from AI.
What is an AI hallucination?
A hallucination is when a generative AI system produces false or invented information and presents it as fact, such as the six non-existent court cases ChatGPT supplied in the Mata v. Avianca case. Hallucinations happen because language models generate plausible text rather than looking up verified facts.
Is a company responsible for what its AI chatbot says?
Courts and tribunals increasingly say yes. In Moffatt v. Air Canada (2024), a Canadian tribunal rejected the airline’s argument that its chatbot was a separate legal entity, found negligent misrepresentation and ordered Air Canada to pay damages.
What was the most expensive AI failure?
By market impact, Google’s Bard launch error: Alphabet lost $100 billion in market value on the day an inaccurate answer in its promotional video was reported, By direct losses on this list, Zillow’s pricing algorithm led to a write-down of about $304 million and the closure of its home-buying business, and the Arup deepfake fraud cost about $25.6 million.
Can AI failures be predicted?
Some can. Biased training data, untested edge cases and missing human review are visible before launch if someone looks for them. Others are harder, because AI performance is uneven: Stanford’s 2026 AI Index notes that models capable of Olympiad-level mathematics still misread analog clocks about half the time. Testing on your own real-world tasks is the most reliable way to find weaknesses early.
How can AI failures be prevented?
Test models on real, representative data before launch, keep a person accountable for consequential decisions, verify generative AI output against trusted sources, limit what AI agents can access, monitor results after launch, and have a plan to switch the system off quickly if it misbehaves.
The verdict
The history of AI failures is mostly a history of human decisions: what data was used, where a model was deployed, whether anyone checked its output and who was accountable when it went wrong. The technology keeps improving, but so does the number of incidents, because AI is now in far more places. The organizations that avoid headlines treat AI as powerful but fallible: they test it on reality, keep people responsible for outcomes and verify what it says.
Related guides: AI trends in 2026 · AI predictions · AI in education · AI in cybersecurity · Cyber security trends · Best machine learning courses
