Bankrupt by AI
All case files
On the recordFined ₩103M · Apr 2021

Bad automation

ScatterLab and the chatbot built on other people’s messages

ScatterLab · Consumer AI · South Korea

Korea’s “Lee Luda” chatbot was trained on 9.4 billion private KakaoTalk messages its users never agreed to hand over. The regulator’s fine was the first time the country’s privacy law reached an AI.

₩103M

PIPC fine · first AI case

9.4B

KakaoTalk messages used

600k

people, without consent

Training-data provenance treated as a deployment liability.

The failure was upstream of the model — in the data it was built on.

The record

  • In April 2021, South Korea’s Personal Information Protection Commission fined ScatterLab ₩103.3 million (~$93,000) for eight violations of the privacy act — described as the first case applying the law to an AI system. (PIPC / korea.kr; FPF; The Register)
  • ScatterLab had trained its “Lee Luda” chatbot on roughly 9.4 billion KakaoTalk messages from about 600,000 users of its other apps, without proper consent. (PIPC announcement, 2021-04-28)
  • The chatbot had already been taken offline in January 2021 after it produced hate speech and appeared to leak personal information. (FPF; The Register)

A likeable bot, an unlikeable dataset

Lee Luda was a hit — a chatty, personable AI persona that drew a large young audience within weeks. The problem was never the personality. It was the provenance of the data underneath: roughly 9.4 billion private KakaoTalk messages, collected from some 600,000 people through ScatterLab’s other apps, repurposed to train the bot without the consent that use required.

The bot also surfaced hate speech and what looked like real personal details, and was pulled in January 2021. But the regulator’s action reached past the visible failures to the invisible one.

The first time the law reached an AI

The PIPC’s fine — ₩103.3 million across eight violations — was described as the first application of Korea’s privacy law to an AI system. Small in money, large in principle: it established that where an AI’s training data comes from is a live legal liability, not a backstage engineering detail.

Consent gathered for one purpose does not silently extend to training a consumer product on the same messages. That is the kind of upstream decision that is cheap to get right at the start and, as ScatterLab found, impossible to retrofit once the model ships.

The lesson

An AI’s training-data provenance is a deployment liability, not a backstage detail — often the first thing a regulator can reach, and the hardest thing to retrofit consent onto.

How we’re reading this

The ₩103.3M fine is small in absolute terms; its weight is as the first application of Korea’s privacy law to an AI system. We cite the regulator’s findings for the data-provenance violations; the hate-speech and leak incidents were widely reported at the bot’s January 2021 shutdown.

Sources

  1. 01
  2. 02
The pattern, anonymizedGood automation, bad automation

Compiled from public filings, court records, company statements and reputable press. Figures are attributed to their sources; allegations are labeled as such. Not legal or investment advice.