Your Family Photos Trained AI Without Your Permission

Billions of photos posted to social media became AI training data without consent. What happened — and what it means for your family's images.

KeepSaiQ Editorial8 min read

The photographer didn't know they were contributing to a dataset. They were posting a birthday photo of their daughter — a child blowing out six candles, a living room half-decorated behind her. The image was public, because that was the default setting on the platform, and because they weren't thinking about defaults. They were thinking about the cake.

Somewhere between 2021 and 2023, that photograph, along with billions like it, was scraped from the internet and incorporated into a training dataset for generative AI. The family was not contacted. No one asked for permission. Publicly posted images were, legally and technically, available. A photograph of a six-year-old at her birthday became one data point in a model that would learn to generate photorealistic images of people and scenes that never existed.

This is not hypothetical. It happened. The question now is what it means.

How the datasets were built

The LAION-5B dataset — assembled by the Large-scale Artificial Intelligence Open Network and used as the primary training foundation for Stable Diffusion and several other widely deployed image generation systems — contains approximately five billion image-text pairs collected from publicly accessible websites. The images came from wherever publicly shared images live: Flickr, Instagram, blogs, news archives, forums, and social platforms of every kind.

LAION's researchers were not bad actors in the conventional sense. They were building an open-access training resource for the scientific community, working in a well-established tradition of large-scale web crawls for machine learning research. They published their methodology and made the dataset freely available — which is precisely how it came to train commercial products used by hundreds of millions of people.

The dataset is not an anomaly. It represents how most large-scale AI training data was assembled during the decade from roughly 2012 to 2022: web-scale scraping, automated filtering, and wide distribution to whoever wanted it. The images of ordinary family life that populated the public internet during those years — children's birthdays, summer vacations, holiday dinners, school recitals — were, by virtue of being publicly accessible, candidates for inclusion. Not because anyone targeted family photos specifically, but because family photos were everywhere online, and everywhere online was being scraped.

What the terms of service actually said — and didn't say

Platforms are quick to note that users agreed to their terms. Those terms, they observe, grant the platform broad rights to use uploaded content. Technically, this is true.

What the terms did not say — and what no ordinary user could have inferred — was that "use your content to improve our services" might encompass making those images available to third parties building commercial AI products, or permitting the kinds of web scraping that assembled training datasets. The gap between what users understood they were consenting to and what was technically permitted has become the subject of multiple ongoing lawsuits in the United States and Europe.

The Electronic Frontier Foundation has extensively documented this problem: consent obtained for one purpose does not constitute meaningful informed consent for a categorically different purpose discovered years later. The principle is basic and well-established in privacy law. Checking a box agreeing to terms that were tens of thousands of words long and made no mention of AI training is not meaningful consent to AI training. Courts are beginning to agree, though slowly.

The legal framework is catching up, but the datasets are already built. Retroactive consent or meaningful opt-out is, for most practical purposes, no longer possible for images scraped before 2023.

The Federal Trade Commission has been examining platform data practices for years, including the breadth of data licenses users unwittingly grant and whether platforms' use of that data constitutes unfair or deceptive practices. The EU AI Act, which entered into force in 2024, includes provisions specifically requiring AI providers to document training data sources and consent frameworks — an acknowledgment, written into law, that the consent practices of the preceding decade were inadequate.

Children are the most exposed

Consider the specific situation of a child photographed throughout their early years by parents who shared those photos online. The child was two when their parents created a Facebook account to share images with distant grandparents. They were seven when the family switched to Instagram. They are now sixteen. During those fourteen years, hundreds or thousands of photographs accumulated across multiple platforms. Many were posted publicly or semi-publicly — the default setting the parent never thought to change.

That child, at no point, could have meaningfully understood or consented to any of this. Not at two, not at seven, not at the ages when scraping was actively happening. And yet their visual likeness — what they looked like at various stages of childhood — may now exist in one or more training datasets, permanently.

The implications extend beyond abstract privacy violations. AI image generation systems trained on photographs of real people can, when given a small set of reference images of a specific person, produce synthetic photographs of that person in contexts they never occupied. The same capability that allows AI to generate convincing images of public figures exists, in principle, for children whose likenesses are in training data. Research on non-consensual synthetic imagery has documented serious real-world harms in cases involving adults; for children who can neither consent nor protect themselves, the exposure surface is deeper.

Regulatory frameworks specifically protecting children's data are emerging across jurisdictions, but they apply prospectively. They shape what will happen with images shared going forward. Children photographed in the years when scraping was unregulated have no practical remedy.

Why deleting photos from social media doesn't fix this

After news coverage of AI training data practices became widespread, many families deleted images from social media. This response is understandable. It is also largely ineffective for images already scraped.

Training datasets are not server-side libraries maintained in sync with source platforms. They are static collections, assembled at a point in time, distributed across many servers, often forked and re-shared repeatedly within the research community. A dataset assembled in 2021 that included your family's photos continues to exist — in full, on whatever infrastructure it was distributed to — regardless of whether you have since deleted those photos from the original platform.

Some AI companies have created opt-out mechanisms, and some dataset maintainers have established removal request processes. These address the margins of the problem rather than its substance. The core training datasets behind deployed AI systems are assembled and trained. The models have already learned from the data. Removing images from a dataset after training does not unlearn what the model incorporated.

This is the honest difficulty: for images already publicly shared and already scraped, the exposure has occurred. The meaningful decisions available to families are about future images — where they live, what infrastructure holds them, and whether that infrastructure is designed to make content accessible to the world or to hold it safely within the family.

The structural difference that actually protects families

The reason family photos ended up in AI training datasets is not primarily a story about bad actors or even bad intentions. It is a story about structure.

Social media platforms are designed to make content maximally discoverable and accessible. Discovery and accessibility are the value they sell to advertisers and, through network effects, to users. An image posted publicly on a platform built for public sharing is, by design, available to anyone capable of reading the web — including automated web crawlers assembling training datasets. The scraping that built LAION-5B targeted publicly accessible images specifically because they were publicly accessible.

A private family archive — closed to external indexing, held on infrastructure with no commercial incentive to share data or make it accessible — operates according to a fundamentally different logic. The goal is not maximizing accessibility; it is the opposite. Content that is never public is not subject to the same exposure mechanism.

This distinction is not primarily technical. It is architectural. A platform built with privacy as its foundational purpose, rather than as an optional setting requiring active configuration, does not expose families to the same risks because it is not trying to accomplish the same thing. One system is built to share; the other is built to hold.

Families choosing where their photographs live are, in effect, choosing between these architectures. For images already scraped, that choice comes too late. For the images being taken today — the birthdays and graduations and ordinary Tuesday dinners that will be the family archive of thirty years from now — the choice is still open.

Where photographs live determines what is possible with them. That decision, made once at the level of platform and practice, shapes everything that follows.

Sources & further reading

  1. LAION-5B: A new era of open large-scale multi-modal datasets (LAION)
  2. Electronic Frontier Foundation — Artificial Intelligence
  3. FTC — Protecting Consumer Privacy
  4. EU AI Act — European Commission regulatory framework for AI

Frequently asked questions

Were my family's photos really used to train AI?

Quite possibly. The LAION-5B dataset — the training foundation for Stable Diffusion and several other major image generators — was assembled by scraping billions of publicly accessible images from across the internet, including social media platforms. If you posted photos publicly (or with default semi-public settings) between roughly 2005 and 2022, those images may be in one or more AI training datasets. Some tools allow you to search for your images in known datasets, though coverage is incomplete.

What did I agree to when I joined social media?

Social media terms of service typically grant platforms broad, royalty-free licenses to use uploaded content to improve their services. What those terms did not say — and what most users never imagined — was that 'improving our services' could include licensing data access to AI labs or permitting scraping that allows others to build commercial AI products. The Federal Trade Commission and European regulators are examining the gap between what users understood they were consenting to and what was technically permitted in the fine print.

Why are children's photos a specific concern?

Minors cannot meaningfully consent to anything, including uses of their likeness that didn't exist when they were photographed. A child photographed at age four in 2016 had no capacity to agree that their image could become part of an AI training dataset. Beyond the legal question of children's data rights, there is a practical harm: AI systems trained on children's photos can produce synthetic images of those specific children in contexts that never occurred, a capability that raises serious child safety concerns as these tools become more accessible.

Can I get my family's photos removed from training datasets?

In most cases, removal is extremely difficult or practically impossible. Some AI companies offer opt-out mechanisms, and some dataset maintainers have created removal request processes. But many training datasets have been distributed to hundreds of researchers and companies; removing data from every downstream copy is functionally intractable. The EU AI Act's provisions on training data transparency apply prospectively — they shape future datasets, not the ones already assembled and already used to train deployed AI systems.

Does it matter whether I set photos to public or private?

Yes, substantially. Images shared publicly on social media platforms — and many platforms default to public — are accessible to web crawlers and can be scraped without further permission. Images stored in a genuinely closed system, inaccessible to external indexing, are not exposed to the scraping operations that built major training datasets. The distinction between public-facing and structurally private storage is one of the most consequential choices families can make about their photo archives going forward.