Mutiny Labs

Your family's digital life · Part 2 of 3 · 7 min read

What happens to the photos you share?

What the research establishes about children's photographs, AI training, and the limits of deletion.

In this series · 3 parts
  1. 1. A practical privacy reset for your family.
  2. 2. What happens to the photos you share?
  3. 3. The conversation to have before a threatening message.

Then somebody actually looked.

↳ the one dataset open enough to audit

LAION-5B is a dataset of 5.85 billion image-URL and caption pairs, released in March 2022 by LAION e.V., a German non-profit. LAION states that its datasets “only contain links and metadata” and that it “has never distributed image content itself.” It is also the reason we know any of this: LAION was open enough to audit. The privately-held training sets have never been examined by anyone outside the companies that own them.

10 June 2024 · Human Rights Watch

Brazil

Researcher Hye Jung Han found that the dataset contained links to 170 photographs of children from at least ten Brazilian states, sourced from personal blogs, photo and video sites, school presentations, and hospital birth records. Many carried identifying detail: children’s names, the names of hospitals, timestamps.

⚠ not stolen, scraped from where it was posted
July 2024 · Human Rights Watch

Australia

A second report found links to 190 photographs of children from every Australian state and territory, including stills taken from YouTube videos that had privacy settings applied. Some entries carried full names, ages, and the names of the preschools the children attended.

The report also raised something a privacy setting cannot address: photographs of First Nations children, and cultural protocols around images of people who have died, protocols that permanent inclusion in a training set makes impossible to honour.

⚠ privacy settings did not prevent inclusion
two numbers that matter more than the totals:

By September 2024, Human Rights Watch had identified 362 Australian and 358 Brazilian children across 360 photographs in the dataset. And they reviewed less than 0.0001% of it, which HRW notes makes their findings likely to be a significant undercount. The small number is the alarming one.

On what the material enables, HRW is direct: “malicious actors have used LAION-trained AI tools to generate explicit imagery of children using innocuous photos.” That is a statement about the class of tools, not about these particular children.

Separately, in December 2023, the Stanford Internet Observatory reported 3,226 dataset entries of suspected child sexual abuse material, 1,008 of which were externally validated by the Canadian Centre for Child Protection. LAION withdrew the dataset within days and published a cleaned version, Re-LAION-5B, in August 2024, removing 2,236 links.

Sources: Human Rights Watch, 3 September 2024 · Stanford Internet Observatory (David Thiel), 20 December 2023
“AI models that were trained on the earlier dataset cannot forget the now removed images.”
Human Rights Watch, on the cleaned-up dataset

The links were deleted. Removing a URL does not alter model copies already trained on it. Owners can retrain from scratch, or attempt machine unlearning, but NIST distinguishes the two and warns that practical approximate methods can leave recoverable information. Stable Diffusion’s own model card confirms it was trained on LAION-5B subsets, and there is no practical mechanism to recall every downloaded copy of those weights. That is the problem, stated precisely.

The account, and the child in the photo.

↳ the reassurance is true, and the gap in it is easy to miss

Meta is worth its own passage, because three different things get blurred together in most coverage. Here they are kept apart.

Category oneTraining data

The Llama 4 model card states the model was pretrained on “a mix of publicly available, licensed data and information from Meta’s products and services. This includes publicly shared posts from Instagram and Facebook and people’s interactions with Meta AI.” Meta’s newsroom has said the same since September 2023.

The Llama 3.x model cards say only “publicly available online data,” so this is a statement about Llama 4, not about every model Meta has released.

Category twoAds and personalization

On 1 October 2025, Meta announced that from 16 December 2025 a person’s interactions with its AI would be used to personalize content and advertising, in what it described as most regions.

Reporting by TechCrunch and Reuters characterized the change as having no opt-out and noted carve-outs for the UK, the EU and South Korea. That characterization is theirs, not Meta’s.

Category threeWhere regulators landed

After noyb filed complaints in eleven countries on 6 June 2024, Meta paused its EU training plans on 14 June 2024, then resumed EU training on 27 May 2025 on a legitimate-interests basis with an objection form.

In Brazil, the ANPD ordered a suspension on 2 July 2024 on penalty of R$50,000 per day, and lifted the suspension on 30 August 2024 on a compliance plan. Meta committed not to use under-18 account data in Brazil “until a definitive decision is taken by ANPD.”

↳ the distinction the whole thing turns on:

Meta states it does not use data from accounts of under-18s to train its AI models. But that exclusion attaches to the account that posted the content, not to who appears in it. Asked about this at an Australian Senate committee hearing in September 2024, Meta’s Global Privacy Policy Director confirmed that Meta does not scrape accounts of under-18s but does use “public photos posted by people over 18”, which, on Meta’s own account, includes photos of children posted publicly by parents, relatives, schools or businesses.

At that same hearing of the Senate Select Committee on Adopting AI, Senator David Shoebridge put it to Meta that unless a user had consciously set their posts to private, the company had “decided that you will scrape all of the photos and all of the texts from every public post on Instagram or Facebook since 2007.” Meta’s Global Privacy Policy Director, Melinda Claybaugh, answered: “Correct.”

Trained models do not have a delete key.

↳ why “just take it down” is not the answer it sounds like

Every instinct says there must be a way to take it back: a deletion request, a right to be forgotten, a button. For data sitting in a database, there often is. For knowledge absorbed into a trained model, the research says otherwise.

NIST’s AI 100-2e2025, published in March 2025, separates exact unlearning (retraining the model without the data) from the approximate methods used in practice, which it notes “remain vulnerable to adversarial attacks, including inversion attacks.” Peer-reviewed work presented at ICLR 2025 found models retained 21% of the knowledge they were meant to forget at full precision, and 83% of it after 4-bit quantization; separate ICLR 2025 work showed targeted relearning can restore removed knowledge. Carlini and colleagues extracted 109 near-copies of training images from diffusion models (USENIX Security 2023), and researchers have extracted more than 10,000 verbatim training examples from ChatGPT for roughly $200 in compute.

The European Data Protection Board reached the same place from the legal side. Its Opinion 28/2024 observes that personal data “may still remain ‘absorbed’ in the parameters of the model… which may ultimately be extractable,” and concludes that “AI models trained on personal data cannot, in all cases, be considered anonymous.”

And where weights are published openly, there is no practical mechanism to retrieve copies already in third-party hands. Removing the public copy of a photograph does not guarantee removing its influence from every model already trained on it, and that is most true where the weights have been distributed rather than kept behind an API.

the honest shape of this page:

There is usually no simple, universal delete operation for a model that has already been trained or distributed. Retraining and model replacement are possible, approximate unlearning is imperfect, and which options exist depends on the model and on who controls its copies. No checklist changes that. What the rest of this page is for is the part that is still open: the identity-fraud surface is genuinely reducible, the sextortion script is genuinely defeatable by a conversation held in advance, and the takedown right is real and enforceable. Those three are worth your Saturday. Grieving the first part is not.

Edited September 29, 2026. Research and source dates are retained; this edit is not a fresh review of every statistic or legal development.