Support & Education Written for parents No shame in it

The kids are
already in the dataset.
Here is what that means.

Most of what is written about children and the internet is written to frighten you into buying something. This is not that. It is the measured record of what happened to family photographs when machine learning arrived, what teenagers are actually facing on their phones this year, and the handful of things that genuinely help.

Nothing here is an argument for shame. Parents posting their children are being proud in public, and teenagers seeking attention online are doing what teenagers have always done. The villain is the pipeline: a set of defaults nobody chose, that turns ordinary family life into permanent, machine-readable material.

Skip to what helps ↓
Every number on this page carries its source and its caveat.
Including the famous ones we think are overstated. Where a figure is a projection, a commissioned survey, or a count of reports rather than incidents, we say so on the spot instead of at the bottom.

Start with what
can be counted.

↳ the volume, and the money

A first day of school. A hospital bracelet. A birthday with the school name in the background. Individually harmless, collectively a training corpus, and separately, a starter kit for identity fraud against a person who will not check their credit for another fifteen years.

973
photographs of the average child posted online before their fifth birthday.
An often-cited 2015 estimate from a commissioned survey of 2,000 UK parents (Nominet / Parent Zone). Treat it as an order of magnitude, not a measurement.
915,000
US children Javelin estimated were victimized by identity fraud in the year to July 2022, at an average household cost of $1,128.
Javelin Strategy & Research, 2022 wave. This is a commercial study; we could not inspect the underlying report or its methodology, so read it as an estimate rather than a measured count. In Javelin’s 2021 study of 5,000 households, 73% of child victims personally knew the perpetrator.
9,605
identity-theft reports filed with the FTC in 2024 by consumers aged 19 and under.
FTC Consumer Sentinel, 2024. This is a report count, not an estimate of incidence: identity theft involving children can remain undiscovered for years, often until the child applies for credit.
the projection everyone quotes, quoted honestly:

In May 2018, Barclays forecast that sharenting would account for two-thirds of identity fraud facing young people by 2030, at £667 million a year. Barclays did not publish the underlying methodology, and the figure is a projection rather than measured data.

We include it because you will see it everywhere, usually stripped of that caveat. The direction it points is consistent with the measured figures above. The precision it implies is not earned.

Source: Barclays research reported by the BBC, 21 May 2018 · bbc.co.uk

Then somebody
actually looked.

↳ the one dataset open enough to audit

LAION-5B is a dataset of 5.85 billion image-URL and caption pairs, released in March 2022 by LAION e.V., a German non-profit. LAION states that its datasets “only contain links and metadata” and that it “has never distributed image content itself.” It is also the reason we know any of this: LAION was open enough to audit. The privately-held training sets have never been examined by anyone outside the companies that own them.

10 June 2024 · Human Rights Watch

Brazil

Researcher Hye Jung Han found that the dataset contained links to 170 photographs of children from at least ten Brazilian states, sourced from personal blogs, photo and video sites, school presentations, and hospital birth records. Many carried identifying detail: children’s names, the names of hospitals, timestamps.

⚠ not stolen, scraped from where it was posted
July 2024 · Human Rights Watch

Australia

A second report found links to 190 photographs of children from every Australian state and territory, including stills taken from YouTube videos that had privacy settings applied. Some entries carried full names, ages, and the names of the preschools the children attended.

The report also raised something a privacy setting cannot address: photographs of First Nations children, and cultural protocols around images of people who have died, protocols that permanent inclusion in a training set makes impossible to honour.

⚠ privacy settings did not prevent inclusion
two numbers that matter more than the totals:

By September 2024, Human Rights Watch had identified 362 Australian and 358 Brazilian children across 360 photographs in the dataset. And they reviewed less than 0.0001% of it, which HRW notes makes their findings likely to be a significant undercount. The small number is the alarming one.

On what the material enables, HRW is direct: “malicious actors have used LAION-trained AI tools to generate explicit imagery of children using innocuous photos.” That is a statement about the class of tools, not about these particular children.

Separately, in December 2023, the Stanford Internet Observatory reported 3,226 dataset entries of suspected child sexual abuse material, 1,008 of which were externally validated by the Canadian Centre for Child Protection. LAION withdrew the dataset within days and published a cleaned version, Re-LAION-5B, in August 2024, removing 2,236 links.

Sources: Human Rights Watch, 3 September 2024 · Stanford Internet Observatory (David Thiel), 20 December 2023
“AI models that were trained on the earlier dataset cannot forget the now removed images.”
Human Rights Watch, on the cleaned-up dataset

The links were deleted. Removing a URL does not alter model copies already trained on it. Owners can retrain from scratch, or attempt machine unlearning, but NIST distinguishes the two and warns that practical approximate methods can leave recoverable information. Stable Diffusion’s own model card confirms it was trained on LAION-5B subsets, and there is no practical mechanism to recall every downloaded copy of those weights. That is the problem, stated precisely.

The account, and
the child in the photo.

↳ the reassurance is true, and the gap in it is easy to miss

Meta is worth its own passage, because three different things get blurred together in most coverage. Here they are kept apart.

Category oneTraining data

The Llama 4 model card states the model was pretrained on “a mix of publicly available, licensed data and information from Meta’s products and services. This includes publicly shared posts from Instagram and Facebook and people’s interactions with Meta AI.” Meta’s newsroom has said the same since September 2023.

The Llama 3.x model cards say only “publicly available online data,” so this is a statement about Llama 4, not about every model Meta has released.

Category twoAds and personalization

On 1 October 2025, Meta announced that from 16 December 2025 a person’s interactions with its AI would be used to personalize content and advertising, in what it described as most regions.

Reporting by TechCrunch and Reuters characterized the change as having no opt-out and noted carve-outs for the UK, the EU and South Korea. That characterization is theirs, not Meta’s.

Category threeWhere regulators landed

After noyb filed complaints in eleven countries on 6 June 2024, Meta paused its EU training plans on 14 June 2024, then resumed EU training on 27 May 2025 on a legitimate-interests basis with an objection form.

In Brazil, the ANPD ordered a suspension on 2 July 2024 on penalty of R$50,000 per day, and lifted the suspension on 30 August 2024 on a compliance plan. Meta committed not to use under-18 account data in Brazil “until a definitive decision is taken by ANPD.”

↳ the distinction the whole thing turns on:

Meta states it does not use data from accounts of under-18s to train its AI models. But that exclusion attaches to the account that posted the content, not to who appears in it. Asked about this at an Australian Senate committee hearing in September 2024, Meta’s Global Privacy Policy Director confirmed that Meta does not scrape accounts of under-18s but does use “public photos posted by people over 18”, which, on Meta’s own account, includes photos of children posted publicly by parents, relatives, schools or businesses.

At that same hearing of the Senate Select Committee on Adopting AI, Senator David Shoebridge put it to Meta that unless a user had consciously set their posts to private, the company had “decided that you will scrape all of the photos and all of the texts from every public post on Instagram or Facebook since 2007.” Meta’s Global Privacy Policy Director, Melinda Claybaugh, answered: “Correct.”

What it looks like from
a teenager’s phone.

↳ this one is not slow, and not abstract

The dataset problem is slow and abstract. This one is neither. It arrives as a direct message, it moves in hours, and it is aimed at children who have done nothing wrong except exist online in the ordinary way.

Financial sextortion

The script, and the scale

The FBI reported in January 2024 that between October 2021 and March 2023 it received more than 13,000 reports of financial sextortion of minors, involving at least 12,600 victims, mostly boys, and at least 20 deaths by suicide, between October 2021 and March 2023.

Reports to NCMEC have climbed steeply: 10,731 in 2022, 26,718 in 2023, more than 36,000 in 2024, and more than 50,000 in 2025, roughly 137 a day. NCMEC’s public totals are rounded lower bounds, so the year-on-year change is best described as roughly two-fifths rather than an exact rate. NCMEC has said that at least 36 teenage boys have died by suicide cumulatively since 2021.

2022: 10,731 reports 10,731 2022 2023: 26,718 reports 26,718 2023 2024: more than 36,000 reports 36,000+ 2024 2025: more than 50,000 reports 50,000+ 2025
Reports received by NCMEC, not unique victims or confirmed incidents. A 2024 law newly required platforms to report this category, which inflates the trend line.
⚠ read these as reports, not incidents
Caveat that belongs with the numbers: the 2024 REPORT Act newly required platforms to report this material, and NCMEC attributes much of the increase to that change. NCMEC counts reports, not victims or confirmed incidents. Sources: FBI, 16 January 2024; NCMEC.
Nudify apps

Built from a school photo

The Internet Watch Foundation reported in March 2026, on 2025 data, that 491 of its reports contained AI-generated child sexual abuse material, a 154% increase. Those reports contained 8,029 criminal images and videos. Of the AI videos specifically, it found 3,443, against 13 the year before. Of the AI videos, 65% were Category A, its most serious classification. Separately, 97% of the AI-generated images depicted girls. Those are two different denominators and we are keeping them apart.

In a Thorn survey published in March 2025 of 1,200 young people aged 13 to 20, one in eight said they personally knew someone who had deepfake nudes made of them while under 18, and one in seventeen said it had happened to them.

⚠ four separate measures, four separate denominators
ISD reported in July 2026 that mainstream platforms sent more than 5.7 million referral visits to the top ten nudify sites between December 2025 and March 2026. The Center for Democracy & Technology found 12% of students (n=1,030, 2025) had heard of deepfake intimate imagery depicting someone at their school.
and then, years later, somebody searches:

Kaplan found in 2023 that 28% of admissions officers surveyed (n=205) had checked an applicant’s social media, and 67% considered doing so “fair game.” An Express Employment / Harris Poll survey of 1,002 hiring decision-makers in December 2022 found 70% of companies research candidates on social media.

An honest note about this entire category: almost every published figure in this area comes from an opt-in industry survey commissioned by a company with a product to sell. The precise percentage is genuinely uncertain. The direction is not.

Trained models do not
have a delete key.

↳ why “just take it down” is not the answer it sounds like

Every instinct says there must be a way to take it back: a deletion request, a right to be forgotten, a button. For data sitting in a database, there often is. For knowledge absorbed into a trained model, the research says otherwise.

NIST’s AI 100-2e2025, published in March 2025, separates exact unlearning (retraining the model without the data) from the approximate methods used in practice, which it notes “remain vulnerable to adversarial attacks, including inversion attacks.” Peer-reviewed work presented at ICLR 2025 found models retained 21% of the knowledge they were meant to forget at full precision, and 83% of it after 4-bit quantization; separate ICLR 2025 work showed targeted relearning can restore removed knowledge. Carlini and colleagues extracted 109 near-copies of training images from diffusion models (USENIX Security 2023), and researchers have extracted more than 10,000 verbatim training examples from ChatGPT for roughly $200 in compute.

The European Data Protection Board reached the same place from the legal side. Its Opinion 28/2024 observes that personal data “may still remain ‘absorbed’ in the parameters of the model… which may ultimately be extractable,” and concludes that “AI models trained on personal data cannot, in all cases, be considered anonymous.”

And where weights are published openly, there is no practical mechanism to retrieve copies already in third-party hands. Removing the public copy of a photograph does not guarantee removing its influence from every model already trained on it, and that is most true where the weights have been distributed rather than kept behind an API.

the honest shape of this page:

There is usually no simple, universal delete operation for a model that has already been trained or distributed. Retraining and model replacement are possible, approximate unlearning is imperfect, and which options exist depends on the model and on who controls its copies. No checklist changes that. What the rest of this page is for is the part that is still open: the identity-fraud surface is genuinely reducible, the sextortion script is genuinely defeatable by a conversation held in advance, and the takedown right is real and enforceable. Those three are worth your Saturday. Grieving the first part is not.

Start with the addresses
in your own house.

↳ yours first, then the others, with permission

An adult’s compromised account can expose the things a family audit is actually about: the recovery channels for everyone else’s logins, the school portal, medical records and the photo library. A reused password on a parent’s email reaches all of them. So this is a sensible place for a family audit to begin, and consent is the other half of doing it properly.

Put an address into the panel. We check it against Have I Been Pwned, the public breach index maintained by security researcher Troy Hunt, and show you what turned up.

the order matters: check your own address first, so you know what the result feels like before you sit down with anyone else.

Then, with their permission, help each person in your house check theirs.

Not behind their back. A teenager who finds out you ran their address without asking learns that you are someone to hide things from, which is the exact opposite of what this page is trying to build. Ask, do it together, and let them type it in.

Breach lookup

Check an address

Start with your own. Old addresses count too: the ones from 2009 are usually the worst.

Privacy: the address is checked against Have I Been Pwned through our own server. We do not save it in application storage or analytics, and we do not add it to any list. Being straight with you about the limit: it does travel through our server to HIBP, and transient operational logs may exist somewhere in that delivery chain, which is not something we can promise away. If you would rather avoid that entirely, look the address up on their own site instead. Full detail in our privacy policy.

Breach data powered by Have I Been Pwned.

The household list

One line per person, one address at a time. Tick them off as you go. The point is not the results: it is that by the end of it, everyone in the house has had the conversation about passwords once, on a calm afternoon, instead of in an emergency.

  • Your own main address, and every old one you can still remember.
  • Your partner’s, with them at the table.
  • Each child’s, with their permission, and with them typing.
  • The shared family address, if you have one: school forms, deliveries, the pediatrician.
  • Any grandparent whose email you already help manage.

Six things worth
doing this weekend.

↳ the reducible part

Six things worth doing this weekend

None of this undoes what is already absorbed. All of it reduces what gets added next, and the identity-fraud surface is genuinely reducible.

  1. Audit the old public albums. Birthdays, first days of school, hospital photos, sports rosters. You are looking for the combination that powers identity fraud: a full name, a birthdate, a school or hospital, a face.
  2. Reset the defaults, without over-trusting them. Move family posts to private audiences and lock down the accounts. Worth knowing: Human Rights Watch found stills from YouTube videos that had privacy settings applied, so treat this as reducing exposure rather than eliminating it.
  3. Strip metadata before posting. Photos carry capture times and often precise locations. Most phones can turn location off for the camera; most platforms strip some of it, and you should not rely on which.
  4. Keep the identifiers out of the caption. The school name, the teacher’s name, the street, the birthdate, the full legal name. The photo is rarely the risk on its own; the caption is what makes it usable.
  5. Talk to your kids about the sextortion script, before it arrives. A stranger becomes friendly fast, moves the conversation to images, then turns and demands money within hours. The one thing that has to land: they will not be in trouble, the demand only escalates if paid, and they should tell an adult and keep the messages.
  6. Know that the takedown right is real now. The TAKE IT DOWN Act (S.146, P.L. 119-12) was signed on 19 May 2025 and criminalizes publishing non-consensual intimate imagery, including AI “digital forgeries.” Covered platforms must remove reported material, and known identical copies, within 48 hours of a valid request. The FTC enforces it, and the platform compliance deadline of 19 May 2026 has passed: this is live.

If it has already
happened.

↳ the numbers to have before you need them

Two things are true at once. The first is that this is more common than the silence around it suggests, and the family it happened to almost certainly did nothing careless. The second is that there are specific people whose actual job this is, and reaching them early changes how it goes.

the last thing, and the only one we would argue about:

The pipeline described on this page was not built by parents and cannot be closed by them. Nothing on a checklist rebalances a system where the default is collection and the burden of objecting falls on a nine-year-old’s guardian. What is in your hands is the next photograph, the next caption, and the conversation you have before the message arrives. That is a smaller thing than the problem deserves, and it is still worth doing on Saturday.

Support & Education · en español

Estafas en Puerto Rico.

The family-facing companion, written in Spanish for you and for your parents: the schemes running on this island right now, the five things that will never happen, and a family code word worth agreeing on tonight.

↳ the one to send to the grandparents

Léelo →
Support & Education · companion page

Who’s watching.

Where a family’s data goes after it leaves the house: brokers, the location trail, and the interpersonal section on stalkerware and shared-account monitoring, which is the part that comes up most often in a household.

↳ jump straight to the interpersonal section

Read it →
Support & Education · printable

The family exposure audit.

Page 9 of the workbook is this page on paper: a room-by-room, account-by-account audit you can fill in with a pen at the kitchen table, with no login and no download form.

↳ fourteen Letter-size pages, print it double-sided

Open it →
↳ there is nothing to buy on this page

This one is not
a pitch.

We build software for a living, and none of it would have helped the families in the reports above. This page exists because the material is hard to find stated honestly, with its caveats attached, in one place. Send it to whoever needs it. If you want the rest of the shelf, it is all open and none of it requires an email address.