The internet has always offered artists a slightly rotten bargain.
Put your work online and people may discover you. Keep it offline and protect it perfectly, but accept that nobody can hire, follow or share what they cannot see.
Generative AI has made that bargain worse. Now an artist can say “no” clearly, choose a platform designed around that refusal, and still watch their work disappear into a dataset by the million.
That is what happened to Cara, the portfolio and social platform built for artists who do not want their work used for AI training without permission. Beginning August 13, the site was targeted by three scraping incidents. Cara’s own account lists roughly 12 million images in the first scrape, about nine million image links and associated information in the second, and 123,000 images plus user data in the third.
The numbers are startling. The more important detail is that Cara had already said no.
Its terms prohibit unauthorized scraping, and the platform says it uses rate limits and anti-bot measures. It was created specifically because artists wanted somewhere to share work without quietly donating it to somebody else’s model.
Apparently, a clearly posted boundary on the internet is still treated less like a locked door and more like a decorative sign.
A clever tool built around a grim assumption
The strangest part of the story is what happened next.
The person responsible for the first large scrape, identified publicly as Heft, apologized and deleted the dataset. According to both Cara and WIRED, he then began helping founder Jingna Zhang identify weaknesses in proposed defences.
Now they are working on Lantern, an open-source tool designed to help artists find their work inside publicly available AI datasets.
Lantern creates what its developers describe as a one-way fingerprint of an image without retaining the original artwork. It can compare that fingerprint against newly published datasets and notify the artist when it finds a match. The alert includes a link to the dataset, giving the creator something concrete to investigate and potentially challenge through a removal request or takedown notice.
That is genuinely useful. Most creators cannot spend their evenings downloading enormous datasets and searching them one image at a time. Detection turns a vague fear—“my work is probably in there somewhere”—into evidence with a location attached.
It is also a little horrifying.
Lantern begins from the assumption that preventing every scrape may be impossible. Instead, it helps artists discover the intrusion after their work has already left the platform. It is less like a lock and more like a smoke alarm: valuable, possibly essential, and sounding only once something has started burning.
[AD PLACEMENT]
The collaboration itself is unusual enough to make a good internet redemption story. Someone causes harm, listens to the people affected, changes course and helps build a remedy. We could use more endings like that.
But the next scraper may not apologize. The next dataset may not be public. The next company may already have copied, filtered and mixed the material into a training pipeline before an artist receives an alert.
One person’s change of heart cannot be the enforcement layer for an entire creative economy.
The burden keeps landing on the artist
The modern creator-protection toolkit is becoming impressively technical.
Artists can add machine-readable opt-out signals, alter images with protective tools, limit resolution, block known bots, monitor datasets and file takedown requests. Researchers continue probing those defences, sometimes to strengthen them and sometimes to demonstrate how easily they can be bypassed. A University of Texas at San Antonio research summary, for example, described a method capable of weakening protections used by tools such as Glaze and Nightshade, underscoring the arms-race nature of technical safeguards.
Every new defence may help. Together, they reveal something backwards about the system.
The artist makes the work, publishes it, declares the rules, installs the protection, monitors for violations, proves ownership and then pursues whoever ignored the boundary. The party that wants millions of training images can simply automate the collection.
We have turned consent into a setting and enforcement into unpaid labour.
Cara’s experience also punctures the lazy advice that artists who dislike scraping should simply stop posting online. An online portfolio is not a frivolous extra for a working illustrator, photographer or designer. It is a storefront, résumé and community space. Telling creators to disappear if they want control is like telling a shopkeeper that the only reliable theft prevention is never opening the store.
[AD PLACEMENT]
The better answer will require more than increasingly elaborate tricks applied to individual images.
Dataset publishers could be required to document where material came from and how consent was obtained. Hosting services could respond consistently to datasets assembled in violation of site rules. AI developers could keep auditable records and honour standardized opt-outs. Lawmakers could decide that a creator’s refusal has meaning before the work is copied, rather than only after an expensive legal argument.
Those are policy choices, not features Lantern can quietly ship in an update.
For now, the project may give artists something they have rarely had in the AI training debate: visibility into where their work went. That matters. You cannot challenge a dataset you cannot find.
But we should be careful not to mistake detection for consent or notification for control.
A healthy creative internet will probably need good smoke alarms. It should also have fire codes.
