Since early 2023, photographer Jinna Chan and a small team of volunteers have developed Cara, an image-sharing social media platform and online portfolio app for artists. The platform has attracted approximately 1.5 million users, many of whom share a common concern: preventing their artwork from being used without permission to train artificial intelligence models. Cara also offers artists a way to promote their work while reducing their dependence on Big Tech platforms and their data-collection practices.
Cara provides artist-focused protections, including Glaze, a tool that alters image data to make artwork more difficult for AI systems and image scrapers to imitate. However, completely preventing web scraping remains extremely difficult. In August, beginning on the 13th, Cara experienced three major scraping incidents that sent server costs soaring and alarmed creators who had moved from platforms such as Instagram, where content may be made available to Meta for AI training.
The first incident came to light after those responsible published a 12-terabyte archive containing approximately 12 million Cara artworks—the platform’s entire publicly accessible image library—on the subreddit r/DefendingAIArt. “It was a fun project,” Reddit user MandarinDawnPoppy994 wrote in a since-deleted post, adding that the archive cost less than $10 to create.
“We found out about it through users tagging us,” Chan told WIRED. The scraper reportedly promoted the archive and encouraged others to use the data, triggering heated debates across AI forums about the ethics of collecting and redistributing artists’ work. “It feels very targeted and very harmful,” Chan said, noting that large-scale scraping is often defended as technically legal because “the law hasn’t caught up” with protections against mass data collection. Chan is also involved in two ongoing class-action lawsuits brought by visual artists—one against Stability AI, Midjourney and others, and another against Google—alleging that copyrighted artwork was used to train image-generation tools.
In a surprising development, the individual who archived Cara’s artwork later expressed regret and agreed to work with Chan on an open-source tool designed to help protect artists from unauthorized scraping and AI exploitation.
Other scrapers, however, continued to target Cara despite the platform’s limited resources and security challenges. Several AI supporters opposed the attacks, while some Cara users were reportedly encouraged by MandarinDawnPoppy994 to carry out what Chan described as a “copycat” scraping operation.
The second scraper collected 8.5 million Cara links, along with metadata such as usernames, titles and tags. The information was uploaded to Hugging Face, an AI development platform. After users requested its removal, Hugging Face issued a statement. The company said it had asked user “CaptiveDreamer” to remove personal metadata but could not remove the URLs because no artwork was hosted on Hugging Face. Instead, the links pointed to copies published by the artists on Cara. Hugging Face added that additional copyright reports based on the same circumstances would not change its decision.
On August 22, a third scraper reportedly downloaded 123,000 images from Cara, as well as text posts and user bios containing potentially sensitive personal information. The material was shared through an academic torrent. Chan subsequently launched a GoFundMe campaign with a $120,000 goal for legal expenses. The funds are intended to support potential cybersecurity and copyright strategies to protect Cara and its artist community. As of Thursday, the campaign had raised more than $100,000, while Cara continued seeking additional legal support.
Source: www.wired.com


