He Collected Their Artwork for AI Use. Now He’s Partnering to Create a Tool to Assist Them.

Since the beginning of 2023, photographer Jingna Zhang and a dedicated team of volunteers have been working hard to sustain an image-sharing social media and portfolio application known as Cara. To date, approximately 1.5 million artists have joined the platform. What enticed them to use Cara? A collective resistance to the unauthorized appropriation of their creations for AI model training and a wish to showcase their artwork while steering clear of exploitation by major tech companies.
However, while Cara successfully filters out AI-generated images and provides protective features—such as Glaze, a tool designed to obscure the style of images captured by scrapers to thwart AI imitation—preventing scrapes entirely proves exceedingly difficult. Just this month, starting on August 13, Cara experienced three significant scrapes, leading to increased server costs and causing concern among creators who transitioned from platforms like Instagram, where all content is readily accessible to Meta for training data.
The first incident was revealed when the individual responsible shared a 12-terabyte archive containing 12 million works from Cara—essentially its entire library of publicly accessible images—on the subreddit r/DefendingAIArt. “It was a fun project,” the user, MandarinDawnPoppy994, stated in his now-deleted post, mentioning that the endeavor cost him under $10.
“We actually learned about it through our users tagging us,” Zhang tells WIRED, noting that the scraper was “boasting and seeking others to collaborate on the dataset on Reddit,” igniting a heated discussion across AI-related forums regarding the ethics behind his actions. “I just feel it’s targeted and very hurtful,” she adds, highlighting that “laws have not kept pace” with protections against such data harvesting, allowing scrapers to often argue that their actions are technically legal. (Zhang is involved in two ongoing class actions filed by visual artists, one against Stability AI, Midjourney, and others, and another against Google, claiming that the companies’ image generator tools were trained on their copyrighted materials.)
In an unexpected twist, however, the individual who collected all the artwork from Cara would ultimately regret his actions and agree to work with Zhang on a new open-source tool aimed at safeguarding artists.
Meanwhile, other scrapers continued to exploit Cara’s vulnerabilities and limited resources. While some AI advocates objected to targeting Cara, a few appeared to be emboldened by MandarinDawnPoppy994 to perform what Zhang describes as “copycat” attacks.
A second scraper extracted around 8.5 million links from Cara, along with metadata such as usernames, titles, and tags, and uploaded these to Hugging Face, an AI development platform. After receiving a barrage of takedown requests, Hugging Face issued a statement indicating that while it would notify the user, “CaptiveDreamer,” to remove personal metadata, it could not act on the URLs since “no copies of the artworks are hosted here,” and the links “direct to the versions published by the artists on Cara.” The company concluded that “further copyright reports on the same basis will not alter this outcome.”
Finally, on August 22, a third scraper acquired 123,000 images from Cara, along with text posts and user bios that contained personal information, and shared everything on a site named Academic Torrents. In response, Zhang initiated a GoFundMe for legal expenses, setting a goal of $120,000, stating that the funds would be used to explore all possible strategies for defending Cara through cyber and copyright laws. As of Thursday, she has raised over $100,000, and she mentions that Cara is actively seeking additional legal assistance.
