Showing posts with label open source. Show all posts
Showing posts with label open source. Show all posts

Sunday, August 23, 2026

Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway; Forbes, August 23, 2026

Rob Salkowitz, Forbes; Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway

"Scrapers defend their actions

Artists’ claims to own and control their own work online are disputed by individuals, groups and commercial entities who believe that advancing the progress of AI entitles them to any and all data they can obtain, regardless of consent. This was apparently the motivation of the original scraper, who posted an archive of nearly 12 million images from Cara to the sub-Reddit r/DefendingAIArt under the handle “MandarinDrawnPoppy994.”

“Scraping is necessary to develop good models. It’s like building a highway – some houses must be demolished, but in the end everyone benefits,” the poster wrote in a thread titled “[AMA] I scraped all of Cara.”

Zhang says she and others reached out to the original scraper and eventually prevailed on him to take down the post. She adds, “not only did the first scraper delete the dataset, but he has turned around now to offer help, and we’re now co-creating an open source tool separate from Cara that will help people check if they've been scraped in new datasets in the future.”

Unfortunately, that was not the end of the problem. Several days later, the site was scraped again by a different actor. This time the data was posted on Hugging Face, a hub of resources for AI developers rooted in the open source community, by a poster under the name “Ioannis/Captive Dreamer.”

After some people reported the post to Hugging Face Trust and Safety, the team responded that “we have reviewed these [copyright reports] carefully. Because no copies of the artworks are hosted here, and because the URLs [in the dataset] point to the copies the artists published on Cara, there is nothing hosted on Hugging Face that we can disable through our notice and takedown process. This is not a judgement about who owns the works (the artists do); it is about what is stored on our servers. Further copyright reports on the same basis will not change this outcome.”

Hugging Face did not respond to a request for further comment for this story.

Third attack in 10 days

Now on Saturday, August 22, Zhang says the site was scraped for a third time, with the perpetrator taking just 123K images, but also a second data set that includes users’ text posts, bio and information they share on the site. That archive has been posted on Academic Torrents, a site that “was established to meet the demands of science in the age of big data” by providing data for researchers, according to its “About” page."

Sunday, April 26, 2026

Devious New AI Tool “Clones” Software So That the Original Creator Doesn’t Hold a Copyright Over the New Version; Futurism, April 26, 2026

  , Futurism; Devious New AI Tool “Clones” Software So That the Original Creator Doesn’t Hold a Copyright Over the New Version

"The advent of generative AI continues to undermine the very concept of copyright, from entire books shamelessly ripping off authors to tasteless AI slop depicting beloved characters going viral on social media. The sin is foundational: all today’s popular AI tools were built by pillaging copyrighted material without permission.

Even software isn’t safe. As 404 Media reports, a new tool dubbed Malus.sh — pronounced “malice,” to give a subtle clue where this is headed — uses AI to “liberate” a piece of software from existing copyright licenses, essentially creating a “clean room” clone that technically doesn’t infringe on the original code’s copyright."

Friday, April 24, 2026

DeepSeek’s Sequel Set to Extend China’s Reach in Open-Source A.I.; The New York Times, April 24, 2026

 Meaghan Tobin and , The New York Times; DeepSeek’s Sequel Set to Extend China’s Reach in Open-Source A.I.

"DeepSeek released its models as open source, which means others can freely use and modify them. By contrast, OpenAI and Anthropic kept their leading models proprietary. The episode demonstrated that an open-source system could perform almost as well as closed versions. In the months that followed, Chinese firms released dozens of other open-source models. By the end of 2025, these models made up a significant share of global A.I. usage.

On Friday, DeepSeek released a preview of V4, its long-awaited follow-up model, which it intends to open source. The new model excels at writing computer code, an increasingly important skill for leading A.I. systems. It significantly outperformed every other open-source system at generating code, according to tests from Vals AI, a company that tracks the performance of A.I. technologies.

DeepSeek released its new model just days after Moonshot AI, another Chinese start-up, introduced its latest open-source model, Kimi 2.6. While these systems trail the coding capabilities of the leading U.S. models from Anthropic and OpenAI, the gap is narrowing.

The implications are meaningful. Using A.I. to write code is faster and frees up human programmers to focus on bigger issues. It also means people can use DeepSeek’s latest release to power A.I. agents, which are personal digital assistants that can use other software applications on behalf of office workers, including spreadsheets, online calendars and email services."