AI Is Going Just Great

Category

Copyright / Data

Training-data lawsuits, scraped archives, artists vs. models, and the legal frontier of synthetic content.

← All categories

  1. August 2026

  2. ·todayConcerningMinor

    Marvel's Spider-Man: Brand New Day Artbook Flagged for Possible AI-Generated Concept Art

    analyticsinsight.net

    Fans argue that a premium artbook should clearly identify any artificial intelligence involvement in its creative material.

    One image in the official 208-page collector's artbook for Spider-Man: Brand New Day is drawing scrutiny after fans spotted warped taxis, distorted buildings, strange windows, and a damaged fire escape in a New York street scene titled "Big City, Little Spider." The image credits concept artist Scott McInnes. Marvel has not confirmed AI involvement.

    The backlash spread across X, Reddit, and ResetEra after The Artbook Collector flagged the artwork on August 18, the same week the book shipped. The timing is uncomfortable for Marvel: the studio recently downsized its visual development team, and it already drew criticism in 2023 for using AI in Secret Invasion's opening title sequence. Fans are calling for transparency about AI use in credited creative work, particularly in premium, paid publications.

    Copyright / DataJobs / Workforce
  3. ·2d agoSadModerateamazon

    Amazon warehouse destroys rare books for AI training; its logo is a T. rex shredding a book

    cybernews.com

    It upsets me that there are ways to do this without destroying the book, but destroying the book is cheaper. Just awful.

    An investigation by 404 Media found that Amazon operates a warehouse, code-named VGT3, where employees scan rare and out-of-print books for AI training data and then discard the physical copies. The facility's logo is a Tyrannosaurus rex poised to destroy a book. Workers in online forums worried the warehouse would shut down without a steady supply of old books; it remains open, with rare books brought in bulk for destruction.

    Removing bindings makes high-speed scanning cheaper, and some legal interpretations hold that copying a book and destroying the original skirts copyright claims. Amazon said it "purchases books through commercial channels to help develop and improve the products and services [its] customers use." Anthropic was separately reported to have destroyed millions of books to train Claude.

    Copyright / DataReal-World Impact
  4. ·1w agoEmbarrassingModerate

    SMU graduate student loses $2 million book deal after agent says he gave conflicting answers about AI use

    keranews.org

    "I'm just a boy in my 20's. I don't understand why anyone is witch-hunting me."

    Jerry Falade, a PhD student at Southern Methodist University, lost a publishing deal worth more than $2 million after his agent said he gave conflicting answers about using AI to write parts of his crime novel. Minotaur Books had acquired Call Me, I'll Hide the Body in a competitive auction that drew 13 other publishers.

    Falade's agent, Marc Gerald, told The Bookseller that another editor raised concerns halfway through the submission process. Falade denied the allegations on social media. "I'm just a boy in my 20's," he wrote. "I don't understand why anyone is witch-hunting me." Publishing consultant Jane Friedman described an industry operating in "a world of both confusion and suspicion. No one really trusts anyone anymore." The U.S. Copyright Office ruled in 2025 that AI-generated content can only be copyrighted with "sufficient expressive elements" from a human, and prompts alone do not qualify. The incident follows Hachette's cancellation of a horror novel in March over similar AI accusations.

    Real-World ImpactCopyright / Data
  5. ·1w agoInfuriatingModerateamazon

    Twitch Quietly Opts All Streamers Into Amazon AI Training, Buries the Opt-Out

    kotaku.com

    "If it was opt-in, nobody would opt-in. Um, that's honestly the answer. So it's going to be on by default." — Mike Minton, Twitch CPO

    Twitch began using streamer VODs, livestreams, clips, chat logs, and channel images to train Amazon generative AI models, defaulting every account to opt-in and tucking the opt-out inside a security settings submenu. The change went live on August 12 without any email notice to creators; Twitch's head of community later explained that "not everyone reads their emails" and that word-of-mouth gets information "further."

    During a stream addressing the backlash, chief product officer Mike Minton explained the opt-out-by-default design with disarming candor: "If it was opt-in, nobody would opt-in. Um, that's honestly the answer." He also confirmed Twitch has no plans to show users which channels are opted in or out, couldn't answer what happens to your chat messages when you appear in someone else's opted-in stream, and said he personally would not opt out if he were a streamer. User feedback on Twitch's own platform was described as running "quite clear" against the move.

    Copyright / DataReal-World Impact
  6. ·2w agoConcerningMajor

    EU AI Act Enforcement Begins, Requiring AI Transparency Labels and Copyright Policies

    helpnetsecurity.com

    The most advanced models "create risks on an entirely new scale."

    On 2 August 2026, the European Commission's AI Office and national authorities began enforcing the EU AI Act, with transparency rules now requiring chatbots to identify themselves as automated systems, deepfakes to carry labels, and machine-generated content to include machine-readable marks. Companies that skip these obligations face fines of up to €15 million or 3% of global annual turnover, whichever is higher.

    Providers of general-purpose AI (GPAI) models face the most immediate scrutiny: they must document training data, publish summaries of content used to train their models, and maintain a copyright policy. The Commission has already named OpenAI, Anthropic, and Google as companies whose relationship with European regulators could grow more complicated under the new powers. Not everything kicks in at once — rules for high-risk AI systems are delayed until late 2027 at the earliest — but a ban on AI-generated non-consensual sexually explicit content and CSAM takes effect in December 2026.

    Safety FailureCopyright / Data
  7. July 2026

  8. ·1mo agoIronicModerateanthropic

    AI companies are buying up pre-2022 printed books because the internet is too full of AI-generated text

    404media.co

    "AI company destroys two million books" is not a headline that generates sympathy.

    ISBNdb, a book database company, is now offering bulk book acquisition services for AI companies, pitching pre-2022 printed books as ideal training data because they are "structurally guaranteed" to be free of AI-generated text. The company handles orders of 1,000 to 1 million books at a time and requires NDAs to keep AI companies' identities hidden — its own site notes that "'AI company destroys two million books' is not a headline that generates sympathy."

    The pitch arrives as the web fills with AI-generated content, which can cause "model collapse" when used as training data. Booksellers on platforms like Alibris and Biblio report historic spikes in bulk purchases with "no rhyme or reason" — varied topics, disregard for price, and a focus exclusively on books with ISBNs. Internal Anthropic documents revealed in a copyright lawsuit detailed plans to scan millions of books and destroy them in the process; a federal judge ruled the copying was fair use specifically because the physical books were destroyed. One bookseller told 404 Media that rare, foreign-language, and low-circulation books are being swept up in these purchases, and if they're pulped during scanning, "they'll be even harder to obtain."

    Hype vs RealityCopyright / Data
  9. ·1mo agoInfuriatingMajoropenai

    New York Times Claims OpenAI Concealed Evidence of Copyright Infringement Detection Tools in Lawsuit

    techcrunch.com

    "If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it." — Ian B. Crosby, lead counsel for plaintiffs

    The New York Times and The Daily News have accused OpenAI of hiding evidence in their ongoing copyright lawsuit, alleging the company misrepresented its ability to search training data and chat logs. A court-ordered deposition of OpenAI data privacy engineer Vinnie Monaco allegedly revealed that OpenAI had already conducted internal searches of its training corpus for copyrighted journalism — and had amassed a database of roughly 78 million de-identified ChatGPT conversations to assess its own infringement exposure. OpenAI had previously argued that such searches were technically burdensome and privacy-sensitive.

    Further compounding the allegations, the plaintiffs claim OpenAI built a "Bloom" filter under an internal initiative called "Project Giraffe" to detect and log output regurgitation shortly after the lawsuit was filed — then allegedly deleted billions of ChatGPT outputs in violation of a court preservation order, submitted a heavily redacted 20-million-log sample the court itself called "unusable," and substituted millions of logs in the requested sample. The NYT and Daily News are now asking the judge to sanction OpenAI, bar it from using the chat log sample as evidence, and compel it to pay legal fees. OpenAI denied the allegations, framing the move as a privacy attack on users by plaintiffs with a weakening case.

    Copyright / DataCorporate Drama
  10. June 2026

  11. ·1mo agoInfuriatingMajoranthropic

    Anthropic Claims Alibaba Used 25,000 Fake Accounts and 28.8 Million Exchanges to Illicitly Distill Claude

    tomshardware.com

    25,000 fake accounts and 28.8 million exchanges — Anthropic says Alibaba ran an industrial-scale operation to distill Claude from April to June 2026.

    Anthropic has accused China's Alibaba of running a large-scale, covert operation to "distill" its Claude AI models without authorization. According to Anthropic, the effort involved roughly 25,000 fake accounts and 28.8 million exchanges with Claude, carried out between April and June 2026 — essentially using Claude's outputs at massive scale to train or improve a competing model.

    Model distillation via fake accounts is a known risk in the AI industry, but the alleged scope here is striking: nearly 29 million exchanges over just a few months amounts to an industrial-grade extraction effort. Anthropic has not yet disclosed what legal or technical remedies it is pursuing.

    Security / AbuseCopyright / Data
  12. — end of timeline —