News / #copyright Tag Copyright 41 articles archived under #copyright · RSS Sign in to follow arXiv — Machine Learning research 2d ago SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features arXiv:2608.10709v1 Announce Type: new Abstract: Quantization-Aware Training (QAT) enables the deployment of quantized models with minimal accuracy degradation. However, in practical scenarios, training labels are often unavailable due to privacy, copyright, or cost constraints.… 5 arXiv — NLP / Computation & Language research 2d ago Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a… 28 arXiv — Machine Learning research 18d ago DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection arXiv:2607.22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain… 25 arXiv — NLP / Computation & Language research 18d ago Agentic Evaluation of Copyright Law Compliance arXiv:2607.21799v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including… 6 Hacker News — AI on Front Page community 23d ago 'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling Article URL: https://www.techradar.com/vpn/vpn-privacy-security/vpns-are-lawful-technical-tools-says-eu-court-in-landmark-anne-frank-copyright-ruling Comments URL: https://news.ycombinator.com/item?id=48997221 Points: 354 # Comments: 61 32 Hacker News — AI on Front Page community 23d ago Judge approves $1.5B Anthropic settlement for pirated books used to train Claude Article URL: https://apnews.com/article/ai-anthropic-copyright-settlement-claude-books-bartz-74b140444023898aeba8579b6e9f0d63 Comments URL: https://news.ycombinator.com/item?id=48996652 Points: 281 # Comments: 200 10 Ars Technica — AI news-outlet 23d ago Anthropic’s $1.5B copyright settlement approved; only 350 authors opted out Anthropic blocks authors from opting out of $1.5B settlement at last minute. 37 r/LocalLLaMA community 23d ago Anthropic claims local models are stealing from it, meanwhile it pays $1.5B for theft $1.5 billion settlement largest known payout in U.S. copyright case Case part of a wave of lawsuits from copyright holders against AI companies Some authors and publishers opted out and continue separate cases against Anthropic   submitted by   /u/Terminator857 [link]… 18 r/LocalLLaMA community 24d ago Anthropic got sued for using copyrighted books for LLM training   submitted by   /u/Informal-Trouble2183 [link]   [comments] 19 TechCrunch — AI news-outlet 24d ago Anthropic’s landmark $1.5B copyright settlement is approved The final approval settles one case, but it doesn't resolve the broader issue of using copyrighted works to train AI models. 31 TechCrunch — AI news-outlet 1mo ago Google faces another AI training lawsuit from major publishers Hachette, Cengage, Elsevier, and other publishers allege that Google trained its AI on copyrighted works without the necessary permissions. 9 r/LocalLLaMA community 1mo ago Grok Build CLI uploads your whole repo — full git history + .env secrets — to xAI's cloud, and the opt-out doesn't stop it (wire-captured) I ran Grok Build CLI (v0.2.93) through mitmproxy. It uploads your entire repo as a git bundle (full history) to xAI's Google Cloud — independent of what you open. With the prompt literally "do not read or open any files," a file I planted came back verbatim when I git clone -d… 5 arXiv — NLP / Computation & Language research 1mo ago Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks arXiv:2607.07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data.… 4 arXiv — Machine Learning research 1mo ago AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate arXiv:2607.08337v1 Announce Type: new Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of… 21 TechCrunch — AI news-outlet 1mo ago New York Times says OpenAI hid evidence in ChatGPT copyright trial News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, escalating their lawsuit with a new motion for sanctions. 16 Ars Technica — AI news-outlet 1mo ago OpenAI faked inability to search training data, hid billions of logs, NYT says OpenAI may be sanctioned for hiding, deleting ChatGPT logs in NYT copyright fight. 24 arXiv — Machine Learning research 1mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 arXiv — Machine Learning research 1mo ago Amplifying Membership Signal Through Chained Regeneration arXiv:2606.31991v1 Announce Type: new Abstract: The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enforcement. Current membership (MIA) and dataset inference (DI) attacks often rely on one-shot… 7 arXiv — NLP / Computation & Language research 1mo ago Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law arXiv:2606.31250v1 Announce Type: new Abstract: Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards:… 36 arXiv — Machine Learning research 1mo ago CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence arXiv:2606.27683v1 Announce Type: new Abstract: Edge devices increasingly invoke large language models (LLMs) through API services for context aware edge intelligence, while edge generated data may be collected to improve LLMs and may introduce sensitive, copyrighted, harmful,… 11 arXiv — NLP / Computation & Language research 1mo ago Position: The Term "Machine Unlearning" Is Overused in LLMs arXiv:2606.27379v1 Announce Type: new Abstract: Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This position paper… 15 Ars Technica — AI news-outlet 1mo ago NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI NYT shifts OpenAI/Microsoft copyright claims after SCOTUS ruling against Sony. 20 arXiv — NLP / Computation & Language research 1mo ago Output Vector Editing for Memorization Mitigation in Large Language Models arXiv:2606.18767v1 Announce Type: new Abstract: Large language models memorize and reproduce sequences from their training data, creating privacy, copyright, and security risks. Existing neuron-level mitigation methods equate editing with zeroing out neuron activations, but the… 24 llama.cpp releases dev-tools 1mo ago b9684 [SYCL] Add conv_3d ( #24691 ) add conv_3d optimize update ops.md restore test script rm unused code rm copyright notes macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu… 15 Hacker News — AI on Front Page community 2mo ago H.R. 6028 would fundamentally change the U.S. Copyright Office Article URL: https://www.eff.org/deeplinks/2026/06/congress-just-rushed-through-disastrous-copyright-office-overhaul Comments URL: https://news.ycombinator.com/item?id=48484496 Points: 209 # Comments: 65 25 r/MachineLearning community 2mo ago ICML rejected paper visibility [D] If ICML conference paper is rejected and no one opts-in or opts-out to keep the reviews visible, will the reviews be visible to everyone? There was clear instruction that only papers with at-least 1 opt-in AND zero opt-out options will be visible. None of the authors selected… 7 arXiv — Machine Learning research 2mo ago Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path arXiv:2606.07271v1 Announce Type: new Abstract: Understanding what generative models retain from training data remains challenging, with implications for copyright and privacy. Beyond verbatim reproduction, models can encode subtler traces of their training data that never… 12 TechCrunch — AI news-outlet 2mo ago Publishers will be able to opt out of AI Search, thanks to new regulation U.K. regulators are requiring Google offer a tool allowing website publishers to opt-out of generative AI search features. The option will be tested in the UK then rolled out globally. 10 arXiv — Machine Learning research 2mo ago Geometric Erasure by Contrastive Velocity Matching in Rectified Flows arXiv:2606.00140v1 Announce Type: new Abstract: While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. To address these challenges, concept erasure… 21 arXiv — NLP / Computation & Language research 2mo ago Divergence Decoding: Inference-Time Unlearning via Auxiliary Models arXiv:2605.31293v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently memorize sensitive training data thereby creating significant privacy and copyright risks. Addressing these risks, i.e., removing such knowledge from an existing model checkpoint, has proven… 34 arXiv — Machine Learning research 2mo ago Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature Differences arXiv:2605.26756v1 Announce Type: new Abstract: Diffusion models can unintentionally memorize training samples, raising concerns about privacy and copyright. While recent methods can detect memorization, they often rely on global or model-specific signals and provide limited… 30 arXiv — NLP / Computation & Language research 2mo ago Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data arXiv:2605.24842v1 Announce Type: new Abstract: This paper examines how the labour of translators has been transformed into foundational data capital for the age of artificial intelligence (AI). Translation memories (TM) and parallel corpora preserve a one-to-one correspondence… 6 Hacker News — AI on Front Page community 2mo ago Show HN: Auto-identity-remove – Automated data broker opt-out runner for macOS Article URL: https://github.com/stephenlthorn/auto-identity-remove Comments URL: https://news.ycombinator.com/item?id=48178184 Points: 282 # Comments: 112 15 arXiv — NLP / Computation & Language research 2mo ago Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training arXiv:2506.01732v3 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which… 11 Ars Technica — AI news-outlet 3mo ago Authors fight for higher payouts from Anthropic’s $1.5B copyright settlement Lawyers accused of rushing historic settlement to seize $320 million in fees. 8 arXiv — NLP / Computation & Language research 3mo ago To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model arXiv:2605.14291v1 Announce Type: cross Abstract: The rapid advancement of Large Vision-Language Models (LVLMs) is increasingly accompanied by unauthorized scraping and training on multimodal web data, posing severe copyright and privacy risks to data owners. Existing… 8 arXiv — Machine Learning research 3mo ago Inference-Time Machine Unlearning via Gated Activation Redirection arXiv:2605.12765v1 Announce Type: new Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning seeks to remove the influence of a targeted forget set while preserving model… 10 arXiv — NLP / Computation & Language research 3mo ago Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical… 17 Vercel — AI dev-tools 4mo ago Team-wide Zero Data Retention and prompt training controls now on AI Gateway AI Gateway now supports Zero Data Retention (ZDR) at the team level, removing the need to configure opt-outs or reach agreements with each provider individually. It routes requests only to providers where ZDR agreements are in place, with support for Anthropic, OpenAI, Google,… 35 Smol AI News news-outlet 7mo ago not much happened today **Stanford paper** reveals **Claude 3.7 Sonnet** memorized **95.8% of Harry Potter 1**, highlighting copyright extraction risks compared to **GPT-4.1**. **Google AI Studio** sponsors **TailwindCSS** amid OSS funding debates. **Google** and **Sundar Pichai** launch **Gmail Gemini… 21 Eugene Yan research 28mo ago Task-Specific LLM Evals that Do & Don't Work Evals for classification, summarization, translation, copyright regurgitation, and toxicity. 9