Sunday, September 06, 2026

How confusing are these arguments?

https://thenextweb.com/news/seattle-times-newsday-sue-openai-microsoft-copyright-paywall-bypass-gpai-code-of-practice-copyright-chapter-measure-1-2

Two more newsrooms join the case against OpenAI and Microsoft

The Seattle Times and Newsday have sued OpenAI and Microsoft, alleging the companies methodically scraped articles in a way that bypasses paywalls. The EU’s general-purpose AI code commits signatories not to circumvent access restrictions, naming subscription models and paywalls specifically.

The Seattle Times and Newsday have jointly sued OpenAI and Microsoft over the training of AI models on their journalism. The two called generative AI “a snake eating its own tail“, Engadget reported.

The complaint alleges the companies were methodically scraping news articles in a way that bypasses paywalls. That is a claim about how the material was obtained, not only about what was done with it.

It says the result offers readers an AI-generated alternative to the articles themselves, cutting traffic and digital advertising revenue.

It also alleges the models hallucinate, attributing false information to the two outlets, and that copyright management information was stripped from articles.





I was thinking along these lines…

https://www.researchgate.net/profile/Nikolaos-Polatidis/publication/413673690_AI_is_as_Good_as_the_Data_it_Learns_From_A_Review_of_Dataset_Needs_Across_AI_Applications/links/6a8ffd2f0f1ada50a9d362dd/AI-is-as-Good-as-the-Data-it-Learns-From-A-Review-of-Dataset-Needs-Across-AI-Applications.pdf

AI is as Good as the Data it Learns From: A Review of Dataset Needs Across AI Applications

Artificial intelligence systems do not rely on data in the same way. Classical machine-learning methods are commonly developed from structured, task specific records, whereas deep-learning systems make greater use of large image, text, audio, video, and sensor collections. Generative and multimodal models extend this dependence further by combining large-scale pretraining corpora with smaller datasets for instruction tuning, human preference modelling, alignment, and safety evaluation. This review examines how dataset requirements change across major AI paradigms, including rule-based systems, classical machine learning, deep learning, natural language processing, computer vision, recommender systems, reinforcement learning, robotics, generative AI, and multimodal AI. The comparison focuses on data modality, supervision, scale, preparation, quality, representativeness, provenance, privacy, copyright, and governance. Across these paradigms, the evidence shows that dataset suitability cannot be judged by size alone: the value of data depends on how well it represents the target task, learning process, and deployment environment. The review therefore develops a cross-paradigm view of data requirements and identifies practical open problems that remain unresolved as AI systems become larger, more heterogeneous, and more widely deployed.



No comments: