1 min readfrom TechCrunch

Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft

Our take

The legal landscape surrounding AI training data continues to evolve. Following similar actions, *The Seattle Times* and *Newsday* have filed lawsuits against OpenAI and Microsoft, alleging the unauthorized use of their journalistic content to train AI models. These suits highlight growing concerns about copyright and fair use in the rapidly advancing field of artificial intelligence. For further insight into AI agent behavior and related developments, explore our article, "OpenAI confirms ‘wiki incident’…"
Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft

The escalating legal battles surrounding AI training data continue to intensify, with the Seattle Times and Newsday joining the growing list of news organizations suing OpenAI and Microsoft. This latest wave of litigation underscores a critical tension at the heart of the AI revolution: the balance between innovation and intellectual property rights. These lawsuits aren’t simply about compensation; they represent a fundamental challenge to the current model of large language model (LLM) development, which has, until now, largely operated under the assumption that publicly available data is fair game for training purposes. We’ve previously explored OpenAI’s responses to data concerns, including their acknowledgement of a recent "wiki incident" [OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure] and the broader discussions around responsible AI practices as predicted by thought leaders [Presentation: A Few Predicted Talks From QConAI 2030]. The core argument these news organizations are making is that their copyrighted content – articles, reports, and investigations representing significant investment and journalistic expertise – has been scraped and utilized to train these powerful AI models without consent or compensation.

The implications of these lawsuits extend far beyond the immediate financial stakes for OpenAI and Microsoft. A ruling in favor of the news organizations could fundamentally reshape the landscape of AI training, potentially requiring developers to secure licenses or permissions for using copyrighted material. This could significantly increase the cost and complexity of developing LLMs, potentially slowing down innovation in the short term. Conversely, a ruling against the news organizations could embolden AI developers to continue operating under the current data-scraping model, further eroding trust in the industry and potentially incentivizing a race to the bottom in terms of data sourcing practices. The release of GPT-6 Astra [GPT-6 Astra: What’s Actually New in OpenAI’s New Frontier Model] highlights the relentless pace of advancement in this space, demonstrating just how quickly these models are evolving and the urgency with which these legal and ethical questions need to be addressed. It’s a complex situation, complicated by the inherent difficulty in tracing the precise origins of data used to train these massive models.

The legal arguments are nuanced, centering on questions of fair use, copyright infringement, and the transformative nature of AI training. While AI developers argue that training LLMs constitutes a transformative use of copyrighted material – creating something new and distinct from the original works – news organizations contend that the use is essentially a commercial exploitation of their intellectual property. The courts will need to grapple with these complex issues, potentially establishing new legal precedents that will govern the development and deployment of AI technologies for years to come. The debate also highlights the broader challenge of adapting existing legal frameworks to accommodate rapidly evolving technologies. Copyright law, for example, was largely conceived before the advent of AI and may require significant revisions to adequately address the unique challenges posed by LLMs. This situation isn’t unique to news organizations; artists, authors, and other creators are also increasingly scrutinizing how their work is used in AI training.

Looking ahead, the resolution of these lawsuits will likely shape the future of AI development and the relationship between AI developers and content creators. Will we see a shift towards more transparent data sourcing practices and the implementation of licensing agreements for copyrighted material? Or will the industry continue to rely on large-scale data scraping, potentially leading to further legal challenges and ethical concerns? The ongoing discussions surrounding data governance and responsible AI are crucial, and the outcome of these legal battles will undoubtedly influence the direction of these conversations. A key question to watch is whether alternative training methods, such as synthetic data generation, will gain traction as a way to mitigate the reliance on copyrighted material and reduce the potential for legal disputes.

Two more news organizations are suing OpenAI and Microsoft over the supposed use of their journalism to train AI.

Read on the original site

Open the publisher's page for the full experience

View original article