data scraping

data scraping on Beyond Market Intelligence: a running collection of 4 stories we have gathered and hand-picked because they are worth your time. Every post here touches on data scraping in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around data scraping, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft
TechCrunch

Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft

The legal landscape surrounding AI training data continues to evolve. Following similar actions, *The Seattle Times* and *Newsday* have filed lawsuits against OpenAI and Microsoft, alleging the unauthorized use of their journalistic content to train AI models. These suits highlight growing concerns about copyright and fair use in the rapidly advancing field of artificial intelligence. For further insight into AI agent behavior and related developments, explore our article, "OpenAI confirms ‘wiki incident’…"

Is it legal to train AI models on copyrighted books? It’s complicated
TechCrunch

Is it legal to train AI models on copyrighted books? It’s complicated

The legality of training AI models on copyrighted books presents a complex and evolving challenge. Many published authors, often unknowingly, have contributed to the datasets powering AI tools now poised to impact their profession. The question of whether this constitutes infringement is at the heart of ongoing debate. While the situation seems inherently problematic, definitive legal answers remain elusive. For deeper insights into related discussions surrounding AI and investment, explore our article, "Will the DOJ’s investigation into a16z spook other VCs?".

Hack suggests AI music generator Suno scraped YouTube for training data
TechCrunch

Hack suggests AI music generator Suno scraped YouTube for training data

Recent allegations suggest AI music generator Suno may have utilized improperly sourced training data. A security breach, involving the unauthorized access of Suno’s source code via an employee’s credentials, revealed a process of scraping audio from YouTube spanning decades. This raises significant concerns about copyright and data ethics within the rapidly evolving AI landscape. For a deeper dive into the challenges of AI agent validation, see our recent article, "Stripe Benchmark Shows AI Agents Build Integrations but Struggle with Validation."

Google faces another AI training lawsuit from major publishers
TechCrunch

Google faces another AI training lawsuit from major publishers

Google is facing a significant legal challenge as major publishers—including Hachette, Cengage, and Elsevier—file a lawsuit alleging unauthorized use of copyrighted material to train its AI models. This action highlights the growing tension surrounding AI development and intellectual property rights. Publishers assert that Google leveraged copyrighted works without securing proper permissions, raising questions about fair use and data sourcing. For a contrasting perspective on AI applications, explore "The founder of Hinge raised $18M to build a new AI dating service, Overtone."