Quartz
Subscribe
Quartz
Subscribe
Edition
Business News
A.I.
Technology
Money & Markets
Leadership
Lifestyle
Latest

Get Quartz in your inbox

Free daily briefing on global business news.

Business News
AirlinesAutomobilesFoodPharmaceuticalsPolitics & GovernmentRetail & EcommerceSpace & AerospaceEarnings
Technology
A.I.ComputingConsumer TechSpace & AerospaceEarnings
Money & Markets
Economic IndicatorsMarketsPersonal FinanceEarnings
Lifestyle
Cars & BikesCollectingEntertainmentFood & Fine DiningHealth and FitnessReal EstateTravel
Quartz

Global business news for a smarter world

Topics

  • Business News
  • Money & Markets
  • Tech & Innovation
  • Generation A.I.
  • Lifestyle
  • Leadership

Products

  • Daily Brief
  • Weekly Digest
  • Member Benefits
  • Quartz Pro

Legal

  • Sitemap
  • About
  • Accessibility
  • Privacy
  • Terms of Service
  • Advertising

© 2026 Quartz Media, Inc. All rights reserved.

A.I.

OpenAI may have broken YouTube rules by training ChatGPT on 1 million hours of video

OpenAI and other tech companies are facing difficulty collecting enough data to train massive AI models

By Maxwell Zeff·2 min read·Updated April 8, 2024
Add QZ to Google

OpenAI reportedly transcribed more than one million hours of YouTube videos to train GPT-4, according to The New York Times on Saturday. The report comes just days after YouTube CEO Neal Mohan said transcribing YouTube videos for AI training would be a “clear violation” of its policies in a Bloomberg interview.

“When a creator uploads their hard work to our platform, they have certain expectations. One of those expectations is that the terms of services is going to be abided by,” said Mohan in an interview with Bloomberg last week. “But it does not allow for things like transcripts or video bits to be downloaded.”


The New York Times report alleges that OpenAI team members, including President Greg Brockman, personally helped collect the YouTube videos, according to sources. The article details how OpenAI, and many tech companies, are facing difficulty collecting enough data to train massive AI models. OpenAI allegedly used Whisper, its AI transcription software, to collect more data to train GPT-4, the latest and greatest model underlying ChatGPT.


OpenAI and Google $GOOGL did not immediately respond to Gizmodo’s requests for comment.


The New York Times report could have massive implications for OpenAI and Google’s ongoing battle at the forefront of generative AI development. Google is unlikely to go quietly if OpenAI is using its content to make ChatGPT even greater. However, the company has made no such allegations yet. In a statement to The Verge this weekend, a Google spokesperson merely said he’s “seen unconfirmed reports” about OpenAI’s training.

YouTube’s terms of service prohibit any user from downloading its content, including the use of botnets or scrapers, unless they have clear permissions from the company. YouTube also prohibits utilizing its content for any “independent” uses of its service.

OpenAI’s Chief Technology Officer, Mira Murati, said she was “not sure” whether YouTube videos were used to train her company’s text-to-video AI model Sora when asked by The Wall Street Journal in March. The New York Times report mentions nothing about Sora, or actual YouTube bits themselves. However, her hesitancy to answer this question directly leads to greater speculation.


The New York Times, itself, is in a copyright battle with OpenAI at the moment. OpenAI and Meta $META are also being sued by a number of authors and content houses for training their AI on copyrighted works.

If these reports are true, it could raise entirely new questions about copyright law in the AI world. Most copyright complaints around AI have been brought by small publishers, but Google could add some real weight behind this fight if it chooses to partake. It would also present a way for Google to slow down OpenAI, which is undoubtedly winning the AI race at the moment.


A version of this article originally appeared on Gizmodo.

Daily Brief

The essential business news, delivered fresh every morning.

Join 500,000+ readers who start their day with Quartz.

By subscribing, you agree to our Terms of Service and Privacy Policy.

Related

Cloud ComputingVerizon lands a $1 billion-plus dark fiber deal with Google for AI data centers
Politics & GovernmentTrump vows new tariffs on the E.U. after Brussels fines Google $1 billion
Politics & GovernmentTrump rolls out new forced-labor tariffs on 60 countries as trading partners push back
A.I.Samsung and SK Hynix are set to announce major memory chip deals with U.S. tech firms
A.I.Meta is upgrading its AI assistant to automate recurring tasks and daily briefings