Leaked materials shared with 404 Media suggest that AI music company Suno assembled a vast training library by copying audio and lyrics from platforms including YouTube Music, Deezer, Genius, Pond5, Jamendo, Freesound, the International Music Score Library Project, and podcast feeds delivered through RSS. The hacker who accessed the company’s systems also claimed to have obtained user information linked to hundreds of thousands of Suno customers, along with Stripe-related payment data.
The documents offer an unusually detailed look at how an AI music platform may have gathered the raw material used to build its models. Suno has already been at the center of major legal fights with the recording industry, which alleges that the company trained its systems on millions of copyrighted recordings. In court filings, Suno previously acknowledged that its models were trained on “essentially all music files of reasonable quality” available on the open internet, amounting to “tens of millions” of recordings. The company has argued that such training qualifies as fair use, and at least one of those cases has already been settled.
What the leaked files appear to show
While the lawsuits established the broad scope of Suno’s training practices, the hacked data appears to reveal more about where the music came from and how it was collected. The Recording Industry Association of America had accused Suno of copying songs directly from YouTube, and the code reviewed by 404 Media reportedly supports that claim.
The leaked material includes source code dated to 2023 and 2024, with comments describing scraping pipelines and the size of several datasets. One file reportedly references sources such as “genius_hq,” “youtube_music,” “freesound,” “jamendo,” “imslp,” “deezer,” and “ytm_tagged,” while noting that non-music content would be filtered out. A file labeled “youtube_music” allegedly recorded the ingestion of 2,013,545 music clips.
Another file contained internal notes on dataset sizes measured in hours. Those figures reportedly included:
- 113,879 hours from YouTube Music
- 17,615 hours from Genius-related material
- 410 hours from Freesound
- 19,514 hours from IMSLP
- 3,726 hours from Jamendo
- 62,117 hours from Pond5 music
- 12,287 hours from Deezer
- 152,162 hours from a dataset called ytm_tagged
- 103 hours from MuseScore lyrics
Taken together, those figures point to an enormous catalog spanning many decades of listening time. The code also reportedly included routines aimed at isolating vocal tracks, including searches for acapella versions of songs on YouTube.
Streaming sources, proxies, and podcast scraping
According to the report, some of the code suggests Suno used proxy infrastructure from Bright Data, a company known for selling scraping tools and data services, to pull music from YouTube. Additional code indicated that the company also targeted spoken-word material. Using a service called PodcastIndex, Suno allegedly identified around 420,000 podcasts with at least five episodes of 30 minutes or more, and sought to download roughly one million hours of podcast audio.
It remains unclear from the leaked files exactly how Suno copied material from every source listed in the code. But some of the reported targets raise immediate questions. Pond5, owned by Shutterstock, is a paid stock-media marketplace with millions of music tracks and sound effects. If the internal figures are accurate, Suno may have copied a substantial portion of that catalog. Genius presents a different issue: it does not directly host full songs in the usual sense, but it does integrate with licensed streaming services and provides lyrics and song-related content.
The legal and copyright backdrop
In one court filing, Suno said its training data included nearly all reasonably accessible music files on the open web, while respecting paywalls and password protections, along with related textual descriptions. The RIAA’s complaint framed that process in harsher terms, alleging that Suno copied vast quantities of the world’s most popular recordings and fed those copies into its AI models so the system could produce outputs that mimic human-made songs. The trade group also accused the company of using unlawful stream-ripping from YouTube while bypassing anti-copying safeguards.
More broadly, this dispute reflects a growing shift across the AI industry. Companies developing generative tools increasingly do not deny using copyrighted material during training. Instead, many now argue that such use is legally protected under fair use or similar doctrines. Earlier reporting by 404 Media described similar mass copying practices involving other AI firms, including Nvidia and Runway ML.
Recent reporting from The Atlantic also described music datasets commonly used in AI training that rely on links to tracks hosted on YouTube or Spotify. Developers then use automated tools to retrieve the audio, sometimes bypassing the parts of the platforms designed to drive advertising revenue, subscriptions, or creator compensation. It is not known whether Suno relied on any of those specific databases.
Suno’s response and the breach claims
In a statement cited by 404 Media, a Suno spokesperson said the company’s AI models were trained on publicly available music files and related metadata accessible on third-party sites across the open web. The spokesperson also said Suno discovered a limited security incident in November 2025, investigated it quickly, and concluded that the exposure mainly involved outdated source code no longer in use. The company further said no sensitive personal data had been compromised and emphasized that it does not have access to customers’ full credit card numbers through Stripe.
Suno also said that, based on the limited customer information believed to be involved, individual notices were not required under applicable privacy laws. The company did, however, submit a training-data disclosure required under California law.
The hacker, identified as ellie.191, told 404 Media that access came through a single employee account compromised using the Shai-Hulud worm, a supply-chain attack that enabled the theft of GitHub and cloud credentials. The hacker said they obtained a customer list containing email addresses and/or phone numbers, as well as Stripe-related payment details depending on how users signed up. A sample of affected users reportedly confirmed that they had registered with phone numbers and said they were never informed of a breach. The hacker said there was no larger ideological motive, adding simply: “I like hacking everything.”
Questions about originality
Suno has also tried to address criticism that its system can produce songs too similar to existing music. A spokesperson said the company’s goal is to help users make original music rather than replicas of someone else’s work. According to Suno, it does not intentionally use artist names as a training metadata category and has invested in safeguards meant to reduce impersonation and other abuse, while developing methods for identifying AI-generated content.
Still, skeptics point to the central contradiction in the company’s position: a model trained on enormous quantities of commercial music may be designed for originality, yet still inherit the stylistic fingerprints of the recordings it consumed. That tension has made Suno a focal point in the wider debate over whether generative AI is a tool for creativity, a machine for imitation, or both.
The controversy has only been sharpened by an earlier remark from Suno founder and CEO Mikey Shulman, who said on a podcast last year that he believes most people do not enjoy much of the time they spend making music. In light of the company’s legal battles and the newly reported leak, that comment now lands in a much more contentious debate about authorship, labor, and the future of music itself.
Comments