M3 Studios NewsM3 StudiosHouston, TX
Back to M3 Studios News

Was Your Music Used to Train AI? Now You Can Search

M3 StudiosSpring, TX5 min readAugust 1, 2026

You can now type your own name into a database and find out whether your recordings were used to train generative AI music models. The Atlantic published a searchable tool called AI Watchdog, built on research by reporter Alex Reisner, that opens up four of the major music datasets circulating in AI development. Two of them hold roughly 12 million and 9 million tracks. The other two each hold more than 100,000. You can search by artist name, by song title, or by ISRC, and the results span major catalogues, independent releases and underground production alike. For a Houston artist who has been arguing about AI in the abstract for two years, this is the first tool that answers the question about your own music specifically.

Reisner reports the datasets have already been downloaded thousands of times. Pinning down exactly who used them is not possible from the outside, but Google and Stability AI have both acknowledged drawing on at least some of the material in published research papers. That combination, wide distribution plus partial corporate acknowledgement, is what makes a search result worth documenting rather than just worth being angry about.

The detail almost every write-up is missing

Here is the mechanic that reframes the whole story. These datasets are largely not collections of audio files. Several are distributed as lists of links to tracks hosted on platforms like YouTube and Spotify. Developers then run automated tools to pull the audio down at scale, and Reisner notes that in some cases those tools bypass logins, advertisements and other mechanisms that exist specifically to generate revenue for creators. In his words, such tools violate the terms of service of those platforms.

Read that twice, because it means the injury has two layers rather than one. The first is the training use itself, which is the part currently being fought over in court. The second is quieter and more concrete: the audio was harvested through the very platforms that are supposed to pay you per play, using methods designed to skip the ads and the sign-ins that generate that payment. Your song was pulled and the play did not count.

A dataset built from links is not a copy of your catalogue. It is a set of instructions for taking your catalogue, repeatedly, from the platforms that owe you money for each one.

Why the ISRC is the part that decides whether you can even look

The tool accepts three search keys, and the difference between them matters more than it sounds. An artist name search depends on how your name was entered by whoever uploaded the track. A title search depends on exact spelling and collides with every cover and remix that shares a name. The ISRC is the only key that identifies a specific recording without ambiguity, because that is what it was designed to do.

Which produces an uncomfortable but useful conclusion. An artist who registered ISRCs properly and kept a list can audit an entire catalogue in a few minutes. An artist who never assigned codes, or who let a distributor assign them and never recorded which was which, is now trying to search for recordings they cannot precisely name. Metadata discipline, the least glamorous administrative task in independent music, has quietly become a rights-enforcement tool. If your codes are not in order, the ISRC and UPC breakdown covers what they are and how to get them assigned correctly.

What a search result actually is, and what it is not

A hit in one of these datasets is evidence that a recording appeared in a compilation circulating among AI developers. It is not proof that a specific company trained a specific model on your specific song, and it is not a legal claim on its own. Treating it as one, firing off demand letters based on a screenshot, is how an artist spends money to accomplish nothing.

What it is worth is documentation, done now, while the record is fresh. Search your name and every ISRC you control. Screenshot each result with the date visible. Note which dataset the hit came from, because they are not interchangeable and some carry their own licensing terms. One of the sources, the Free Music Archive dataset, is available for personal listening but requires licensing for commercial use, which means the terms attached to a track can matter as much as its presence.

Keep that file with your registrations rather than in a phone camera roll. Litigation, licensing schemes and settlement funds all eventually require someone to prove which recordings they own and when they learned about a use. Artists who can produce a dated, organized record are in a different position from artists who remember being upset in July.

What the audit looks like on a real catalogue

Put numbers on it so the task stops feeling abstract. An independent artist with six years of releases might have thirty to fifty recordings carrying ISRCs, including singles, album tracks, features and the remix nobody remembers agreeing to. Searching that catalogue is a single sitting, because each code is a paste and a look, and the output is a short list of which recordings appear where and which do not appear at all.

That list is more informative than the headline that sent you looking. If the hits cluster on the releases that got the most streaming traction, you learn that the harvesting tracked popularity, which is what you would expect from tools pulling from platform links. If a track appears that you never uploaded to a major platform, you have learned something about how it circulated. And if an unreleased or privately shared recording appears, that is a materially different situation from a commercially released one, and it is worth documenting with particular care and showing to an attorney rather than posting about.

The artists who will struggle are the ones whose catalogue lives across three distributors, two abandoned artist names and a folder of untitled bounces. That is not a moral failing, it is the normal condition of a working independent career, and it is fixable in one afternoon of list-building that pays off every time a registration, a claim or a licensing scheme requires an inventory. Making that list is the actual assignment here. The search is the easy half.

The pattern this fits into

Every serious fight in the music business has followed the same shape. A new technology makes copying cheap, the copying happens at scale before the law arrives, and the artists who eventually get paid are the ones whose ownership was documented before the payout mechanism existed. That was true of mechanical royalties for physical copies, it was true when streaming arrived and left a black box of unmatched money behind, and it is true here.

The difference this time is speed. The datasets exist, the downloads already happened, and now there is a public search box. That compresses the useful window down to the next few weeks of housekeeping. Registration, split documentation and clean codes are the entire toolkit, and registration is the step most independent artists skip while assuming the copyright exists automatically, which it does, in a form that is much harder to enforce.

What to do this week, in order

Pull your ISRC list, or build one from your distributor's back catalogue export if you never kept it. Search AI Watchdog by every code and by your artist name and any alternate spellings you have released under. Screenshot and date whatever comes back, note the dataset, and file it with your copyright registrations and split sheets. Then check that anything unreleased or unregistered gets a code and a registration before it goes out, since the point of the exercise is that the next release should be documented from day one.

Nothing in that list requires a lawyer or a budget. It requires an afternoon and a habit. If the administrative side of a catalogue is the part that has never been handled properly, the publishing and royalty material in the M3 Studios education library covers the registration and metadata chain end to end, and future sessions are worth finishing with the codes and paperwork attached rather than promised.

As of July 29, 2026. Dataset contents and tool availability can change. A search result is information, not legal advice, and specific claims should be reviewed by a qualified attorney.

Frequently asked questions

How do I check if my music was used to train AI?

Use The Atlantic's AI Watchdog tool, which makes four major music datasets circulating in AI development searchable. You can search by artist name, by song title, or by ISRC. The ISRC gives the most reliable result because it identifies one specific recording, while name and title searches depend on how the metadata was entered.

What is in the AI music training datasets?

Four datasets identified in Alex Reisner's reporting for The Atlantic. Two contain roughly 12 million and 9 million tracks. The other two each contain more than 100,000. They span major label catalogues, independent releases and underground production, and Reisner reports they have been downloaded thousands of times.

Are the datasets actual audio files?

Largely no. Several are distributed as lists of links to tracks hosted on platforms such as YouTube and Spotify. Developers use automated tools to download the audio at scale, and Reisner reports that in some cases those tools bypass logins and advertisements, which he notes violates the terms of service of those platforms.

Does finding my song in a dataset mean I can sue?

No. A hit shows a recording appeared in a compilation circulating among AI developers. It is not proof that a particular company trained a particular model on that recording, and it is not a legal claim by itself. The practical value is documentation: dated screenshots, the dataset name, and a record kept alongside your copyright registrations.

Which AI companies have acknowledged using this material?

Google and Stability AI have both acknowledged drawing on at least some of the material in published research papers. Beyond that, the datasets' wide distribution makes it difficult to determine from the outside exactly who has used them.

Follow M3 Studios for the money and rights mechanics Houston artists actually use: Instagram @metamusicmediainc, TikTok @metamusicmediainc, YouTube @metamusicmediainc. Questions: info@metamusicmedia.com. Finish the next one with the codes and paperwork attached, in Spring, TX: metamusicmedia.com/pages/book-your-session.

  1. The Atlantic, AI Watchdog, the searchable databases of material identified in generative AI training, reporting and research by Alex Reisner. https://www.theatlantic.com/category/ai-watchdog/
  2. MusicTech, "Was your music used to train AI? This free tool will tell you," June 2026 (the four datasets and their approximate sizes, the thousands of downloads, the Google and Stability AI acknowledgements, the lists-of-links distribution method, the terms-of-service point, and the Free Music Archive licensing note). https://musictech.com/news/music/is-your-music-used-to-train-ai-tool/
  3. Resident Advisor, "A new tool allows artists to check whether their tracks are used in AI," June 2026 (the tool's launch and search options). https://ra.co/news/85456
  4. United States Copyright Office, registration basics for sound recordings and musical works. https://www.copyright.gov/registration/
  5. International ISRC Agency, ISRC handbook and code assignment. https://isrc.ifpi.org/
Continue readingMore from M3 Studios News
Apple Music
Aug 1, 2026

Why Apple Music Raised Prices, and Where the New Dollar Actually Goes

blue dot fever
Jul 31, 2026

Blue Dot Fever: Why Tours Keep Dying in Music's Biggest Year Ever

City's Initiative
Jul 31, 2026

Houston Arts Grants: Two Rounds Open August 28

© Meta Music Media Inc · M3 StudiosM3 Studios logo mark, Houston recording, mixing and visual production studio, Meta Music Media Inc↑ Back to top
M3 Studios News · Houston, TX