GetTogether.community content is used to train LLMs

Question

GetTogether.community content is used to train LLMs

cassidyjames opened this issue 2 years ago · 0 comments

You can confirm that 8.4k tokens were scraped from GetTogether.community by CommonCrawl and are included in Google's C4 dataset. It's likely that other LLMs have scraped and will continue to scrape user-generated content from GetTogether.community to train their proprietary large language models.

This can be discouraged for CommonCrawl and ChatGPT with the proper robots.txt inclusion:

User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /