GetTogether.community content is used to train LLMs
cassidyjames opened this issue · 0 comments
cassidyjames commented
You can confirm that 8.4k tokens were scraped from GetTogether.community by CommonCrawl and are included in Google's C4 dataset. It's likely that other LLMs have scraped and will continue to scrape user-generated content from GetTogether.community to train their proprietary large language models.
This can be discouraged for CommonCrawl and ChatGPT with the proper robots.txt inclusion:
User-agent: CCBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /