The Great Data Harvest and Moving Away From Github
Some time ago, back in November 2025, I sought out a post in the Community Discussion forums of GitHub asking about their policy regarding what data they use in their training set for Copilot. For some time, they’ve been explicit about using publically available source to train their models. It’s being done, and as far as I’m aware, It’s considered standard practice to use any publically published material, even if it was originally pirated.