The Great Data Harvest and Moving Away From Github



Some time ago, back in November 2025, I sought out a post in the Community Discussion forums of GitHub asking about their policy regarding what data they use in their training set for Copilot.

For some time, they’ve been explicit about using publically available source to train their models. It’s being done, and as far as I’m aware, It’s considered standard practice to use any publically published material, even if it was originally pirated. The assumption the industry is running under is that anything publically availble falls under fair use. No matter the license.

The legality around this is a question that’s still not been fully figured out in any of the world’s jurisidictions. I’ve only seen a couple of cases here and there dealing with the topic, and as things stands the knife may end up falling on either side of the debate.

But, coming back to the Community Discussion: What I was curious about is their public policy regarding use of private repositories. Considering the current economic incentives and the race to be first to build a full software developer replacement bot, those private repos must look like really appealing feed for the algorithmic beast to chew on. Sadly, there is no clear answer to find. The only explicit user they mention are the Enterprise level users. But any standard pro or smaller business are left in the dark.

With this in mind, since some time ago, I’ve decided to migrate away from GitHub for business and personal use. Transferring all my data to the only known safe haven: self hosting on my own hardware.

The value of most types of data has suddenly increased, in the past it was primarily data that would tie people together and reveal behaviour, wants and needs that people feared would be harvested for ads, but this new era of constructing automata through assimilation and mimicry has made almost all of our output the target for data colonisation.

Some people may not find the idea of having their expressions and creation poorly cloned without consent with the goal of replacing you as jarring as I. I’d argue that reflects poorly on the value of their work.

But I believe we should, and people increasinly will feel this.

Listening to Karen Hao describe the desperation in which people are being made reduntant by automated systems only to later end up in a new job where their task is to feed the same machine their knowledge to increase its capability to do the same to their collegues is a very grim look in to the current state of some industries.

I suspect we’re heading in to an era of an internet increasingly barren, Its content replaced by worthless facsimile, as people flee the eye of Altman seeking the one ring. Talk about a solid self-fullfilment of the Dead Internet Theory. I don’t really see any other recourse for people who wish to avoid exploitation. That is at least until the day we’ve figure out a new social contract for what we are owed for our efforts in society.

So, goodbye Cloud. Never really like you in the first place, but I begrudgingly accepted you out of convenience. Now the calculus feels completely different. As it stands, there are few companies I feel any business owner can trust not to sell the data for training. With how legislation is leaning and legal interpretation is being made, It truly is a case of a data free-for-all. If you have it, you can train on it, no matter how you got it.

Visual-Structural template by Lukas Orsvärn.