logoalt Hacker News

kstrauser • today at 2:37 AM • 3 replies • view on HN

Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.

Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.

I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.


Replies

schacon • today at 6:17 AM

So, I haven't worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it's never been a real problem afaik.

But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"

All parts of this are incorrect.

You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.

tanoku • today at 9:35 AM

> I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is.

All the details on how GitHub's infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no "secrets" or "rumors" here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.

- https://github.blog/engineering/architecture-optimization/in... - https://www.youtube.com/watch?v=Ri8hSZNKzu4 - https://github.blog/engineering/building-resilience-in-spoke... - https://github.blog/open-source/git/counting-objects/ - https://www.youtube.com/watch?v=DY0yNRNkYb0 - https://cursor.com/blog/git-at-any-scale

hedora • today at 3:14 AM

I'd expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and "try to prove this true or false; here's a github PAT".

➕ show 1 reply