this post was submitted on 05 Oct 2026
88 points (92.3% liked)
Programming
28764 readers
452 users here now
Welcome to the main community in programming.dev! Feel free to post anything relating to programming here!
Cross posting is strongly encouraged in the instance. If you feel your post or another person's post makes sense in another community cross post into it.
Hope you enjoy the instance!
Rules
Rules
- Follow the programming.dev instance rules
- Keep content related to programming in some way
- If you're posting long videos try to add in some form of tldr for those who don't want to watch videos
Wormhole
Follow the wormhole through a path of communities !webdev@programming.dev
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I tried to implement my own git from scratch. And when I got to choose a hashing algorithm, I reached the exact same conclusion as the author pretty fast.
The hashing algorithm doesn't really matter as long as it's good enough to prevent accidental collisions. And if there is a collision, it is extremely easy to check.
If you are about to create a new object and an object with that hash already exists, you compare the 2 objects. If they are the same, nothing happened. If they are different, you just found a collision. Throw some obscure error that will only be witnessed once in the lifetime of the universe and be done with it. Tell the user to change a single bit of the input and be done with it.
The purpose of object hashes in git was never to provide security. Security is achieved through other means.
Let's explicit this: you can, and should, sign your commits if you want security (no commit tampering). You need to:
gpg --full-generate-keygit config --global user.signingkey [public key ID]git config --global commit.gpgsign trueNote that I used
--globalhere but you can do without to sign on a per-project basis.Also
gpgUX is terrible, to find the[public key ID]use the commandgpg --list-keys, it should looks like:Last, if you want to save/backup you keys the default location of
gpgfiles is~/gnupgor you can export keys withgpg --export [key ID] > path/to/file.keyfor the public part andgpg --export-secret-keys [private key ID]for the private part. Both export can have a-aor--armorargument to output as base64 text instead of raw binary.[private key ID]can be found withgpg --list-secret-keys.EDIT: after writing this I checked https://git-scm.com/book/en/v2/Git-Tools-Signing-Your-Work and TIL you can sign tags too.
EDIT 2: you can also use your ssh key to sign ->
git config --global gpg.format ssh(I am less inclined to do this as you won't have subkeys and expiration date but this would be simplier indeed)If you want to avoid the gpg mess, you can actually also configure gut to sign with your SSH key.
Worth to keep in mind though that the
gpgsignature basically only signs the (git) hashes of the contained objects, so the choice of hash function is indeed important for security.~~Indeed it is still relevant~~. I am wondering what does it sign for tags.
EDIT: Are you sure about that? signing the hash felt logical but digging about it this seems a false assumption. git is signing the whole commit object.
Yes, but the "commit object" is just a bunch of metadata that refers to the tree by its hash:
And the tree in turn refers to the files (or subfolders) by their hashes:
If you can produce a hash collision, you can therefore have two repositories with the same signed commit but different file content.
Yes you are right, damned! Why git isn't signing the diff (patch)? This is so wrong on so many level, signing already hash the content, we can sign multiple GiB without any issue.
Hm, I wouldn't call it wrong on many levels, I think it's quite elegant. It's a form of a Merkle Tree, and knowing the hash of a commit does not just allow you to verify its content, but also the whole history up to that point (as the commit contains the hash of its parent, which in turn contains hashes of its content and parents). I guess that's not too unimportant if you consider scenarios where you pull code from distributed repositories (like forks), as you can ensure that the parts of the history you know are actually what they claim to be. And if you accept hash-then-sign as being secure, this is just as secure for signed commits (assuming the hash function is secure).
I think the only unfortunate part here is that git ended up using SHA-1 as a hash function, which 20 years ago might or might not have been a reasonable choice -- I don't know how "broken" it was regarded back then, or how widespread SHA-256 use was.
Even if it's not the purpose it can fulfill it. Hashes are used for all kinds of things like traceability or for pinning versions. They are used in a context where people depend on that a hash always has the same code behind it and SHA1 can not guarantee that anymore. A better hashing algo is needed to keep the way how people use git secure.
It's not enough to say "you shouldn't use it that way". We have to accept the reality of how git is used and adapt.
If someone has access to write to a repo you are downloading code from, and you don't trust that someone, you shouldn't be running the code in that repo.
The only "security" vulnerability is: I have audited and verified that the code in commit 05682ab36ca. And I will use only that version.
What you do then is: download that commit and store it in a local repo, and package it so you can distribute it to your clients.
Auditing external software is tremendous effort, even if it is open source. Using your local repo instead of downloading that commit from GitHub every time is very little extra effort in comparison.
And if you really need it. You can always hash with sha256 yourself and verify that the commit with that sha1 still produces the same sh256. So you can keep downloading it each time, just need to store the sha256.
You are exactly doing what I described. You are saying "your using it wrong" instead of acknowledging that's how people use it and and then improving that. You might not need this feature because you are "using it correctly" but that doesn't matter. For the lived reality of many people sha256 is the correct move.
I've never heard of anyone relying on git's sha1 hashed for security
I have. In some languages you can use git repos as dependencies. It's good practice to use hashes instead if tags because tags are not immutable. SHA1 hashes are better than tags but better hases would be even better.
Also traceability. You can build a chain from your commit over ci to the artifact if you do it right. If you can change the files to a hash you can destroy that trust chain. Sha256 makes that impossible.
People use the hashes in ways you might not. It wasn't meant as a trust anchor but it became one. And that is now being addressed. I also think that the git maintainers have put a lot of thought into backward compatability and making the change as painless as possible.
Good practice != Something that ensures security. As I said, it's an even better practice to just host your own fork if what you want is security. The reason to use git hashes is mostly so you don't have to trust the author following semver in there version number. It's a matter of ensuring that your code will always compile, not a security feature.
Have you read the article? Your arguments derive from an assumption of "sha1 is mutable, sha256 is immutable". First of all, every hashing algorithm is going to have collisions, that's an unavoidable "feature" of hashes. And sha1 is in no way "mutable", it takes a great amount of effort (and money) to generate a collision. Furthermore, those collisions are not arbitrary. You have to calculate them beforehand. As the article says: what is more likely? Paying 40k€ in compute time to generate a single collision? Or just paying an open source maintainer 40k€ to let you do 1 new commit that most people are going to download anyway?
Yes I read the article. I don't agree with all the conclusions.
Now is the best time to change to a new algo. SHA1 is not completely broken yet. It most probably will be at some point.
Also AFAIK they will have compatibility tools for SHA1 repos in git. I don't see the big deal with the path they chose.
We will probably nor agree on that one. I think the hashes are important and that they are mutable right new with effort and later without. It's okay that we don't agree though. The git maintainers have made their decision anyways.