A cryptographic hash function takes an input of any size (a password, a paragraph, a ten-gigabyte video file) and produces a fixed-length string of characters called a digest. Feed it the same input twice and you get the same digest every time. Feed it two inputs that differ by a single character and the digests come out looking completely unrelated, with roughly half the bits flipped. That property is called the avalanche effect, and it's what makes hashes useful: a tiny change anywhere in the input should be undetectable by looking at the input alone but glaringly obvious in the output.
Hashing only goes one direction. There's no key, no inverse function, nothing to "decrypt." Given a digest, the only way to find an input that produces it is to guess inputs and hash them until one matches, which is exactly what a password cracker does, and exactly why hashing a password directly, with no salt and no deliberate slowness, is a bad idea even with a strong algorithm behind it. That last point matters more than most guides give it credit for, and we'll come back to it.
The three algorithms, by the numbers
MD5, SHA-1, and SHA-256 differ mainly in output size and internal construction, and the output size alone tells you a lot about how hard each one is to brute-force blindly.
- MD5: 128-bit output, shown as 32 hex characters. Designed in 1991 by Ron Rivest. Fast, small, and the first of the three to fall.
- SHA-1: 160-bit output, 40 hex characters. Designed by the NSA, published in 1995. Deprecated by NIST in 2011, formally broken by a real demonstrated attack in 2017.
- SHA-256: 256-bit output, 64 hex characters. Part of the SHA-2 family, published in 2001. Still holds up against every known practical attack, which is why it's the current baseline for security-sensitive hashing.
A bigger output space isn't the whole story. MD5 wasn't broken because 128 bits became brute-forceable in the naive sense; it was broken because researchers found structural flaws in its internal compression function that let them construct collisions far faster than brute force ever could. Size matters, but the internal design is what actually held or didn't.
What "cryptographically broken" actually means
A hash is considered broken, in the cryptographic sense, when someone finds a way to produce a collision (two different inputs that hash to the same output) faster than brute force says should be possible. For MD5, that's not a theoretical worry anymore. Researchers Xiaoyun Wang and colleagues published a practical collision-finding method for MD5 in 2004, and the attack only got faster from there. By the early 2010s, chosen-prefix collisions, where an attacker picks the meaningful content of both files in advance and then appends garbage that forces a matching hash, could be generated on ordinary hardware in seconds. Is MD5 still safe to use? For anything an adversary might want to forge, no, and the clearest proof of that showed up in the wild, not just in a research paper.
The Flame malware, discovered in 2012, used a forged Microsoft digital certificate to sign itself as legitimate Windows code. The forgery relied on an MD5 collision attack against Microsoft's own certificate authority infrastructure: an advanced, previously unknown variant of the chosen-prefix technique, likely developed by a nation-state actor. That's the practical difference between "theoretically weak" and "broken." MD5's weakness was used to compromise real code-signing trust in a real, high-profile incident, not just demonstrated in a lab.
SHA-1 took longer to fall but got there too. In February 2017, Google and CWI Amsterdam published SHAttered, the first publicly demonstrated SHA-1 collision: two different PDF files that hash to the identical SHA-1 digest. Producing it took roughly 9,223,372,036,854,775,808 SHA-1 computations (yes, that's the actual number Google published) and around 6,500 CPU-years plus 100 GPU-years of compute, run in parallel over a few months on Google's infrastructure. That's a lot less than a naive brute-force collision search would need, which is precisely the point. The mathematical shortcut, not the raw compute, is what made SHA-1 broken rather than merely old.
Why MD5 and SHA-1 still show up everywhere
Here's the part that trips people up: collisions matter enormously in some contexts and not at all in others. Whether MD5 vs SHA-256 (or SHA-256 vs SHA-1) matters for your use case comes down to one question. Is there an adversary deliberately trying to construct a matching input, or are you just trying to catch accidental damage?
If you're checking whether a file downloaded correctly, or deduplicating files on a backup server, or comparing a version you saved yesterday against today's copy, nobody is trying to trick you. Random bit-flip corruption doesn't care about MD5's structural weaknesses; it just needs any change to produce a different hash, and MD5 still does that reliably. That's why plenty of download pages, backup tools, and internal deduplication systems still quote an MD5 checksum next to a file. It's fast, it's short, and there's no attacker in the threat model.
Git is the example people bring up most, and it's worth being precise about it: Git historically used SHA-1 to name every commit, tree, and blob by the hash of its content. That's content-addressing, a way to detect accidental changes and let identical content collapse to the same object, not a defense against a forger. Git has since moved to hardened SHA-1 (which detects the specific structure SHAttered-style attacks rely on) and has an opt-in SHA-256 mode, but the original design choice made sense for its actual threat model at the time. A checksum for accidental corruption and a hash defending against a motivated forger are different jobs, even when the same algorithm gets used for both.
When SHA-256 (or better) is non-negotiable
Once there's an adversary in the picture who benefits from forging a match, MD5 and SHA-1 are off the table and SHA-256 becomes the minimum. Digital signatures, TLS certificates, and code-signing all fall here, anywhere a forged collision would let someone impersonate a legitimate file, site, or piece of software. Browsers stopped trusting SHA-1 certificates years ago for exactly this reason, and any CA still issuing them would be a serious problem.
Passwords deserve a paragraph of their own, because "use SHA-256 for passwords" is advice that sounds right and is still wrong. SHA-256 is fast, and that's a feature for file integrity but a liability for passwords, because a fast hash lets an attacker with a stolen database try billions of guesses per second against it. What password storage actually needs is a slow, salted, purpose-built algorithm: bcrypt, scrypt, or Argon2. These deliberately cost real time and memory per guess, which is irrelevant when you're hashing one password at login but devastating to an attacker trying to hash billions. So the accurate takeaway isn't "SHA-256 good, MD5 bad." It's that general-purpose hash functions, even secure ones, were never the right tool for password storage in the first place.
Try it
GlaeKit's Hash Generator computes MD5, SHA-1, and SHA-256 digests from any text, entirely in your browser via the Web Crypto API. Nothing is uploaded anywhere.
Frequently asked questions
Is MD5 still safe to use for file checksums?
For detecting accidental corruption or comparing files where nobody is trying to fool you, yes. MD5 still reliably changes output when input changes. It's not safe anywhere an attacker could benefit from crafting a file that matches a target hash, since practical collision attacks against MD5 have existed since the mid-2000s.
Why is SHA-1 considered broken?
In 2017, Google and CWI Amsterdam publicly demonstrated the first real SHA-1 collision (nicknamed SHAttered): two distinct PDF files with identical SHA-1 hashes. The attack needed far less compute than a brute-force search would require, which is the actual definition of "broken" in cryptography: a shortcut exists that shouldn't.
Should I hash passwords with SHA-256?
No, not on its own. SHA-256 is secure but fast, and fast is the wrong property for password storage since it lets an attacker with a leaked database try enormous numbers of guesses per second. Use a slow, salted algorithm designed for passwords instead: bcrypt, scrypt, or Argon2.
What does "collision" mean in hashing?
A collision is two different inputs that produce the identical hash output. Collisions are mathematically inevitable for any hash function once you look at enough inputs, since the input space is infinite and the output space isn't. A hash is "broken" specifically when someone finds a way to construct a collision on purpose, faster than random chance would predict.
Is SHA-256 vs SHA-1 the only real comparison that matters today?
For new security-sensitive work, largely yes. SHA-256 (or the SHA-3 family, or SHA-512) is the practical baseline, and SHA-1 shouldn't be chosen for anything new. MD5 and SHA-1 both still turn up in legacy systems and non-security roles like checksums, so recognizing them and knowing when they're actually a problem still matters day to day.