Machines can authenticate with long random keys stored in software or hardware. People need a practical way to control those credentials or prove their identity. They might remember a password, carry an authenticator, or unlock a device with a fingerprint.
Passwords do not require a dedicated authentication device, and they can be replaced when forgotten. Their weaknesses follow from how people choose and use them: short passwords are guessable, reused passwords expose several accounts, and a stolen password can be entered by anyone.
Three related terms are easy to misuse:
-
Identification is claiming an identity, as when typing a username.
-
Authentication verifies that claim.
-
Authorization determines what an authenticated party is allowed to do. We will cover authorization when we cover access control.
Three Kinds of Evidence
Evidence of identity falls into three categories, known as authentication factors:
-
Something you know, such as a password or a PIN.
-
Something you have, such as a phone, a hardware token, or a smart card.
-
Something you are, such as a fingerprint or a face. These are biometrics.
Multi-factor authentication (MFA) requires factors from different categories. A password followed by a code from a phone is two factors. A password followed by a security question is one factor used twice, since both are things the user knows, and an attacker who can phish one can usually phish the other.
Sending the Password
The most direct protocol sends the password to the server, which checks it. The Password Authentication Protocol (PAP) was defined in 1992 for dial-up connections using the Point-to-Point Protocol (PPP), and it did just what it says. The client sent a username and password in the clear, and the server compared them with its records.
On an unprotected link, PAP exposes the password to anyone who can observe the traffic. A web password form uses a similar approach. The browser sends the password, and the server checks it. The form travels over HTTPS, where Transport Layer Security (TLS) encrypts the traffic and authenticates the server. This resembles PAP’s password submission, but it is not the PAP protocol.
TLS protects the password in transit, but the server still receives it. A compromised login service can capture passwords as users enter them, regardless of how stored passwords are protected. Malware on the user’s device or a lookalike login site can also capture the password before it reaches the intended server.
Challenge-Handshake Authentication
The Challenge-Handshake Authentication Protocol (CHAP), defined alongside PAP in 1992 and revised in 1996, applies challenge-response authentication to passwords. The server sends a random challenge. The client returns a hash of an identifier, the password, and the challenge, computed with MD5. The server computes the same hash and compares. MD5 is a cryptographic hash function that was widely used in the 1990s. Researchers showed in 2004 that collisions could be found quickly, and it is no longer considered secure.
The password never crosses the network. A recorded response is useless for a later login, because the challenge changes each time. CHAP also let the server issue new challenges at intervals while the link was up, to confirm that the same party was still connected.
CHAP requires the server to have the password, or an equivalent secret, in usable form so that it can compute the expected response. It cannot verify a response using only a one-way password hash. By contrast, a server receiving a password over TLS can check it against a stored hash, as described below.
CHAP also permits offline guessing. An eavesdropper can record the challenge and response, compute responses for likely passwords, and look for a match. Preventing replay does not prevent guessing a weak secret, and CHAP by itself does not encrypt the later data traffic.
Microsoft’s variant, MS-CHAPv2, was widely used for virtual private networks and enterprise Wi-Fi. In 2012, researchers demonstrated that a captured exchange could be attacked with work comparable to searching the 56-bit DES key space. Microsoft acknowledged the weakness and recommended protecting these exchanges inside a secure tunnel or using a stronger authentication method.
Storing Passwords
In December 2009, an attacker used a SQL injection flaw to copy the user database of RockYou, a company that made applications for social networks. The database exposed 32 million users’ passwords in plain text. The leaked passwords became a widely used dictionary for password-cracking tools.
A service that needs only to verify passwords should store a one-way hash rather than a recoverable password. The basic idea is to compute \(h = H(\text{password})\) when the password is set. At login, the server hashes the submitted password and compares the result with \(h\).
Preimage resistance does not protect a password that an attacker can guess. A stolen hash lets the attacker compute \(H(\text{guess})\) for likely passwords and compare the results without reversing the hash function.
Online and Offline Guessing
Password guessing takes two forms, and the defenses against them are different:
-
Online guessing submits guesses to the login service. The service can slow it down with rate limits, lock accounts after repeated failures, and log every attempt. Note that having the service lock an account can become a way for an attacker to perform an availability attack against a user.
-
Offline guessing uses a stolen hash or another captured value that can verify guesses, such as a CHAP response. The attacker tests guesses on its own hardware, outside the service’s rate limits and logging.
The storage defenses that follow are all aimed at offline guessing.
Dictionary Attacks
Attackers do not guess at random, since people do not choose at random. A dictionary attack tests likely passwords first: common words and names, keyboard patterns, passwords from earlier breaches such as the RockYou list, and variations produced by rules that apply common substitutions, capitalize the first letter, or append a year or an exclamation point.
The approach works because people’s choices have changed little over decades. In 1979, Robert Morris and Ken Thompson of Bell Labs, AT&T’s research laboratory, studied 3,289 Unix passwords and found that 86 percent were short strings, dictionary words, names, or other easily searched categories. In the 32 million RockYou passwords leaked in 2009, the most common choice was “123456,” used by about 290,000 accounts. Analyses of recent breaches keep finding the same patterns, and younger users choose passwords much like everyone else.
Precomputed Tables and Rainbow Tables
If every site hashed passwords with the same function and nothing else, an attacker could hash an entire dictionary once and store the results. Cracking a stolen file would then require nothing more than looking up each hash in the table, and the same table would work against every site.
Storing every candidate and its hash takes considerable space. A rainbow table trades some lookup computation for less storage. It keeps enough information to reconstruct groups of candidates when needed. Introduced in 2003, the technique works when a site hashes each password directly and the passwords come from a limited set.
LinkedIn hashed each password directly. In June 2012, about 6.5 million of its SHA-1 password hashes appeared online. An attacker could hash each candidate password once and compare it against the entire stolen collection.
Salt
Preventing reuse of precomputed hashes requires a different hashing input for each stored password. Unix has done this since at least 1979, when Bob Morris and Ken Thompson described a design that combined each password with a random value and repeated the hash computation to make each guess more expensive.
A salt is a random value generated when a password is set and stored alongside its hash. A conceptual example is \(H(\text{salt} \parallel \text{password})\), where \(\parallel\) means concatenation. Real password hashing functions take the password and salt as separate inputs, along with cost settings. At login, the server uses the stored salt and settings to check the submitted password.
Different salts make identical passwords produce different stored hashes. An attacker therefore cannot spot password reuse by comparing hashes, or use one precomputed table against the whole database. Each guess must be evaluated separately for each salt. The attacker can still reuse a list of candidate passwords, but not the hash computations.
A salt is not secret, and it does not slow down a guess against a single account. An attacker targeting one user reads the salt from the stolen file and hashes guesses with it. The defense against that is a slow hash function.
For local accounts on many Linux systems, /etc/shadow stores the password hashes. A hash field beginning with $y$ uses yescrypt, which Debian 11 adopted as its default. The field also records the cost settings and salt. Existing accounts can retain older formats until their passwords change.
Slow Hash Functions
Cryptographic hash functions are designed to be fast, and for password storage that is a defect. The legitimate server computes one hash per login. An attacker computes one per guess, and a fast function helps the attacker far more than the server.
A 2022 benchmark measured about 22 billion SHA-256 computations per second on one Nvidia RTX 4090 graphics card. At that rate, testing all \(26^8\), or about 209 billion, eight-letter lowercase strings would take under ten seconds. This is an estimate for raw SHA-256 on that hardware, not for a dedicated password hashing function.
A password hashing function is deliberately expensive to compute, with a cost that can be raised as hardware improves. Three are in wide use:
-
bcrypt (1999) repeats an expensive key setup from the Blowfish cipher. Each increment of its cost parameter roughly doubles the work. The same benchmark measured 184,000 guesses per second at cost 5. Raising the cost to 10 multiplies the work by 32, suggesting about 420 days for the full lowercase search on that card. Likely passwords could still be found much sooner.
-
scrypt (2009) requires a configurable amount of memory for each hash. An attacker running many guesses in parallel must supply memory and memory bandwidth for them, increasing the hardware cost.
-
Argon2 won the Password Hashing Competition, a public selection process, in 2015. It has adjustable time, memory, and parallelism settings. The Open Worldwide Application Security Project (OWASP), a software security nonprofit, recommends Argon2id for new password storage.
Password-Based Key Derivation Function 2 (PBKDF2), standardized in 2000, commonly repeats HMAC for a configurable number of iterations. It does not require large amounts of memory, so it offers less resistance to parallel hardware than a suitably configured scrypt or Argon2. It remains useful where compliance requirements call for an approved implementation of this method.
A service might choose settings that take about 100 milliseconds on its own hardware, then measure the effect under realistic login loads. Attackers can use faster or parallel hardware, so that server-side timing does not imply the same delay for each attacker’s guess.
Credential Stuffing
People reuse passwords across sites. Credential stuffing takes username and password pairs leaked from one site and tries them against other sites, relying on that reuse.
Credential stuffing is an online attack that does not require a flaw in the target’s password storage. A strong password hash cannot prevent a login made with the correct password.
In October 2023, 23andMe, a consumer genetic testing company, disclosed unauthorized access to customer accounts. Its December disclosure attributed access to 0.1 percent of accounts to credentials reused from other breaches.
Those accounts could view information that other users had shared through DNA Relatives and related features. The attackers used that access to collect millions of additional profiles. A person’s information could therefore be exposed through a relative’s compromised account even if the person’s own password was unique.
Password Spraying
Per-account lockouts limit repeated guesses against one account. Password spraying tries a few common passwords, such as “Winter2025!” or “Welcome1,” against many accounts. Keeping each account below the lockout threshold lets the attacker continue searching for users who chose one of those passwords.
In late November 2023, a group that the U.S. and UK governments attribute to Russia’s foreign intelligence service, tracked by Microsoft as Midnight Blizzard, used password spraying to break into a legacy test account at Microsoft. The account was not used for production and did not have multi-factor authentication. From there, the attackers reached email accounts belonging to members of Microsoft’s senior leadership. Microsoft detected the intrusion in January 2024.
Mitigations
Each password attack has its own defenses, and no single measure covers all of them:
| Attack | Defenses |
|---|---|
| Eavesdropping | Encrypting the channel with TLS, and never sending a password over an unprotected link |
| Theft of the password file | A salt for every account, and a slow password hashing function with a high cost setting |
| Dictionary guessing | Rejecting common and previously breached passwords when a user sets one |
| Online guessing and spraying | Rate limits across accounts and source addresses, not only per account, and detecting failures spread across many accounts |
| Credential stuffing | Unique passwords, multi-factor authentication, breached-password checks, and detection of unusual logins |
Password managers help users generate and store a different random password for each site. A breach of one site’s password database then does not supply a working password for the others.
Checking against breach data does not require sending a full password or hash to a third party. Have I Been Pwned, a breach notification service, accepts the first five hexadecimal characters of a password’s SHA-1 hash. It returns matching hash suffixes, and the site completes the comparison locally. SHA-1 serves as a lookup identifier here, not as the site’s password storage function.
Password policies have also changed. Many organizations required mixtures of uppercase letters, digits, and symbols, and forced changes every 90 days. Such rules encouraged predictable choices, such as “Password1!” followed by “Password2!” three months later.
In July 2025, the U.S. National Institute of Standards and Technology (NIST) published SP 800-63B-4, its updated authentication guidelines. The password requirements include:
-
At least 15 characters for a password used on its own, and at least 8 for a password used as one factor of multi-factor authentication.
-
No composition rules requiring particular mixtures of character types.
-
No required periodic changes. Require a change when there is evidence of compromise.
-
A check of every new password against a list of common, expected, and compromised passwords.
These guidelines govern covered federal systems and also inform other organizations’ policies. At authentication assurance level 2, which requires two factors, they require offering a phishing-resistant method. At level 3, phishing resistance is mandatory. They also restrict the use of codes delivered through the telephone network because of risks such as number reassignment and interception. Rutgers does not pay attention to the “no required periodic changes” recommendation.
Next: Part 5: Beyond Passwords
Lecture 4: Part 1 | Part 2 | Part 3 | Part 4 | Part 5 | Part 6 | Appendix
Lecture 4 Study Guide | List of terms