You write a regex maybe once a month, if that. Long enough between uses that you forget whether it's {2,4} or {2:4}, whether a lookahead is (?= or (?:, and whether you need to escape a dot inside a character class. Then you spend ten minutes on Stack Overflow relearning something you already learned three times before. This page exists to shortcut that: the building blocks, a set of ready-to-use patterns for the things people actually need regex for, and the gotchas that make a pattern silently wrong or, worse, hang your tab.
This is a regex cheat sheet, not a tutorial on finite automata. If you want to understand why regex works the way it does, there are good books for that. If you want to remember what \b means at 4pm on a Friday, keep reading.
The building blocks
Almost every pattern you'll write is some combination of these five ideas: what kind of character you're matching, where in the string you're anchored, how many times something repeats, how you group and capture pieces, and occasionally a lookahead or lookbehind to match based on context without consuming it.
| Syntax | Meaning |
|---|---|
\d | A digit, 0-9 |
\w | A "word" character: letters, digits, underscore |
\s | Whitespace: space, tab, newline |
. | Any character except a newline (unless the s flag is set) |
^ / $ | Start / end of the string (or of each line, with the m flag) |
\b | A word boundary, between a \w character and a non-\w character |
* / + / ? | Zero or more, one or more, zero or one |
{n,m} | Between n and m repetitions ({3} exactly three, {3,} three or more) |
(...) | A capture group: matched text you can extract or back-reference |
(?:...) | A non-capturing group, used for precedence without saving the match |
a|b | Alternation, matches a or b |
(?=...) | Positive lookahead: must be followed by this, but it isn't part of the match |
(?<=...) | Positive lookbehind: must be preceded by this, without consuming it |
Lookahead and lookbehind get overused. They're the right tool when you need to match based on what surrounds text without including that surrounding text in the result, say, a number followed by "px" where you don't want "px" in your capture. For most day-to-day matching, a plain group does the job with less mental overhead.
Common regex patterns, ready to use
These are the patterns that come up constantly enough to be worth keeping somewhere. Test each one against your actual data before shipping it: every "universal" pattern below has edge cases, and the notes explain where.
Regex for email validation
^[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}$
Good enough for form validation. It is not, and cannot be, a complete implementation of the email address spec (more on that below).
URL matching
https?:\/\/[\w.-]+(?:\.[a-zA-Z]{2,})+[\w\-._~:/?#[\]@!$&'()*+,;=%]*
Covers the common case: scheme, host, optional path and query string. Real URLs allow characters this doesn't anticipate, so treat it as "find probable links in text," not a strict validator.
Phone number (US-style, flexible formatting)
^\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}$
Matches (555) 123-4567, 555.123.4567, and 5551234567. International numbers need a different pattern entirely; formats vary too much to fake with one regex.
Extracting a hex color code
#([A-Fa-f0-9]{6}|[A-Fa-f0-9]{3})\b
Put the six-digit alternative first. Regex alternation tries options left to right and stops at the first match, so if the three-digit form came first, it would grab just the first three characters of a six-digit code.
Matching a date in YYYY-MM-DD format
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$
Validates the month is 01–12 and the day is 01–31. It doesn't know February has 28 days, or that April doesn't have 31; regex can't count that way. If the date needs to be real, parse it with a date library after the format check passes.
Trimming whitespace
^\s+|\s+$
Replace matches with an empty string to strip leading and trailing whitespace. Most languages have a built-in trim() now, so you rarely need this yourself, but it's handy inside a larger find-and-replace pass over a whole document.
Splitting on commas outside quotes
,(?=(?:[^"]*"[^"]*")*[^"]*$)
Splits a CSV-style line on commas, but skips over commas sitting inside quoted fields like "Smith, John". It works using a lookahead that checks whether there's a balanced number of quotes remaining after the comma. For anything beyond simple one-line CSV, use a real CSV parser: quoted newlines and escaped quotes will break this fast.
Mistakes that quietly break your regex
Most regex bugs aren't syntax errors. The pattern compiles fine and runs fine, it just matches the wrong thing, or matches nothing at all, or takes down the page. Three come up more than the rest.
Greedy vs. lazy quantifiers. By default, * and + are greedy: they grab as much as they possibly can, then backtrack only if the rest of the pattern fails to match. Run <.+> against <b>bold</b> and it matches the whole string, from the first < to the last >, not just <b>. Add a ? after the quantifier, <.+?>, to make it lazy, matching as little as possible instead, which gets you just <b>. Most "my regex matched way too much" bugs are exactly this.
Forgetting to escape special characters. Characters like . ( ) [ ] { } * + ? ^ $ | all mean something to the regex engine. Write 3.14 to match a decimal number and the . will happily match any character, so "3x14" passes too. Escape it as 3\.14 when you mean a literal dot. The same applies to unescaped parentheses meant as literal text: they'll silently start a capture group instead.
Catastrophic backtracking. This one isn't cosmetic; it can hang the browser tab or the server process entirely. It happens when a pattern has nested quantifiers over overlapping character sets, something like (a+)+b tested against a long string of a's with no trailing b. The engine tries every possible way of splitting the a's between the inner and outer +, and that number of combinations grows exponentially with string length. A string of 30 a's can already take seconds; add a few more characters and it's effectively forever. I've watched a form validation regex like ^([a-zA-Z0-9]+)*$ freeze a page solid on a 40-character pasted input, with the CPU pegged on a single thread, because the browser's regex engine has no timeout of its own. The fix is almost always the same: replace the nested repetition with a single quantifier over a more specific character class, so there's exactly one way to parse a match instead of thousands.
On email validation specifically: the "perfect" email regex is a myth, and not because nobody's tried hard enough. RFC 5322, the actual spec for email address syntax, permits quoted local parts, escaped characters, and comments embedded inside the address; the official grammar runs to dozens of production rules. A regex that fully implements it is thousands of characters long and still debated on Stack Overflow. In practice, a simple pattern that checks for "something@something.something" catches the typos that matter, and anything past that should be verified by actually sending a confirmation email, which is the only real proof the address exists.
Try it
GlaeKit's Regex Tester lets you paste a pattern and test text side by side, with matches highlighted live and capture groups listed out. Everything runs in your browser using JavaScript's built-in regex engine, and nothing is uploaded.
Frequently asked questions
What's the difference between greedy and lazy matching?
A greedy quantifier (*, +) matches as much text as it can before backtracking to satisfy the rest of the pattern. A lazy quantifier (*?, +?) matches as little as it can, only expanding when forced to. Use lazy matching whenever you want the shortest match between two delimiters, like text inside a single pair of tags.
How do I match a literal dot instead of "any character"?
Escape it with a backslash: \.. Unescaped, . matches any single character except a newline, which is almost never what you want when you're trying to match a decimal point, a file extension, or an abbreviation.
What's a capture group?
Parentheses around part of a pattern, like (\d{4})-(\d{2})-(\d{2}), create capture groups: the matched text inside each pair is saved and can be referenced afterward as $1, $2, and so on, or accessed programmatically from the match result. Use (?:...) instead when you need the grouping for structure but don't need to extract that piece.
Why does my regex hang or freeze the page?
This is catastrophic backtracking, usually caused by nested quantifiers like (a+)+ or (a*)* matching against input that almost, but doesn't quite, satisfy the pattern. The engine tries an exponential number of ways to split the match before giving up. Rewrite the pattern so there's only one way to parse a given match, typically by removing the nested repetition or tightening the character class.
Should I use one giant regex to validate an email address?
No. RFC 5322's full grammar for valid email syntax is far more permissive and complex than most people expect, and a regex that implements all of it is impractical to write or maintain. A simple pattern that catches obvious typos, combined with actually sending a confirmation email, covers real-world needs better than chasing spec completeness.