Regular Expressions Tutorial with Examples: grep, sed, Python and Cheat Sheet
What is a regular expression #
A regular expression (regex) describes a pattern of text. Instead of searching for the word error, you can search for "a line that starts with a date, then the word ERROR or WARN, then a number". You use regex to find text, validate input and replace parts of it.
Every example below can be run in a terminal. Create a file to practice:
cat > app.log <<'LOG'
2026-10-01 10:15:02 INFO user=ana ip=10.0.1.15 action=login
2026-10-01 10:16:40 ERROR user=bob ip=10.0.2.200 action=upload size=5242880
2026-10-01 10:17:11 WARN user=ana ip=192.168.0.7 action=logout
2026-10-02 08:01:33 ERROR user=eve ip=10.0.1.99 action=login
LOG
Which regex flavor #
Tools implement different dialects. The basic syntax is the same, the extras are not.
| Flavor | Used by | Notes |
|---|---|---|
| POSIX BRE | grep, sed (default) |
+, ?, |, () and {} need a backslash |
| POSIX ERE | grep -E, sed -E, awk |
The one used in most shell examples here |
| PCRE | grep -P, PHP, many editors |
Has \d, lazy quantifiers, lookahead |
| RE2 | Go, Terraform, Prometheus | No lookahead or backreferences, linear time |
| Python / JavaScript | re, RegExp |
Close to PCRE |
The building blocks #
| Pattern | Matches |
|---|---|
abc |
the text abc |
. |
any one character (except a newline) |
[abc] / [^abc] |
one of a, b, c / any character except them |
[a-z], [0-9] |
a range |
\d \w \s |
digit, word character, whitespace (PCRE, Python, RE2). In POSIX use [0-9], [[:alnum:]_], [[:space:]] |
* + ? |
0 or more, 1 or more, 0 or 1 |
{n} {n,} {n,m} |
exactly n, at least n, between n and m |
^ $ |
start and end of the line |
\b |
word boundary |
( ) |
group and capture |
a|b |
a or b |
\. |
a literal dot (escape metacharacters with a backslash) |
grep: search #
grep ERROR app.log # plain text
grep -E 'ERROR|WARN' app.log # alternation
grep -E '^2026-10-01' app.log # lines that start with a date
grep -E 'user=(ana|eve)' app.log # group with alternation
grep -E 'size=[0-9]+' app.log # one or more digits
grep -oE 'ip=[0-9.]+' app.log # -o prints only the match
grep -oE '[0-9]{1,3}(\.[0-9]{1,3}){3}' app.log # IPv4 addresses
grep -vE 'INFO' app.log # -v inverts
grep -c ERROR app.log # count the lines
-o is the most useful flag: it turns grep into an extractor.
sed: replace #
sed -E 's/ERROR/FAIL/' app.log # first match in each line
sed -E 's/ip=[0-9.]+/ip=REDACTED/g' app.log # g = every match
sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/' app.log # 2026-10-01 -> 01/10/2026
sed -E 's/^[0-9-]+ [0-9:]+ //' app.log # remove the timestamp
sed -nE '/ERROR/p' app.log # print only matching lines
sed -E -i.bak 's/INFO +/INFO /' app.log # edit the file, keep a backup
In the replacement, \1, \2... are the groups captured by the parentheses, and & is the whole match.
Python #
import re
line = "2026-10-01 10:16:40 ERROR user=bob ip=10.0.2.200 action=upload"
m = re.search(r"(?P<level>ERROR|WARN) user=(?P<user>\w+) ip=(?P<ip>[\d.]+)", line)
if m:
print(m.group("level"), m.group("user"), m.group("ip")) # ERROR bob 10.0.2.200
print(re.findall(r"\d+", line)) # every group of digits
print(re.sub(r"ip=[\d.]+", "ip=x", line)) # replace
pattern = re.compile(r"^\d{4}-\d{2}-\d{2}") # compile once when reused
print(bool(pattern.match(line))) # True
Always use raw strings (r"...") so Python does not interpret the backslashes. Named groups (?P<name>...) make the code readable.
Greedy and lazy #
Quantifiers are greedy: they take as much as possible.
echo '<b>one</b> and <b>two</b>' | grep -oP '<b>.*</b>' # one match: from the first <b> to the last </b>
echo '<b>one</b> and <b>two</b>' | grep -oP '<b>.*?</b>' # two matches: <b>one</b> and <b>two</b>
.*? is lazy: it takes as little as possible. When you can, use a more precise pattern such as [^<]* instead.
Lookahead (PCRE and Python) #
grep -oP 'user=\K\w+' app.log # \K drops what came before: ana, bob, ana, eve
grep -oP '\d+(?= bytes)' file.txt # (?= ) matches only if followed by " bytes"
POSIX and RE2 do not support them: rewrite with a capture group.
Real-world patterns #
| Goal | Pattern |
|---|---|
| IPv4 address (shape only) | ^([0-9]{1,3}\.){3}[0-9]{1,3}$ |
| CIDR block | ^([0-9]{1,3}\.){3}[0-9]{1,3}/[0-9]{1,2}$ |
| Semantic version | ^v?(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)$ |
| AMI id | ^ami-[0-9a-f]{8}([0-9a-f]{9})?$ |
| AWS account id | ^[0-9]{12}$ |
| AWS ARN | ^arn:aws:[a-z0-9-]+:[a-z0-9-]*:[0-9]{12}:.+$ |
| S3 bucket name | ^[a-z0-9][a-z0-9.-]{1,61}[a-z0-9]$ |
Date YYYY-MM-DD |
^[0-9]{4}-(0[1-9]|1[0-2])-(0[1-9]|[12][0-9]|3[01])$ |
| Kubernetes name (DNS label) | ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ |
An IPv4 pattern that checks 0-255 per octet is long and error-prone. For validation, check the shape with a regex and the range with code or a library.
Good habits #
- Anchor validation patterns with
^and$, or a test will pass forabc123xyzwhen you wanted only123. - Escape the dot:
10.0.1.1also matches10a0b1c1. Write10\.0\.1\.1. - Do not parse everything with regex: HTML, JSON and YAML have parsers. Use jq for JSON.
- Beware of catastrophic backtracking with nested quantifiers such as
(a+)+$in PCRE engines. RE2 is safe against this. - Test interactively with
grep -oEon sample data before putting a pattern in production.
Next steps #
See how infrastructure tools use regex in regex in Terraform, Ansible and Prometheus.