Regular Expressions Tutorial with Examples: grep, sed, Python and Cheat Sheet

· 3 min read · DevOps Tutorials

What is a regular expression #

A regular expression (regex) describes a pattern of text. Instead of searching for the word error, you can search for "a line that starts with a date, then the word ERROR or WARN, then a number". You use regex to find text, validate input and replace parts of it.

Every example below can be run in a terminal. Create a file to practice:

cat > app.log <<'LOG'
2026-10-01 10:15:02 INFO  user=ana ip=10.0.1.15 action=login
2026-10-01 10:16:40 ERROR user=bob ip=10.0.2.200 action=upload size=5242880
2026-10-01 10:17:11 WARN  user=ana ip=192.168.0.7 action=logout
2026-10-02 08:01:33 ERROR user=eve ip=10.0.1.99 action=login
LOG

Which regex flavor #

Tools implement different dialects. The basic syntax is the same, the extras are not.

Flavor Used by Notes
POSIX BRE grep, sed (default) +, ?, |, () and {} need a backslash
POSIX ERE grep -E, sed -E, awk The one used in most shell examples here
PCRE grep -P, PHP, many editors Has \d, lazy quantifiers, lookahead
RE2 Go, Terraform, Prometheus No lookahead or backreferences, linear time
Python / JavaScript re, RegExp Close to PCRE

The building blocks #

Pattern Matches
abc the text abc
. any one character (except a newline)
[abc] / [^abc] one of a, b, c / any character except them
[a-z], [0-9] a range
\d \w \s digit, word character, whitespace (PCRE, Python, RE2). In POSIX use [0-9], [[:alnum:]_], [[:space:]]
* + ? 0 or more, 1 or more, 0 or 1
{n} {n,} {n,m} exactly n, at least n, between n and m
^ $ start and end of the line
\b word boundary
( ) group and capture
a|b a or b
\. a literal dot (escape metacharacters with a backslash)
grep ERROR app.log                          # plain text
grep -E 'ERROR|WARN' app.log                # alternation
grep -E '^2026-10-01' app.log               # lines that start with a date
grep -E 'user=(ana|eve)' app.log            # group with alternation
grep -E 'size=[0-9]+' app.log               # one or more digits
grep -oE 'ip=[0-9.]+' app.log               # -o prints only the match
grep -oE '[0-9]{1,3}(\.[0-9]{1,3}){3}' app.log   # IPv4 addresses
grep -vE 'INFO' app.log                     # -v inverts
grep -c ERROR app.log                       # count the lines

-o is the most useful flag: it turns grep into an extractor.

sed: replace #

sed -E 's/ERROR/FAIL/' app.log                          # first match in each line
sed -E 's/ip=[0-9.]+/ip=REDACTED/g' app.log              # g = every match
sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/' app.log   # 2026-10-01 -> 01/10/2026
sed -E 's/^[0-9-]+ [0-9:]+ //' app.log                   # remove the timestamp
sed -nE '/ERROR/p' app.log                               # print only matching lines
sed -E -i.bak 's/INFO +/INFO /' app.log                  # edit the file, keep a backup

In the replacement, \1, \2... are the groups captured by the parentheses, and & is the whole match.

Python #

import re

line = "2026-10-01 10:16:40 ERROR user=bob ip=10.0.2.200 action=upload"

m = re.search(r"(?P<level>ERROR|WARN) user=(?P<user>\w+) ip=(?P<ip>[\d.]+)", line)
if m:
    print(m.group("level"), m.group("user"), m.group("ip"))   # ERROR bob 10.0.2.200

print(re.findall(r"\d+", line))              # every group of digits
print(re.sub(r"ip=[\d.]+", "ip=x", line))    # replace
pattern = re.compile(r"^\d{4}-\d{2}-\d{2}")  # compile once when reused
print(bool(pattern.match(line)))             # True

Always use raw strings (r"...") so Python does not interpret the backslashes. Named groups (?P<name>...) make the code readable.

Greedy and lazy #

Quantifiers are greedy: they take as much as possible.

echo '<b>one</b> and <b>two</b>' | grep -oP '<b>.*</b>'      # one match: from the first <b> to the last </b>
echo '<b>one</b> and <b>two</b>' | grep -oP '<b>.*?</b>'     # two matches: <b>one</b> and <b>two</b>

.*? is lazy: it takes as little as possible. When you can, use a more precise pattern such as [^<]* instead.

Lookahead (PCRE and Python) #

grep -oP 'user=\K\w+' app.log           # \K drops what came before: ana, bob, ana, eve
grep -oP '\d+(?= bytes)' file.txt       # (?= ) matches only if followed by " bytes"

POSIX and RE2 do not support them: rewrite with a capture group.

Real-world patterns #

Goal Pattern
IPv4 address (shape only) ^([0-9]{1,3}\.){3}[0-9]{1,3}$
CIDR block ^([0-9]{1,3}\.){3}[0-9]{1,3}/[0-9]{1,2}$
Semantic version ^v?(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)$
AMI id ^ami-[0-9a-f]{8}([0-9a-f]{9})?$
AWS account id ^[0-9]{12}$
AWS ARN ^arn:aws:[a-z0-9-]+:[a-z0-9-]*:[0-9]{12}:.+$
S3 bucket name ^[a-z0-9][a-z0-9.-]{1,61}[a-z0-9]$
Date YYYY-MM-DD ^[0-9]{4}-(0[1-9]|1[0-2])-(0[1-9]|[12][0-9]|3[01])$
Kubernetes name (DNS label) ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$

An IPv4 pattern that checks 0-255 per octet is long and error-prone. For validation, check the shape with a regex and the range with code or a library.

Good habits #

  • Anchor validation patterns with ^ and $, or a test will pass for abc123xyz when you wanted only 123.
  • Escape the dot: 10.0.1.1 also matches 10a0b1c1. Write 10\.0\.1\.1.
  • Do not parse everything with regex: HTML, JSON and YAML have parsers. Use jq for JSON.
  • Beware of catastrophic backtracking with nested quantifiers such as (a+)+$ in PCRE engines. RE2 is safe against this.
  • Test interactively with grep -oE on sample data before putting a pattern in production.

Next steps #

See how infrastructure tools use regex in regex in Terraform, Ansible and Prometheus.

#Regex #Linux #Grep #Sed #Python