Skip to content

ይህ ትምህርት ገና ወደ አማርኛ አልተተረጎመም፤ ስለዚህ በእንግሊዝኛ ቀርቧል። የእንግሊዝኛውን ገጽ ክፈቱ

Regular expressions: the basics

4 min read

Professional programmers say it plainly: regular expressions are hard to read, hard to check, and easy to get subtly wrong. They're also the best tool for one job: finding text by its shape. This lesson teaches a small, useful core; lesson 6.5 adds groups and how to test what you write.

Finding text by its shape

To find every phone number in a message, find would need the exact text. You only know the shape: a run of digits. A regular expression describes shapes:

import re

msg = "Call Hana on 0911234567 or Abel on 0712345678."
print(re.findall(r"\d+", msg))

In the editor, press Escape then Tab to move on.

A (a regex) is a of text. \d means "any digit" and + means "one or more of the thing before", so \d+ is "a run of digits". re.findall returns every piece of the text that fits, as a list. The re module comes with Python; import it like random in lesson 3.5.

The r before the quotes makes a : a string where \ is just a backslash. You'll see why below.

The core symbols

These eleven cover most real patterns. Each "Fits" column is text the whole pattern fits exactly.

SymbolMeansPatternFits
\dany decimal digit (0 to 9, or one like ٣ from lesson 6.3)\d\d12
\wa letter, digit, or _, in any script\w+Selam, ሰላም
\sa space, tab, or newlineETB\s\dETB 5
.any character except a newlineb.tbat, b t
+one or more of the thing before\d+7, 250
*zero or moreab*a, abbb
?zero or one: optionalcolou?rcolor, colour
{n}, {n,m}exactly n, or n to m\d{4}, \d{2,3}2018; 12, 123
[...]one of these characters[79]7, 9
^the start of the text^09the 09 at the start of 0911234567
$the end of the textbirr$the birr at the end of 250 birr
import re

print(re.findall(r"\d+", "Bus 12 and 7"))
print(re.findall(r"\w+", "Selam, ሰላም!"))
print(re.findall(r"colou?r", "color or colour"))
print(re.findall(r"\d{4}", "Year 2018, room 12"))

In the editor, press Escape then Tab to move on.

Change the last pattern to \d{2} and predict what it finds before you run it.

Now the raw string. In a normal string, \ starts an escape, like \n for a new line. \d isn't one, so Python keeps it, but warns you:

match, search, and fullmatch

findall gives every piece that fits. To find just the first one, re.search returns a object, or None if nothing fits. .group() gives the text it found:

import re

found = re.search(r"\d+", "ETB 250")
print(found)
print(found.group())

In the editor, press Escape then Tab to move on.

span=(4, 7) says the match starts at index 4 and stops before 7, like a slice. Now re.match, which sounds like it does the same job. Predict:

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

import re
print(re.match(r"\d+", "ETB 250"))
Pick the output

Python prints

None

re.match only tries at the very start of the text. The text starts with E, not a digit, so there's no match.

If you expected the same match as search, so does nearly everyone. re.match only looks at the start of the text. A third function, re.fullmatch, fits only if the pattern covers the whole text:

import re

print(re.fullmatch(r"\d{4}", "12345"))
print(re.fullmatch(r"\d{4}", "2018"))

In the editor, press Escape then Tab to move on.

So search finds the first place the pattern fits, anywhere; match tries only at the start; fullmatch needs the whole text.

ጥያቄ

You want the first amount in "Paid ETB 250, then ETB 80". Which call finds it?

When nothing fits, all three return None, and .group() on None fails:

መልመጃ

Add up a receipt

findall returns strings, even when they're digits: look at the quotes in ['45', '30', '15']. Find every amount in the receipt, convert each one with int(), and keep a running total.

Output
Total: 90

Your code

import re

receipt = "Coffee 45, bread 30, tea 15"
# Find every amount and add them up

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

import re

receipt = "Coffee 45, bread 30, tea 15"
total = 0
for amount in re.findall(r"\d+", receipt):
    total = total + int(amount)
print("Total:", total)

ዋና ዋና ነጥቦች

  • A regular expression is a pattern for the shape of text. Write it as a raw string, r"...", or Python warns about \d.
  • The core: \d \w \s . for kinds of characters, + * ? {n} for how many, [...] for a choice, ^ $ for the start and end.
  • findall returns every piece that fits, as a list of strings.
  • search fits anywhere, match only at the start, fullmatch the whole text. They return a match or None, so check before calling .group().