Regular expressions: groups and testing
Lesson 6.4 gave you the core symbols and the three ways to try a pattern. This lesson takes parts out of a match, shows the one surprise that catches everyone, and builds a real pattern the careful way: by testing it on text that should fail.
Groups: taking out parts
Parentheses mark a : a part you want to take out on its own. Predict:
What will this print?
Decide before you look. Guessing wrong is how this sticks.
import re
print(re.findall(r"(\d+)-(\d+)", "10-20 30-40"))Python prints
[('10', '20'), ('30', '40')]When a pattern has groups, findall returns the groups, not the whole match: one tuple per match, one item per group.
With search, .group(1) is the first group and .group() is still the whole match.
On lesson 6.2's payment message, one pattern replaces the find and slice work (\. is
a real full stop). The Mawj bot reads payment messages with patterns like this.
import re
sms = "Paid ETB 250.00 to Selam Coffee. Transaction ID: DF12GH34."
found = re.search(r"ETB (\d+\.\d+)", sms)
print(found.group(1))In the editor, press Escape then Tab to move on.
Groups are numbered by their opening parentheses, left to right. Ask for one the pattern doesn't have and Python says so:
Greedy patterns
One more surprise. Predict both lines:
What will this print?
Decide before you look. Guessing wrong is how this sticks.
import re
print(re.findall(r"<.*>", "<a><b>"))
print(re.findall(r"<.*?>", "<a><b>"))Python prints
['<a><b>'] ['<a>', '<b>']
.* takes as much as it can and still fit, so it runs to the last >. Adding ? makes it take as little as it can.
* and + are
:
they take as much text as they can. *? and +? take as little as they can. On the
payment message, to (.*)\. gives Selam Coffee. Transaction ID: DF12GH34, because
.* runs on to the last full stop; to (.*?)\. stops at the first and gives
Selam Coffee.
Building and testing a phone pattern
Ethiopian mobile numbers start with 09 (Ethio Telecom) or 07 (Safaricom Ethiopia),
then eight more digits: 0[79]\d{8}. From abroad, +251 replaces the 0. | means
"or", and (?:...) groups without capturing, so findall still returns whole numbers:
(?:\+251|0) +251 or 0 (\+ is a real plus sign)
[79] then 7 or 9
\d{8} then exactly eight digitsA pattern is only as good as its tests, and the useful tests are the numbers that should fail:
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
tests = ["0911234567", "+251712345678", "0811234567", "091123456",
"09112345678", "+2510911234567", "0911 234 567", "O911234567"]
for number in tests:
print(number, bool(re.fullmatch(PHONE, number)))In the editor, press Escape then Tab to move on.
The first two should pass. Every other one is a real typing mistake: a prefix no operator
uses, a digit short, a digit extra, +251 with the 0 kept, spaces, and a capital O
typed for a zero. (08 is kept for a third operator. If one starts using it, the pattern
becomes [789], and that test changes.) Change fullmatch to search and run it again:
the two long ones pass, because search finds ten good digits inside them.
Quick check
Which function checks that a whole string is one phone number, and nothing more?
To pull numbers out of a message, findall with the same pattern:
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
msg = "Call Hana on 0911234567 or Abel on +251712345678, not 0811234567."
print(re.findall(PHONE, msg))In the editor, press Escape then Tab to move on.
When startswith, in, or split can do the job, use them instead.
number.startswith(("09", "07")) from lesson 6.1 needs no explaining; a regex always does.
Exercise
Check typed phone numbers
Print each typed number with valid or invalid after it, using the phone pattern
and re.fullmatch.
0911234567 valid
0811234567 invalid
+251912345678 valid
0712 345678 invalidYour code
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
typed = ["0911234567", "0811234567", "+251912345678", "0712 345678"]
# Print each number with valid or invalidIn the editor, press Escape then Tab to move on.
Show a solutionHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
typed = ["0911234567", "0811234567", "+251912345678", "0712 345678"]
for number in typed:
if re.fullmatch(PHONE, number):
print(number, "valid")
else:
print(number, "invalid")Key takeaways
- Parentheses make a group:
.group(1)takes out the first;findallreturns the groups, one tuple per match. - Groups are numbered by their opening parentheses; asking for one that isn't there is
IndexError: no such group. *and+are greedy and take as much as they can;*?and+?take as little as they can.- Test a pattern on text that should fail, typos included, and check whole strings with
fullmatch. - Prefer
startswith,in, orsplitwhen they do the job.