ይህ ትምህርት ገና ወደ አማርኛ አልተተረጎመም፤ ስለዚህ በእንግሊዝኛ ቀርቧል። የእንግሊዝኛውን ገጽ ክፈቱ
Project: Text analyzer
You'll build a text analyzer: give it a paragraph, in English or Amharic, and it counts
characters, words, and sentences, finds the most common words and the longest one, and
pulls out every phone number. It brings together the whole module, plus the tally plan
from lesson 4.4 and sorted with a key from lesson 5.5.
Here's the finished report for a short, made-up paragraph:
Characters: 145
Words: 26
Sentences: 4
Top words:
coffee 3
and 2
bread 2
the 2
is 2
Longest word: coffee
Phones: 0911234567, +251712345678Keep these names; every milestone reuses them. As in lesson 5.8, the functions return
values and only report prints.
| Name | Its one job |
|---|---|
text | The paragraph being analyzed |
words | text.split(): every word as typed |
clean_word(word) | Return the word without punctuation at its ends, in lowercase |
word_counts(text) | Return counts, a dict from each clean word to how often it appears |
sentence_count(text) | Return the number of sentences, ended by . or ። |
find_phones(text) | Return a list of the phone numbers in the text |
report(text) | Print the report |
Milestone 1: clean_word and the word count
Coffee. and coffee should count as the same word. strip() can take the characters
to remove: word.strip(".,") removes dots and commas from both ends, and leaves the
middle alone. Include ። and the Ethiopic comma ፣, so Amharic works too.
መልመጃ
clean_word passes every check
Write clean_word(word): strip .,!?።፣ from the ends and return the word in
lowercase. The starter returns the word unchanged, so the checks report each word
it gets wrong.
Your code
def clean_word(word):
# Strip . , ! ? ። ፣ from both ends; return it in lowercase
return word
checks = [("Coffee.", "coffee"), ("ቡና።", "ቡና"), ("Hana,", "hana"), ("7.", "7")]
for word, expected in checks:
if clean_word(word) != expected:
print(f"Wrong for {word}: got {clean_word(word)}, expected {expected}")
text = "Selam Coffee opens at 7. Hana orders coffee and bread."
words = text.split()
print(f"Words: {len(words)}")
print("All checks done")In the editor, press Escape then Tab to move on.
መፍትሄውን አሳይHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
def clean_word(word):
return word.strip(".,!?።፣").lower()
checks = [("Coffee.", "coffee"), ("ቡና።", "ቡና"), ("Hana,", "hana"), ("7.", "7")]
for word, expected in checks:
if clean_word(word) != expected:
print(f"Wrong for {word}: got {clean_word(word)}, expected {expected}")
text = "Selam Coffee opens at 7. Hana orders coffee and bread."
words = text.split()
print(f"Words: {len(words)}")
print("All checks done").lower() does nothing to ቡና, as you saw in lesson 6.1, so the same line works for
both languages.
Milestone 2: word_counts and the top five
Numbers and phone numbers shouldn't count as words in the top five. isalpha() is
True when every character is a letter. Predict:
ውጤቱን ገምቱ
Decide before you look. Guessing wrong is how this sticks.
print("0911234567".isalpha())
print("ቡና".isalpha())Python prints
False True
Digits aren't letters, so the phone number is False. Ge'ez letters are letters to Unicode, so ቡና is True: isalpha isn't only for English.
word_counts is the tally plan, with one if to keep only real words. Then sorted
with a key puts the most common first: lambda pair: pair[1] compares each pair by its
count, as in lesson 5.5.
መልመጃ
The five most common words
Finish word_counts(text) using the tally plan, then print the top five like the
output below.
coffee 3
and 2
bread 2
the 2
is 2Your code
def clean_word(word):
return word.strip(".,!?።፣").lower()
def word_counts(text):
counts = {}
# Tally each clean word that isalpha()
return counts
text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
"The coffee is fresh and the bread is fresh too. "
"Call 0911234567 or +251712345678 to order.")
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
for word, count in top:
print(word, count)In the editor, press Escape then Tab to move on.
መፍትሄውን አሳይHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
def clean_word(word):
return word.strip(".,!?።፣").lower()
def word_counts(text):
counts = {}
for word in text.split():
word = clean_word(word)
if word.isalpha():
counts[word] = counts.get(word, 0) + 1
return counts
text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
"The coffee is fresh and the bread is fresh too. "
"Call 0911234567 or +251712345678 to order.")
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
for word, count in top:
print(word, count)ጥያቄ
fresh appears twice too. Why isn't it in the top five?
Milestone 3: sentences in two scripts
A sentence ends with . in English and ። in Amharic. Replace one with the other
(lesson 6.1), then split, strip, and drop the empty pieces (lesson 6.3):
መልመጃ
sentence_count passes every check
Write sentence_count(text) so it counts sentences ended by either full stop.
Your code
def sentence_count(text):
# Treat ። like ., then count the pieces that aren't empty
return 0
checks = [("One. Two.", 2), ("ሰላም ነው። ቡና ደርሷል።", 2), ("Selam. ሰላም።", 2), ("No full stop", 1)]
for text, expected in checks:
if sentence_count(text) != expected:
print(f"Wrong for {text}: got {sentence_count(text)}, expected {expected}")
print("All checks done")In the editor, press Escape then Tab to move on.
መፍትሄውን አሳይHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
def sentence_count(text):
text = text.replace("።", ".")
return len([s for s in text.split(".") if s.strip()])
checks = [("One. Two.", 2), ("ሰላም ነው። ቡና ደርሷል።", 2), ("Selam. ሰላም።", 2), ("No full stop", 1)]
for text, expected in checks:
if sentence_count(text) != expected:
print(f"Wrong for {text}: got {sentence_count(text)}, expected {expected}")
print("All checks done")text = text.replace(...) only moves the name inside the call; the caller's string
never changes. It's honest to say what this misses: a price like 250.00 would count
its dot as a sentence end. For this project's texts, that's fine.
Milestone 4: find_phones and the report
find_phones is one line with lesson 6.5's pattern. report puts everything together.
The longest word is max with key=len over the clean words in counts, so a phone
number can't win.
መልመጃ
The whole analyzer
Write find_phones(text) and report(text) so the output matches the report at
the top of this page.
Your code
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
# Paste clean_word, word_counts, and sentence_count here
def find_phones(text):
pass
def report(text):
pass
text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
"The coffee is fresh and the bread is fresh too. "
"Call 0911234567 or +251712345678 to order.")
report(text)In the editor, press Escape then Tab to move on.
መፍትሄውን አሳይHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
def clean_word(word):
return word.strip(".,!?።፣").lower()
def word_counts(text):
counts = {}
for word in text.split():
word = clean_word(word)
if word.isalpha():
counts[word] = counts.get(word, 0) + 1
return counts
def sentence_count(text):
text = text.replace("።", ".")
return len([s for s in text.split(".") if s.strip()])
def find_phones(text):
return re.findall(PHONE, text)
def report(text):
words = text.split()
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
print(f"Characters: {len(text)}")
print(f"Words: {len(words)}")
print(f"Sentences: {sentence_count(text)}")
print("Top words:")
for word, count in top:
print(f" {word} {count}")
print(f"Longest word: {max(counts, key=len)}")
print(f"Phones: {', '.join(find_phones(text))}")
text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
"The coffee is fresh and the bread is fresh too. "
"Call 0911234567 or +251712345678 to order.")
report(text)coffee and orders both have six letters; max returns the first it meets. Forget
import re and Python tells you exactly what's missing:
Milestone 5: Amharic, and a payment message (stretch)
Run report on an Amharic paragraph. Nothing in the program has to change: split,
strip, isalpha, len, and the phone pattern all work on Ge'ez text.
import re
PHONE = r"(?:\+251|0)[79]\d{8}"
def clean_word(word):
return word.strip(".,!?።፣").lower()
def word_counts(text):
counts = {}
for word in text.split():
word = clean_word(word)
if word.isalpha():
counts[word] = counts.get(word, 0) + 1
return counts
def sentence_count(text):
text = text.replace("።", ".")
return len([s for s in text.split(".") if s.strip()])
def find_phones(text):
return re.findall(PHONE, text)
def report(text):
words = text.split()
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
print(f"Characters: {len(text)}")
print(f"Words: {len(words)}")
print(f"Sentences: {sentence_count(text)}")
print("Top words:")
for word, count in top:
print(f" {word} {count}")
print(f"Longest word: {max(counts, key=len)}")
print(f"Phones: {', '.join(find_phones(text))}")
text = "ሰላም ቡና ቤት ጠዋት ይከፈታል። ቡና እና ዳቦ አለ። ቡና ለማዘዝ 0911234567 ይደውሉ።"
report(text)In the editor, press Escape then Tab to move on.
Last, the job the Mawj bot does: read a payment message and return the amount and the transaction ID together, as a tuple (lesson 4.3).
መልመጃ
parse_sms
Write parse_sms(sms) that returns (amount, reference) as strings, or None when
either is missing. Use re.search with one group each.
('250.00', 'DF12GH34')
NoneYour code
import re
def parse_sms(sms):
# Search for the amount after "ETB " and the reference after "ID: "
pass
print(parse_sms("You paid ETB 250.00 to Selam Coffee. Transaction ID: DF12GH34. Thank you."))
print(parse_sms("Your balance is ETB 1,200.00."))In the editor, press Escape then Tab to move on.
መፍትሄውን አሳይHide the solution
This is one way to solve it, not the only one. If yours prints the same thing, it works.
import re
def parse_sms(sms):
amount = re.search(r"ETB (\d+\.\d+)", sms)
reference = re.search(r"ID: (\w+)", sms)
if amount and reference:
return (amount.group(1), reference.group(1))
return None
print(parse_sms("You paid ETB 250.00 to Selam Coffee. Transaction ID: DF12GH34. Thank you."))
print(parse_sms("Your balance is ETB 1,200.00."))The second message has no ID, and its amount has a comma, so \d+\.\d+ doesn't fit
1,200.00 either. Real messages vary like this; test your pattern on every kind you
can find, especially the ones that should fail.
ዋና ዋና ነጥቦች
- Small functions that return values, and one
reportthat prints, make each part easy to check. - Clean words before counting:
strip(".,!?።፣")andlower()work in both scripts. - The tally plan plus
sorted(..., key=lambda pair: pair[1], reverse=True)gives the most common words; ties keep their first order. - Treat
።like.to count sentences in both languages. - A regex with groups turns a payment message into a tuple; test it on messages that should fail.