Skip to content

ይህ ትምህርት ገና ወደ አማርኛ አልተተረጎመም፤ ስለዚህ በእንግሊዝኛ ቀርቧል። የእንግሊዝኛውን ገጽ ክፈቱ

Project: Text analyzer

6 min readፕሮጀክት

You'll build a text analyzer: give it a paragraph, in English or Amharic, and it counts characters, words, and sentences, finds the most common words and the longest one, and pulls out every phone number. It brings together the whole module, plus the tally plan from lesson 4.4 and sorted with a key from lesson 5.5.

Here's the finished report for a short, made-up paragraph:

Output
Characters: 145
Words: 26
Sentences: 4
Top words:
  coffee 3
  and 2
  bread 2
  the 2
  is 2
Longest word: coffee
Phones: 0911234567, +251712345678

Keep these names; every milestone reuses them. As in lesson 5.8, the functions return values and only report prints.

NameIts one job
textThe paragraph being analyzed
wordstext.split(): every word as typed
clean_word(word)Return the word without punctuation at its ends, in lowercase
word_counts(text)Return counts, a dict from each clean word to how often it appears
sentence_count(text)Return the number of sentences, ended by . or ።
find_phones(text)Return a list of the phone numbers in the text
report(text)Print the report

Milestone 1: clean_word and the word count

Coffee. and coffee should count as the same word. strip() can take the characters to remove: word.strip(".,") removes dots and commas from both ends, and leaves the middle alone. Include ። and the Ethiopic comma ፣, so Amharic works too.

መልመጃ

clean_word passes every check

Write clean_word(word): strip .,!?።፣ from the ends and return the word in lowercase. The starter returns the word unchanged, so the checks report each word it gets wrong.

Your code

def clean_word(word):
    # Strip . , ! ? ። ፣ from both ends; return it in lowercase
    return word

checks = [("Coffee.", "coffee"), ("ቡና።", "ቡና"), ("Hana,", "hana"), ("7.", "7")]
for word, expected in checks:
    if clean_word(word) != expected:
        print(f"Wrong for {word}: got {clean_word(word)}, expected {expected}")

text = "Selam Coffee opens at 7. Hana orders coffee and bread."
words = text.split()
print(f"Words: {len(words)}")
print("All checks done")

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

def clean_word(word):
    return word.strip(".,!?።፣").lower()

checks = [("Coffee.", "coffee"), ("ቡና።", "ቡና"), ("Hana,", "hana"), ("7.", "7")]
for word, expected in checks:
    if clean_word(word) != expected:
        print(f"Wrong for {word}: got {clean_word(word)}, expected {expected}")

text = "Selam Coffee opens at 7. Hana orders coffee and bread."
words = text.split()
print(f"Words: {len(words)}")
print("All checks done")

.lower() does nothing to ቡና, as you saw in lesson 6.1, so the same line works for both languages.

Milestone 2: word_counts and the top five

Numbers and phone numbers shouldn't count as words in the top five. isalpha() is True when every character is a letter. Predict:

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

print("0911234567".isalpha())
print("ቡና".isalpha())
Pick the output

Python prints

False
True

Digits aren't letters, so the phone number is False. Ge'ez letters are letters to Unicode, so ቡና is True: isalpha isn't only for English.

word_counts is the tally plan, with one if to keep only real words. Then sorted with a key puts the most common first: lambda pair: pair[1] compares each pair by its count, as in lesson 5.5.

መልመጃ

The five most common words

Finish word_counts(text) using the tally plan, then print the top five like the output below.

Output
coffee 3
and 2
bread 2
the 2
is 2

Your code

def clean_word(word):
    return word.strip(".,!?።፣").lower()

def word_counts(text):
    counts = {}
    # Tally each clean word that isalpha()
    return counts

text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
        "The coffee is fresh and the bread is fresh too. "
        "Call 0911234567 or +251712345678 to order.")
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
for word, count in top:
    print(word, count)

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

def clean_word(word):
    return word.strip(".,!?።፣").lower()

def word_counts(text):
    counts = {}
    for word in text.split():
        word = clean_word(word)
        if word.isalpha():
            counts[word] = counts.get(word, 0) + 1
    return counts

text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
        "The coffee is fresh and the bread is fresh too. "
        "Call 0911234567 or +251712345678 to order.")
counts = word_counts(text)
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
for word, count in top:
    print(word, count)

ጥያቄ

fresh appears twice too. Why isn't it in the top five?

Milestone 3: sentences in two scripts

A sentence ends with . in English and ። in Amharic. Replace one with the other (lesson 6.1), then split, strip, and drop the empty pieces (lesson 6.3):

መልመጃ

sentence_count passes every check

Write sentence_count(text) so it counts sentences ended by either full stop.

Your code

def sentence_count(text):
    # Treat ። like ., then count the pieces that aren't empty
    return 0

checks = [("One. Two.", 2), ("ሰላም ነው። ቡና ደርሷል።", 2), ("Selam. ሰላም።", 2), ("No full stop", 1)]
for text, expected in checks:
    if sentence_count(text) != expected:
        print(f"Wrong for {text}: got {sentence_count(text)}, expected {expected}")
print("All checks done")

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

def sentence_count(text):
    text = text.replace("።", ".")
    return len([s for s in text.split(".") if s.strip()])

checks = [("One. Two.", 2), ("ሰላም ነው። ቡና ደርሷል።", 2), ("Selam. ሰላም።", 2), ("No full stop", 1)]
for text, expected in checks:
    if sentence_count(text) != expected:
        print(f"Wrong for {text}: got {sentence_count(text)}, expected {expected}")
print("All checks done")

text = text.replace(...) only moves the name inside the call; the caller's string never changes. It's honest to say what this misses: a price like 250.00 would count its dot as a sentence end. For this project's texts, that's fine.

Milestone 4: find_phones and the report

find_phones is one line with lesson 6.5's pattern. report puts everything together. The longest word is max with key=len over the clean words in counts, so a phone number can't win.

መልመጃ

The whole analyzer

Write find_phones(text) and report(text) so the output matches the report at the top of this page.

Your code

import re

PHONE = r"(?:\+251|0)[79]\d{8}"

# Paste clean_word, word_counts, and sentence_count here

def find_phones(text):
    pass

def report(text):
    pass

text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
        "The coffee is fresh and the bread is fresh too. "
        "Call 0911234567 or +251712345678 to order.")
report(text)

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

import re

PHONE = r"(?:\+251|0)[79]\d{8}"

def clean_word(word):
    return word.strip(".,!?።፣").lower()

def word_counts(text):
    counts = {}
    for word in text.split():
        word = clean_word(word)
        if word.isalpha():
            counts[word] = counts.get(word, 0) + 1
    return counts

def sentence_count(text):
    text = text.replace("።", ".")
    return len([s for s in text.split(".") if s.strip()])

def find_phones(text):
    return re.findall(PHONE, text)

def report(text):
    words = text.split()
    counts = word_counts(text)
    top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
    print(f"Characters: {len(text)}")
    print(f"Words: {len(words)}")
    print(f"Sentences: {sentence_count(text)}")
    print("Top words:")
    for word, count in top:
        print(f"  {word} {count}")
    print(f"Longest word: {max(counts, key=len)}")
    print(f"Phones: {', '.join(find_phones(text))}")

text = ("Selam Coffee opens at 7. Hana orders coffee and bread. "
        "The coffee is fresh and the bread is fresh too. "
        "Call 0911234567 or +251712345678 to order.")
report(text)

coffee and orders both have six letters; max returns the first it meets. Forget import re and Python tells you exactly what's missing:

Milestone 5: Amharic, and a payment message (stretch)

Run report on an Amharic paragraph. Nothing in the program has to change: split, strip, isalpha, len, and the phone pattern all work on Ge'ez text.

import re

PHONE = r"(?:\+251|0)[79]\d{8}"

def clean_word(word):
    return word.strip(".,!?።፣").lower()

def word_counts(text):
    counts = {}
    for word in text.split():
        word = clean_word(word)
        if word.isalpha():
            counts[word] = counts.get(word, 0) + 1
    return counts

def sentence_count(text):
    text = text.replace("።", ".")
    return len([s for s in text.split(".") if s.strip()])

def find_phones(text):
    return re.findall(PHONE, text)

def report(text):
    words = text.split()
    counts = word_counts(text)
    top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)[:5]
    print(f"Characters: {len(text)}")
    print(f"Words: {len(words)}")
    print(f"Sentences: {sentence_count(text)}")
    print("Top words:")
    for word, count in top:
        print(f"  {word} {count}")
    print(f"Longest word: {max(counts, key=len)}")
    print(f"Phones: {', '.join(find_phones(text))}")

text = "ሰላም ቡና ቤት ጠዋት ይከፈታል። ቡና እና ዳቦ አለ። ቡና ለማዘዝ 0911234567 ይደውሉ።"
report(text)

In the editor, press Escape then Tab to move on.

Last, the job the Mawj bot does: read a payment message and return the amount and the transaction ID together, as a tuple (lesson 4.3).

መልመጃ

parse_sms

Write parse_sms(sms) that returns (amount, reference) as strings, or None when either is missing. Use re.search with one group each.

Output
('250.00', 'DF12GH34')
None

Your code

import re

def parse_sms(sms):
    # Search for the amount after "ETB " and the reference after "ID: "
    pass

print(parse_sms("You paid ETB 250.00 to Selam Coffee. Transaction ID: DF12GH34. Thank you."))
print(parse_sms("Your balance is ETB 1,200.00."))

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

import re

def parse_sms(sms):
    amount = re.search(r"ETB (\d+\.\d+)", sms)
    reference = re.search(r"ID: (\w+)", sms)
    if amount and reference:
        return (amount.group(1), reference.group(1))
    return None

print(parse_sms("You paid ETB 250.00 to Selam Coffee. Transaction ID: DF12GH34. Thank you."))
print(parse_sms("Your balance is ETB 1,200.00."))

The second message has no ID, and its amount has a comma, so \d+\.\d+ doesn't fit 1,200.00 either. Real messages vary like this; test your pattern on every kind you can find, especially the ones that should fail.

ዋና ዋና ነጥቦች

  • Small functions that return values, and one report that prints, make each part easy to check.
  • Clean words before counting: strip(".,!?።፣") and lower() work in both scripts.
  • The tally plan plus sorted(..., key=lambda pair: pair[1], reverse=True) gives the most common words; ties keep their first order.
  • Treat ። like . to count sentences in both languages.
  • A regex with groups turns a payment message into a tuple; test it on messages that should fail.