Skip to content

ይህ ትምህርት ገና ወደ አማርኛ አልተተረጎመም፤ ስለዚህ በእንግሊዝኛ ቀርቧል። የእንግሊዝኛውን ገጽ ክፈቱ

Unicode and Amharic text

5 min read

Python handles Amharic as well as English, once you know what a string really is. By the end of this lesson you'll count Ge'ez letters correctly, know why Amharic takes more space on disk, split Amharic sentences, and read Ge'ez numerals.

One letter, one number

Predict the length of Selam written in Ge'ez letters.

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

print(len("ሰላም"))
Pick the output

Python prints

3

Each fidel is one character to Python, so ሰላም is three characters, just like it's three letters on the page.

If you said 6, you split each fidel into its consonant and vowel (ሰ is s + ä). If you said 9, you'd heard that Amharic takes more space. It does, but not in characters; that's the next section.

Almost every script people write today, Ge'ez included, is part of one standard, , which gives each character its own number, its . To Python, a is one code point, and a string is a sequence of code points. ord() gives a character's code point; chr() goes the other way.

print(ord("ሰ"))
print(hex(ord("ሰ")))
print(chr(0x1200))
start = 0x1230
print("".join([chr(n) for n in range(start, start + 7)]))

In the editor, press Escape then Tab to move on.

Unicode charts write code points in hexadecimal (base 16), so ሰ is U+1230. hex() shows a number in hex, and in code you type a hex number with 0x in front: 0x1230 is the same number as 4656. The Ge'ez letters (Unicode calls them Ethiopic) start at U+1200 with ሀ. The last line prints the seven code points from U+1230: the ሰ family, in fidel chart order. Change start to 0x1208 for the ለ family.

Characters are not bytes

A computer saves and sends text as bytes, not characters. Predict:

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

print(len("ሰላም".encode()))
Pick the output

Python prints

9

In UTF-8, each Ge'ez character becomes 3 bytes, so three characters are 9 bytes.

A is a number from 0 to 255, the unit of files, memory, and networks. Turning characters into bytes needs a rule, an . The one almost everything uses is : English letters take 1 byte each, Ge'ez letters take 3. .encode() gives the bytes (in UTF-8 unless you name another encoding), and .decode() turns them back into a string.

data = "ሰላም".encode()
print(data)
print(len("Selam".encode()))
print(data.decode())

In the editor, press Escape then Tab to move on.

The b means bytes. Each \x and the two hex digits after it are one byte, so \xe1\x88\xb0 is ሰ. What matters: a limit counted in bytes, like a 30-byte database field, fits 30 English letters but only 10 Ge'ez letters, whatever len() says about the string.

ጥያቄ

Which is true?

Amharic sentences and order

Amharic ends a sentence with ። (አራት ነጥብ, four dots), not .. They're different characters, so split on ።. Predict the result:

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

print("ሰላም ነው። ቡና ደርሷል።".split("።"))
Pick the output

Python prints

['ሰላም ነው', ' ቡና ደርሷል', '']

split only cuts. The space after the first ። stays at the start of the second piece, and the text after the last ። is an empty string.

If you picked the tidy list, you expected split to clean up. It only cuts. Strip each piece and keep only the ones that aren't empty (an empty string is falsy, from lesson 2.4), with a comprehension from lesson 4.7:

text = "ሰላም ነው። ቡና ደርሷል።"
sentences = [s.strip() for s in text.split("።") if s.strip()]
print(sentences)
print(len(sentences))
print(text.split("."))

In the editor, press Escape then Tab to move on.

The last line splits on . and finds nothing to cut.

sorted() compares strings code point by code point. The Ethiopic block is laid out in the traditional fidel order, so ሀ comes before ለ, and ለ before መ:

words = ["መጽሐፍ", "ለምለም", "ሀገር"]
print(sorted(words))

In the editor, press Escape then Tab to move on.

That's all code point order promises. It isn't a full dictionary sort, so don't rely on it where the exact order matters.

Ge'ez numerals

Ge'ez has its own numerals: ፩ ፪ ፫ for 1, 2, 3. In lesson 3.5 you checked isdigit() before calling int(). Predict:

ውጤቱን ገምቱ

Decide before you look. Guessing wrong is how this sticks.

print("፩".isdigit())
Pick the output

Python prints

True

Unicode lists ፩ as a digit, so isdigit() says True. It checks for any digit, not only 0 to 9.

So isdigit() says yes. But int() reads only decimal digits: 0 to 9, and digits from scripts that write numbers the same place-value way, like Arabic-Indic ٣. Ge'ez numerals don't work that way, so ፩ isn't one:

GEEZ is a constant (the capitals say so), and you use it like any other dict:

GEEZ = {"፩": 1, "፪": 2, "፫": 3, "፬": 4, "፭": 5, "፮": 6, "፯": 7, "፰": 8, "፱": 9}
print(GEEZ["፫"] + GEEZ["፬"])

In the editor, press Escape then Tab to move on.

This covers one to nine. Ge'ez has no zero: tens and hundreds have their own signs (፲ is ten, ፳ twenty, ፻ a hundred), so bigger numbers need more than a lookup.

When one letter is two code points

In other scripts, one letter on screen can be more than one code point. é can be one code point, or e followed by an accent code point that joins onto it. Inside a string, \u and four hex digits write a character by its code point:

import unicodedata
one = "\u00e9"
two = "e\u0301"
print(len(one), len(two))
print(one == two)
print(unicodedata.normalize("NFC", two) == one)
print(len("👍🏽"))

In the editor, press Escape then Tab to move on.

Both look like é, but == compares code points. unicodedata.normalize("NFC", ...) rewrites text into one standard form so they compare equal. A thumbs-up with a skin tone is 2 code points too. Just remember: len() counts code points, not what the eye sees.

መልመጃ

Words and sentences in Amharic

Count the words and the sentences in this short paragraph. Words are separated by spaces; sentences end with ።.

Output
Words: 10
Sentences: 3

Your code

text = "ዛሬ ቡና ጠጣሁ። ከዚያ ወደ ቢሮ ሄድኩ። ሥራው ብዙ ነበር።"
# Count the words, then the sentences

In the editor, press Escape then Tab to move on.

መፍትሄውን አሳይ

This is one way to solve it, not the only one. If yours prints the same thing, it works.

text = "ዛሬ ቡና ጠጣሁ። ከዚያ ወደ ቢሮ ሄድኩ። ሥራው ብዙ ነበር።"
words = text.split()
sentences = [s for s in text.split("።") if s.strip()]
print(f"Words: {len(words)}")
print(f"Sentences: {len(sentences)}")

ዋና ዋና ነጥቦች

  • Unicode gives every character in every script a number, its code point. ord() and chr() go between them.
  • A Python string is a sequence of code points, so len("ሰላም") is 3.
  • Bytes are not characters: in UTF-8, a Ge'ez letter is 3 bytes. .encode() and .decode() go between them.
  • Split Amharic sentences on ።, then strip and drop the empty pieces. sorted() follows code point order.
  • "፩".isdigit() is True, but int("፩") fails; map Ge'ez numerals with a dict.