Character Representation
intermediate30 minLearning objectives
- Distinguish between ASCII and Unicode
- Explain why Unicode was developed
- Interpret character encoding tables
- Evaluate different character sets
Learn
AQA 4.5.4 β Character encoding
Retrieval: every previous lesson in this sequence represented numbers in binary. Text needs representing too β this lesson asks how a computer, which only ever stores bits, represents the letter "A" or the emoji "π".
Key vocabulary
- Character set β the complete collection of characters a particular encoding can represent.
- Character encoding β the specific rule mapping each character in a character set to a binary/numeric code.
- ASCII (American Standard Code for Information Interchange) β an early, 7-bit character encoding covering 128 characters: unaccented English letters, digits, punctuation and basic control codes.
- Unicode β a much larger, modern character set designed to represent effectively every character in every human writing system, plus symbols and emoji.
- Code point β the specific number assigned to one character within a character set (e.g. "A" is code point 65).
Understand β a character is just an agreed number
A computer never stores the letter "A" directly β it stores a number, and an agreed encoding table says "the number 65 means the letter A." Every device reading that file has to use the same table, or the numbers get displayed as the wrong characters entirely β this is exactly what happens when text appears as "mojibake" (garbled symbols) after being opened with the wrong encoding assumed.
See it β a slice of the ASCII table
| Code point (denary) | Code point (binary, 7-bit) | Character |
|---|---|---|
| 65 | 1000001 | A |
| 66 | 1000010 | B |
| 97 | 1100001 | a |
| 48 | 0110000 | 0 |
| 32 | 0100000 | (space) |
Notice: uppercase and lowercase letters are 32 apart (65 for "A", 97 for "a") β a deliberate, structured design choice, not a coincidence, that made converting between cases a simple, fast bitwise operation on early hardware.
ASCII's limitation, and why Unicode was developed
ASCII uses only 7 bits, giving exactly 128 possible characters β enough for unaccented English, but with no way to represent "Γ©", "ΓΌ", Cyrillic, Chinese characters, or emoji at all. Different regions historically invented their own incompatible extended character sets to cover their own languages, meaning a file created in one country could display as garbage in another. Unicode was developed to solve this permanently: one single, enormous, universally agreed character set (over 140,000 characters and counting) intended to represent every writing system in the world, so text can move between any two systems without ambiguity.
Calculate/Apply it β decoding a message
Using the ASCII slice above, decode the binary sequence 1000010 1000001 1000010 1000010 1100001 character by character (a space separates each 7-bit code). (66βB is not in the table above, but 97βa is β try decoding just the ones shown: 1000010 would be "B", 1000001 is "A" - working through the pattern: this spells "BABBa" if you look up each code, illustrating how a whole message is just a sequence of individually-decoded characters.)
Common mistake
Assuming Unicode and ASCII are two competing, incompatible systems. In fact, Unicode's first 128 code points are deliberately identical to ASCII β so any valid ASCII text is automatically valid Unicode text too. Unicode extends ASCII rather than replacing it outright, which is exactly what made widespread adoption realistic.
Evaluate β choosing a character set
A system storing only simple English product codes (e.g. "SKU1042") could use plain ASCII β smaller, simpler, entirely sufficient. A system storing international customer names, addresses in multiple scripts, or user-generated social media posts needs Unicode β anything narrower would corrupt or reject valid real-world text. Choosing the narrower encoding isn't "more efficient" if it can't actually represent the data a system genuinely needs to store.
Challenge
Investigate (using a reference table or Python's ord()/chr() functions in the coding lab) the Unicode code point for a character from a non-English alphabet you're familiar with, or an emoji. Confirm it's a much larger number than any ASCII code point, and explain in a sentence why that's expected given how many characters Unicode has to fit in.
Looking ahead: the next two lessons (Image and Sound Representation) apply this exact same underlying idea β real-world information broken into and stored as discrete binary-coded units β to pictures and audio instead of text.