ScholaFly

CS02-02 Computer Science Watch

Character sets and ASCII

Subscribe on YouTubeLike this lesson on YouTube

Watch on YouTube

In this lesson

In this video you'll learn about character sets and ASCII for GCSE Computer Science, with worked examples and the mistakes examiners report. By the end you'll be able to explain what a character set is, describe how characters are represented in binary using ASCII, and relate the number of bits per character to how many characters can be represented - naming which board uses 7 bits and which uses 8.

What it covers

  1. 1:17 What a character set actually is
  2. 4:38 Reading the table both ways
  3. 7:18 Bits per character
  4. 9:26 Exam technique

Key words

About this video

GCSE Computer Science - Character sets and ASCII | Units and file sizes 2/6 (2026/27 exams)

In this video you'll learn about character sets and ASCII for GCSE Computer Science, with worked examples and the mistakes examiners report.

By the end you'll be able to explain what a character set is, describe how characters are represented in binary using ASCII, and relate the number of bits per character to how many characters can be represented - naming which board uses 7 bits and which uses 8.

For: AQA, Edexcel, OCR GCSE Computer Science
Watch first: CS01-02 Denary to binary and back

Specifications: AQA 8525 3.3.5, Edexcel 1CP2 2.2.1, OCR J277 1.2.4

Video code: CS02-02 - search YouTube for "ScholaFly CS02-02" to come straight back to this video.

Videos in this chapter:
CS02-00 — Units, characters and file size - Intro
CS02-01 — Bits, bytes and the units of storage
CS02-02 — Character sets and ASCII
CS02-03 — Character codes run in order
CS02-04 — Unicode
CS02-05 — Working out the size of a text file
CS02-06 — The limits of a fixed number of bits

#CharacterSetsAndASCII #GCSEComputerScience #ComputerScience

For more, visit ScholaFly: https://scholafly.com

Read the transcript

Every letter you have ever typed left the keyboard as a number. Capital A goes out as sixty-five. Capital B goes out as sixty-six. Nothing in the electronics forces those particular numbers. They come from an agreement people wrote down in nineteen sixty-three, and your message arrives as words only because the machine at the other end holds the same copy. And that agreement had to settle one thing first: how many bits to spend on a single character. Spend too few and most of the world's alphabets will not fit. Spend too many and every piece of text you own gets bigger.

You are on video two of six in Units, characters and file size. Every number in this one gets counted in bits, so if that is not solid yet, start with C S oh two, oh one, Bits, bytes and the units of storage.

First, what a character set actually is. In this topic the definition is the content, and it is marked word by word. A character set is all the characters a computer can represent, and every one of them is given its own unique binary code. Two words in that sentence do the work. The first is all: the set holds every character the machine can handle, letters, digits, punctuation, the space, and the invisible ones like a new line. The second word is unique. No two characters share a code, and no code appears twice anywhere in the set. That has to be true or the whole thing collapses. If the code eighty-nine sat on two different rows, a machine reading eighty-nine out of a file could not decide which character to put on your screen. So here is your handle for this video: the set is the box, and the code is one row inside the box. Point at a single row and you have named one character and its code. Draw a boundary round every row and that whole thing is the set. Three students have each written a sentence saying what a character set is. Only one of the three is right. A says it is a method for turning letters into binary. B says it is all the characters a computer can represent, each with its own unique code. C says it is the letter U and its code, eighty-five. Pick one. I'll wait. It is B, and B is the only one carrying both of the words that matter: all the characters, each with its own unique code. A describes a process instead of a thing, which is the commonest way this sentence goes wrong. A character set turns nothing into anything. It is an agreed list, and the machine looks characters up in it. It is not encryption, and it is not a programming language. Encryption hides meaning on purpose, while a character set publishes it so every machine agrees. C is true, but it is one row out of the table. It names a character and its code, which is not the same thing as naming the set. Give the box, never the row, and get the words all and unique into the sentence you write down.

The table itself comes next - ASCII, the American Standard Code for Information Interchange - and no board expects you to remember a single number in it. The part you need is printed for you in the question. What is examined is whether you can read it in both directions. Here is the set on screen. First direction: give the code for capital P. Run down the character column to P, then read straight across to the number. Capital P sits on the row holding eighty, so the code for capital P is eighty. Other direction now. Find the character that has the code seventy-four. This time you go down the code column first, then read back across to the character. Pause it there and work it out. I'll wait. Seventy-four is sitting on the row for capital J, so the character with code seventy-four is capital J. Give both halves when you answer: the number is the code, and the symbol is the character it stands for. Here is the sharp one. Lower case n has the code one hundred and ten, and capital N has the code seventy-eight. Those are not two spellings of one letter. They are two different characters, and the set gives every character its own row, so they hold two different codes. Which means a capital and a small letter are never swappable in an answer. Write the exact one the question asked for. There is an order running through those numbers, and C S oh two, oh three, Character codes run in order, does the arithmetic with it. Search ScholaFly C S oh two, oh three to find it. One table, read two ways, and the capital and the small letter are always two separate rows.

That box has a size, and its size is decided by how many bits you hand to one character. So, two counting questions together. How many different characters can a seven-bit code hold, and how many can an eight-bit code hold? Try it on paper before I show you. I'll wait. Seven bits give one hundred and twenty-eight different characters. Eight bits give two hundred and fifty-six. Add one bit and the count doubles, because every pattern you already had now comes in two versions: one ending in a zero, one ending in a one. So n bits give two to the power of n different codes. That relationship is the examined part, not the numbers on the rows. The extra bit buys room, and only room: one hundred and twenty-eight more rows on top of the original set, for accented letters and extra symbols. That doubled version has a name of its own, Extended ASCII, and it becomes important in a minute, because the boards handle it differently. Two hundred and fifty-six rows still will not hold every writing system on the planet. The much bigger set that does is Unicode, and C S oh two, oh four, Unicode, is the video for it. Bits per character sets the size of the box, and every bit you add doubles how many characters can live inside it.

Right, the exam side of this, and it opens on a line of genuinely good news. O C R reported on a twenty twenty-three Computer systems question about character sets, and said this about the answers. Many responses accurately identified that a character set stores all the characters. So if the word all was already in your sentence, the core of this is understood, and that puts you with most students on it. What is left is the second word, and the number. The number is where your board matters more than anywhere else in this topic. Three boards, three true statements, and you only need the one that is yours. A Q A examines ASCII as a seven-bit code. Their report on the twenty twenty-four Computing concepts paper puts it plainly. For the calculation many students incorrectly thought that ASCII is an 8-bit code. The reason is the one from a minute ago. Extended ASCII is not required for that specification, so seven is the version being examined there. Edexcel also examines seven, and its mark scheme refuses eight. Their report on the twenty twenty-four Principles of Computer Science paper says this. Whilst many candidates did give the correct answer there were also a significant number who gave '8' as the response, which is the number of bits used by extended ASCII. O C R goes the other way, and this one is a decision about the exam rather than a fact about ASCII. Their report quotes the specification itself. The specification for J277 states that in the exam ASCII will be described as having 8-bits to avoid confusion between ASCII and extended-ASCII, which are not differentiated in the specification. So the answer is not seven, and it is not eight. It is seven for A Q A, seven for Edexcel, and eight for O C R, and each of those is correct on its own paper. Which makes the working rule a short one. Never let that number travel on its own. Say it with the board attached, and look at the name printed on the front of your paper.

Before the end, here is the whole video pulled back into four lines. A character set is all the characters a computer can represent, each with its own unique code. The set is the box, and the code is one row. The table is given to you, and you read it both ways: character across to code, or code back across to character. Bits per character fixes how many rows fit. Seven bits, one hundred and twenty-eight. Eight bits, two hundred and fifty-six. n bits, two to the power of n. And the number wears a board badge: seven for A Q A and for Edexcel, eight for O C R, because O C R's papers do not separate ASCII from Extended ASCII.

Quick bit of bookkeeping, and it is entirely for your benefit. A thumb on a video says you could read that code table both ways without help, so revision week can go straight past it. If bits per character is still wobbly, leave the thumb off and park this in a playlist for the weekend, since it reads very differently once the table is not new. Then carry on to C S oh two, oh three, Character codes run in order, which puts the order inside those codes to work.

Next in the chapter: Character codes run in order.

For more, visit scholafly.com, or watch the next video.

Related terms

For: AQA GCSE 8525, Edexcel GCSE 1CP2, OCR GCSE J277

On the specification

BoardSpecStatement
AQA GCSE 85253.3.5Character encoding
Edexcel GCSE 1CP22.2.1Explain how computers encode characters using 7-bit ASCII.
OCR GCSE J2771.2.4Data storage - Numbers
For teachers

This GCSE Computer Science lesson teaches character sets and ASCII. By the end, students should be able to explain what a character set is, describe how characters are represented in binary using ASCII, and relate the number of bits per character to how many characters can be represented - naming which board uses 7 bits and which uses 8. It works through four worked examples and the mistakes examiners report.