Bits and positional notation
A bit has two possible values, 0 and 1. An eight-bit byte has 256 possible patterns. Meaning comes from an agreed interpretation: the same byte might represent an integer, part of a character, or an instruction. In binary, positions from right to left have weights 1, 2, 4, 8 and so on. Thus 1101 means 8 + 4 + 1 = 13. Hexadecimal uses weights 1, 16, 256; digits A–F mean 10–15. Each hexadecimal digit corresponds to four bits, so 0x2D is 0010 1101, or 45 decimal.
To encode a nonnegative integer, repeatedly divide by 2, record the remainders, then read them in reverse. Leading zeroes change the chosen width, not the unsigned value.
What decimal value is binary 10110?
Worked solution
16 + 4 + 2 = 22.
Width, range and overflow
With w bits there are 2^w patterns. Unsigned integers range from 0 to 2^w−1; eight bits allow 0–255. Adding one to 255 requires nine bits. A fixed-width system must define what happens: reject, wrap, saturate or report an error. These are different behaviours, not interchangeable repairs. Our converter rejects values outside its range.
Python integers grow as needed, subject to memory, so Python's 255 + 1 does not itself wrap at eight bits. To model eight-bit wrapping explicitly, use (value + 1) % 256. Storage formats and hardware interfaces still impose widths even when your language's integer type does not.
How many unsigned values fit in six bits?
Worked solution
64 values: 0 through 63.
Signed integers and byte order
Two's complement assigns the highest bit weight −2^(w−1), with ordinary positive weights for the others. At eight bits, 11111111 is −128+127 = −1; 10000000 is −128. The range is −128 to 127, not −127 to 127. Decode an unsigned pattern u by subtracting 256 when its top bit is set.
A multibyte value also needs byte order. For 0x1234, big-endian storage puts 12 then 34; little-endian puts 34 then 12. This is a byte-order choice, separate from signedness. When exchanging data, specify width, signedness and byte order together.
Interpret eight-bit 11111110 as unsigned and signed.
Worked solution
Unsigned 254; two's complement 254−256 = −2.
Text is not a byte sequence until encoded
Unicode assigns code points to text symbols; UTF-8 encodes those code points into bytes. ASCII characters take one UTF-8 byte, while many other code points take several. The string A中 has two code points but four UTF-8 bytes. Python str represents text; bytes represents a byte sequence. Encode when crossing a byte-oriented boundary, and decode using the agreed encoding when reading.
A displayed character can comprise multiple code points, such as a base letter plus a combining accent. Therefore code-point length is not always the number of visible characters. Splitting encoded text at an arbitrary byte boundary can cut a multibyte sequence and make decoding fail. Store and transmit the encoding agreement as part of the format.
Why can len(text) differ from len(text.encode('utf-8'))?
Worked solution
The first counts code points; the second counts bytes, and a code point can need multiple bytes.
Floating point and exact alternatives
A floating-point format stores a sign, a scaled significand and an exponent. Finite binary fractions have denominators that are powers of two. Decimal 0.1 is 1/10, so its binary expansion repeats and finite storage rounds it. On usual Python platforms, 0.1 + 0.2 differs slightly from 0.3. This is a representation issue, not evidence that addition is arbitrary.
Compare approximate measurements with tolerances chosen for the problem, using both relative and absolute tolerance near zero. For exact money amounts, consider integer minor units or Decimal constructed from strings. Decimal also has a finite precision context: it is not exact for every possible division. Fractions can preserve rational values exactly at the cost of growing numerators and denominators.
Is binary 0.5 exact? Is binary 0.1 exact with finite bits?
Worked solution
0.5 = 1/2 is exact; 0.1 = 1/10 is not.
Common misconceptions
- Bit patterns need an interpretation; a leading 1 does not imply negativity without a signed format.
- Byte length, code-point count and visible character count are different.
- Changing the printed precision does not change the stored float.
Lab setup
Download each script and run it in a terminal with Python 3.11 or later: python m02_binary.py. On Windows, py -3 is an alternative; on some systems use python3. The labs use only the standard library. Predict the result before running, then complete the variations. Run without -O so assertions remain enabled. Outputs below were captured by the builder. Code and output are identical in both language editions.
Lab 1 — Build a converter
Follow the remainders through the loop. The script checks the round trip and explicitly rejects overflow.
"""Base conversion and fixed-width signed interpretation."""
def binary(value, width=8):
if not 0 <= value < 2 ** width:
raise ValueError("value does not fit unsigned width")
bits = []
for _ in range(width):
bits.append(str(value % 2))
value //= 2
return "".join(reversed(bits))
for value in [0, 13, 127, 128, 255]:
bits = binary(value)
signed = value if value < 128 else value - 256
assert int(bits, 2) == value
print(f"{bits}: unsigned={value:3}, signed={signed:4}, hex={value:02x}")
try:
binary(256)
except ValueError:
print("256 rejected at width 8")
else:
raise AssertionError("overflow must be rejected")
00000000: unsigned= 0, signed= 0, hex=00
00001101: unsigned= 13, signed= 13, hex=0d
01111111: unsigned=127, signed= 127, hex=7f
10000000: unsigned=128, signed=-128, hex=80
11111111: unsigned=255, signed= -1, hex=ff
256 rejected at width 8
- Predict the output for 42.
- Use width 4 and test 15 and 16.
- Explain why the signed interpretation changes at 128.
Worked solution
42 is 00101010 and hexadecimal 2a. At width 4, 15 is 1111; 16 is rejected. The eight-bit sign bit has weight −128, so patterns at or above 128 represent u−256 when interpreted as signed.
Lab 2 — Inspect bytes and rounding
Run the text round trip and compare float and Decimal. The truncated text case deliberately verifies an error.
"""Separate text from bytes, and approximate numbers from decimal arithmetic."""
from decimal import Decimal
from math import isclose
text = "A中"
encoded = text.encode("utf-8")
print("code points:", len(text), "bytes:", len(encoded), "hex:", encoded.hex())
assert encoded.decode("utf-8") == text
print("float:", 0.1 + 0.2, "exact equality:", 0.1 + 0.2 == 0.3)
assert isclose(0.1 + 0.2, 0.3, rel_tol=1e-12, abs_tol=1e-12)
exact = Decimal("0.1") + Decimal("0.2")
print("decimal:", exact)
assert exact == Decimal("0.3")
try:
encoded[:2].decode("utf-8")
except UnicodeDecodeError:
print("truncated UTF-8 rejected")
code points: 2 bytes: 4 hex: 41e4b8ad
float: 0.30000000000000004 exact equality: False
decimal: 0.3
truncated UTF-8 rejected
- Replace the text with ASCII-only ABC.
- Try encoding the original text as ASCII.
- Compare Decimal('0.1') with Decimal(0.1). Explain the difference.
Worked solution
ABC has three code points and three UTF-8 bytes. ASCII cannot encode 中, so it raises UnicodeEncodeError. Constructing Decimal from a float preserves that float's already rounded value; constructing from '0.1' preserves the intended decimal input.
Exercises with worked solutions
Try before opening the solution. ★ applies an idea; ★★ combines ideas; ★★★ asks for design or proof.
Write 45 in eight-bit binary and hexadecimal.
Worked solution
45=32+8+4+1, so 00101101. Group into 0010 and 1101: 0x2D.
Give the unsigned range of ten bits.
Worked solution
There are 2^10=1024 patterns, giving 0–1023.
Encode −5 in eight-bit two's complement.
Worked solution
256−5=251, whose binary pattern is 11111011. Decode it as −128+123=−5.
Store 0x0102 in two bytes in both byte orders.
Worked solution
Big-endian: 01 02. Little-endian: 02 01. Both encode 258 under their respective agreement.
Why does slicing the first two bytes of A中 fail to decode?
Worked solution
A takes one byte and 中 takes three. The slice retains A and only the first byte of 中, leaving an incomplete UTF-8 sequence.
Which of 1/8 and 1/10 has a terminating binary expansion?
Worked solution
1/8=0.001 binary terminates because 8 is a power of two. Reduced denominator 10 includes a factor of five, so 1/10 repeats.
Design storage for prices with exactly two decimal places and no fractional cent.
Worked solution
Store integer cents: 12.34 becomes 1234. Sum cents exactly; divide and format only at display boundaries. Specify currency, maximum amount and rounding policy for operations such as tax.
A record contains an unsigned count up to 1000. Specify width, signedness and byte order.
Worked solution
Ten bits suffice mathematically. A practical byte-aligned field can use two unsigned bytes, big-endian, accepting only 0–1000. Reject reserved/out-of-range values rather than assuming every 16-bit pattern is valid.
Self-check quiz
Choose an answer for feedback; reset to retry. A text answer key is available without JavaScript.
How many patterns in a byte?
Eight-bit 10000000 as two's complement?
Does Python int automatically wrap at 255?
What does UTF-8 encode?
Which fraction is exact in finite binary?
Which construction preserves decimal input 0.1?
Answer key
- A — Eight independent bits give 2^8 patterns.
- B — The highest bit has weight −128.
- C — Python integers are not limited to eight bits.
- A — Text and its byte representation are distinct.
- B — Four is a power of two.
- C — The string avoids an intermediate binary approximation.
Guided reading
- Python floating-point tutorial — Read the representation-error explanation; describe when a tolerance is appropriate.
- Python Unicode HOWTO — Read encodings and decoding errors; distinguish str and bytes.
Review and the next step
Encode 29, decode an eight-bit negative value, and explain why A中 needs four UTF-8 bytes. Choose a representation for a catalogue's book count and price. Next, Module 03 builds a program that uses these values.
Key terms
| Term | Meaning |
|---|---|
| Bit | A binary digit with two possible values. |
| Two's complement | Signed interpretation with a negative highest-bit weight. |
| Encoding | An agreed mapping from values to representations. |
| Rounding | Choosing a representable value near the intended value. |