Lobsters retrocomputing - 24 Aug 2026

Page 1 of 3

[TOC]

For generations, large systems have been built in layers, with well-defined interfaces between each layer. This design allows any old layer to be replaced with a newer, better implementation, without disrupting the rest of the system. Eventually, though, the once-good design becomes too outdated and the whole lot must be replaced with a new design and a new implementation that better fits current needs.

But in recent years, computing technologies have advanced at a ferocious pace and IT systems have become immensely complex. To keep pace with the rapidly changing worldwide market, the IT industry takes the expedient approach: it builds newer, more complex software technologies atop outmoded, simple technologies, instead of taking the correct - but costly, time-consuming, and disruptive - approach of replacing old technologies with newer, better ones. This situation is due in part to the economics and practices of modern software development and in part to the deeply enmeshed nature of some older, shortsightedly planned, hastily developed technologies.

The progressively deep and complex layering of software technologies moves programmers farther and farther away from the hardware, so much so that many experienced programmers today no longer understand something as fundamental as the instruction execution cycle or system reset sequence. Surely, a software developer can have a profitable career working on social media and line of business applications in JavaScript, without ever having peered inside a computer case. But software development is more than captivating user interfaces, elegant GraphQL requests, undocumented microservices, and neck-deep layers of NPM packages. To be effective, a programmer must, at the very least, possess an accurate mental model of how a computer executes programmes.

I wrote this article for programmers who work exclusively with software, far above hardware. My goal here is modest; I aim to explain the most basic aspects of computing: how a computer is built from disparate subsystems, how these subsystems collectively perform simple calculations, and how complex system behaviour emerges from combinations of just two symbols, a 0 and a 1.

But who really cares about such lowly matters? Every IT practitioner should, I say. Recognising that things happen when they do is a sign of intelligence, a commonplace phenomenon within Animalia. But the drive to understand why things are as they are - that is a mark of intellect, a quality unique to Homo sapiens.

Bears, birds, and many other species can plan for the immediate future and the welfare of their spawn. But we humans, unlike our fellow animals, possess a unique ability to plan far into the future and take actions that are designed to benefit the many generations of progeny. But to exercise this ability, we are obliged to assume the burden continually to observe, record, analyse, calculate, and predict. To lighten this weighty responsibility, we have over millennia amassed a collection of ingenious tools and techniques. Thousands of years ago, we used stones to build computers that can measure the passage of seasons, so that we may plan a bountiful harvest thereby saving the tribe from starvation, come winter. Today, we use silicon, stones in a different guise, to make supercomputers that can predict the paths of winter storms, so that we may save the lives of millions of coastal residents.

The term "computer" can be applied broadly to any mechanical means that can be used to ease the task of calculation. So, beans and bones are as much "computing devices" as are slide rules and supercomputers. And until the 1940s, a "computer" was the person who performed numerical computations using the slide rule. In this article, "computer" means the modern digital computer. With all its hardware sophistication (and software sophistry) the modern digital computer is still a mechanical aid for calculating. Yet, simple calculations like additions and subtractions, when mixed with human ingenuity, achieves marvellous, and seemingly miraculous, feats like weather prediction simulation, air traffic control, global positioning satellites, artificial intelligence (AI), and even Minecraft.

I shall now describe how a computer jacks itself up from twiddling bits to playing chess: from bits to logic and arithmetic operations; from arithmetic operation to comparison relation; from comparison to looping; then to algorithms and finally AI.

bits and bytes - A digital computer can manipulate only binary digits (bits). A bit can be either 0 or 1, the two stable states of the flip-flop. A flip-flop represents the 0 and the 1 digital logic symbols using two different voltage levels, a low (typically $0.0\ V$ DC) and a high (traditionally $5.0\ V$, but $3.3\ V$ in modern devices). The flip-flop stores a logic symbol. It stands at the voltage-bit boundary.

An example of a simple, modern computer is the 1980s home computer. Such a computer is called the "8-bit machine", because it manipulates data 8 bits, or one byte, at a time. As a consequence of this byte-at-a-time processing scheme, the CPU can access a particular bit only through the enclosing byte.

In the early 1970s when the microprocessors first emerged, they were rather costly. So, the 4 bit nybble (also spelled nibble) was used as the datum on some early machines. But every modern CPU uses the 8-bit byte as the basic unit of data. The byte is also known as the word on an 8-bit machine. But the size of the word matches the CPU's native datum size. So, on a 16-bit, 32-bit, or 64-bit computer, a word is respectively those quantities.

These days, even the smallest of computers, like Raspberry Pi or iPhone, employ 64-bit CPUs. But the 8-bit CPU is most assuredly not gone, nor is it forgotten; the popular Arduino single-board microcontroller platform uses a modern, sophisticated 8-bit CPU. Indeed, an 8-bit machine is the simplest possible full-fledged computer. So, we focus on the 8-bit architecture in this article. This keeps the presentation simple, yet relevant.

Throughout history, humans have used base-12 and base-10 number systems. And computer scientists routinely use base-16 and base-8, in addition to base-2. Ordinarily in number systems, the larger the base, the greater the representation power. Why, then, digital computers use the miserly base-2? As most things in engineering, the answer is the cost of complexity.

In electronics, the 0 symbol of the binary system is commonly represented with the $0.0\ V$ DC and the 1 symbol with the $3.3\ V$. The hardware can distinguish between these two voltage levels reliably and efficiently, because there are only two possibilities - either a low or a high - and the two levels are adequately separated so as to avoid ambiguity. Why not use more than two states, say with the base-8 system? Well, it turns out that squeezing several, closely-spaced voltage levels into the same overall range decreases the voltage separation between the levels, which increases the likelihood of level detection errors, which, in turn, must be compensated with increased hardware complexity and commensurate rise in production cost. A more complex hardware also consumes a greater amount of energy during operation. On the other hand, broadening the overall voltage range may lower level detection error rates somewhat, but it, too, raises the machine's energy consumption. So, modern digital implementations use the binary logic system and a relatively small DC voltage range $[0.0, 3.3]\ V$, for efficiency, simplicity, and economy.

A boolean is a logical value that can be either a $False$ or a $True$. A boolean is the fundamental operational unit of mathematical logic. Logic lies at the foundation of modern mathematics. A boolean value is represented quite naturally inside the CPU as an individual bit: the $False$ with the 0 and the $True$ with the 1. This bit-boolean boundary is where the voltage level, a lowly bit quantity stored inside a flip-flop, accedes to the truth value, a lofty mathematical concept of the mind. Here, the $0.0\ V$ DC level, or the 0 bit, becomes the logical $False$, and the $3.3\ V$, or the 1 bit, becomes the logical $True$.

Logical operators ($$, $$,$$, etc.) that manipulate and combine boolean values are directly implemented inside the CPU using logic gates constructed of transistors. Logic gates are commonly implemented using only NAND gates, due to the functional completeness property that allows a combination of NAND gates to implement arbitrary boolean expressions. Using just one gate type simplifies the design and manufacture of chips and reduces their cost.

Traditionally, a byte was the number of bits needed to represent an alphanumeric symbol. At one time, 6 bits, 7 bits, and 9 bits were used as the byte datum, alongside 8 bits. But since the 1980s, all computer architectures have converged on the 8-bit byte. There are several reasons why the 8-bit byte became the de facto standard. The fact that 8 is a multiple of 2 has something to do, no doubt. Electrical engineers and computer scientists have an affinity for the numbers that are multiples of 2, especially those that are the powers of 2. And the 8-bit byte's range of $2^8 = 256$ distinct values is the smallest range with the capacity to represent all the commonly used symbols. These symbols were later codified into the American Standard Code for Information Interchange (ASCII) character set. This byte-alphabet boundary is where a mere collections of bits transmogrify into numbers and letters, the units of symbolic thought.

numbers and letters - Just as a decimal (base-10) digit can represent $0$ up to $9$, a binary (base-2) digit can represent 0 up to 1. The decimal number system requires multiple digits to represent a number greater than $9$. Likewise, the binary number system requires multiple digits to represent a number greater than 1. Hence, decimal $0$ is represented as binary 0, $1$ as 1, $2$ as 10, $3$ as 11, $4$ as 100, and so on. The minimum value a byte can represent is 00000000 (decimal $0$), and the maximum value it can represent is 11111111 (decimal $255$). Thus, a byte can represent $2^8 = 256$ different values. Numbers greater than $255$ must be represented by adjoining two bytes, which can represent $2^{16} = 65,536$ distinct values. An 8-bit CPU manipulates data one byte at a time. So, to load a 16-bit quantity from memory, the CPU must access the memory twice. Since the 1940s, the dawn of the electronic computer, numbers have been represented and manipulated using many different formats. But all modern computers now use 2's complement format to represent signed integers and IEEE 754 floating-point format to approximate real numbers.

To represent character symbols, computers use ASCII and Unicode standard formats. ASCII was adopted as the standard binary format in the 1960s, with the advent of the modern digital computer. The numbers, letters, and myriad symbols that appear on the American standard keyboard are mapped to the ASCII symbol table. Although the original ASCII specification used 7-bit characters, the modern version uses 8-bits to represent one character, so it can represent up to 256 different alphanumeric symbols. For instance, the word "hi" in ASCII binary representation requires two bytes: 01101000 for h, followed by 01101001 for i. But tens of thousands of symbols are needed to represent all the languages in the world. The worldwide personal computer revolution in the 1980s kickstarted various efforts to represent complex characters in multiple bytes, which culminated in the Unicode standard of the 1990s. The ASCII standard 8-bit character format has now been subsumed into the Unicode standard as UTF-8.

Note that even an advanced, modern CPU supports only the following data types: 1-bit boolean, 8-bit integer (byte), 16-bit integer, 32-bit integer, 64-bit integer, 32-bit single-precision floating-point, 64-bit double-precision floating-point, and 64-bit pointer (memory address). The character value is represented as a Unicode unsigned integer quantity as described above, and a string value is represented as a sequence of character values. All other data types supported by high-level programming languages are represented in hardware as a combination of pointers and byte sequences. Now that we know how numbers are represented in the hardware, we are ready to discuss arithmetic operations that manipulate those numbers.

arithmetic and logic - The computer represents numbers and letters using standard numeric formats like 2's complement, IEEE 754, and Unicode. But representation is one thing, computation is quite another matter. The most basic arithmetic operation that a digital computer performs is the addition. The CPU adds numbers using a circuit with an unimaginative name - the adder. Two binary integers are added just like two decimal integers. In the decimal number system, $9 + 1 = 10$, because adding $1$ to the maximal value that a decimal digit can represent, namely $9$, exceeds the capacity of a digit, and hence must be represented by two digits. Similarly, in the binary number system, 1 + 1 = 10 (decimal $2$), because adding 1 to the maximal value that a bit can represent, namely 1, exceeds the capacity of a bit, and hence must be represented by two bits. When bits are assembled into an 8-bit quantity, this quantity can represent a total of $2^8 = 256$ different values from 00000000 (decimal $0$) to 11111111 (decimal $255$). This is analogous to an 8-digit decimal quantity being capable of representing $10^8 = 100,000,000$ different values (from $00,000,000$ to $99,999,999$).

But what about negative numbers? In the 2's complement format, the left-most bit of a binary quantity represents the negative sign: 0 is positive, and 1 is negative. Hence, 00000001 is +1, and 11111111 is -1. Because this format reserves the most significant bit for sign representation, the largest positive integer an 8-bit quantity can represent is 01111111 (decimal $127$) and the smallest negative number is 10000000 (decimal $-128$). An 8-bit singed integer, therefore, can represent 128 negative values, the $0$ which is conventionally treated as a positive value, and 127 other positive values, giving a total of 256 different values. In other words, an 8-bit quantity, singed or unsigned, can represent $2^8 = 256$ different integer values. So, an unsigned byte can represent decimals in the closed interval $[0, 255]$, and a signed byte can represent decimals in the closed interval $[-128, +127]$. By using the 2's complement representation, the integer adder unit can compute subtraction using the same procedure used to compute addition: $7 + (-9) = -2$.

Inside the CPU, a dedicated hardware unit, called the arithmetic logic unit (ALU), performs integer arithmetic (add, sub, mul, div) and boolean logic (not, and, or, xor, shift, rotate). Hence, the ALU is the core part of the CPU, and it is the ALU that gives the CPU its computational prowess. Traditionally, the 8-bit CPUs performed floating-point arithmetic in software, because including specialised, floating-point hardware would have made the machines uneconomical for the home market of the time. Even more powerful 16-bit CPUs of that era, Intel 8086 and Motorola 68000 for example, also lacked floating-point hardware, but they could be augmented with an optional (and expensive) add-on floating-point unit (FPU) that supported floating-point arithmetic (fadd, fsub, fmul, fdiv). The FPU was then known as the floating-point coprocessor. Modern CPUs, starting with the 32-bit Intel 80486 processor, have integrated, on-chip FPU hardware.

Now, we know how the CPU represents boolean values, characters, and numbers, and how it manipulates those values using arithmetic and logic operations. Next, we shall examine how conditions and loops are implemented in the hardware.

decision and repetition - A sequence of bits can represent letters and numbers. And the adder unit can add and subtract numbers. But programmes must also make decisions based on conditions:

...

jump a < b csq alt

csq: ; consequent code block

...

alt: ; alternative code block

...

For clarity and concision, the above code is written in pseudo assembly language. It implements the if-then-else conditional branch, which is supported by all high-level programming languages. The mathematical relation a < b, where a and b are integers, represents a condition that a is less than b. It turns out that any general decision can be implemented using comparison operations. Navigational decisions, for example, can be implemented as follows: turn left if a < b; go straight if a = b; turn right if a > b.

When comparing a and b, the CPU first computes the subtraction a - b, then checks if the result is negative, zero, or positive. A negative means a < b; a zero means a = b; a positive means a > b. By following one execution path or another on the basis of this comparison result, the CPU accomplishes conditional branching. This comparison-decision boundary is where the comparisons inside the ALU take on the semblance of the decisions of the mind.

Computers are designed to perform the same set of instructions repeatedly, tirelessly. A repetitive execution of a block of assembly code, also known as looping, is accomplished by jumping back to the top of the block, based on the conditional branch located at the bottom of the block:

... ; before loop

n <- 10 ; define maximum loop index

i <- 0 ; initialize loop index

loop: ; start of loop

... ; perform iteration i of loop

i <- i + 1 ; increment loop index

jump i < n loop end

end: ; end of loop

... ; after loop

Above, integer i is the loop index and n is the maximum number of iterations. These variables are initialised prior to loop entry. The index i is incremented after each iteration. While i < n, the CPU jumps back to the top of the loop to the label loop:. But when i = n, the loop termination condition is met, so the CPU jumps out of the loop to the label end:, and continues executing the code that follows the loop. Traditionally, we do not use variables in assembly language as we do in a high-level language; instead we use registers and memory addresses, which are expressed in hexadecimal values. But in this article, we use variables for clarity.

We now know how the CPU performs decisions and repetitions. Next, we shall see how the CPU can compute sophisticated algorithms using the above-mentioned simple operations: arithmetic and logic, decisions, and repetitions.

algorithm and intelligence - A CPU equipped with arithmetic-logic operations and branching-looping instructions is capable of universal computations. Such a CPU is said to be Turing complete. For example, exponentiation $b^e = b \cdot b \cdot\ ...\ \cdot b$ is computed as $e$ repeated multiplications of $b$, and factorial $n! = (n - 0) \cdot (n - 1) \cdot\ ...\ \cdot 2 \cdot 1$ is computed as $n$ repeated multiplications of $n - i$, where $i$ ranges over the open interval $[0, n)$. The exponentiation algorithm uses looping and multiplication, and the factorial algorithm uses looping, subtraction, and multiplication. A more complicated algorithm, $sin(a)$ where $a$ is an angle in radians, can be computed using Taylor series expansion with a sufficiently large number of iterations $n$:

$$

sin(a) = \sum_{i=0}^{n} \frac{(-1)^i}{(2i+1)!} a^{2i+1}

$$

Next page | More Lobsters retrocomputing | Headlines

Original: https://amenzwa.github.io/stem/ComputingHistory/How8bitComputersWork/