Lobsters retrocomputing - 24 Aug 2026

Page 2 of 3

The sine algorithm and its constituents, the exponentiation and the factorial, use looping, addition, subtraction, multiplication, and division. In this way, arbitrarily complex computations can be constructed from combinations of simple ones.

Algorithms like these are implemented as functions that take inputs and return results. Functions may be invoked using the call instruction from anywhere in the code. A function call, just like a branch, alters programme execution flow. But there is a difference: upon encountering a branch the CPU jumps immediately to the designated code block and the execution continues from there; but before entering the function body code block, the CPU must assemble the input values for the function to use and also must arrange for the function to return its result to the calling code. In this way, when the function finishes, execution control returns to the caller, along with the function's result.

More sophisticated functions, like sorting, searching, and graph algorithms, are implemented by composing simpler functions. Complicated AI algorithms like chess and Go use sorting, searching, and graph algorithms. This algorithm-intelligence boundary is where a semblance of higher order thinking emerges from the seething primordial foam consisting of just 0s and 1s.

By its very name, "artificial intelligence" is not intelligence, but an imitation thereof. Thinking, feeling machines have not been invented, because humans do not yet grasp the inner workings of human intellect. There are two broad categories of AI algorithms: rule-based AI and connectionist AI. Rule-based AI is used when the problem is well understood and the solution can be expressed as a small set of clearly defined rules. Board game algorithms are prime examples of rule-based AI. Connectionist AI, more popularly known as artificial neural networks (ANNs), is used when the problem is poorly understood, has a large set of potential solutions which makes it impracticable to describe it in terms of rules, and a vast amount of data can be collected beforehand. Autonomous vehicle navigation, crowd behaviour modelling, and similar open-ended problems are solved using ANNs. I explain these concepts in more detail in my article How Artificial Intelligence Works.

In this section, we explore the software aspects of computing, in particular the software that is the closest to hardware - the machine code. Machine code consists of only 0s and 1s. It is generated from assembly language, which is represented in human-readable mnemonics, like add, sub, jump, call, and so on. There is one-to-one correspondence between machine instructions and assembly mnemonics. That is, an assembly programme can be assembled into its equivalent machine code programme, and the machine code programme can be disassembled back into the original assembly programme.

Every CPU family implements its unique set of machine instructions. This is known as the instruction set architecture (ISA) of the CPU. A typical ISA comprises instructions for clock and peripheral configuration, for arithmetic and logic, for memory access, and for input and output. The ISA is the programmer-visible aspects of the hardware.

In a typical 1980s 8-bit home computer system, memory was a scarce resource, so the ISA had to provide ways for programmers to access memory efficiently. Consequently, instructions that access memory on these CPUs were numerous and intricate. A typical 8-bit CPU, for example, supports many different memory addressing modes: immediate, absolute, relative, implied, indirect, indexed, and combinations thereof. The complicated details of these addressing modes are not relevant to our discussions. But recognise that this complexity enabled a low-end CPU to do as much work as necessary using as few bits of instruction representation as possible. This increased code density of the ISA and thus reduced memory requirements of programmes.

The only abstraction mechanisms that assembly language offers are the ability for the programmer to allocate memory locations for data storage and to organise code into callable procedures. And the most important procedures on an 8-bit home computer were the interrupt service procedures and the system service procedures, which were all stored in the ROM. The ROM also held the ubiquitous BASIC interpreter. Interrupt service procedures are handlers for I/O, software errors, and hardware exceptions. System service procedures are analogues of system calls in modern operating systems.

When the 8-bit machine boots up, the CPU performs a system-wide hardware reset. Hardware reset clears CPU's internal registers (accumulator, general purpose registers, programme counter, etc.) and also clears peripherals' registers (configuration registers, data transfer registers, etc.). This puts the whole system in the well-defined initial state. The CPU then executes the reset interrupt service procedure. The CPU knows where to find this procedure, because the entry address of the procedure is conventionally stored in the ROM at a well-known address, such as 00000000,00000000 or 11111111,11111111. The reset procedure first checks the health of RAM and I/O devices. It then completes the system initialisation by performing a few housekeeping tasks, like setting up the interrupt service vector table in RAM, putting status register bits to appropriate initial values, and so on. The vector table is a list of addresses for various service procedures for handling timer, I/O, software error, hardware exception, and other interrupts.

Once the system has been fully initialised, the reset procedure starts the BASIC interpreter. On the classic 8-bit home computer, the BASIC interpreter served not only as the high-level language interpreter, it also served as a programme editor and an operating system command line interpreter that provides access to system services, such as loading programmes from external storage devices.

While home users invariably used BASIC, commercial software were typically written in assembly. A programmer-defined procedure in assembly language is just a sequence of instructions grouped under a label. The following pseudo-assembly procedure (a function, actually) doubles its integer argument:

...

dbl: ; double the input integer supplied in accumulator A

mul 2 ; multiply value in A with 2, store result in A

return ; pop (SR+retaddr) off stack at SP, load PC with retaddr

...

This procedure can be invoked as follows:

...

A <- 13 ; load A with argument value 13

call dbl ; push (retaddr+SR) on stack at SP, load dbl's addr in PC

A -> mem ; store result in A to RAM location mem

...

The above pseudo assembly code is stripped to its essentials so as to highlight how procedures are defined and invoked in assembly, using addresses, programme counter (PC), stack pointer (SP), accumulator A, call instruction, and return instruction. To call the dbl procedure, the caller passes $13$ as the argument to dbl via the accumulator A, and dbl returns its result value, $26$, to the caller, also via A. The call instruction automatically pushes onto the stack the return address and the contents of the status register (SR). The stack is a designated portion of the RAM dedicated for use by the called procedure. The 16-bit return address is the address of the instruction immediately after the call dbl, which is A -> mem. The 8-bit SR holds the state of the CPU. Next, the CPU jumps to the first instruction in dbl, which is mul 2. When the dbl procedure finishes, the return instruction at the end automatically loads the SR and the PC with corresponding contents popped off the stack. The CPU then jumps to the A -> mem instruction, and programme execution continues from there. The above code sequence would be written in C as follows:

...

int dbl(int a) { return 2 * a; }

...

mem = dbl(13);

...

Every time a procedure is called, requisite memory for use by the called procedure is allocated on the stack and the current state of the CPU (the contents of its registers) is also pushed onto the stack. When the procedure is finished, the state preserved on the procedure's stack frame is used to restore the CPU to its pre-call state, and the stack frame is deallocated to reclaim the memory. If a called procedure calls another procedure, more memory will be consumed to enable that subsequent call. Hence, a runaway recursive call will quickly exhaust the stack memory, and will throw the stack overflow exception.

Although the BASIC language is far removed from the underlying processor, special language constructs provide direct read/write access to memory and registers inside the CPU and the peripherals. This gave hobbyist programmers efficient and precise ways to initialise, configure, and control peripherals, like graphics card, sound card, and storage controller, without having to descend into assembly. Some BASIC implementations even accept assembly instructions embedded inside the BASIC programme.

One way of programatically interacting with peripherals is memory-mapped I/O. A processor that implements the memory-mapped I/O scheme allocates portions of the address space to ROM, RAM, and I/O registers. This method allows the programmer to access both memory and peripherals uniformly using memory access instructions along with sophisticated addressing modes. This approach is simple, effective, and convenient. But because significant portions of the address space are allotted to ROM and I/O, the total addressable RAM is much reduced. All 8-bit machines were equipped with 16-bit address bus capable of addressing $2^{16} = 65,536$ bytes (64 KB). Early machines were equipped with only 16 KB or 32 KB of RAM. In those minimalist configurations, the convenience of memory-mapped I/O outweighed the loss of usable RAM addresses. The popularity of memory-mapped I/O is ever greater, now; the ARM processor, arguably the most successful processor architecture in history, uses this scheme in its Cortex-M line of modern microcontrollers. And because the Cortex-M's 32-bit address bus has $2^{32} = 4,294,967,296$ bytes (4 GB) of address space and most microcontrollers are equipped with only a few hundred KB of ROM and a few tens of KB of SRAM, the loss of a few addresses to memory-mapped I/O devices is hardly noticeable.

A classic 8-bit CPU handles I/O in two ways: synchronous and asynchronous. In synchronous I/O mode, the CPU would send a block of data to the peripheral, say a storage device, and would wait for this much-slower device to complete receiving the data block, before sending another block to it. The time the CPU spent waiting for the peripheral could not be put to profitable uses, such as handling another device or performing calculations.

In asynchronous I/O mode, the CPU cues up a chunk of data transfer by filling a data buffer in RAM, issues an asynchronous-write instruction, and immediately moves on to perform other tasks. When the peripheral has completed receiving the block of data, it sends an interrupt signal to the CPU. When an interrupt is received by the CPU, it completes the currently executing instruction, but immediately jumps to the corresponding service procedure for handling the interrupt. The service procedure then cues up another chunk of data and immediately returns to the interrupted code to resume execution. The interrupt service procedure is invoked just like a normal procedure, as mentioned above. Indeed, the reset service procedure mentioned earlier is the service procedure associated with the reset interrupt.

There are two types of interrupts: maskable interrupt request (IRQ) and non-maskable interrupt request (NMI). Some interrupt service procedures are critical to the continued functioning of the CPU, and hence they must be attended to immediately and, hence, are designated as NMIs. But most interrupts, especially I/O, are maskable IRQs. IRQs are assigned priorities, based on their relative importance. A higher priority IRQ (say, a keyboard input request) may interrupt the execution of a lower priority IRQ (say, a disc write request) service procedure. In this situation, the CPU will push the state of the disc write service procedure onto the stack, and jump to the keyboard input service procedure. The reset signal is the quintessential NMI. Hence, no matter what the CPU is doing, when the reset signal arrives, it must immediately perform hardware reset, and invoke the reset service procedure.

The only data types a classic 8-bit CPU can natively support are 1-bit flag, 8-bit byte, and 16-bit address. For 16-bit (or 32-bit) integer values, the CPU performs byte operations two (or four) times, and it emulates floating-point values in software. On the other hand, a modern 32-bit processor, like the Cortex-M, supports natively 1-bit flag, 8-bit byte, 16-bit half-word, 32-bit word, and 32-bit address. And if the microcontroller is equipped with a floating-point hardware, as in the case of the Cortex-M4F processor on the RobotDyne Black Pill board, it performs 32-bit single-precision floating-point operations in hardware. The Cortex-M4F, however, must perform 32-bit operations twice for 64-bit double-word integer and emulate in software 64-bit double-precision floating-point operations, since its FPU can only handle single-precision floating-point operations.

The C programming language supports char (8-bit), short (16-bit), int (32-bit), long (64-bit), float (32-bit single-precision), and double (64-bit double-precision) data types, plus 32-bit or 64-bit pointer, depending on the processor architecture. C's native data types closely match those supported by the underlying hardware. Hence, C is the lowest (the closest to hardware), high-level (above assembly) language in popular use. Indeed, many statements in C can either be translated directly into assembly or they can be expressed as combinations of a few mnemonics. Yet, C is far more convenient to use than assembly language, because of its readability and its ability efficiently to support high-order organisational constructs, like data structures, functions, and modules. Modern optimising C compilers can produce code that out performs hand-crafted assembly in most situations. For this reason, embedded programming is done in C, now.

As capabilities of processors have grown rapidly over the past few decades, even minuscule microcontrollers can now cope adequately with the processing overhead inherent in higher-level languages like C++, Python, TypeScript, Rust, and OCaml. The growth of hardware capabilities will inevitably bring about the explosion of software complexity. As software complexity increases, type safety, modularity, composability, and other niceties these higher-level languages offer become necessities. So, in a not-so-distant future, modern high-level languages will become more prevalent in embedded software development.

To keep the discussions in this section short and simple, we shall talk about a fictional computer with minimal functionality: it has a CPU with 16-bit-wide address bus and 8-bit-wide data bus; it has a memory system with ROM and RAM; it executes a programme stored in ROM, and the programme can access the RAM; it can perform simple calculations like addition and subtraction; but, it has no keyboard input, no display of any kind, and no permanent storage. We will not discuss real 8-bit home computers of the 1980s; simple as they were, they still were laden with real-life complexities that can muddy the presentation with pragmatic details.

An implementation of the instruction set architecture in hardware is called the microarchitecture. A component diagram of a simple CPU microarchitecture that can perform arithmetic and logic operations, execute loops, call procedures, and handle interrupts is shown in the diagram below. It consists of the following components: control unit (CU) which comprises control logic, timing logic, and instruction decoder (ID); interrupt controller (IC); instruction register (IR); status register (SR); arithmetic logic unit (ALU); accumulator (A); general purpose register (B); programme counter (PC); stack pointer (SP); internal data bus to which these components are attached; buffer circuit that isolates the internal data bus from the external data bus on the motherboard; and address lines that connect directly to the external address bus on the motherboard. ROM, RAM, and I/O peripherals are located outside the CPU on the motherboard, and so are the address bus and the external data bus.

It is often said that the CPU is the brain of the computer. If that were so, the CU is the cerebrum; with the aid of ID, IC, and SR, the CU controls all activities that take place inside the CPU. Many circuitries inside the CPU as well as those in external components are synchronous circuits. These components must be driven by the master clock signal in order for them to operate in synchrony. Timing signals go to every component in the system. To reduce clutter, timing signal paths are not shown in the above diagram. A typical classic 8-bit home computer ran at the clock frequency between 1 and 2 MHz. This clock signal is supplied by an external crystal oscillator.

The CPU chip is plugged into a socket on the motherboard, and the socket pins are connected to clock line, interrupt line, reset line, power line, ground line, address bus lines, and data bus lines. The buses connect the CPU chip to ROM chips, RAM chips, video interface, audio interface, keyboard interface, and other peripherals located around the motherboards. The unidirectional address bus is 16 bits wide, so the CPU can address 64 KB of memory. Address signals travel from the CPU to the address decoders inside the peripherals. The bidirectional data bus is 8 bits wide. Data signals travel in both directions between the CPU and the peripherals. The external data bus is connected to the internal data bus via a tri-state digital buffer. This buffer electrically isolates the internal data bus from the external data bus. It connects the two buses only when the CPU instructs it to do so. And it also controls the I/O data flow direction. The buffer also isolates each peripheral from the external data bus, so that when a peripheral is inactive, it does not interact electrically with the bus. When the CPU wishes to communicate with a peripheral device, the CU causes that device to become active and all other peripherals to become inactive, thereby leaving the data bus for the exclusive use of the CPU and the chosen device. This chain of events constitutes bus arbitration.

The primary function of the CPU is to run programmes. Programmes are sequences of machine instructions. The CPU runs the programme by fetching, decoding, and executing the first instruction, then the next, and the next, and so on. The very procedure executed after the CPU powers up is the reset interrupt service procedure. This procedure, along with other interrupt service procedures are stored in the ROM. When an interrupt occurs during normal operation, the CPU suspends the currently executing procedure, and runs the interrupt service procedure. Once the interrupt has been serviced, the CPU resumes executing the suspended procedure. We shall now examine these operational stages, individually.

Next page | Previous page | More Lobsters retrocomputing | Headlines

Original: https://amenzwa.github.io/stem/ComputingHistory/How8bitComputersWork/