Lobsters retrocomputing - 24 Aug 2026
Page 3 of 3
For the CPU to execute an instruction, it must first fetch the instruction from memory. System service procedures and interrupt service procedures are stored in the ROM and user programmes are stored in the RAM. The address of the instruction to be fetched must already be held in the PC. At the start of the fetch stage, the CU transmits the instruction address held in the PC along the address bus. The address decoder of the ROM (or that of the RAM) recognises this address as being one of its own, retrieves the contents of the addressed byte, and transmits the contents along the data bus. The CU knows that the transmitted byte is an instruction, so it loads the instruction into the IR.
The CU now proceeds to the decode stage. The CU instructs the ID to interpret the contents of the IR. The circuitry inside the ID recognises the instruction type: arithmetic operation, logic operation, conditional branch, procedure call, memory access, I/O access, or whatever.
Having decoded the instruction, the CPU enters the execution stage. If the instruction is a read-from-RAM, say one that uses absolute addressing, the CU knows that the instruction comprises three bytes - one byte for the opcode and two bytes for the operand address. The CU uses the operand address to retrieve the addressed byte from RAM, and loads the contents into accumulator A. The ALU can now use the value in accumulator A as an operand. The CU then increments the PC by 3, so the PC now holds the address of the next instruction.
If the instruction is an add operation, the CPU instructs the ALU to add the 8-bit contents of accumulator A and register B, and puts the sum into accumulator A. An addition may generate a carry. A subtraction, on the other hand, may yield a zero or a negative. The CU sets the carry, zero, or negative bit in the SR, depending on the result in accumulator A.
If the instruction is a conditional branch, say one that uses relative addressing and with a loop termination condition i < 10 where i is the loop index, the CU increments the loop index and subtracts 10 therefrom. If the result in accumulator A is not 0, the loop must continue. Then, the CU decrements the PC by the jump distance specified in the branch instruction. At the start of the next fetch stage, the CU will execute the first instruction at the top of the loop, thereby iterating through the loop once more. If the result in accumulator A is 0, however, the CU increments the PC by 2, so the PC now holds the address of the instruction that follows the end of the loop. The CU will fetch that instruction at the start of the next fetch stage.
When the CPU receives an interrupt signal from a peripheral, the IC checks if the interrupt is an IRQ or an NMI. If the interrupt is an IRQ, the IC checks if it has sufficient priority to interrupt the currently executing procedure. If the interrupt is an NMI, however, the IC informs the CU to handle the interrupt, immediately.
Before handling an interrupt, the CU finishes the currently executing instruction first, and pushes the current values of the SR and the PC onto the stack. This effectively suspends the currently executing procedure. Next, the CU loads the PC with the address of the interrupt service procedure. The CU knows where to fetch the service procedure's address, because the address is stored in the interrupt vector table, indexed by the interrupt number. The interrupt vector table is located in ROM at 00000000,00000000 (or, 11111111,11111111, depending on the CPU design). So, if the interrupt number is 1, its service procedure address is stored in the vector table at the following location:
(00000000,00000000 base) + (interrupt number 1) * (2 bytes/address) =
00000000,00000010 holds low byte of handler address
00000000,00000011 holds high byte of handler address
Since every address is two bytes in length, the low byte value of the service procedure's address is at ROM location 00000000,00000010 (decimal $2$) and the high byte value of the address is at ROM location 00000000,00000011 (decimal $3$). The values of these two bytes together form the address of the service procedure for interrupt number 1, and this address is loaded into PC. The CU will execute the first instruction of the handler procedure, starting at the next fetch stage.
Upon reaching the interrupt return instruction at the end of the handler procedure, the CU pops the saved PC and SR values off the stack and loads the saved PC value into the PC and saved SR value into the SR, thereby restoring the suspended procedure's state. The suspended procedure resumes execution at the start of the next fetch stage.
I explained in this article how a computer is built, how it boots up, how it executes instructions, how it performs simple calculations, how these calculations when combined exhibit behaviours as complex as playing chess. My main goal here was to give software developers a better understanding of how the computer hardware executes programmes, with the hope that this low-level understanding may make them more mechanically sympathetic programmers. This hardware execution model I presented also sheds light on the inner workings of virtualisation technologies, hardware emulators, virtual machines, and microcontrollers.
As much as possible, I avoided descending into the low-level details like radiation shielding, clock signal phases, clock skew, signal degradation during bus transfer, serial and parallel I/O protocols, power regulation, heat dissipation, circuit layout, chip fabrication, etc. I also shunned modern, advanced topics like pipelining, branch prediction, vector processing, SIMD, VLIW, multi-core processors, on-chip instruction and data caches, on-chip high-speed buses, NUMA, etc. Those interested in modern processors and microcontrollers should read textbooks like Computer Architecture: A Quantitative Approach and The Designer's Guide to the Cortex-M Processor Family: A Tutorial Approach.
- 6502.org
- This site is the most comprehensive repository of 6502 resources. Its books page lists many of the classic publications on general assembly programming, as well as those that focus on the 6502.
- The Visual 6502
- This web application visualises a 6502, as it executes code. You can run the CPU at full speed or single step it. The single-stepping execution mode is particularly enlightening.
- KIM-1 Simulator
- The KIM-1 was the most popular trainer for the 6502. Many people learn microprocessor programming on trainers, back in the day. This web application simulates that learning experience.
Previous page | More Lobsters retrocomputing | Headlines
Original: https://amenzwa.github.io/stem/ComputingHistory/How8bitComputersWork/