Lobsters retrocomputing - 10 Sep 2026
Decoding the NEC V20 Microcode
reenigne's decoding of the 8088 microcode in 2020 opened the doors for extremely accurate emulation of the 8088 CPU.
Although I had added support for the NEC V20 in MartyPC, the V20 core was not cycle-accurate in terms of the V20's actual timings. It was an 8088 in a V20's clothing - a copy-paste of my 8088 core with V20 instructions tacked on.
This was not an ideal case, but the prospect of making my V20 core cycle-exact without the microcode seemed like it might be a discouraging slog of trial and error.
So why not get the microcode, then?
I recently commissioned InfoSecDJ to take die photography of an NEC V20 CPU (actually a second-source V20 fabricated by Sharp, but a V20 nonetheless). He did an excellent job.
| The NEC V20 CPU die - InfoSecDJ |
This photomosaic is extremely high resolution - 5.6 Gigapixels to be exact, an astonishing 70478x80672 resolution - too large to even fit in the JPEG image format!
You can see the entire thing at full resolution here.
The rectangular region just below the center of the die is the main microcode ROM.
| The V20 Microcode ROM block |
The ROM array is 258x116, containing 29,928 bits. It's evenly divisible by 29, which we know is the microcode word length of the V20, so that's a good sign. But it implies 1032 microcode words, when we were only expecting 1024. That's a bit odd, isn't it? We'll figure out the reason for that a bit later on.
Here's a close-in crop of the ROM array:
| Microcode bits, zoomed in |
The bright, horizontally running traces are part of the chip's metal layer. The yellowish dots are interconnects that connect the metal layer with the polysilicon layer beneath it. Note the vertical bars behind the metal layer - and note that there is occasionally a gap in the polysilicon on either side of each interconnect.
These gaps form a transistor - with the presence of a transistor indicating a 1 bit. I'll highlight the 1 bits to make that a bit easier to see.
| Defining bit locations with MaskRomTool |
MaskRomTool allowed me to draw the rows and columns that defined the locations of the bits. Unfortunately, I found that its bit-detection methods were based on thresholding, and there wasn't enough of a difference in contrast between a bit and a non-bit to make this an effective detection mechanism. Notice the bit histogram is very compressed toward the far axis. This was not going to work. Perhaps the thresholding technique would have worked better without the bright metal layer, but I didn't want to ask InfoSec DJ to attempt removing it. Another approach was needed.
| The quick and dirty classification tool |
Shown here is a '1' bit. Can you spot it by the transistor behind the metal layer? The buttons ended up being extraneous - you just need to hit either 1 or 0 on your keyboard to classify the bit. In theory, you could do this 29,928 times and you'd have the job done in a few hours. I had originally intended for this to be my backup method in case the CNN training didn't work out - I had a few friends willing to volunteer to help, and the JSON logs that the "Bit Voter" produces can be merged to support distributed work with consensus. Fortunately, this was not needed.
This is what a training run looks like.
[Epoch 01] train: loss=0.6945 acc=0.7273 f1=0.0164 | val: loss=0.6924 acc=0.7876 f1=0.0000
val precision=0.0000 recall=0.0000 cm=[[178, 0], [48, 0]]
[Epoch 02] train: loss=0.6860 acc=0.7151 f1=0.3826 | val: loss=0.6257 acc=0.7965 f1=0.0729
val precision=0.5000 recall=0.0394 cm=[[178, 0], [46, 2]]
[Epoch 03] train: loss=0.3751 acc=0.8914 f1=0.7213 | val: loss=0.3900 acc=0.7655 f1=0.6327
val precision=0.4661 recall=1.0000 cm=[[125, 53], [0, 48]]
[Epoch 04] train: loss=0.1159 acc=0.9523 f1=0.9139 | val: loss=0.0495 acc=0.9912 f1=0.9773
val precision=0.9773 recall=0.9773 cm=[[177, 1], [1, 47]]
[Epoch 05] train: loss=0.0251 acc=0.9945 f1=0.9888 | val: loss=0.0460 acc=0.9867 f1=0.9744
val precision=0.9514 recall=1.0000 cm=[[175, 3], [0, 48]]
[Epoch 06] train: loss=0.0319 acc=0.9933 f1=0.9802 | val: loss=0.0438 acc=0.9823 f1=0.9659
val precision=0.9350 recall=1.0000 cm=[[174, 4], [0, 48]]
[Epoch 07] train: loss=0.0185 acc=0.9945 f1=0.9212 | val: loss=0.0274 acc=0.9956 f1=0.9891
val precision=0.9792 recall=1.0000 cm=[[177, 1], [0, 48]]
[Epoch 08] train: loss=0.0141 acc=0.9956 f1=0.9913 | val: loss=0.0271 acc=0.9956 f1=0.9891
val precision=0.9792 recall=1.0000 cm=[[177, 1], [0, 48]]
[Epoch 09] train: loss=0.0101 acc=0.9978 f1=0.9940 | val: loss=0.0447 acc=0.9867 f1=0.9735
val precision=0.9488 recall=1.0000 cm=[[175, 3], [0, 48]]
[Epoch 10] train: loss=0.0110 acc=0.9967 f1=0.9907 | val: loss=0.0437 acc=0.9912 f1=0.9773
val precision=0.9773 recall=0.9773 cm=[[177, 1], [1, 47]]
Early stopping: no val F1 improvement >= 0.0 for 3 epoch(s).
Best val F1: 0.9891
If you have a CUDA-capable GPU, training is rather quick - this only took a few minutes.
The idea is that we want to maximize our accuracy - but reaching 1.0 may not be feasible, and maybe not even desirable (there's a thing called overfitting). Sometimes going on for longer just makes things worse, so we end training if we're not seeing a steady improvement.
The output of the training is a neural network model - we can then use this model to run inference on an entire input data set. Inference is just a fancy word for applying our model to actually do what we trained it to do - predict whether a given image contains a 0 bit or 1 bit.
Before we move on - a quick note to head off any potential controversies. CNNs loosely fall in the broader scope of AI from a computer science perspective, but we are not using "AI" in the modern, controversial sense that typically refers to a large language model (LLM).
When we run an inference pass, we get a confidence score for each pixel. We can use this confidence score to mark bits the model is less confident about, under some specific threshold (I used < 99% here). Here's the result of the first run, with ambiguous bits colored red:
| The first output of our bit-classification CNN |
I took all the ambiguous bits and manually sorted them back into the training folders, then re-ran the training, repeating until I got this result:
| The final "good enough" inference run |
This was pretty good - only 4 bits remain ambiguous, and it was faster just to manually verify them than to train another model.
Great, we have our 29k microcode bits and we saved hours of tedious manual labor (in exchange for hours of writing a training script in Python, but at least that is reusable!).
We still have to turn this rectangular blob of bits into a list of 29-bit microcode words. In other words, we need to reorganize the bitmap until it is 29x1032 instead of 258x116. How exactly to go about doing that is not obvious, but we can put it aside for the moment until we've decoded the matching decoder PLA.
The decode or "activation" PLA sits above the main microcode ROM block, with some intermediate circuitry sandwiched in between.
| The decode PLA |
| The decode PLA, zoomed in |
| The final PLA extraction |
| The 4:1 microcode column multiplexers |
Between the two blocks of microcode, a 4-way multiplexer allows two logical inputs to select one of the four words currently activated by the decode PLA. These inputs are pulled from the two low-order bits of the microcode program counter.
| The NEC V20 microcode word format from court documents |
| Working on V20 word extraction |
| Decoding the V20 Microcode with an Excel Spreadsheet |
The Main Decoding Effort
Unlocking the Group Decode ROM
| The NEC V20's Group Decode ROM PLA |
Decoding the first few lines of the GDR show us some familiar patterns.
01 111100?? 00100000001000 PREFIXES f0,f1,f2,f3
01 1111010? 00100000001000 HLT,CMC f4,f5
01 1111?0?? 00100000000100 f0,f1,f2,f3,f8,f9,fa,fb
01 1111??0? 00100000000010 f0,f1,f4,f5,f8,f9,fc,fd
01 1111?0?0 00100000000001 f0,f2,f8,fa
01 1111?100 00100000000001 HLT,STD f4,fc
01 00001111 00100000000001 EXT PFX 0f
01 001??110 00100000000000 26,2e,36,3e
01 01100100 00100000000001 REPNC 64
01 0110010? 00100000001000 REPX 64,65
01 0100???? 00000000011100 INC/DEC 40,41,42,43,44,45,46,47,48,49,4a,4b,4c,4d,4e,4f
01 1111?11? 01000100011011 GRP f6,f7,fe,ff
The True V20 Microcode Word Format
| The final NEC V20 microcode word format |
V30
In the image below, within the indicated circle, one side or arm of the metal structure was cut depending on the CPU type being fabricated. This meant both CPUs could share the same microcode mask. This also explains the discrepancy in microcode word count originally noticed!
| A zoomed-in view of the V20/V30 selection circuitry |
More Lobsters retrocomputing | Headlines
Original: https://martypc.blogspot.com/2026/09/decoding-nec-v20-microcode.html