Lobsters retrocomputing - 13 Sep 2026

Page 7 of 8

This then informs how we'll set up our bitbanger. A half-duplex system will suffice for downloading, since the sender will wait for us to indicate receipt between packets, and that will let us concentrate entirely on receiving until a full packet is obtained. The absolutely fastest speed we can receive at is generally limited by how quickly we can clock data bits into an accumulator from the receive line. If we connect the receive line to the least-significant input of one of the GPIO registers (we'll use the leftmost for convenience), we can do it in 23 cycles:

getabit MACRO

in b,(71) ; 11 cycles

rr b ; 8 cycles (rotate b bit 1 into carry)

rra ; 4 cycles (rotate carry into a little-endian)

; = 23 cycles

endm

The Z80's clock speed (in any of these systems) does not neatly divide into any standard bitrate, but theoretically 23 cycles per bit gives us a maximum possible transfer speed of ((315 000 000/88)/23) =~ 155632.4 bits per second. That suggests you might be able to get 115200bps with an unrolled loop, but at speeds this fast time required for other tasks starts to be a concern, such as storing to memory, checking how many bytes have been received, and branching back to get another, all of which together will certainly be more than 23 cycles.

The other problem is the time required to sense the start bit, because this can occur at any moment, and hardware UARTs generally end up repeatedly snooping the line at some multiple of the bitrate to ensure they won't miss one. We, on the other hand, can't even check for a start bit at just twice 115200bps. In fact, the fastest we can check for a start bit (a zero) is

startbt in a,(71) ; 11 cycles

rra ; 4 cycles

jp c,startbt ; 10 cycles taken or not

; = 25 cycles

which because of the unavoidable branch is actually longer than the time to clock in a data bit!

However, these numbers do suggest that half that speed, i.e., 57600bps, is plausible. Flipping the equation around, that gives us a relatively generous ((315 000 000/88)/57600) =~ 62.1 cycles per bit, long enough to do our housekeeping tasks on each byte, and our tight startbit loop can poll the line at ((315 000 000/88)/25) =~ 143181.8 bits per second (a familiar number to some of you), which is at least twice the data rate and should be sufficient for the sort of continuous data transfer we'd experience receiving a data packet. We will target this speed. (Note from the future: an early draft used in a,(c) in the startbit loop, which is a 12-cycle instruction. This single extra cycle reduced the startbit poll rate to 137674.8bps, and at that speed multiple bytes got missed and/or corrupted. We are probably only just fast enough to make this work.)

Parenthetically, VZ-300 owners in the audience will now have asked if this will work for them. If we substitute its lower clock speed at 57600bps, we get ((17734475/5)/57600) =~ 61.5 cycles per data bit and a maximum startbit poll rate of ((17734475/5)/25) == 141875.8 bits per second exactly. Because you can sample a little bit faster but never slower, we would need a separate version for the VZ-300; the same code will not work reliably on both. Sorry! That will be the subject of a future article.

Note that by making receive fast, we made transmit slower: unless we occupy the least-significant bit of another GPIO register, which seems rather wasteful, the next fastest position is the second-to-least significant bit. This snippet needs no less than 31 cycles to send the next bit of a character stored in a register other than the accumulator (here we'll use B):

putbit MACRO

xor a ; clear accumulator and flags (4 cycles)

rr b ; rotate low bit into carry (8)

rla ; rotate carry into low bit (4)

rla ; rotate up one more bit (carry must be zero) (4)

out (71),a ; 11 cycles

; = 31 cycles

endm

If we had to completely guard that GPIO register from interfering with any other GPIO pins on the same register, it would be even longer because we would need to read the current state and then do the bitmasks. Mercifully we'll just refuse to support that, and as 31 cycles is still well within our 62 cycle maximum per bit, she'll be right.

Now that we've gamed out how we'll connect our serial wires, we next need to figure out which pin on the BennVenn expansion header goes with which GPIO line. As I mentioned earlier, a minor gripe is that the lines are not necessarily wired in order, so it's a good idea to verify what pin is where. This is again most easily done with one of Ben's proto boards because you can continuity-test the back of the pin connector and any of the pin rows to see how they are connected (or, if you want, just wire directly to the marked lines on the proto board itself, but the angle makes it a bit less elegant). Find ground, then find pins one and two. I marked my findings with a

TextaSharpie.

If you don't have his board, you can still figure it out a little less conveniently with a voltmeter. Ensure all pins are set to output and turned off (something like FORI=68TO73:OUTI,0:NEXT will do from BASIC). The only live lines at that point should be ground, 3.3V, 5V and 9V. Find ground first, which you might do by checking for continuity with the ground test point Ben provides on the board, then use that as your common to find the voltage pins. Mark those; the rest are GPIO. Turn on pins one and two individually (OUT 71,1 or OUT 71,2) and look for voltage.

Having done so, our serial port will be the very cheap, easily available and extremely flexible HW-597 USB-to-TTL converter, based on the CH340. There are buckets of these things on eBay from many manufacturers and they quietly reproduce in my desk drawer like hamsters. Which I don't mind, because it works at 3.3V or 5V, it connects to your host directly with USB, and pretty much every modern operating system has built-in drivers for it (at least both my M1 MacBook Air and my Raptor Talos II running Fedora do). Jumper it for 3.3V operation as shown here, then connect the ground to the ground pin, transmit - relative to the host, not the VZ200 - to pin 1, and receive to pin 2. This is how the wiring looks on mine.

Since we will be drawing power from the connected host and not the BennVenn, don't connect the 3.3V line. Instead, for the programs below, ensure the HW-597 is already plugged into your host (such as with a USB extension cable) and showing bright status LEDs before powering on the VZ-200, or it may try to unsuccessfully power itself from the other lines and get a little daft.

To test your connection, a simple program in the Github repo (BITS) toggles the screen colour as it sees activity on the receive line. This is easiest to watch at a slow bit speed of around 150 baud or so. Here, I hooked it up to the Talos II, ran picocom -b150 /dev/ttyUSB0 (adjust for the path to your device), and just banged on the T2's keyboard. If you get alternating flashes of green and orange on the VZ while you do so, then your receive line at least has basic connectivity. A more thorough test is now to accept entire bytes. This program displays any byte it gets from the receive line (ASCII), effectively one half of a very slow terminal program. We'll use the internal ROM routine to display a character, which will get us scrolling for free, and then force the update instead of waiting for the next IRQ - which is disabled anyway to make sure our timing remains precise. (Typing only upper case characters works; lower case shows as symbols.)

This program runs at a sedate 300bps, and the reason is because the VZ ROMs are written for space efficiency, not time efficiency, at least to any extent they're efficient at all. 300 baud gives us an apparent surfeit of cycles using our formula - 11931 cycles per bit - but we may well need all of them since we've really got no idea how long it can take the ROM routines to do any arbitrary screen update. We won't be using the ROM much for our data blaster program, but a general purpose terminal emulator would have to consider a proper solution to achieve faster speeds. This is something else we might revisit in a future article.

The other purpose of this ASCII test program is to mock up how we'll write the fast serial loader. Despite the fact we have over 10,000 cycles between bits at 300bps and could easily have written each bit we read as a subroutine call to save memory, I still inlined each clocked-in bit using a macro because we necessarily need to at 57.6kbps - among other things, each CALL is 17 cycles and the RET to return from it is 10, which would consume almost half our CPU budget by themselves. I'd also like to observe, again with my usual biases showing, that cycle counting isn't nearly as much fun on the Z80 as it is on the 6502. Most opcode tables will fortunately collapse the whole Z80 T-state and M-state business into a single unified cycle count, but unlike the 6502 where there are instructions with execution times of 2, 3, 4, 5, 6 or 7 cycles (so you can easily make a busywait from any combination), the Z80's cycle time options start at 4 and go as high as 23, skipping many numbers, and many of the smaller cycle times require specific conditions like not taking a branch. Having considered our little half-terminal program, here's what I settled on for 57.6kbps, written as AS macros:

getabit MACRO

in b,(c) ; 12 cycles

rr b ; 8 cycles (rotate b bit 1 into carry)

rra ; 4 cycles (rotate carry into a little-endian)

; = 24 cycles

endm

topwait MACRO

ld (ix+0),b ; 19 cycles

endm

botwait MACRO

topwait

endm

getbit MACRO

topwait

getabit

botwait

; 19 + 24 + 19 = 62

endm

From our previous maximal case I turned the in b,NN instruction into a slightly slower in b,(c), which burns an additional cycle, but means we have 38 cycles left over of our 62 which we can split exactly between two ld (ix+N),b instructions of 19 cycles each. (This also lets us possibly alternate between multiple connected serial devices by changing C, but one catastrophe at a time, I always say.) The separate top and bottom waits are for situations where we have an odd number of cycles left over and need to have different wait times; consider this future expansion for the VZ-300.

When the stop bit arrives, we need to dump the byte into a buffer and get ready for the next one in the same 62 cycles, since we expect the other end will be ready to fire the next start bit at us immediately. To make an interesting and vaguely useful display (as well as not requiring additional memory), the screen itself would seem like a good place, but this also imposes some constraints: we only have 512 bytes there (i.e., 32x16), some of which we also need for indicating status, meaning our received packets should really be no larger than 256 or 384 bytes to allow for a transmission log and other useful info. The protocol we select should have packets no larger than that, be easy to implement (because I'm lazy), and be something that pretty much everything can speak. While we've seen Xmodem-1K or Xmodem-CRC implemented other places (like The Newsroom's Wire Service variant), I just decided to go with good old O.G. Xmodem. That contains 132-byte packets and is easy to write and checksum, and any errors over USB between your host computer and the VZ200 would undoubtedly be from bad bit framing rather than line noise which the default checksum algorithm should detect. While it overruns memory a bit at the end, this is largely irrelevant for just loading something we intend to immediately execute.

All that preamble yields us a stop bit stanza like this:

; stow character in screen buffer during stop bit time

; unfortunately we don't have enough cycle headroom to do

; a running checksum, so we have to do it after the fact

ld (ix+0),a ; 19 cycles

ld (ix+0),a ; 19 cycles

inc ix ; 10 cycles

dec l ; 4 cycles

jp nz,datapak ; 10 cycles

; 19 + 19 + 10 + 4 + 10 = 62

Here we use the IX index register as a pointer into screen memory and the L register as the packet length countdown. A double-store of the same location onscreen once again soaks up 38 cycles, then the increment and decrement, then the branch. A nice thing about the JP instruction, which is absolute instead of relative, is that the conditional branch form requires 10 cycles regardless of whether it's taken or not, so this entire stanza always consumes precisely 62 cycles as well. We use that instruction a lot in the cycle-exact portions so that we always have predictable CPU time.

Once we get a full packet, we know the sender won't do anything until we reply, so we can relax our timing and validate the packet at leisure, copy it into the correct place in memory and send the ACK for the next one. We send bytes using the same send-bit route I showed you before or a trivial variation, padded to 62 cycles per bit and also inlined. We accept .VZ-format files in this loader, so we have special handling for the first packet to make sure it has a generally correct format and note the type and memory address, which is where the rest of this packet and subsequent packets will be copied to. If it doesn't, then we send CAN and force the sender to abort.

By contrast, the start bit is handled with the same 25-cycle code I showed you before, because this is the fastest way we can be sure we won't miss one. But this also means we have no way of checking the keyboard nor implementing a timeout: there is no spare time to count cycles or scan for keys, and the only free-running timer is the VDG end-of-frame IRQ which will totally mess up our timing if that runs, so during the entire transaction IRQs are disabled as well. The program therefore assumes your sender is up and ready to go the moment it starts executing. On startup it fires off the initial NAK and waits, possibly forever if the other end never gets the signal until you reset the VZ200. (We display a message to alert you that no transmission has yet been received, which is immediately overwritten on-screen by the .VZ metadata.)

Let's see our loader in action. On the host side, to send the program to the VZ200 you can use any terminal program that speaks Xmodem (and just about everything does), just as long as it automatically starts the transfer as soon as the initial NAK arrives. With both my MacBook Air laptop and my Raptor Talos II workstation, I use lsx with the usb2ppp tool from BURLAP, which we earlier used to tunnel PPP over a serial line for the Brother GeoBook, but can be used to run pretty much any program over a serial port. Here's an example, substituting the path to the HW-597 that your machine uses (e.g., Fedora Linux on my Raptor Talos II uses /dev/ttyUSB0):

% usb2ppp /dev/cu.usbserial-110 57600 lsx -b ascii.vz

opening /dev/cu.usbserial-110

setting up for serial access

setting flags on serial port fd=3

starting process lsx

subprocess pid= 64301

Sending ascii.vz, 2 blocks: Give your local XMODEM receive command now.

We start the loader, which you can run as a separate program on its own and it will load anything that does not encroach on its default location at $e000. (We'll find an even better spot for it in just a minute.) Since our ASCII half-terminal program loads and executes from $8000, this is no problem; it takes up three blocks (there's apparently an off-by-one bug in lsx), so it's a good quick test of the machinery. The loader immediately sends NAK and we're off to the races.

Since we're assaulting screen memory every 62 cycles without regard for the VDG, there are accordingly "snow" artifacts everywhere, but we don't (and indeed can't) care. We have read our metadata, which we display on the top line (file type and starting address). Below that onscreen is an image of the current packet followed by a connection log of dots for successfully validated packets or an X for one that failed. This is a grab from a much longer transfer but it gives you the general idea. At the end of transmission, we automatically execute the file if it's a binary, or print a RUN you can just hit RETURN on to run a BASIC program.

Bytes Sent: 384 BPS:20

Transfer complete

subprocess terminated

restoring terminal settings

Ta-daaaaa!

In memory of Dick Smith's food company I've christened it the BFL, short for Bush Food Loader, a joke I otherwise refuse to explain (image from The Australian, paywalled link).

Now we want to make it part of the system. To convert it to "firmware" takes advantage of a specific feature Ben built into the SD card reader for easier updates: the ability to load and run a new system "ROM" directly from the card.

The default memory map for the VZ200 puts the system ROMs between $0000 and $3fff, reserving the space from $4000 to $67ff for ROM cartridges, though a cartridge could technically take over any address range above the TOM (that's how the RAM expanders worked, after all), and we'll come back to that point later. For the system to recognize code at $4000 or $6000 as part of a cartridge, a sequence $aa $55 $e7 $18 is required in that order, and execution then starts at $4004 or $6004. As shipped to you Ben's device not only fills in RAM above the TOM, it also fills RAM in from $4000 to $67ff and puts its own code there with that sequence, effectively "slushware" (a la the DECmate II, and we'll use the same term here since it's not really ROM). This code is run by the system ROM on startup like a "cartridge," because that's what it looks like to the system ROM, and this code is what does the RAM test and initializes SD card access. Once the system is up, it then does this:

The file VZDOS.VZ it's loading from the card is "magic." If present, it will be used to temporarily replace the slushware at runtime; for a couple extra seconds spent loading it you don't have to mess around with burning it to the cartridge. But this binary is not signed or checksummed, nor does the load check if it's even a new copy of the slushware - any program will serve as long as the .VZ header loads it to $8000 but the code is written to execute from $4004 (as the onboard image would). The slushware then places a little trampoline copy routine at $a000 and runs that to copy the 8K from $8000 to $9fff to $4000, overwriting the old slushware, and jump into the new one.

Next page | Previous page | More Lobsters retrocomputing | Headlines

Original: http://oldvcr.blogspot.com/2026/09/a-dick-smith-vz200-without-dick-smith.html