Lobsters - 02 Oct 2026
Page 3 of 3
Whenever you reuse an existing format for a purpose it wasn't designed for, you'll probably need to add conventions on top of it, and every convention you add eventually reserves some part of the input space. So it's worth asking early: what can my users no longer express?
The cheat sheet
Here's the whole thing in one table. The left column is the change you make to the filesystem (for example, in a RUN step), the right column is what ends up in the layer's tar archive to represent it. Remember that .wh.* entries always live next to the path they affect, in the same parent directory, and only have an effect on earlier layers:
| Filesystem change | Layer representation |
|---|---|
| Add foo | tar entry foo |
| Modify foo | new, complete tar entry foo |
| Delete file foo from an earlier layer | empty file .wh.foo |
| Delete directory dir/ (and everything in it) | empty file .wh.dir |
| Replace the inherited contents of directory dir/ | dir/.wh..wh..opq, plus the new children of dir/ |
| Regular file literally named .wh.foo | can't be represented unambiguously |
So, is it a hack?
I promised you we'd come back to this.
On one hand: yes, it's a little hacky. Reserving magic file names leaks a protocol convention into what otherwise looks like an ordinary filesystem namespace. A layer entry doesn't mean what tar says it means, and you only find out if you know the rules. And it produces funny edge cases, like the ones we just saw, where an innocent COPY deletes a different file, but only after the image travels somewhere else.
And it turns out I'm not the only one who thinks so. At the beginning of this article, I said that my take on tar and whiteouts was just an intuition. Well, image-spec#24 gave me some actual history: according to people involved in the spec, the .wh. scheme was inherited from AUFS, the union filesystem early Docker was built on. As one of the maintainers put it: "the original image code was just based on how AUFS did things because AUFS was the only real union filesystem at the time". Another one was even more blunt: "The .wh. is a silly approach."
On the other hand: it's also extremely pragmatic, and I'd go as far as calling it elegant. OCI keeps using plain, standard tar archives. Any tool that can read tar can read a layer. Generating a layer is just writing a tar. And with one tiny naming convention, you get deletion semantics without inventing a brand new archive format and all the tooling that would come with it. The price is that you can't have files starting with .wh. in your images, which... let's be honest, is a price very few people will ever notice. (Again, if you consider me a statistically valid sample: it took me over 10 years to find out, and not because of an actual bug, but because I accidentally started reading the spec!)
That's pretty much the argument that won in that same thread: "Is there a realistic use case for distributing .wh. files, other than packing up a container runtime into an image?", followed by the observation that there are already "millions of container images using this approach". Changing it would mean every implementation having to support two formats forever, just to unlock a handful of weird file names.
Personally, I lean towards: it feels hacky, but I really admire how simple and practical it is. Boring technology, plus a little bit of protocol glue.
Wrapping up
All of this started because I was trying to understand SOCI. I went looking for answers about lazy-loading compressed container layers, and somehow ended up learning about magic .wh.* files and building a repository full of intentionally broken Dockerfiles. Pretty standard rabbit-hole-driven development, I suppose. (And yes, the SOCI stuff is coming, keep an eye on AWS Bites!)
I started with what felt like a fairly mundane question: "how does Docker delete a file?". The answer turned out to involve magic file names, opaque directories, an entire reserved namespace of file names, and a Docker image that changes its filesystem after being exported and imported again. Not bad for one innocent question!
If I had to condense everything into a mental model, it would be this: layers aren't miniature filesystems. They are ordered filesystem changesets. The filesystem a container sees is what you get after interpreting all those changesets, in order, following the OCI rules. Additions and modifications are plain tar entries, deletions are whiteouts, and that convention means that a regular file whose name starts with .wh. can't survive the round trip through an OCI image.
I still think tar was a brilliant choice. It's simple, boring, ubiquitous technology, and boring technology tends to survive. Whiteouts are the little bit of protocol glue needed to stretch it beyond what tar knows how to express on its own.
Now I'm curious to hear from you:
- Did you already know about whiteouts? If so, congratulations, you were several layers ahead of me!
- Do you know why BuildKit lets these names through without complaining?
- Have you run into other weird consequences of how OCI layers work?
- And the big one: elegant bit of Unix pragmatism or a hack we've all learned to live with? Feel free to disagree with me!
You can find me on Bluesky, and the full experiment (with all the scripts, logs and version details) is on GitHub at lmammino/broken-dockerfile. PRs with more cursed experiments are very welcome.
Further reading
- OCI Image Specification: Image Layer Filesystem Changeset: the source of truth for everything about changesets and whiteouts.
- Docker docs: storage drivers and the overlay2 driver.
- Linux kernel docs: Overlay Filesystem, if you want to see how whiteouts and opaque directories work at the filesystem level.
- SOCI snapshotter: the thing that started this whole rabbit hole.
P.S. This article sparked lots of interesting comments on lobste.rs, so go and have a read if you want to dig even deeper. A special mention goes to david_chisnall, who shared tons of insights, most of which I was totally unaware of. What a legend!
Previous page | More Lobsters | Headlines
Original: https://loige.co/hidden-design-compromises-of-docker-layers/