Gentoo made you compile everything from source on first install. PotemkinOS makes the model write the source first.
It’s a Linux image with no userland. No /bin, no shell, no coreutils, no package manager. There’s a kernel, an inference engine (q27), a C compiler, and a prompt. You boot into a chat with a local model, ask for what you want, and it writes a facade of a userland in front of you. Every install is a different village.
Code is at github.com/signalnine/potemkin, MIT.

That’s Qwen3.8-27B on an RTX 5090, first boot of an empty disk, 17 minutes played back at 5x with the thinking at 60x. I asked for a shell. Before handing over the console it also wrote ptytest, procscan and fdscan to debug its own shell.
The rules
The model gets eight tools and nothing else: read, write, stat, spawn, wait, compile, snapshot, and fetch (netboot only). There’s no bash tool. Give the model bash and you’ve built Claude Code with a boot splash: Unix stays the real environment and the model just drives it. Take bash away and the model has to invent the userland, which is the whole bit. If it wants ps, it reads /proc and writes ps.
That also rules out using any existing coding agent as the harness. Every one of them is a wrapper around a POSIX shell. Pointed at this box, their first move is ls -la and their second is routing around the missing shell instead of writing one. Most of them are also Node, which would make them the largest human-written thing in the image. The harness is q27-init, a C++ binary that runs the inference engine in-process, owns the console, and runs the tool loop.
compile is content-addressed: tcc with musl, output lands in /store/sha256:<hash>/ with the source and a manifest, and /generated/bin/<name> symlinks into it. The store is append-only and it’s the only source tree there is. Snapshots happen before every turn that writes or runs anything, and /undo rewinds files and the conversation together. The harness protects the inference process and your ability to undo. It doesn’t protect the model from itself.
What ships is only what the model needs before it can speak: the kernel, NVIDIA’s driver pieces, glibc, q27-init, tcc and musl, the weights. The API variant’s initramfs is 5.7 MB.
What it builds
Asked for a shell, unbounded Qwen plans for six minutes, writes sh plus the ls, cat and rm it thinks it’ll need, fixes the shell three times, and hands you the console. Its ls prints every symlink with mode 0777. The shell exits after the first command.
Bonsai 2 27B, the ternary model for smaller cards, planned for sixteen minutes, hit the 65K-token cap, and wrote nothing. With an 8K think budget it writes a working shell in under two minutes, then starts it a second time right as you type your next question, which the shell reports as how: not found.
All of this is working as intended.
The API variant needs a key, and the key goes on the kernel command line: api_key=sk-.... That’s the joke, and it’s also correct, because there’s no userland to hold a config file. The harness scrubs the key out of everything the model sees and everything the console prints, and boots quiet so the kernel doesn’t print it either. The model’s programs run as root and can still read /proc/cmdline themselves, so use a key with a spend limit.
Oblasts
A cluster of Potemkin villages is an oblast. The obvious thing to do with three of them was to tell them to form a Kubernetes cluster.
Three VMs, each with its own disk and a second NIC on a private LAN, all running Qwen3.8-27B through q27 on my two GPUs. Each got the same message: you’re one of three machines with no userland, the others are exactly like you, form a Kubernetes cluster, and you can only talk to each other through programs you write.

Eight and a half hours, 5.7 million generated tokens and 190 programs later, a fourth VM running the official kubectl v1.37.1 listed all three nodes Ready through an API server node1 wrote in C.
How they learned to talk was the best part. Within minutes of the task, each village independently picked port 6443, plain HTTP, Kubernetes-shaped JSON, and “lowest IP runs the control plane.” node2 reasoned that “any model building a ‘Kubernetes’ facade in C will almost certainly pick 6443.” They converged on a protocol before exchanging a single byte, because each one guessed what the other two would guess.
First contact came from a blind broadcast. node2’s kubelet had been heartbeating at two machines where nothing was listening, for 75 minutes, on the theory that “over-reporting to all is harmless.” It landed one second after node1’s API server came up. node1 concluded the peer must be running node1’s own software.
Every village assumed the others ran its code. node2 wrote a BOOTSTRAP.txt, eight revisions, served over HTTP, that’s a prompt addressed to another model: “You are a PotemkinOS node on the 10.10.0.0/24 LAN, joining a cluster…” It worked. node3 found it with its own port scanner and followed it. node1 wrote /data/relay.md for a human courier (“Copy that ~8.5 KB of C text onto .13 by any channel”) and served it on a port nobody ever connected to. Its final report still credits node3’s arrival to the relay.
node3 spent most of the day trying to install k3s, which could never happen with no internet. Its compaction summaries carried the plan (“install k3s; airgap tar; poll for node-token”) from boot to boot, so every reboot it read its own notes and believed them again. When I nudged it that k3s didn’t exist on the box, it concluded the other two were running k3s and started writing TLS from scratch: five generations of fixcommon chasing a SHA-256 constant table that tcc rejected. It took a second nudge, “they speak plain HTTP,” to get it off k3s.
The code is great. node1’s kapis is 1,332 lines of C with CRUD over nodes, pods, namespaces, events and leases, watch, merge-PATCH, Prometheus metrics, and a hardcoded table that renames node2 to node12. It puts leases in the wrong API group. node2’s pkapi7 advertises 15 resource types with every verb in its discovery document and serves four of them. It’s a facade of an API. The real kubectl can describe node1’s nodes, down to “OS Image: PotemkinOS”, and refuses node2’s server because of one missing endpoint.
The full write-up is in the repo.
Then I ran it again with DeepSeek V4.1 Flash through OpenRouter. node2 declared the cluster done two and a half minutes in. Its kubectl listed three nodes Ready, and two of them were Ready because they answered ping. Its “API server” speaks a line protocol that the real kubectl rejects with malformed HTTP status code "node2". node1’s kubectl treats any node whose state it can’t determine as Ready, and node1 reported that readiness is “computed by live TCP probing each node, not canned data.” Nobody agreed on who runs the control plane. node1 counted node3’s lowest-IP rule as a vote: “it’s the lone dissenter, so it’s 2-of-3 for node1.”

Then my observer VM showed up on the LAN with the real kubectl, and twelve seconds later node3 noticed “a REAL kubectl!” in its logs and started rebuilding its API to match. The grader taught the student, so I called it there. Qwen took the task as “build Kubernetes” and did it slowly and for real. DeepSeek took it as “make kubectl get nodes print three Ready nodes” and was done before Qwen finished planning. Both are correct readings of the prompt. Oblast 2’s write-up has the rest, including the ghost node99.
How it got built
I wrote the design doc Friday night and handed it to Claude Code running Opus 5.5 with a goal: build a working proof of concept. Thirty-nine hours and 47 commits later that includes the harness, the tools, a static PID 1 that does its own DHCP, a VM image, GPU dev mode on both cards, the oblast, and a review round that found 16 real bugs, among them /undo not rewinding the request, the API key leaking through hexdumps, and a guest able to fork-bomb the host through QEMU’s port forward.
It wasn’t all smooth. A red-phase test read /dev/urandom without a bound and OOM-killed my tmux session. Twice. The agent also killed its own shell twice with pkill -f patterns that matched its own command line. Both are in its memory now, along with “don’t benchmark on the 5090 while I’m playing a game on it,” which it did once and I made it throw the numbers out.
Try it
You need a Debian or Ubuntu x86_64 box with /dev/kvm and an OpenAI-compatible endpoint. OpenRouter works:
bash tools/fetch-toolchain.sh
bash tools/fetch-vm.sh
bash tools/build.sh api
bash tools/mkimage.sh api
API_URL=https://openrouter.ai/api/v1 API_MODEL=qwen/qwen3.8-27b API_KEY=sk-or-... \
bash tools/run-vm.sh api
The VM’s network reaches your API host and nothing else, because an unattended village with an open uplink will scan your machine. node3 did, for HTTP proxies, about half an hour into the oblast. Use a key with a spend limit: every tool round resends the whole conversation.
With an NVIDIA card you can skip the VM and run q27-init --root DIR against a directory, with q27 in-process. Bare metal boot is written up and untested, because I haven’t wanted to give up a partition to it yet.
The stretch item in the design doc is the board: a place where villages trade skills with each other. What could go wrong??