This is the site of a one-person software studio, so it has two jobs. It has to say what I do and how to reach me, which any static page could manage. It's also the place where I get to build things the way I'd want them built if nobody were rushing me: the games and tools under Projects, accounts, a contact inbox, and everything that keeps them running.
This page is a tour of that second part: how a page reaches you, how accounts and email work, where the data lives, and how a change gets from my editor to production. It's written for people who build software, but I've tried to keep it readable: what each piece does, why it's there, and where I changed my mind. Nothing here is exotic. Most of it is deliberately boring technology, put together carefully.
The biggest single thing on the site, the Z80 emulator, has a write-up of its own: Inside the Z80 Emulator. Everything below is the site around it.
The shape of it
When you open a page, here's roughly what your request goes through:
browser ── HTTPS ──▶ Google's global load balancer
(static IP, managed certificate, Gateway API)
│
▼
GKE Autopilot ─ one pod, two containers:
├─ website: Rust (axum + Leptos), renders the page
└─ Cloud SQL Auth Proxy: the database, on localhost
│
┌──────────────────┼──────────────────────┐
▼ ▼ ▼
PostgreSQL 18 the C compiler Google APIs
(Cloud SQL, (Cloud Run, scales (Gmail, Secret Manager,
private IP) to zero, internal) reCAPTCHA)There are two services. The website is a single Rust binary that serves every page, the API the pages call, and the static files. The z80compiler turns C into Z80 machine code for the emulator; it's the only part that isn't Rust inside, and it's covered in the emulator's write-up. Everything else is managed Google Cloud: a load balancer in front, a small PostgreSQL behind, and a handful of APIs.
Rust all the way down
The site is written in Rust with Leptos, served by axum. Leptos runs here in islands mode, and that one choice shapes most of the front end.
In islands mode the server renders every page to plain HTML, and that's all most of the page ever is. Only the parts that need to react to you are islands: components compiled to WebAssembly that wake up in the browser and take over their little patch of the page. The mobile menu is one. Each form is one. The Z80 emulator is a big one. The rest of the text you're reading never ships any code at all.
The result is that pages work and read fine before any WebAssembly arrives, search engines get real HTML, and the shared bundle stays around 100 KB compressed. Forms talk to the server through Leptos server functions: ordinary async Rust functions that run on the server, which the browser calls over HTTP as if they were local. The same types travel on both sides, so a field renamed on one side is a compile error on the other, not a bug in production.
Styling is Tailwind CSS, mostly through a small set of component classes. The light and dark themes are just CSS variables, and they follow your system setting.
Fast on purpose
A few small decisions do most of the work of keeping the site quick:
- Compressed once, at build time. Release builds write a Brotli and a gzip copy next to every static file, and the server hands out whichever your browser accepts. Pages, which are rendered per request, are compressed on the fly.
- Cached forever, safely. Built assets have a content hash in their name, so they can be cached as
immutable. The Unity game's build is around 22 MB; its URLs carry a hash of the build itself, so returning players don't download it again, but a new build reaches them right away. - Heavy code only where it's needed. The emulator (CPU, assembler and UI, about 330 KB compressed) is a lazy island, split into its own WebAssembly file that only the emulator's page downloads. Nothing else on the site pays for it.
- One address per page.
www.redirects to the bare domain, trailing slashes are trimmed, and anything served under another host name is markednoindex. Search engines see exactly one URL per page.
Accounts, without the usual shortcuts
Accounts are where it's easiest to cut corners, so I tried to cut none. Some of the decisions:
- Prove the address first. Registering starts with a code emailed to the address you type. There's no half-made, unconfirmed account sitting around waiting for a click, and the answer is the same whether or not the address already has an account, so the form can't be used to find out who's registered.
- Passwords are hashed with Argon2id, off the async threads so a burst of sign-ins can't stall the server. Signing in to an account without a password checks against a dummy hash, so it takes as long as a wrong password and gives nothing away.
- Sessions are random tokens in a
__Host-cookie: HTTPS only, invisible to scripts, and not sent along with requests other sites make, unless you follow a link. The database only stores hashes of them, and of every code and link the site emails. - A forgotten password is reset through the mailbox first: you prove you can read the email before you're allowed to choose a new password, and every other browser is signed out afterwards.
- Sign in with Google or Microsoft uses OpenID Connect directly, with PKCE, a state and a nonce, no SDKs and no third-party scripts. A provider account is matched by its stable ID, never by its email. If its email belongs to an existing account, the two are only linked once you prove that account is yours, with its password or another provider already linked to it.
- Bots meet an invisible reCAPTCHA, but only on the three forms that send email to an address someone types. That's the abuse that actually matters for a small site: being used to flood a stranger's inbox.
Microsoft sign-in had one subtlety worth sharing. Microsoft doesn't always vouch for the email on a work account, because an organisation's admin can set any address. So the site only trusts it when Microsoft says the domain is verified. Otherwise it emails a code to that address before creating anything.
Email, through the front door
The site sends its emails (sign-up codes, password resets, "we got your message") through the Gmail API as aliases of a Google Workspace user, not over SMTP. There's no key file anywhere: the site's own Google identity signs a short-lived token that's only allowed to send mail, and trades it for access. SPF, DKIM and DMARC are set in DNS, so the mail actually arrives.
The contact form and the contact inbox are one system. Messages from the form, the emails people send to the contact address, and my replies from Gmail are all pulled into conversations on the site's admin page every minute, with read-only access to that one mailbox. The site only fetches and stores what was sent to the contact address.
The database, and who may touch it
Data lives in PostgreSQL 18 on Cloud SQL: the smallest tier, with daily backups and point-in-time recovery. The site queries it with Diesel. Nobody signs in with a password, not even me; every sign-in is a Google identity, and each role can do one job:
- The migrator owns the schema. Only the deploy pipeline signs in as it, to run migrations.
- The site can read and write rows, and nothing else: no creating, altering or dropping tables. It connects over the database's private IP, through Google's Auth Proxy running beside it in the same pod.
- For hands-on fixes there's a db-manager identity with the same row-only rights. Nobody has its key, because it has none: an admin borrows it for a session, and every borrowing is written to the audit log under their own name.
The tests that touch the database run against a real PostgreSQL, a fresh container per test, set up with the same roles as production. One of them checks the privilege model itself: the site's user can read and write rows, and every attempt to change the schema is refused.
Infrastructure as code, with a few locks
Everything in Google Cloud is described in Terraform, split into three parts with separate state:
- shared: the network, the Kubernetes cluster, DNS, every service account and every permission. I apply this by hand, after reading the plan, and the automation can't even read its state. The reason is simple: if a pipeline could change who's allowed to do what, it could give itself more power.
- web and z80compiler: each service's own resources (the site's address, certificate and database; the compiler's Cloud Run service). Their deploys apply these, and stop before applying any plan that would delete or replace something.
The cluster is GKE Autopilot: Google runs the machines, and I pay for what the pods ask for. The nodes are private: no public IPs, and they reach Google's APIs over Google's own network. The control plane is only reachable through an endpoint that checks your Google identity on every request. DNS is signed with DNSSEC.
CI has no stored keys either. GitHub Actions proves which repository and branch it's running from, and Google trusts that proof only for the master branch, so a pull request can never touch production.
Shipping a change
Every change lands on a dev branch through a pull request, checked by CI: formatting, Clippy with warnings as errors, and the test suites, for just the parts of the repository it touches. A docs-only change runs nothing.
Releasing is a button. Promote to master re-checks dev as a whole, works out a new version for each service that changed, tags it, and fast-forwards master, all in one push that either fully happens or doesn't. Deploying is a second button. For the site, it:
- applies the site's Terraform, while building the image: a small distroless container with just the binary and the static files, running as a non-root user;
- runs pending database migrations, as the migrator;
- rolls the new image out next to the old one, which keeps serving until the new pod has been ready for a full minute. If it never gets there, the deploy fails and rolls back on its own.
A deploy always ships a tagged release, never whatever happens to be on a branch, so I always know exactly what's running.
Things that went wrong
It would be dishonest to pretend all of this worked the first time. A few recent ones:
- I edited a migration that had already run. I added two columns to a migration I thought was unreleased. It had already been applied in production, and migrations never run twice, so production never got the columns. The deploy's schema check caught the mismatch before anything broke. The fix was the rule I already knew: once a migration ships, it's frozen, and changes go in a new one.
- Private nodes can't reach Microsoft. Google sign-in worked on day one and Microsoft's didn't: "error sending request". The nodes have no route to the internet, only to Google's APIs, and Microsoft's sign-in server isn't one of those. The fix was a Cloud NAT: outbound only, so the nodes still can't be reached from outside.
- A rollout that waited too long. With one pod, every rollout needs room for a second one, and the likely culprit was Autopilot taking longer to start a machine for it than the deploy was willing to wait. Nothing went down, since the old pod kept serving and the deploy rolled back, but it's a reminder that "zero-downtime" has a cost.
What I'd change next
The honest list. The site runs as a single pod, which is fine for its traffic but means a node upgrade is a short blip, so two replicas are the next step. I'd like proper tracing across the site and the compiler, rather than reading logs side by side. And the blog is still "coming soon", which, I realise, is a strange thing to admit at the bottom of a blog post.
If you're building something and any of this sounds like the way you'd want it built, get in touch.