Two agent desks on one server. Both correct. Both readable by each other, because both ran as the same unix user.
We spent five weeks putting two customers on one machine. The isolation was the easy part, and it was not the part that kept us from shipping.
Nothing was misconfigured
Here is the state we started from, and every line of it is correct. The home directory was 750. The credentials file was 600. The runtime's config directory was 700. A reviewer reading that list finds nothing to fix.
It protected nothing. Both desks ran as the same unix user, and POSIX permissions do not separate a user from itself. Every guard was pointed at a boundary that was not there.
That reframed the whole project. We were not going to add an isolation mechanism. The mechanism was already installed and switched off. A per-box user does not introduce a new protection; it gives the existing protections something to act on.
It also merged two pieces of work we had been treating as separate. Isolating two customers and running many desks on one machine turn out to be the same build, wanted for two different reasons.
We tried to break it, and wrote down what happened
A configuration that should isolate and has never been tested against a real cross-read is the bug we ship most often. It reads correctly in review. It fires at nothing. So the acceptance test was not a checklist.
We put two desks belonging to two different brands on one host, planted a credential file in each, and tried to read across. Seven attempts, in every direction we could think of, including through the process table rather than the filesystem. All seven were refused.
The number that matters is not seven. It is three: the reads that succeeded first. Before each denial we had the machine's root user, and the desk that owned the file, read the same file and print the same fingerprint. A denial only means something if you have proved the thing being denied was really there. A test that reads an empty directory and reports "access denied" has measured nothing at all, and looks identical in the log.
Then we broke our own guards on purpose. Three mutations, one at a time: remove the line that pins a desk to its own user, point it at the shared user instead, and delete the config that carries it. Each one refused to start the desk. A guard you have not watched fail is a guard you are guessing about.
The thing that actually blocked us was a missing door
The isolation passed. We still could not ship it, and for a reason that had nothing to do with security.
Our own API could not express the situation. The call that adds a desk to a machine creates it inside the caller's own account, by construction. Two different customers on one machine was not a thing the system could be asked for. Our own acceptance test had to build the pair by hand, outside the product.
Anything a test has to do by hand is not a shipped capability. So the work that unblocked five weeks of isolation engineering was a second endpoint, a consent field, and a switch in the dashboard. Mechanism first, then the plumbing to reach it, and the second half took longer.
What consent looks like when both sides are customers
Hosting is off on every machine, and the only party who can turn it on is the brand that owns the machine. A platform administrator cannot set it for you. Turning it back off stops the next placement and leaves a running desk alone, because silently killing something a second customer depends on is not a setting, it is an outage.
One detail we argued about and got right by accident. When a guest desk runs on your machine, you are told that it is there and how many there are. You are never told which company, which name, or which id. The original specification asked the delete guard to name the guest desk in its refusal, and the shipped code refuses with a count instead. The engineer treated the spec as superseded by a sibling fix that existed to stop exactly that leak. Naming the guest in an error message would have reopened the hole through the error path.
What we did not prove
Two things, stated plainly, because finding them in a footnote later is worse.
No customer build has run end to end on a co-tenanted desk. We have provisioned one, measured it, and torn it down. A desk reporting healthy is not a desk that has finished a job, and we have been caught by that exact gap before: four consecutive builds on this platform once reported success and produced no files.
The isolation measurement was taken on hosts with the Linux ptrace restriction enabled, which is the default on current Ubuntu and Debian images. We measured a host with it switched off separately and it behaved the same way, but not under a full two-customer walk.
So the rule we have been running under since early September is lifted for isolation, and not for everything. One desk per machine is still the default, still fully supported, and still what you get unless you opt in.
The lesson that was not about isolation
The most expensive thing we learned in five weeks had nothing to do with unix users.
Code reaches our own fleet one way, a customer's machine a second way, and the installer a third. A merge satisfies none of them. While we were building isolation, the client installer served a version seven weeks behind the code, and a paying customer's machines sat five releases back. Both were invisible because checking meant logging into fourteen machines.
Both are now watched by a detector rather than by somebody remembering. That is the change worth copying, and it is the one we would have skipped if the isolation work had gone smoothly.