Two strangers' repos, deployed by agents, unaided
We handed two agents two repositories neither had ever seen and let them deploy, with nobody to ask. Both got there. Then we counted what broke.
The claim this platform is built on is that a coding agent can take an application to production on its own. That is easy to demonstrate on our own code, which is exactly why demonstrating it on our own code proves nothing.
So: two repositories found on GitHub, cloned cold, neither of them ours and
neither of them seen before. One a Node service with a Python ML sidecar, the
other a .NET 8 app that needs Postgres. Two agents, each given the gg skill and
nothing else — no runbook, no hints, and nobody to ask.
Both of them got there.
| the repo | outcome | wall clock | gg calls |
|---|---|---|---|
| Node + Python ML sidecar | 1/1, verified serving | 15m20s | 19, zero errors |
| .NET 8 + Postgres | 1/1, verified serving | 24m36s | ~48, half of them polls |
The part worth more than the times: neither agent accepted 1/1 as proof. A
private service on gagarin has no public address, so each one built a probe image,
ran it as a job inside the project, and went and got a real HTTP response out of
the thing it had just deployed. Nobody told either of them to do that. They did
not trust the status line, which is the correct instinct about a status line.
That is the central claim, tested against strangers' code, and it held. Worth saying plainly, because everything below is a list of what went wrong.
Six defects, in forty minutes of agent time
Deploying someone else's unfamiliar app is the best bug-finder we have, and it is not close. Every one of these is ours, not theirs.
A crash reported as SIGSEGV when it was nothing of the kind. An ordinary
unhandled exception in the app surfaced as exit 139, which reads like a corrupt
cross-architecture build — so the agent went looking for an image problem that did
not exist. gg logs had the real stack trace the whole time. Status was lying
with the truth.
A dependency that was declared and did not route yet. A service crash-looped on
a database timeout, with the edge declared and visible in gg status the whole
time. The NetworkPolicy had not converged; it healed itself two minutes later. Our own
diagnosis table attributes that exact symptom to a missing edge, so the
documentation pointed firmly away from the answer.
No way to reach a private service. The strongest signal of the whole exercise, because both agents hit it independently and both invented the identical workaround — build a probe, run it as a job, curl from inside. Two strangers converging on the same workaround is not a missing convenience. It is a missing feature, and it is ours to ship.
Undocumented Postgres privileges. What the role we hand you can and cannot do is decisive for EF Core, and for Rails and Django after it, and the only way to find out is to connect and ask Postgres yourself.
Volume ownership for a non-root container. Undocumented, irreversible, and not something an agent can fix from outside the container. The agent read the situation correctly and deployed the app without the volume it wanted, which is the right call and a bad outcome.
And a sixth, fixed the same day: gg destroy could not delete a job. It routed to
the resource path and printed a raw API endpoint as the workaround, which is the
platform admitting out loud that it has a seam in it.
Five of those six are still open as we publish this. They are written down, they are ours, and none of them is the kind of thing a demo on our own repository would ever have found.
Booting is not demonstrating
The .NET app reached 1/1 and cannot do anything.
It needs six third-party API keys, and they fail lazily — not at boot, but the first time a request actually wants one. Those are a stranger's accounts. No agent can invent them and neither can we.
This matters more than it looks, because "I deployed your project" is a sentence
we were about to go and say to people. It is true for a self-contained repo and
misleading for anything that calls paid APIs, where the honest version is "it
builds and boots, and you would need your own keys" — a much weaker gift. The tell
is greppable, which is the useful part: read the .env.example before you spend
the afternoon.
And one thing this cloud cannot host
A third candidate was examined and not attempted: a realtime voice bot, WebRTC, media over UDP. Signalling is HTTP and would pass our ingress fine. The media would not: a service here gets one TCP port behind an HTTPS ingress, and that repo expects to be handed real ports on the host.
Kubernetes is not the constraint — a Service takes UDP, and plenty of clusters run media on one. The constraint is our own model, and it is a product decision rather than a limit of the substrate. Two true things follow from it. Native UDP is a real feature with a real cost, because ICE wants a port range and a routable candidate address, which is the exact thing a managed platform exists to take away from you. And relayed media already works today, undocumented: egress here is open, so a service can dial an external TURN relay outbound and media flows browser → relay → service with no inbound UDP at all. That sentence belongs in our docs and is not in them yet.
Better to say where the edge is than to find out from somebody who already left.
What we take from it
Forty minutes of agent time bought six defects, a documented ceiling on what we can honestly promise, and a product boundary we had not named. We will run it again — as a product test, not as a marketing stunt, which is what it would quietly become if we let it.
The claim survived. The platform around it picked up five open tasks, which is the trade we would make again.