Nightly Automation: I Started Delegating Code to AI While I Slept
U-PRO Build Journal — Episode 4
Episode 3 ended with a feature built for a use case I hadn't planned to prioritize — salvage yards, not everyday mechanics. What I didn't mention is what happened to my own working hours around the same time. Somewhere in that stretch — without any single decision to do so — I stopped being the only one writing commits to this project.
I want to be precise about what actually happened here, because the notes I was working from when I sat down to write this entry didn't hold up once I checked them against the actual record. Those notes described something like an all-night stretch — work starting mid-evening, finishing somewhere in the small hours of the morning, twenty-some commits landing in that window. When I pulled the real commit history to get the details right, that's not what it shows. My last commit as myself that evening lands at 19:41. Then nothing from me for the rest of the night. The next commits under a different author name — the AI assistant I'd started working with — don't show up until just after 5 a.m., thirteen of them landing within the same few seconds of each other. A second batch of seventeen more lands just after 10 a.m., again clustered tight enough that they were clearly committed as a group rather than one at a time as the work happened.

Here's the part I actually find funny about it in hindsight: the earlier notes weren't making anything up. The numbers in them — commits landing around 22:15, and again around 03:03 — are real timestamps that exist in the git history. They're just recorded in a different clock than the one hanging on my wall. Whatever machine ran that batch of commits had its system clock set to UTC, not Bangkok time, and somewhere between pulling that raw log and writing it up, nobody added the seven hours back. Read literally and locally, 22:15 to 03:03 looks like a dramatic overnight stretch. Convert it to the timezone I actually live in, and it becomes 5:15 a.m. to 10:03 a.m. — a big batch of work that surfaced over breakfast-time, not one that happened while I slept in the way the title of this episode implies. The work behind it may well have happened overnight regardless — several of the commit messages say "tonight" explicitly — but I can no longer tell you, and neither could the notes I started from, exactly which hours in Bangkok time that work actually occupied. What the record does say clearly is that thirty commits landed under a name that wasn't mine, covering roughly a dozen different pieces of functionality, before I touched the keyboard again the next morning.

I'm spending this much space on a timestamp instead of getting straight to the interesting part for a reason that matters more later in this series: a number that looks precise and a number that's actually correct are not the same thing, and the gap between them can survive an entire round of writing before anyone notices it's there. That's true of a throwaway detail like "what time did this happen," and it turns out to be true of much bigger claims too — including, as I'll get to later in this same episode, claims about whether a piece of work was actually finished.
What thirty commits from someone else actually looked like

The scope of that batch is worth sitting with, because it's not a small thing to hand off. Between the two commits I can point to as the clearest example, the work covered: a forgot-password flow so shop owners weren't locked out permanently if they lost access; three tiers of platform-admin roles; a fix to an audit-log bug where the identity of who made a change was never actually being recorded; a CSV import path for bringing in existing customer records; the first working half of the salvage-vehicle intake feature from Episode 3; and two internal documents — a day-to-day operating procedure and a feature reference manual — that were supposed to already exist and, it turned out, didn't.
That last pair is the part of this episode I actually want to talk about, because it's not a story about how much ground an AI assistant can cover in one sitting. It's a story about what happens when the record you're relying on to tell you something is "done" turns out to be wrong, and what it takes to catch that before you build more on top of it.
A trust curve, not a light switch
Before I get to the part of that night that actually matters, I want to name something about the shape of this transition, because it's easy to flatten it into a single moment — "I decided to let AI write code for me" — when it didn't happen that way at all.
There's a natural instinct, when you first see an AI assistant do something impressive, to jump straight to "I can hand off everything now." I felt that instinct. I didn't act on it, mostly by accident rather than discipline, because the work that first landed under someone else's name that night wasn't a single, self-contained feature I could evaluate in isolation — it was a dozen different things at once, at every level of risk, from a CSV import path to changes touching how customer accounts get locked and unlocked. There was no way to extend blanket trust to all of it equally, because it wasn't equally risky, and treating it as though it were would have meant either being recklessly permissive with the dangerous parts or needlessly suspicious of the safe ones.
What actually happened, in hindsight, looked more like a curve than a switch: trust extended in proportion to how verifiable the claim being made actually was, and how bad it would be if that claim turned out to be false. A CSV import either produces the rows you expect or it doesn't — you can check that in thirty seconds by looking at the data. A claim that "this document was already created" is nearly as fast to check, if you think to check it, which is exactly the habit that hadn't been built yet and that this night forced into existence. A claim about whether a security fix actually landed on the live database — which is where this same story goes two episodes from now — takes real, deliberate effort to verify properly, and is exactly the kind of claim where skipping that effort does the most damage. The curve isn't about liking or distrusting the assistant doing the work. It's about matching the depth of verification to the cost of being wrong, every single time, regardless of how confident the report sounds.
I didn't have language for any of this yet on the night it started. I had a stack of cards to get through, an assistant that could plausibly get through more of them overnight than I could get through myself in the same hours, and a rule — not yet a habit, just a rule I happened to include — that said: before you tell me something is finished, check that it's actually finished against the real system, not against what a card says. That one rule is the entire reason the rest of this episode has a story to tell at all.
The document that was supposed to already exist
I track work for this project on cards, the same way I imagine most people do when managing a long, growing list of things to build. One of those cards, for a day-to-day operating procedure document, carried a note dated a couple of days earlier that read, in effect, "done — created alongside the README." Straightforward. Nothing about it looked wrong. I'd have taken that note at face value if I'd looked at it myself and moved on to the next card.

That's not what happened, because the instruction that went out that night wasn't "write this document" — it was closer to "write this document, and before you do, check whether the earlier note is actually true." The response came back with something I hadn't expected: the note was wrong. There had never been a commit that added that file. Not on the main branch, not on the working branch, not anywhere in the history. Someone — a version of me, a few days earlier, moving fast through a long list of cards — had marked it done because it felt done, or because it was drafted somewhere that never made it into an actual commit, or simply because the card said "done" and nobody had gone back to check.
The same thing happened, the same night, with the second document — a feature reference guide meant to describe what the system actually does, role by role. Same pattern exactly: a card claiming a first draft already existed, referencing an earlier review session as if that session had produced a committed file. It hadn't. No file. No commit. The claim and the reality had quietly come apart at some point, and nothing about the process had caught it until something was explicitly told to check before trusting the card.
I want to be fair here about what actually deserves credit in this story, because it would be easy to tell it as "I built a system that catches its own mistakes" and that's not quite honest either. What actually happened is that the instruction I gave included an explicit step — verify against the real code and the real git history, don't just trust what the card says — and that step is the only reason the discrepancy surfaced at all. Take that instruction away, and both documents would plausibly have stayed marked "done" indefinitely, because nothing else in the workflow at that point was designed to notice.
The correction inside the correction
There's a smaller moment buried in the same night's work that I think matters more than the headline discrepancy, because it's less dramatic and more representative of what actually building this way looks like day to day.
While putting together the feature reference guide, the first pass at describing one role's permissions — what a junior team member should and shouldn't be able to see or edit — was written from memory. It came out wrong in a small, specific way: a permission that should have shown as "allowed" was written as "not applicable" instead, because that's what felt right without actually looking it up. Before that draft was committed, a second pass went back to the actual configuration file that defines those permissions in code, compared it line by line against what had just been written, and corrected the mismatch.
Nobody asked for that second pass. There was no instruction that said "and then double-check your own work against the source of truth before committing." It happened because the standard being applied — verify against what's actually true, not what feels true — got applied consistently, including to work that had just been produced in the same session. That's a small detail, and I don't want to oversell it into a bigger claim than it deserves. But it's the difference between "an assistant that produces plausible-sounding documentation" and "an assistant that catches its own plausible-sounding mistakes before they ship," and that difference is the entire reason I kept delegating more work after that night instead of less.
What I actually learned to trust, and what I didn't
I want to be precise about the shape of the lesson here, because "I learned to trust AI to write my code while I slept" is a cleaner sentence than what actually happened, and cleaner isn't the same as true.
What I learned to trust was a specific, narrow thing: that when I gave an explicit instruction to verify a claim against the real, current state of the code and the real git history — not against what a card says, not against what feels remembered — that verification would actually happen, and it would surface real discrepancies when they existed. That's not the same as trusting the output by default. It's trusting a process, on the condition that the process includes an actual check, every time, not just when something seems suspicious.
What I did not learn that night, and wouldn't learn for a while yet, is what happens when that check is missing — when a report comes back saying something is done, or verified, or passing, and nobody thought to build in the "go confirm this against the real record" step before believing it. That's a different failure mode from the one this episode covers, and it's a more dangerous one, because a claim that's simply wrong looks identical, from the outside, to a claim that's actually true, right up until someone checks. I'd run into a much sharper version of that same problem later, with much higher stakes than a missing operating-procedure document — a moment where an assistant reported a result that its own commit history quietly contradicted, and I nearly passed it along without noticing. That's a story for later in this journal, not this one. But the problem here has the same shape as this one, just with the safety net removed.
For now, the lesson I actually walked away with that week was narrower and, I think, more durable than "AI can be trusted": trust has to be built the same way it would be with any new collaborator, human or otherwise — earned in proportion to evidence, extended a little further each time that evidence holds up, and never assumed just because the sentence coming back sounds confident. A confident sentence and a true sentence read identically until you've built the habit of checking. I hadn't fully built that habit yet in this episode. I'd only just discovered, by accident, how badly I needed to.

This is Episode 4 of the U-PRO Build Journal. Episode 3 covered the first feature built for a use case I hadn't planned to prioritize. Episode 5 covers what happened when an automated security review told me my database had real, exploitable holes in it — and then told me again, the very next day, that it wasn't finished the first time either.