
During the first Homelable deployment, both containers were healthy. The automation still stopped.
As I covered in Part 6, the auditor disagreed with how Docker represented a mount and named container capabilities. Fixing that required more than remembering that “the deployment failed.” We needed the revision that created it, the checks that passed, the checks that failed, and the state we had deliberately left behind.
That information went into a checkpoint in the repository. When work resumed, the next step was to correct the audit and inspect the retained deployment. It was not to delete everything and try again.
This became a recurring part of building the homelab with AI. Writing code was one part of the work. Keeping decisions, implementation, and live evidence aligned across sessions took its own process.
I wanted to remain involved
I use AI tools extensively for research, implementation, debugging, and documentation. I also wanted this project to teach me something.
Letting an agent produce a large change and then approving it because the explanation sounded convincing would have defeated that purpose. I wanted to read the code, understand what it would change, run the relevant checks, and interpret the result.
The patch workflow helped enforce that habit. The agent prepared a change; I reviewed and applied it in my checkout, ran validation, and handled the commit and push. We then continued from the resulting revision.
That added friction. Sometimes I was reading a patch when I would rather have been trying the new service. But it gave me a reason to ask why a script needed a privilege, why a cleanup operation selected a particular resource, or what a passing test actually established.
I could still misunderstand something. Human approval only helps when there is something concrete to review and the person approving it pays attention.
Giving the project a memory
A long conversation contains proposals, corrections, abandoned approaches, and statements that were true several weeks ago. I needed a way to resume work without treating all of those as current instructions.
The repository documents had different jobs:
| Document | What it answered |
|---|---|
AGENTS.md |
How should an agent work in this repository? |
| Architecture decision records | Why did I choose this design, and what would justify changing it? |
| Configuration and scripts | What behavior am I trying to deploy? |
| Dated checkpoints | What did I actually verify, and what remains unresolved? |
| Runbooks | How do I carry out or recover an operation? |
| Issues | What work is still actionable? |
Those distinctions helped when the records disagreed. A configuration file described the intended state. A checkpoint recorded an observation at a particular time. Neither automatically proved what was running now.
The checkpoints used explicit statuses such as “Validated,” “Configured but not retested,” “Planned,” and “Deferred.” A script could exist and pass local tests while its live behavior remained unverified.
For the failed Homelable apply, the checkpoint recorded healthy containers alongside the audit failure. It also listed everything we had not tested yet: login, restart behavior, repeat application, backup, and recovery.
That made the stopping point useful. The next session had a defined starting state and a list of claims it still needed to establish.
Recording why a decision changed
The move from disposable guests to a persistent integration host needed more than an updated VM diagram.
The original design had proved useful isolation and cleanup behavior. It also required too much machine setup for every application test. I wanted to preserve both findings.
The replacement architecture decision recorded the failed attempt, explained the new lifecycle boundary, and identified which parts of the earlier decision remained valid. Production isolation and separate identities still applied. Recreating the entire guest for each workload did not.
It also recorded the new costs: a retained host could accumulate drift, and a baseline snapshot did not establish recovery after total host loss.
Without that reasoning, a later agent could reasonably suggest disposable guests again. The decision record gave it the evidence behind the choice, including conditions under which we should reconsider it.
Turning repeated instructions into skills
Some instructions kept coming back.
For repository changes, I wanted a branch name, a complete patch, an application check, validation results, and commit and PR text. I did not want to discover halfway through applying a change that a new file had been omitted.
I captured that process in a reusable patch-handoff skill. The handoff included checking the patch against its expected base revision and distinguishing completed validation from checks I still needed to run.
Later, the blog acquired its own repeated requirements. Mermaid diagrams needed consistent colors and shapes, accessible descriptions, readable mobile sizing, and arrows that represented the actual process. Those became another skill, shared by the homelab and website tasks.
I kept repository-specific operating rules in the repository and reusable procedures in skills. That reduced how much I had to repeat in each prompt, although the instructions still needed maintenance when the workflow changed.
Delegating work with a defined boundary
I used subagents for bounded investigations and separate tasks for work with a different context, such as the website.
A useful assignment named the question, the relevant evidence, and the expected result. “Investigate this network boundary and report the constraints” was easier to evaluate than asking another agent to improve the whole design.
The parent task still had to reconcile the findings. Two agents agreeing did not establish that a configuration worked. Their conclusions needed to survive review and, where appropriate, a live test.
The blog work gave me a practical example of coordination between tasks. This homelab task held the engineering history; the Website task held the rendering implementation. I could approve a diagram here, pass a scoped implementation request across, and review the result in the local website.
That separation was useful, but it created another handoff to manage. The receiving task needed the relevant decision and constraints, not an assumption that it knew everything from the other conversation.
Making completion mean something specific
I tried to keep the work small enough that I could explain what the next step would prove.
For infrastructure changes, that meant a bounded change followed by evidence before advancing. Commands identified where they ran and whether they changed state. A failed operation became a recorded stopping point.
The checks also had different limits. Local tests could establish how a parser handled a fixture. Controller validation could establish whether the tooling worked in its execution environment. Integration could exercise application behavior. Production still needed its own acceptance checks.
The same distinction applied to backups. Creating an archive, inspecting it, and starting an isolated restore were separate results. The documentation needed to say which had happened.
This took more effort than marking a task complete when the command exited successfully. It also made it easier to resume after a failure without accidentally upgrading an assumption into a fact.
Where the process became too heavy
The disposable integration workflow was the clearest example. We built substantial machinery around qualifying and removing a guest before the application had even started. I changed the design when the coordination cost became hard to justify.
The diagrams were a smaller version of the same problem. Achieving consistent styling and readable layouts led to renderer changes, authoring conventions, and repeated fixes around subgraphs.
These were reminders to review the process as critically as the code it produced.
Coming next
The repository now carries the decisions, operating instructions, and verified stopping points needed to continue the project. I still review changes and test the resulting systems; the documents help me work out where to start and what remains uncertain.
The next post returns to the network upgrade: separating the homelab into networks, moving services across them, and checking which connections should still work afterward.