Daemon lifecycle
Install, start, stop, reload, diagnose. macOS launchd + Linux systemd, side-by-side.
The dotagent daemon is one process per user. The OS (launchd on macOS / systemd on Linux) keeps it alive. dotagent never installs itself — dotagent install only generates the unit file; you load it into the OS scheduler with one extra command.
If anything below isn't behaving as expected, the Diagnostics section at the bottom has the canonical "is this thing on?" check.
TL;DR
# 1. Generate the unit file.
dotagent install
# 2. Load it into the OS scheduler (one of these two).
launchctl bootstrap "gui/$(id -u)" ~/Library/LaunchAgents/run.avelino.dotagent.plist # macOS
systemctl --user daemon-reload && systemctl --user enable --now run.avelino.dotagent # Linux
# 3. Verify.
dotagent status
cat ~/.config/dotagent/state/daemon.pidGenerate the unit file
Writes a single unit, regardless of how many agents you have:
macOS
~/Library/LaunchAgents/run.avelino.dotagent.plist
Linux
~/.config/systemd/user/run.avelino.dotagent.service
The unit's ExecStart / ProgramArguments points at the currently-running dotagent binary (resolved via std::env::current_exe). If you move the binary after install, re-run dotagent install so the unit refreshes.
Both unit files set:
Property
macOS (launchd)
Linux (systemd)
Auto-restart
KeepAlive=true
Restart=always
Start at login
RunAtLoad=true
(enable with --now)
Throttle
ThrottleInterval=10
RestartSec=10
stdout capture
StandardOutPath=…run.avelino.dotagent.log
StandardOutput=append:…
stderr capture
StandardErrorPath=…-error.log
StandardError=append:…
The stderr file catches crashes, not activity. The daemon only mirrors its
tracingstream to stderr when stderr is a terminal — otherwise it would grow forever in a file no rotation policy covers. Readlogs/daemon/dotagent.logto see what the daemon is doing; read…-error.logto find out why it died.DOTAGENT_LOG_STDERR=1/0forces the mirror on or off.
Templates live in crates/dotagent-unit-gen/templates/.
Note:
dotagent installaccepts--alland a positionalNAMEfor backwards compatibility with the legacy per-agent install flow. Both are no-ops now — the daemon manages every discovered manifest internally. The CLI prints a notice to remind you.
Start
macOS (launchd)
What this does: registers the plist with launchd's gui/<uid> domain (the user session domain) and immediately spawns the daemon since RunAtLoad=true.
Verify:
brew services start dotagent(coming soon, once the tap publishes) is a friendlier alternative — it runslaunchctl bootstrapfor you.
Linux (systemd user units)
daemon-reload makes systemd re-scan ~/.config/systemd/user/. enable --now does two things: marks the unit to auto-start at the next login AND starts it immediately.
For the unit to actually survive logout, you need lingering enabled (otherwise systemd kills user sessions when you log out):
Verify:
Stop
macOS
bootout is the inverse of bootstrap — unloads the plist from launchd, sends SIGTERM, and removes the daemon from the user domain. The plist file stays put on disk; the daemon just isn't running.
Don't
kill -9the daemon —KeepAlive=truewill respawn it within 10s. Either bootout (above) or uselaunchctl stop:
Linux
Doesn't disable auto-start at next login. To prevent re-spawn at the next boot, also:
Restart
After upgrading the dotagent binary itself, the running daemon is still pointing at the old in-memory binary. Replacing the file doesn't change the running process. Restart it explicitly:
macOS
kickstart -k sends SIGTERM, waits for exit, then re-spawns.
Linux
If you only changed manifests (
agent.tomlfiles) orconfig.toml, usedotagent reload(SIGHUP) instead — it's cheaper.
Boot orphan reap
A daemon that exits through its shutdown path reaps its children and deletes state/supervisor.json. A daemon that is SIGKILLed, panics, or is replaced by launchctl kickstart -k does neither. Its children survive with nobody holding their deadline — the supervisor that would have enforced it died with the daemon. Observed in production as an agent process alive for 26 minutes against a 600-second timeout.
So a starting daemon opens its state store and audit log, and then — before it starts a supervisor of its own — looks at what the last one left behind:
The order matters: the reap runs before the snapshot writer starts, because the writer's first tick would overwrite the only record of those processes.
It refuses to kill on doubt
The OS recycles pids. A pid read off disk with no further proof eventually names somebody else's process, so every record is corroborated against what the OS reports for that pid right now. A record is reaped only when all of these hold:
pid > 1, and not our own
signalling init or ourselves
the record carries pgid and command
a snapshot from a build that predates identity checking
recorded pgid == pid
a record this supervisor did not produce — every supervised spawn does setpgid(0, 0)
the OS still reports the pid
a process that already exited (the common case)
observed pgid == pid
a recycled pid that is not its own group leader
observed start time within 5s of the recorded one
a process that inherited the number after the old daemon died
observed command consistent with the recorded one
a recycled pid running something else entirely
Anything missing, ambiguous, or unparseable is a skip, never a kill. Leaving one orphan alive costs memory; killing the wrong process costs somebody's work. A snapshot that cannot be read or parsed signals nothing at all and is left on disk, since a corrupt registry is exactly when guessing is most expensive.
Confirmed orphans get the same treatment a live deadline expiry gets: SIGTERM to the whole process group, one grace window, then SIGKILL. Each one is logged at warn with the agent, label, pid, pgid, and how far past its deadline it ran, and audited as orphan_reaped at critical — a process that outlived its deadline unsupervised is worth an out-of-band notification, not just a log line. Skips are logged at debug with the reason (already_gone, start_time_mismatch, command_mismatch, not_group_leader, ambiguous_record, unusable_record) and are never audited: refusing to signal is the expected outcome, not an event.
Rollout
pgid and command are new fields in the snapshot. A state/supervisor.json written by an older binary has neither, so every record in it classifies as ambiguous_record and nothing is killed. The reap becomes effective from the second restart after the upgrade — the first one is what writes a snapshot the next boot can actually check.
Reload (SIGHUP)
Reads ~/.config/dotagent/state/daemon.pid, sends SIGHUP to that process. The daemon picks up changes on its next tick:
New manifests in
~/.config/dotagent/agents/Modified manifests (drift is detected and audited)
Updated
config.tomlUpdated plugin binaries (resolution is re-done per tick)
What reload does NOT do:
Doesn't swap the binary. A SIGHUP'd process keeps its in-memory code. For binary swaps, use
restart(above).Doesn't drain in-flight runs. Currently-spawned agents keep running until they exit.
Doesn't cut short an in-flight answer. Retiring persistent instances waits for the slot each one holds, so an instance answering a trigger when the signal lands is retired after that answer is delivered. The wait is bounded by the agent's
timeout_seconds.dotagent reloaditself returns as soon as the signal is sent — it is the daemon-side effect that waits.Doesn't reset the audit chain. The new tick continues from the current
prev_hash.
Errors:
reading ... (daemon not running?)
state/daemon.pid is missing. Daemon isn't running.
sending SIGHUP: No such process
PID file is stale (daemon crashed without Drop cleanup). Start it again.
A reload replaces the whole in-memory config, not just the parts the ingress reads. Retention thresholds and the daily summary's time and destinations were pinned at boot until recently; editing config.toml and reloading now takes effect on the next tick, as the table above says it should.
What wakes the daemon
The daemon sleeps until the earliest thing it has to be awake for. Two independent reasons produce a wake-up, plus a safety cap:
The safety cap is what picks up a manifest dropped into agents/ without a reload — within 30 minutes even if nothing is scheduled.
An inbound trigger is not on this list, and that is the point. A chat message or an MCP tool call is drained by a worker task running beside the loop, so it is answered on its own clock instead of waiting for the loop to come back around:
Triggers were an arm of the loop's select! until recently, which meant a scheduled run with a 20-minute deadline held every queued message for its whole duration. Now only triggers queue behind each other. See Triggers → Serialization.
The daily summary is a wake-up reason in its own right: it is the one scheduled thing the daemon does that no agent schedule accounts for. Before that was made explicit it only landed inside its window by coincidence — the 30-minute sleep cap happened to match the 30-minute grace window, two unrelated constants with nothing connecting them. That coincidence broke the moment a tick overran its own sleep budget, and it never held at all for a time / grace_minutes someone picked themselves.
dotagent tick --dry-run reports the agent next event only, so the real next wake-up can be earlier than what it prints when the summary is closer. See [daily_summary].
The background tasks do not add wake-ups
Two supervisor tasks run beside the loop, and neither polls:
reaper
sleeps until the nearest deadline
parked
snapshot writer
rewrites the file every 2s
parked after one write
Both park on the same signal, which a spawn — or a retime that pulls a deadline in — raises. A daemon supervising nothing therefore holds no timers at all: it is asleep until its next scheduled event, and nothing underneath it is ticking.
This matters on a laptop, where periodic work is charged twice: once for the CPU, and again for keeping the machine out of deep idle. The reaper formerly swept on a fixed 5-second tick (~17k wake-ups a day) and the snapshot writer rewrote a byte-identical [] every 2 seconds (~43k writes a day) to record that nothing had changed.
The reaper losing its tick also made it more precise, not less: it now fires on the deadline instead of up to one tick after it.
Power
Scheduled runs can be held back while the machine is on battery — see [power]. The check happens per dispatch, not per wake-up: the daemon still wakes on schedule, decides nothing is to be dispatched, and goes back to sleep. Deferred runs leave no trace, so plugging the charger back in dispatches the current window as if it had just come due.
Uninstall
Stop the daemon first, then remove the unit:
dotagent uninstall is idempotent — running it twice doesn't error, the second call just prints "nothing to remove".
dotagent uninstall does NOT delete your data. Manifests, heartbeats, audit log, config — all stay in ~/.config/dotagent/. See installation.md for a full wipe.
Diagnostics
"Is the daemon running?"
If either step fails:
File missing → daemon never started (or was stopped cleanly).
PID exists but
psempty → stale pidfile (daemon crashed without theDropguard firing). Just start the daemon again.
Platform-native checks
macOS:
Linux:
"What's the daemon doing right now?"
Tail the structured log:
That is the file to watch. dotagent.log is where the daemon's whole tracing stream goes, and under launchd / systemd it is the only place it goes.
The captured stdout/stderr files are a different thing:
…-error.log is not an activity log. The daemon installs its stderr tracing layer only when stderr is a terminal, because under a service manager stderr is a plain file that nothing rotates — an unbounded log growing next to the rotated one it duplicates. So the error file receives panics, and startup failures that happen before logging is up, and nothing else. A daemon that is running normally leaves it empty, which is the correct outcome and not a symptom.
Override with DOTAGENT_LOG_STDERR=1 (force the mirror on — useful when the unit is rewired to journald, which does rotate) or DOTAGENT_LOG_STDERR=0 (force it off even in a terminal).
Health dashboard:
"What will the daemon do next?"
That next event timestamp is the next agent window. The daemon wakes at whichever comes first between it, the daily summary, and the 30-minute safety cap — see What wakes the daemon.
"Did the audit log break?"
The daemon verifies the chain on startup and emits AuditChainBroken (with notify) if it fails. It checks the live segment — the only file that changes. Manual check:
Past 32MB the log rotates: the live file becomes audit.log.<YYYYMMDDTHHMMSS> and a fresh one opens with a seam (audit_log_rotated) whose prev_hash is the old file's tail hash, so the chain has no gap at the rename. Rotation is normal and verifies clean:
Segments are never deleted by dotagent — the log sweeper in [logs] does not touch them. If you prune them yourself, verification still reads as intact and reports how far back it reaches. What it will not forgive is cutting lines off the head of the live file: that removes the seam too, and the orphaned hash left behind is what fires audit_chain_broken. Details in security/threat-model.md.
Signal reference
The daemon process responds to:
SIGHUP
Wake immediately; re-read manifests + plugins on the next tick. dotagent reload.
SIGTERM
Graceful shutdown. Drops daemon.pid. Audit gets DaemonStopped.
SIGINT
Same as SIGTERM.
SIGKILL
Immediate kill — no Drop runs → stale pidfile. Auto-restart via launchd/systemd.
Don't kill -9 unless you have to — let launchctl bootout / systemctl --user stop do the work.
Common patterns
Run the daemon manually (development)
For debugging the daemon itself, bypass launchd/systemd:
Don't do this while the supervised daemon is also running — they'll both try to write daemon.pid and step on each other.
Run on a non-standard root
The unit file inherits the env var of the shell that ran install — not great for permanence. Better: set the env var inside the unit file directly (EnvironmentVariables for launchd / Environment= for systemd).
Make config / env changes survive restart
If you need persistent overrides (custom DOTAGENT_HOME, custom OTEL_EXPORTER_OTLP_HEADERS, etc.) edit the unit file:
macOS (~/Library/LaunchAgents/run.avelino.dotagent.plist):
After editing, reload:
Linux (~/.config/systemd/user/run.avelino.dotagent.service):
After editing:
The unit file is regenerated by
dotagent install— your manual edits are lost if you re-run it. Long-term, treat the unit file as a generated artifact and keep custom envvars inconfig.tomlinstead where possible.
Related
installation.md— install paths (brew, release, cargo, source)cli.md—install,uninstall,reloadtroubleshooting.md— sintoma → diagnosticobservability.md— log streams + OTelpaths.md— every file the daemon touches
Last updated
Was this helpful?