Deployment Guide¶
This guide covers how system administrators deploy and configure dotvault across an organisation.
Architecture overview¶
dotvault runs as a per-user daemon. Each user has their own instance, their own Vault identity, and their own secrets. The administrator's role is to:
- Set up the Vault infrastructure (KV engine, auth methods, policies)
- Deploy the dotvault binary to machines
- Distribute a configuration file (or Group Policy on Windows)
- Arrange for dotvault to start in each user's session
Vault infrastructure¶
KV engine¶
Enable KVv2 and create the user prefix namespace:
Policies¶
Create a template policy that scopes each user to their own secrets. See KV Engine & Policies for the full policy file.
Auth method¶
Enable and configure at least one auth method. OIDC is recommended for desktop environments as it integrates with existing SSO.
Configuration distribution¶
Linux¶
Place the config file at the system-wide location:
dotvault also checks paths listed in $XDG_CONFIG_DIRS.
Deploy with your existing configuration management (Ansible, Puppet, NixOS, etc.):
macOS¶
Place the config file at:
Deploy via MDM (Jamf, Munki) or configuration management.
Windows¶
Place the config file at:
Or use Group Policy to manage configuration centrally via the registry.
Registry takes precedence
On Windows, if Group Policy registry keys exist at HKLM\SOFTWARE\Policies\goodtune\dotvault, dotvault loads all configuration from the registry and ignores the YAML file entirely. The --config CLI flag is refused while a policy is present unless that policy opts in with a BypassSystemConfig REG_DWORD of 1 (the bypass_system_config: true equivalent); see Windows Group Policy.
Running as a user service¶
systemd (Linux)¶
Upgrading from a manually-created unit
Previous versions of this guide showed an example ~/.config/systemd/user/dotvault.service snippet. If you created one, remove it before enabling the packaged unit — the per-user path shadows /usr/lib/systemd/user/ and your hand-rolled unit (which lacks Type=notify, WatchdogSec, the env-file paths, etc.) will silently take precedence:
rm ~/.config/systemd/user/dotvault.service
systemctl --user daemon-reload
systemctl --user enable --now dotvault.service
Behavioural change to be aware of: services declaring After=dotvault.service now block until dotvault completes its initial sync (the packaged unit uses Type=notify and delays READY=1 until secrets are on disk). The previous hand-rolled unit had no readiness gate, so dependents started in parallel. If a dependent's startup ordering matters to you, this is the change to plan for.
The RPM, DEB, and APK packages all ship a dotvault.service user unit (a Type=notify service with WatchdogSec=120 and the OpenTelemetry-friendly logging settings) at the canonical /usr/lib/systemd/user/ path. dotvault is a per-user daemon — it authenticates to Vault with the OS user's identity and writes secrets into that user's $HOME — so installing it as a system service that runs as root would write to root's $HOME and authenticate to Vault as root, which is almost never what you want. The unit carries ConditionUser=!root, so a root start is skipped rather than run — systemctl status names the failed condition.
Enable per-user once the package is installed:
The daemon watches ~/.dotvault-token itself (via inotify on Linux), so subsequent rewrites of the file (typically from an interactive dotvault login in another shell) trigger an immediate token re-read on the running daemon within seconds — no extra unit to enable. See Config reload for the full mechanism.
Or enable globally for every login session on the machine:
--global enables the unit in every user's session; each user runs their own instance and authenticates with their own Vault identity.
Socket activation (optional)¶
The packages also ship three socket units, installed but not enabled: dotvault-api.socket (the local API socket), dotvault-agent.socket (the SSH agent socket) and dotvault-docker.socket (the Docker volume plugin socket). Without them the daemon binds its sockets itself, and each socket disappears whenever the daemon does — systemctl --user restart dotvault.service is a brief outage for anything borrowing a token, requesting a signature, or starting a container at that moment. With a socket unit enabled, systemd binds the socket and holds the listening fd across daemon restarts, so clients queue in the backlog and are served when the daemon returns; the socket also exists from sockets.target at session start, before the daemon has authenticated.
systemctl --user enable --now dotvault-api.socket
systemctl --user enable --now dotvault-agent.socket # optional, independent
systemctl --user enable --now dotvault-docker.socket # optional, independent
Things to know:
- The config is still the master switch.
api.enabled/agent.enabled/docker.enableddecide whether the surface exists; the socket unit only decides who binds it. If systemd passes a socket the daemon is not configured to serve, the daemon drains it — connections are accepted and closed immediately, so clients fail fast with EOF — and logs a warning naming the mismatch. (Merely closing the daemon's copy would refuse nobody: systemd retains its own listening fd, so clients would hang in a backlog no one accepts.) Enabling the socket unit is not a substitute for enabling the feature. - This is fd-passing, not start-on-demand. The service stays
WantedBy=default.targetand runs regardless of connections — it is syncing files and keeping a token alive. Enable both the.socketand the.service; enabling only one is a half-configured state. - A queued connection is bounded by startup, not by authentication. All three surfaces begin serving before the daemon has a Vault token: the API socket answers an honest
401, the SSH agent answers "no identities" (and refuses a signature), and the volume plugin refuses aMountthat would need Vault with a message naming the cause, until one arrives. That distinction matters because the backlog has no timeout of its own — a daemon that can never obtain a token (no local token, no peer to borrow from) would otherwise leave every client blocked forever on a connection systemd had already accepted. - Owner-only is verified, not assumed. The units set
SocketMode=0600, and the daemon independently checks the inherited socket's filesystem mode and refuses anything wider —SocketModedefaults to0666, so a hand-edited unit that drops the line fails loudly instead of silently exposing the token endpoint to every uid on the box. - The socket unit's path wins. Under activation the socket lives wherever
ListenStream=says; if that differs fromapi.unix.path/agent.unix.path/docker.socket, the daemon logs the divergence and reports the activated path indotvault status. For the volume plugin the adopted path reaches/api/v1/statusand the log, butdotvault statusdials the configureddocker.socketand prints its registration hint from it — so keepdocker.socketand the unit'sListenStream=the same (they agree by default), and remember the.specfile that registers the plugin with the engine must name the unit's path. - The volume plugin gains the least from it. A container already holding a volume never notices a daemon restart — the materialised directory lives under
RuntimeDirectory=dotvault,RuntimeDirectoryPreserve=yeskeeps it across the restart, and the daemon resumes refreshing it. The socket unit only smooths the engine's own calls: adocker runordocker volume createlanding mid-restart queues instead of the engine reporting the plugin unreachable. - Owner-only covers the directory too. Alongside the socket-node mode, the daemon verifies the parent directory is owner-only and owned by the same user (
DirectoryMode=0700in the units) — a socket in a directory someone else can write to can be swapped for an impostor, which no node check catches. - Requires systemd ≥ 227 (
FileDescriptorName=support). On older systemd the fds arrive unnamed, and the daemon drains them with a warning rather than serving them. - Not applicable on macOS (launchd has its own, incompatible mechanism) or Alpine/OpenRC; on those the daemon's self-bind path runs unchanged.
Docker volume plugin registration (automatic, per user)¶
The Linux packages also ship a systemd user-tmpfiles drop-in at /usr/share/user-tmpfiles.d/dotvault-docker.conf. When a user's manager runs systemd-tmpfiles --user --create at login, it creates ~/.local/lib/docker/plugins/dotvault.spec containing unix://$XDG_RUNTIME_DIR/dotvault/docker.sock, which is what registers the Docker volume plugin with that user's rootless Docker engine. Without it every user has to write that file by hand, and the path and contents are both per-user — the home directory and the uid — so a package cannot simply install the file.
Note that this is /usr/share/user-tmpfiles.d, not /usr/lib/tmpfiles.d: the system tmpfiles directories are under /usr/lib, the user ones under /usr/share. Nothing else about the packaging changes — there is no %post scriptlet, and the drop-in is inert on Alpine, which runs OpenRC.
Things to know:
- It runs from
systemd-tmpfiles-setup.servicein the user manager, which — unlike the system unit of the same name — is not symlinked at install time and relies on the distro applying systemd's shipped preset (90-systemd-user.presetdoesenable systemd-tmpfiles-setup.service). Where a distro does not, it simply never runs and nothing says so.systemctl --user is-enabled systemd-tmpfiles-setup.serviceis the check;systemd-tmpfiles --user --createapplies it without a re-login. - An existing spec file is left entirely alone. The drop-in uses tmpfiles'
ftype and:-prefixed modes, and both apply only when the item is created: the contents of an existing spec are never rewritten, and its permissions are never reset (a bare mode field is re-enforced on every run, so a spec a user had tightened to0600would come back0644at each login). A user who wrote their own spec for a customiseddocker.socketkeeps it exactly as they left it. - The path is correct under socket activation too.
dotvault-docker.socket'sListenStream=%t/dotvault/docker.sockand the daemon's own default socket are the same path, so the registered spec is right whichever binds it. It is wrong only when an operator has customised one of them, which is the case the previous point covers. - The spec is created regardless of
docker.enabled. For a user who has not enabled the plugin,docker volume create -d dotvaultthen fails with a connection error instead ofplugin not found. There is no exposure — the socket is0600in a0700directory, and an unconfigured daemon binds nothing — only a less obvious error message. - Opting out uses tmpfiles' vendor-override convention: a same-named symlink to
/dev/nullin a directory of higher precedence. A user doesln -s /dev/null ~/.config/user-tmpfiles.d/dotvault-docker.conf; to opt a whole machine out, put the same symlink at/usr/local/share/user-tmpfiles.d/dotvault-docker.conf. The user search path, highest precedence first, is~/.config/user-tmpfiles.d,$XDG_RUNTIME_DIR/user-tmpfiles.d,~/.local/share/user-tmpfiles.d,/usr/local/share/user-tmpfiles.d,/usr/share/user-tmpfiles.d— note it does not include/etc/user-tmpfiles.d, which serves the system instance only. The installed basename is fixed for exactly this reason. - Podman and rootful Docker are unaffected: Podman is registered in
containers.confand rootful Docker in/etc/docker/plugins, neither of which dotvault writes. See the Docker volumes guide.
Enable lingering if the daemon must outlive a login session
A --user service normally stops when the user's last session ends, and $XDG_RUNTIME_DIR (where the SSH agent and local API socket live) is torn down with it. For a machine people reach over SSH — where a tmux job or the local API socket is expected to survive a disconnect — enable lingering so the user manager keeps running:
Environment-variable overrides (e.g. OTEL_EXPORTER_OTLP_ENDPOINT) can be set via four optional EnvironmentFile= paths referenced by the unit:
~/.config/dotvault/env(preferred for per-user secrets)~/.config/dotvault.env/etc/default/dotvault/etc/sysconfig/dotvault
The system-wide paths are typically world-readable, so the per-user ~/.config/dotvault/env is the right place for anything sensitive (e.g. an OTLP bearer token in OTEL_EXPORTER_OTLP_HEADERS). Create the file with chmod 600; all four are silently ignored if absent.
%h vs ~ in custom unit drop-ins
The packaged unit references the per-user paths as %h/.config/dotvault/env and %h/.config/dotvault.env. %h is systemd's home-directory specifier — equivalent to ~ when you're creating the file at the shell. If you reference the file from a systemctl --user edit drop-in or a custom unit, write %h (or ${HOME}); systemd does not expand ~ in EnvironmentFile= directives, so a literal ~/.config/... would be silently skipped.
The unit hard-codes a couple of system paths that the package owns: ExecStart=/usr/bin/dotvault run, plus the EnvironmentFile= paths listed above. If you install dotvault into a non-standard location (e.g. /usr/local/bin), copy the unit out to ~/.config/systemd/user/dotvault.service and adjust those lines.
Slow initial sync and the systemd startup window
With Type=notify, two different deadlines govern dotvault's lifecycle:
TimeoutStartSec— the pre-READY=1window. systemd waits this long for the daemon to finish auth + initial sync and signal ready. The packaged unit sets it to 300 seconds; the systemd default of ~90s is too tight for resource-constrained hosts (many rules, slow Vault, cold TLS handshake). If the daemon doesn't reachREADY=1in time, systemd marks the start a failure and restarts — causing a boot loop on chronically slow hosts.WatchdogSec— the post-READY=1liveness check. The daemon kicks the watchdog at half this interval after becoming ready; if the kicks stop, systemd restarts the unit. The packaged unit sets it to 120 seconds.
WatchdogSec does not extend the startup window — only TimeoutStartSec does. To raise the startup window (or the watchdog) on a host where the defaults are too tight, use a drop-in:
systemctl --user edit dotvault.service
# Under [Service], one or both of:
# TimeoutStartSec=600
# WatchdogSec=300
TimeoutStartSec=infinity disables the pre-ready timeout entirely if your environment can't bound the first sync.
Note also that anything declaring After=dotvault.service now blocks until the first sync completes — a behavioural change from the previous manually-created unit which had no Type=notify gate.
Hardening and the FUSE mount¶
Do not add systemd sandboxing to this unit, and check any systemctl --user edit drop-in for it. The packaged unit sets no NoNewPrivileges=, RestrictNamespaces=, RestrictSUIDSGID= or LockPersonality=, and that is deliberate: NoNewPrivileges= makes execve ignore the setuid bit and file capabilities for the daemon and every process below it, which disarms the FUSE mount helper and breaks fuse.enabled outright. The daemon asks the kernel to mount directly first, but that needs CAP_SYS_ADMIN, which a user manager cannot grant, so the helper is the only path it has.
The constraint covers the whole seccomp-based class, not those four names — RestrictRealtime=, MemoryDenyWriteExecute=, SystemCallFilter=, ProtectClock=, ProtectKernelTunables=, PrivateDevices= and friends all make systemd imply NoNewPrivileges=yes in a unit that cannot install the filter otherwise, which is every user-manager service. A drop-in adding any one of them re-breaks the mount, and does so with nothing in the log naming the cause.
Little is given up by leaving them out. A per-user service already runs with the user's own authority — it can read ~/.ssh and write ~/.bashrc or ~/.config/systemd/user/ — so a compromised daemon that wants an unconfined process arranges to be run again outside the unit rather than escaping the sandbox. These directives were defence in depth against an exploit that cannot persist, not a boundary. UMask=0077 is not seccomp-based, has no effect on the mount, and is kept.
If you do not use fuse.enabled, the directives are harmless — but the packaged unit is shared, so re-add them only in a drop-in on hosts where you know the filesystem is off.
Upgrading from v0.35.0 or earlier
Units shipped up to v0.35.0 carried those four directives, and fuse.enabled could not have worked under them. A package upgrade replaces the packaged unit, but it will not touch a copy you made at ~/.config/systemd/user/dotvault.service or a drop-in you wrote — both shadow the packaged file. If the filesystem does not mount after upgrading, check those two places first, then systemctl --user daemon-reload.
launchd (macOS)¶
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.goodtune.dotvault</string>
<key>ProgramArguments</key>
<array>
<string>/usr/local/bin/dotvault</string>
<string>run</string>
</array>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>StandardErrorPath</key>
<string>/tmp/dotvault.err</string>
</dict>
</plist>
Deploy to /Library/LaunchAgents/ (all users) or ~/Library/LaunchAgents/ (single user).
Windows Task Scheduler¶
Create a scheduled task that runs at user logon:
$action = New-ScheduledTaskAction -Execute "C:\Program Files\dotvault\dotvault.exe" -Argument "run"
$trigger = New-ScheduledTaskTrigger -AtLogOn
$settings = New-ScheduledTaskSettingsSet -AllowStartIfOnBatteries -DontStopIfGoingOnBatteries
Register-ScheduledTask -TaskName "dotvault" -Action $action -Trigger $trigger -Settings $settings
Or deploy via Group Policy as a scheduled task.
Logging¶
dotvault writes all logs to stderr:
- Text format when stderr is a TTY (interactive use)
- JSON format otherwise (service/daemon use)
Control verbosity with --log-level:
Available levels: debug, info (default), warn, error.
Override the auto-selected format with --log-format:
dotvault run --log-format json # force structured logs
dotvault run --log-format text # force human-readable logs
dotvault run --log-format auto # default — text on TTY, JSON otherwise
This is useful when running under a service manager that captures stderr but is connected to a TTY for debugging, or when forcing structured logs for ingestion into a log collector regardless of how the daemon was launched.
dotvault itself never writes a log file — integrate with your platform's log collection (journald, syslog, Windows Event Log via a wrapper, etc.), or use the OTel mirror described next. On systemd hosts the packaged unit routes stderr to the journal, so the OpenTelemetry collector's journaldreceiver can filter on _SYSTEMD_USER_UNIT=dotvault.service (or _SYSTEMD_UNIT when the unit was enabled with systemctl --global) to pick logs up directly.
Separately, every log record is also mirrored to the OTel LoggerProvider alongside stderr (see Observability below) — this is additive, not a replacement: stderr/journald keeps working exactly as before, so a collector outage or observability.enabled: false never loses the local copy. When observability is enabled, the mirrored records flow to your collector as OTLP log records, which lets the collector fan them out to a file exporter, a syslog/journald forwarder, or — via a collector build that includes a Windows Event Log exporter (not a stock component of the core OTel Collector distribution, so this typically means a custom/contrib collector build) — the Event Log. dotvault ships none of that collector-side configuration; the fan-out target is entirely operator-authored.
Security note: because every record now leaves the process once a collector endpoint is configured, treat observability.endpoint/insecure/headers as part of the logging trust boundary, not just the metrics one — a call site that logs something sensitive is no longer contained to local stderr/journald.
Observability¶
dotvault can export OpenTelemetry metrics and logs to an OTel collector. Each signal is configured in its own nested block — metrics: and logs: — so the two can go to separate backends or one can be switched off independently. Disabled by default; enable with:
observability:
enabled: true # master switch for both signals
export_interval: "15s" # metric export cadence
metrics:
endpoint: "http://127.0.0.1:4317" # http:// = plaintext for the local hop
protocol: "grpc" # or "http/protobuf"
logs:
endpoint: "https://logs.vendor.example"
protocol: "http/protobuf"
headers:
X-Api-Key: "logs-backend-key"
The top-level endpoint / protocol / insecure / headers fields still work as shared defaults that the per-signal blocks layer onto — the model deliberately mirrors the OTel SDK's own env-var convention (generic OTEL_EXPORTER_OTLP_* plus signal-specific OTEL_EXPORTER_OTLP_METRICS_* / _LOGS_*) — but they are deprecated and being retired in stages:
Shared exporter fields are deprecated
Setting observability.endpoint, protocol, insecure, or headers at the top level (with observability enabled) logs a startup WARN naming the fields and increments the dotvault.config.deprecated counter once per field per process start, so a fleet's collector can measure migration progress. Everything keeps working for now; a later release makes the warning louder and 1.0 removes the shared fields. Move the settings into the per-signal metrics: / logs: blocks — or, for values shared across both signals (a common bearer token, a common collector), use the standard OTEL_EXPORTER_OTLP_* environment variables, which remain fully supported. Tracked in #140.
Field semantics: enabled (unset = on whenever observability.enabled is; explicit false switches that signal off — the top-level flag remains the master switch and a per-signal true cannot resurrect a disabled subsystem), endpoint / protocol (non-empty overrides), insecure (set overrides), and headers, which replace the shared map wholesale rather than merging — merging credential maps invites sending one backend's bearer token to the other. An explicitly empty headers: {} therefore means "this signal sends no headers" even when the shared map is populated. Watch the inverse case, too: a signal that overrides endpoint but leaves headers unset inherits the shared map — shared bearer token included — and sends it to the new backend. The daemon warns at startup when it sees that combination; set an explicit per-signal headers: ({} for none) to state the intent and silence it. Enabling observability.enabled with both signals explicitly off is rejected at config load as the contradiction it is.
Endpoint form. Both signals share one endpoint contract, and a full URL with a scheme is the recommended form: the scheme carries the TLS intent (https:// → TLS, http:// → plaintext, on gRPC and http/protobuf alike), and an explicit path is used verbatim — no mount path is assumed, so a vendor route like https://collector.example/tenant-42/v1/metrics works as written. A URL without a path gets the OTLP standard /v1/metrics / /v1/logs appended automatically (http/protobuf; gRPC has no URL path), so https://otel.example alone still routes correctly. Bare host:port also works — the traditional gRPC form — with TLS then governed by the insecure flag; dns:/// targets pass through to the gRPC resolver. An explicit insecure: true forces plaintext even over an https:// scheme — prefer stating the intent in the scheme and leaving the flag alone. The insecure-transport-with-auth-headers warning fires however the plaintext arises, flag or http:// scheme.
Upgraders: gRPC endpoints with an http:// scheme
Earlier releases ignored a URL scheme on gRPC endpoints (stripping it and attempting TLS regardless). The scheme is now authoritative, so a gRPC endpoint written http://… switches from attempting TLS to plaintext. The daemon logs a WARN naming this change whenever it builds a gRPC exporter from an http:// endpoint — use https:// or drop the scheme to keep TLS, or keep http:// if plaintext was the intent all along (the common case for a loopback collector, where the old TLS attempt simply failed).
Metric temporality. metrics.temporality selects the aggregation temporality the metric exporter reports: cumulative (the OTLP default), delta (counters, observable counters, and histograms report per-interval deltas — what Datadog and some other vendors expect), or lowmemory (only synchronous counters and histograms delta). The vocabulary and instrument-kind mapping are exactly those of OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE, which an unset field falls through to. Metrics-only — setting it under logs: is a config error.
Disabling the logs signal leaves the global LoggerProvider on the OTel no-op implementation, so both the Log* helpers and the slog→OTel mirror go quiet for that daemon (stderr/journald logging is unaffected) — metrics.enabled alone gives you series without shipping any log records.
Windows Group Policy
The observability block round-trips through the GPO/registry layer like every other section — author it under SOFTWARE\Policies\goodtune\dotvault\Observability (or generate the values with dotvault reg-import), with the per-signal overrides as Observability\Metrics / Observability\Logs subkeys (tri-state Enabled/Insecure DWORDs, and their own Headers subkeys — a present-but-empty per-signal Headers key means "no headers for this signal", distinct from an absent key meaning inherit). Header values round-trip too (as REG_SZ values under the respective Headers keys), so a .reg export carries the live tokens — treat the artefact as a secret. To keep tokens out of the policy hive and out of any exported config, leave headers empty and set them via the standard OTEL_EXPORTER_OTLP_HEADERS environment variable (through a machine-wide environment policy) instead. See Windows Group Policy for the full registry schema.
The standard OTEL_* environment variables (generic OTEL_EXPORTER_OTLP_ENDPOINT / OTEL_EXPORTER_OTLP_HEADERS, and the signal-specific _METRICS_* / _LOGS_* variants) are honoured by the SDK whenever the corresponding config field is empty, so endpoint and header configuration can live entirely outside the config file. Put credential-bearing values (OTEL_EXPORTER_OTLP_HEADERS) in the per-user EnvironmentFile (~/.config/dotvault/env, mode 0600) rather than a world-readable location — this is the recommended way to share a token across both signals without it appearing in any config artefact, and it is where the shared-field deprecation steers that use case.
Upgraders: dotvault.sync.ticks gained a third outcome value
A sync cycle interrupted part-way through — the ordinary shape of a daemon shutting down — used to be recorded as outcome="ok", because the cycle reported no error. It now reports outcome="cancelled". Two kinds of existing query change meaning: one written as outcome="error" is unaffected in what it counts but no longer sees the whole of "not a clean cycle", and one written as outcome != "ok" now counts every daemon restart as an anomaly. Prefer outcome="error" for alerting, which is the failure rate and excludes shutdowns by construction. Nothing else about the instrument changed, and no other metric is affected.
The exporter emits a bounded set of instruments:
| Metric | Type | Attributes |
|---|---|---|
dotvault.sync.ticks |
counter | outcome={ok,cancelled,error} — cancelled is a cycle the daemon stopped part-way through (a shutdown, almost always), counted apart from both neighbours so it neither inflates the success rate with cycles that never ran every rule nor puts routine shutdowns in the failure rate. A cycle that failed a rule and was then cancelled counts as error: the failure is the actionable half |
dotvault.sync.duration |
histogram | outcome={ok,cancelled,error} — same vocabulary as the counter above |
dotvault.vault.calls |
counter | op={read,write,lookup_self,renew_self}, status |
dotvault.token.renewals |
counter | outcome={renewed,reauth_required,failed} |
dotvault.token.ttl_remaining |
histogram | (no attrs) |
dotvault.token.denylist |
counter | event={denied,suppressed,cleared} — a token Vault rejected entering suppression, a lookup-self skipped because of one, and a suppression dropped after the credential situation changed (token file written, peer socket reconnected, SIGHUP). A rising suppressed against a flat denied is a host sitting without a usable token; before suppression existed that same condition showed up as dotvault.vault.calls{op="lookup_self",status="denied"} climbing at one call per 10s indefinitely. See Tokens Vault has already refused |
dotvault.enrol.attempts |
counter | engine, outcome={completed,error} |
dotvault.web.requests |
counter | route, status_class={1xx…5xx} |
dotvault.config.reloads |
counter | outcome={no_change,applied,error} |
dotvault.sighup.received |
counter | (no attrs) — each SIGHUP forces an immediate ~/.dotvault-token re-read and config reload |
dotvault.config.deprecated |
counter | field — deprecated config fields in active use, one increment per field per process start; sum by field to measure fleet migration progress |
dotvault.build_info |
gauge | version, go_version, os, arch — constant 1, one series per build, following the Prometheus *_build_info convention. version is the v-stripped release semver on tagged builds and dev on untagged/hand-rolled ones (same value as dotvault version), so a dev series in a fleet view means an unofficial build, not missing data. Join other series against it to slice by build — e.g. dotvault.config.deprecated joined by version shows whether deprecated-config stragglers are just old builds. The same identity also rides every series as OTel resource attributes (service.name, service.version, user.name, host.name, os.type, host.arch, process.runtime.*) for backends that surface target_info — see Resource attributes for what user.name and host.name disclose |
Log records¶
log/slog to stderr / journald is still the primary logging path, but every record handled through it is also mirrored to the OTel LoggerProvider — the mirror is additive, not a replacement, so a collector outage or observability.enabled: false never loses the stderr/journald copy. This is what lets a collector fan dotvault's operational log stream (not just deployment-fact records) out to a file exporter, a syslog/journald forwarder, or the Windows Event Log — see Logging above.
One record is emitted directly through the OTel logger rather than via the slog mirror, because it must reach a central collector without ever printing to an end user's terminal:
configuration loaded from Windows Registry (Group Policy); file-based config is ignored— WARN severity, attributepath=<would-be config file>. Fires once per daemon/sync startup on a GPO-managed Windows box. Routing this through slog would print an INFO line on every CLI invocation on a GPO-managed install, which is exactly the noise this record avoids.
Health probes are served on whichever HTTP surfaces the daemon has: the loopback listener when web.enabled: true, and/or the per-user Unix socket when api.enabled: true. A deployment with neither enabled has nothing to probe; enable one of them, or rely on the systemd sd_notify(READY=1) signal instead. The OTel httpcheckreceiver speaks TCP, so it needs web.enabled; a socket-only daemon is probed with curl --unix-socket <path> http://localhost/readyz, which is the way to get a readiness check without opening a port at all.
GET /healthz— liveness, always 200 while servingGET /readyz— readiness, 200 once the daemon is authenticated to Vault AND has completed its initial sync cycle, 503 otherwise. Mirrors thesd_notify(READY=1)contract so a KubernetesreadinessProbeor the OTelhttpcheckreceivernever observes a green daemon before secrets exist on disk. The auth check reflects the cached in-memory token, not a per-probe Vault round-trip; a revoked token flips/readyzback to 503 within the lifecycle check cadence (default 5 min).
Both return JSON and are loopback-only, suitable for the OTel httpcheckreceiver.
Security considerations¶
- File permissions — all managed files are written with
0600. dotvault warns if the config file is group or world writable. - Token security —
~/.dotvault-tokenis written with0600. Secret values are never logged, even at debug level. dotvault uses this dotvault-specific filename rather than Vault's default~/.vault-tokenso a concurrentvaultCLI session cannot clobber the daemon's cached token.
Transitional — upgrading from v0.19.0 or earlier
Releases before v0.20.0 used Vault's default ~/.vault-token. There is no migration: dotvault re-authenticates once on first start, and any token it previously wrote to ~/.vault-token lingers at 0600 until it expires. If dotvault was the only writer of that file, delete it after upgrading. This note will be removed in a future release (around v0.23.0).
- Atomic writes — all file writes use temp file + rename to prevent partial writes.
- Web UI — loopback only, CSRF-protected, strict Content Security Policy.
- Windows — DACL-based permission checks via the Windows Security API.
- systemd sandboxing — the packaged user unit deliberately carries none, because the seccomp-based directives disarm the FUSE mount helper; see Hardening and the FUSE mount for what that does and does not cost.
Config reload¶
SIGHUP is the running daemon's reload trigger. It does two things at once: re-reads ~/.dotvault-token immediately (picking up a token freshly written by dotvault login), and re-runs the configuration loader immediately instead of waiting for the next config-refresh tick.
What a reload can and cannot apply:
- Applied in place — the dynamic sections:
rules,enrolments,sync.interval, andremote_configitself (the overlay fetcher is rebuilt and the refresh cadence re-derived on the next pass; note that a remote document still cannot carryremote_config— the section is local-only). These are the same sections the daemon already re-reads periodically on its config-refresh tick (default: the sync interval; seeremote_config.refresh_interval), whether the change came from an edited local config or the remote overlay. The signal just skips the wait. - Restart required — the static sections:
vault,web,api,agent,observability, and the top-levelbypass_system_configflag. These configure subsystems constructed once at startup (the Vault client, the web listener, the local API socket, the SSH agent, the OTel exporter). A reload that finds them changed logs a warning naming the changed sections; restart the daemon (systemctl --user restart dotvault.service) to apply them.
The packaged systemd unit wires SIGHUP as ExecReload=, so the canonical gesture on Linux is:
This targets the unit's MainPID specifically — preferable to kill -HUP $(pgrep -x dotvault), which would also signal any unrelated dotvault sync or go run ./cmd/dotvault invocation the user happens to be running (SIGHUP's default disposition is to terminate, so those side processes would die). On macOS the equivalent targeted form is launchctl kill SIGHUP gui/$(id -u)/com.goodtune.dotvault.
On Windows, SIGHUP is never delivered to processes. The system-tray icon (installed by both dotvault.exe run and dotvaultw.exe) carries a Reload config menu entry that performs exactly the same token re-read + immediate config reload; static-section changes log the same restart-required warning. Alternatively, changes to the dynamic sections still converge on the next config-refresh tick with no action at all.
Token re-read is automatic on Linux
On Linux the daemon also watches ~/.dotvault-token directly with inotify and re-reads it the moment the file is created or replaced — so when an interactive dotvault login writes a fresh token, the running daemon picks it up within seconds without any signal. This is built into the daemon (internal/tokenwatch); there is no separate unit to enable, and it works regardless of how dotvault was started. Deletes are ignored — the daemon keeps using its current in-memory token until a replacement is written. The watcher is a no-op off Linux; operators who want automatic re-read on macOS should script the launchctl kill form above on a launchd WatchPaths trigger.
Earlier releases shipped a dotvault-token-watch.path user unit that achieved the same nudge by SIGHUP-ing the daemon. It has been removed; the package upgrade deletes the unit files, but an enabled symlink left in ~/.config/systemd/user/ by a previous systemctl --user enable will persist and keep firing a (now redundant, but harmless) SIGHUP. After upgrading, clear it with systemctl --user disable --now dotvault-token-watch.path.