incus-compose implements health checks via a sidecar container called
ic-healthd. Incus has no native healthcheck support, so ic-healthd fills that
role.
ic-healthd is a core component. Every
healthcheck, every restart policy (restart: always | on-failure | unless-stopped), and everydepends_on: { condition: service_healthy }is enforced by this sidecar, not by Incus. If healthd is misconfigured, stopped, or crashing:
- instances are not restarted, and
- the project may fail to come up at all:
incus-compose upwaits forservice_healthydependencies to be reported healthy by healthd. If that status never arrives,upblocks until--dependency-timeout(default 5m;0waits forever) and then fails.Opt out of healthd entirely with
incus-compose up --no-healthd(this also drops the dependency wait);--no-depsskips the wait too. When health, restart, or startup-ordering behavior looks wrong, debug healthd first (see Debugging ic-healthd).
incus-compose up makes sure a healthd is watching the project when any service
declares a healthcheck, has a restart policy other than no, or is depended
on with condition: service_healthy. By default that is one daemon shared by
the whole server; see Scope. It:
user.healthcheck.scope, which is how the daemon
finds it.ic-healthd container if it is not already there, with an Incus
trust token.user.healthcheck.status.sequenceDiagram
participant IC as incus-compose
participant I as Incus
participant H as ic-healthd
IC->>I: mark the project user.healthcheck.scope
IC->>I: create a trust token
IC->>I: create the daemon if missing,<br/>inject token + API URL + marker
I->>H: start
H->>I: register its cert with the token (consumed)
H->>I: discover instances, then open a<br/>lifecycle listener
loop per watched instance
H->>I: exec user.healthcheck.test
I-->>H: exit code
H->>I: write user.healthcheck.status
end
The daemon is running before the regular services start, so service_healthy
dependencies can be evaluated. A project-scoped sidecar is removed by
incus-compose down; the shared daemon is never touched by it.
Health check config and runtime state live in the instance's user.* config
keys. There is no separate config file. ic-healthd reacts to
incus config set/instance create/delete changes as they happen via the Incus
event stream; incus-compose healthd reload remains available to force a full
manual resync.
See the Docker healthcheck docs for the value semantics: https://docs.docker.com/reference/dockerfile#healthcheck
user.incus-compose.managed true
user.healthcheck.enabled true
user.healthcheck.test '["CMD","wget","-q","--spider","http://localhost"]'
user.healthcheck.start_period 10s
user.healthcheck.start_interval 2s
user.healthcheck.interval 10s
user.healthcheck.timeout 5s
user.healthcheck.retries 3
user.healthcheck.status unknown | stopped | starting | healthy | unhealthy
user.healthcheck.restart always | on-failure | unless-stopped
user.healthcheck.ignore true
These keys are visible in incus config show <instance>.
user.healthcheck.status is the only key ic-healthd writes, and nothing else
writes it. All the others are set by incus-compose at instance creation time
and treated as read-only by the daemon.
That split is what makes the status trustworthy: an instance reports what the
daemon last saw, never what another process assumed. So a fresh instance carries
no status at all until the daemon reports one, and list shows it as Unknown
for that moment. The exception is up --no-healthd, where nothing will ever
report: those instances are created with unknown and keep it.
user.healthcheck.stopped is the one signal that goes the other way.
incus-compose stop sets it to say a stop was deliberate; the daemon reads it
and leaves unless-stopped instances alone, and writes the stopped status
itself from the event it sees anyway.
incus-compose pause sets the same marker, because a frozen instance answers no
healthcheck and would otherwise read as one that needs restarting. See
Pausing a watched service.
ic-healthd watches an instance only when it carries
user.healthcheck.enabled: "true". A healthcheck: block or a restart policy
alone is no longer enough - the instance has to say it wants watching.
incus-compose writes the key automatically for every service that declares a
healthcheck: or a restart policy other than no, so you do not normally set
it by hand.
Upgrading from a release before this? Instances created by an earlier version do not carry the key, so they are not watched: their healthchecks do not run and their restart policies are not enforced. Run
incus-compose uponce per project to fix it.upadds config keys an instance is missing without recreating anything, so no--recreateand no downtime is needed. ic-healthd logs a warning for every instance it finds with a healthcheck but no opt-in, so you can see what is affected withincus-compose healthd logs.
user.incus-compose.managed: "true" is a separate key, written on the project
and on every instance incus-compose creates. It records who created the thing,
which matters for other incus-compose features; ic-healthd does not read it and
it has no effect on whether an instance is watched.
Set user.healthcheck.enabled: "false" via x-incus. The service keeps its
healthcheck: block - it simply is not watched:
services:
sidecar-tool:
image: docker.io/example/tool:latest
x-incus:
user.healthcheck.enabled: "false"
user.healthcheck.ignore: "true" also excludes an instance, from discovery and
from every event handler. incus-compose sets it on the ic-healthd sidecar so it
does not watch itself. For a normal service prefer enabled: "false" - it says
the same thing in the same namespace as the rest of the healthcheck config.
One ic-healthd watches any number of projects from a single Incus event listener, so by default there is exactly one on the server:
| Scope | Where it runs | Watches |
|---|---|---|
global (default) |
instance ic-healthd in the Incus incus-compose project |
every project marked user.healthcheck.scope=global |
project |
instance {project}-ic-healthd in the project |
that one project |
The shared daemon gets a project, a bridge (icompose0) and a root disk of its
own, so nothing about how your default project is set up can break it. Both
are created on the first healthd up and neither is removed by healthd down.
up writes the choice to the Incus project as user.healthcheck.scope, and
that stored value then beats both the flag and the compose file:
flowchart TD
S([which daemon watches this project?]) --> P{"user.healthcheck.scope<br/>on the Incus project?"}
P -->|yes| USE[use it]
P -->|no| C{"--healthd-scope given?"}
C -->|yes| USEC[use it]
C -->|no| X{"x-incus-compose.healthd.scope?"}
X -->|yes| USEX[use it]
X -->|no| D[global]
So a project keeps the scope it was brought up with. Changing your mind later
means changing that key and running up again:
incus project set my-project user.healthcheck.scope=project
incus-compose up
up never leaves both daemons on one project: switching to global removes the
project's own sidecar before marking the project, and switching to project
marks it first so the shared daemon lets go before the sidecar appears.
x-incus-compose:
healthd:
scope: project
or incus-compose up --healthd-scope project the first time. Reasons to:
The cost is one container, one certificate and one event listener per project,
and the sidecar's limits.* counting against that project's quota (see
Sizing the sidecar).
Projects last brought up by a version before this carry no
user.healthcheck.scope at all, and the shared daemon only watches projects
that positively carry global. They are therefore invisible to it and keep
running on their own sidecar, with no window where both watch them.
Run incus-compose up per project when you want it moved. That removes the old
sidecar, revokes its certificate and hands the project to the shared daemon.
When keys are missing, ic-healthd falls back to:
| Key | Default |
|---|---|
| start_period | 0s (disabled) |
| start_interval | 5s |
| interval | 30s |
| timeout | 30s |
| retries | 3 |
retries must be greater than 0.
After retries consecutive failures the instance is restarted. The first
restart waits interval * retries; the delay doubles on every further restart,
capped at 5 minutes.
stateDiagram-v2
state "stopped on purpose" as parked
[*] --> stopped: the daemon finds it,<br/>not running yet
stopped --> starting: started, inside<br/>the start period
stopped --> healthy: started, test passes
starting --> healthy: test passes
starting --> unhealthy: start_period elapsed,<br/>test still failing
healthy --> unhealthy: retries consecutive failures
unhealthy --> healthy: test passes again
healthy --> stopped: it stopped
unhealthy --> stopped: it stopped
unhealthy --> restarting: a restart policy is set
restarting --> starting: first delay interval * retries,<br/>doubling, capped at 5m
stopped --> parked: it was stopped on purpose, with<br/>restart unless-stopped
parked --> starting: started again
This is user.healthcheck.status, the verdict you can read with
incus config get. It is not the same thing as the daemon's internal
per-instance state machine (idle/checking/restarting), which tracks what the
scheduler is doing right now - see
ic-healthd Internals - Instance state.
incus-compose does not read or inherit the HEALTHCHECK instruction embedded in
Docker images.
Incus imports OCI images via umoci, which converts the OCI image config into an
OCI runtime spec. The Docker HEALTHCHECK extension is not part of the OCI
image spec and is discarded during that conversion. Fetching it from the
registry at up time would require registry access on every run and fails in
air-gapped environments.
Workaround: Always declare healthcheck.test explicitly in the compose
file:
services:
db:
image: docker.io/postgres:16-alpine
healthcheck:
test: ["CMD", "pg_isready", "-U", "postgres"]
interval: 10s
timeout: 5s
retries: 5
restart: always, on-failure, or unless-stopped without a healthcheck
block is also handled. ic-healthd monitors the instance state and restarts it
when stopped, without running an exec-based test command.
With unless-stopped, instances stopped intentionally
(user.healthcheck.stopped=true, set by incus-compose stop) are not
restarted.
incus-compose pause freezes an instance, which stops it answering any
healthcheck. To the daemon that is indistinguishable from a service that fell
over, so without help it would restart out of the pause on the next interval.
pause therefore sets user.healthcheck.stopped, the marker stop already
uses, and unpause clears it. The daemon sees a stopped instance it was told
about, parks it, and touches nothing until it starts again.
Two consequences worth knowing:
user.healthcheck.status reads stopped. The daemon reports
what it can observe, and it cannot probe a frozen instance.instance-resumed as a start, so after unpause they leave the
instance unwatched until the next resync. Update the daemon
(incus-compose healthd up), or force one with
incus-compose healthd reload.Since: v1.3.0
ic-healthd runs in its own container and must reach the Incus HTTPS API from the inside. Two things are configured:
network - the Incus network (or host bridge) healthd attaches its NIC
to.incus - the Incus API URL healthd connects to.flowchart LR
subgraph P["its project"]
H["ic-healthd<br/>or {project}-ic-healthd"]
end
H -->|NIC| BR["network:<br/>that project's own bridge,<br/>project:network,<br/>or a host bridge"]
BR -->|"IPv4 gateway"| EP["incus:<br/>https://gateway:client-port<br/>or a pinned URL"]
EP --> API[Incus HTTPS API]
Both can be set in the compose file or overridden on the CLI. CLI flags and environment variables take priority over the compose file.
name: my-project
x-incus-compose:
healthd:
# Incus API endpoint healthd connects to.
# Default: `core.https_address` when it names a host, else the host address
# on `network` below with the port incus-compose itself connected on.
incus: https://<ip-of-the-projects-bridge>:8443
# `<project>:<network>` for a managed network, or a plain bridge name.
# We assume the current project if you leave the first part empty.
# Default: the bridge of the project the daemon runs in.
network: :default
| Flag | Environment variable | Compose key |
|---|---|---|
--healthd-incus |
INCUS_COMPOSE_HEALTHD_INCUS |
x-incus-compose.healthd.incus |
--healthd-network |
INCUS_COMPOSE_HEALTHD_NETWORK |
x-incus-compose.healthd.network |
incus-compose healthd up takes the same two options as --incus and
--network.
networkicompose0 for the shared daemon, the project's own default network
for a project-scoped one. Either way healthd can come up before the rest of
the project.<project>:<network> - a managed Incus network, optionally in another
project. A network the compose file declares is created before the sidecar
attaches to it; anything else must already exist.: - a host bridge name (e.g. incusbr0, or a bridge
Incus does not manage such as br0). It must already exist.The bridge's IPv4 gateway is used as the default Incus endpoint, so healthd can
reach Incus over it. A bridge Incus does not manage carries no gateway in its
config, so pair one with an explicit incus below.
incusEmpty (default) - resolved in this order:
core.https_address, if it names a host. 10.0.0.5:8443 is used as
https://10.0.0.5:8443 and nothing below is consulted. A bare :8443
names no host, so it falls through.network, with the port incus-compose connected
on. This is the case core.https_address = :8443 lands in: Incus listens
on all interfaces, so the bridge IP reaches it.Step 2 needs a HTTPS connection - over a unix socket there is no port to
reuse, so set --healthd-incus explicitly. It also needs a managed bridge; on
one Incus does not manage there is no gateway to read and healthd up fails
naming the network rather than guessing an endpoint the sidecar cannot reach.
An explicit URL - used verbatim, e.g. https://10.0.0.1:8443. Combine
with network to pin both the bridge and the endpoint.
empty below means the fallback, i.e. core.https_address names no host.
network |
incus |
Behavior |
|---|---|---|
| default | empty | Own bridge IP + client port (the default) |
| default | URL | Own bridge for the NIC, pinned endpoint |
bridge / project:network |
empty | Different bridge, auto-detected IP |
bridge / project:network |
URL | Different bridge, pinned endpoint |
Whichever daemon watches a project can exec into its instances and start, stop and restart them. What differs is how far that reaches.
A project-scoped sidecar gets a restricted certificate:
The shared daemon is registered unrestricted, deliberately. A restricted
certificate carries a fixed list of projects, and the whole point of the shared
daemon is to pick up projects created after it was registered. Practically it
means a compromised ic-healthd container is a compromised Incus server.
It is one container, running one binary, on an image you control via
--healthd-image, reachable only over the bridge you point it at. If that is
not a trade you want to make, use scope: project (see
Choosing project scope) - every project then gets a
daemon bounded to itself, at the cost of one container each.
The healthd command group manages the sidecar directly without touching
services. Each follows the project's scope, so in a global-scope project they
act on the shared daemon in the incus-compose project:
| Subcommand | Description |
|---|---|
logs [--follow] |
Stream the ic-healthd container log |
reload |
Send SIGHUP to force a full manual resync (rarely needed) |
restart |
Restart the ic-healthd container |
status |
Print the shared daemon's health status key |
up |
Create the sidecar, or replace one running an older image |
down [--force] |
Stop and remove the sidecar |
healthd up accepts --image, --binary, --incus, --network, --scope,
--pull and --timeout. Inside a project it refuses with an error when no
service there requires healthd (no healthcheck, no restart policy, no
service_healthy dependency).
Since: v1.3.0: healthd status.
All of them also run with no compose file in sight, where they act on the shared daemon. That is how you put one on a server before any project exists, and how you look at it afterwards:
incus-compose healthd up # create the shared daemon
incus-compose healthd logs # watch it
healthd up this way marks no project and so watches nothing by itself -
projects opt in on their own up. The others fail with
no ic-healthd is running rather than guessing at a project.
healthd down on the shared daemon stops health checking for every project
using it, so it lists the other projects and asks first. --force skips the
question, and is required when there is no terminal to ask on (CI, scripts).
incus-compose down never touches the shared daemon at all.
Everything the daemon is configured with - debug logging, workers,
restart-workers, x-incus, incus - is injected as environment on the
container when it is created, and a running daemon is never reconfigured in
place. None of it is compared against what the daemon runs either: up replaces
a sidecar only for a newer image (see Sidecar Image), so
changing any of these is a manual recreate:
# verbose logging on
incus-compose healthd down --force
incus-compose --debug healthd up
# and back off again
incus-compose healthd down --force
incus-compose healthd up
--debug is the global incus-compose flag and is inherited by healthd
operations; the others come from x-incus-compose.healthd (see
Sizing the sidecar).
--trace is a level below it, for the lines the daemon emits per Incus event
and per check. They are what you want when a project is not being watched and
you need to see the events arriving, and what you do not want otherwise - on a
busy server they bury everything else. It implies --debug:
incus-compose healthd down --force
incus-compose --trace healthd up
incus-compose healthd logs --follow
With
scope: globalthis restarts the daemon every other project is using, so health checking pauses server-wide for a few seconds.--forceis what says you meant it; drop it to be told which projects are affected and asked first.
incus-compose up --no-healthd
You can run the daemon yourself instead of letting up create a sidecar, and
point incus-compose at it with up --external-healthd /
down --external-healthd. incus-compose then uses healthd features but does not
create or look up a sidecar of its own.
Set it permanently for a project in the compose file instead of passing the flag every time:
x-incus-compose:
healthd:
external: true
--external-healthd and the compose key combine with OR: either one is enough
to turn it on, there is no flag to force it back off for a project that sets it
in the compose file.
For the ic-healthd run flags, the registration handshake, and the local
edit-run-reload loop, see
ic-healthd Internals - Running the daemon directly.
Default image: ghcr.io/lxc/incus-compose/ic-healthd:{version}
Override with --healthd-image flag or INCUS_COMPOSE_HEALTHD_IMAGE env var.
The container is named ic-healthd for the shared daemon,
{project}-ic-healthd for a project-scoped one, and carries two tags:
user.healthcheck.ignore=true, so ic-healthd skips itself during discovery and
every event handler, and user.healthcheck.daemon=true, which incus-compose
uses to locate the sidecar instance (healthd logs/restart/etc.) - ignore
is a general opt-out any instance can carry, so it can't double as the sidecar's
own identifying marker.
Both incus-compose up and incus-compose healthd up upgrade the daemon for
you: when the image you ask for is a newer release than the one it is running,
the sidecar is stopped, removed and recreated from that image. The comparison is
semver and only ever moves forward, so a machine on an older incus-compose
cannot downgrade a daemon shared with everybody else. Tags that are not release
versions - moving tags like latest, and git describe builds - are not
comparable, so those replace on any difference and CI and --healthd-binary
keep rolling.
Updating the daemon on a server therefore needs no compose file and no project:
incus-compose self-update
incus-compose healthd up
The image alias is the only thing that triggers this. A daemon running the image you asked for is left alone however much else differs - see Changing the daemon's settings.
The sidecar runs with limits.cpu: 2 and limits.memory: 256MiB. Change that,
or set any other Incus instance config on it, with x-incus:
x-incus-compose:
healthd:
workers: 256
restart-workers: 64
x-incus:
limits.cpu: 4
limits.memory: 512MiB
workers (128) and restart-workers (32) cap the health checks and the
restarts the daemon runs at once across every project it watches. They are
separate pools because a restart holds its worker far longer than a check does -
see ic-healthd Internals - Worker pools. A
shared daemon watching many projects is the case worth raising them for.
Quota. A project-scoped sidecar lives in your project, so its
limits.cpu/limits.memoryare added to what your services use when Incus checks a project-levellimits.*. Budget for it. The shared daemon lives in its own project and does not count against any compose project, which is one more reason the default scope isglobal.
The first project to bring the shared daemon up supplies its incus, workers,
restart-workers and x-incus; a later project whose healthd block differs is
warned about and otherwise ignored, so one compose file cannot restart the
daemon everybody else is using. To apply new settings, take it down and back up:
incus-compose healthd down --force
incus-compose healthd up
Because healthd drives all health and restart behavior, most "container did not
restart" or "stuck service_healthy" problems are diagnosed from the sidecar.
Work through these in order.
Instances are named <service>-1 (the replica index starts at 1) and live in
the Incus project named after your compose project, so pass --project.
ic-healthd writes its verdict to user.healthcheck.status
(unknown | stopped | starting | healthy | unhealthy):
incus config get web-1 user.healthcheck.status --project <project>
starting that never becomes healthy means the test never passes within the
start period; unhealthy means it failed retries times. An empty value or
unknown on a running instance means no daemon has reported on it at all -
check that the sidecar is running (step 4) and that the instance carries
user.healthcheck.enabled: "true".
All inputs live in user.healthcheck.*. If a key is wrong, healthd behaves
wrong - it never reads the compose file directly:
incus config show web-1 --project <project> | grep -E 'user\.(healthcheck|restart)'
incus-compose healthd logs --follow
Enable debug logging for full per-check detail (failures, retry counts,
inStart transitions, restart delays). The --debug flag is inherited by the
sidecar at creation, so recreate it with debug on (see
Changing the daemon's settings):
incus-compose healthd down --force
incus-compose --debug healthd up
incus-compose healthd logs --follow
The container is named ic-healthd in the Incus incus-compose or
{project}-ic-healthd for a project-scoped one. If it is missing or stopped,
nothing is being monitored:
incus-compose list # the daemon is listed by default (since 1.0.0-rc.1)
incus-compose healthd up # create it if missing
Remember: incus-compose start never (re)starts the sidecar - only up does.
healthd runs user.healthcheck.test via incus exec. Run it yourself to see
why it fails:
incus-compose exec <service> -- sh -c 'wget -q --spider http://localhost; echo exit: $?'
If you change user.healthcheck.* keys directly (instead of via up),
ic-healthd picks them up on its own via the Incus event stream - no action
needed. If you ever suspect it missed something (e.g. after a change made while
its event listener was disconnected and before it reconnected), force a full
resync:
incus-compose healthd reload # sends SIGHUP
incus-compose up hangs or times out on dependenciesIf a service uses depends_on: { condition: service_healthy }, up waits for
healthd to report the dependency healthy before starting the dependent
service. A broken or missing healthd means that status never arrives and up
blocks until --dependency-timeout (default 5m) elapses, then fails.
Confirm the dependency's status with steps 1-3 above; it is likely stuck on
starting or unhealthy.
If you only want to bring the project up without the wait, opt out:
incus-compose up --no-healthd # also stops managing healthchecks/restarts
# or keep healthd but skip the wait:
incus-compose up --no-deps
Sidecar has wrong config (missing --incus/--project flags)?
This can happen when ic-healthd was created by an older version of incus-compose. Recreate it:
incus-compose healthd down --force
incus-compose healthd up
Sidecar not running after incus-compose start?
start never creates or starts the sidecar; only up does. Use
incus-compose healthd up to start it independently.