Appearance
How do we run a fleet change and independently verify the result? #
A fleet change is complete when the resulting device state is correct, not when the job runner says the command was sent.
Use a bounded cohort, define the expected outcome before execution, and verify that outcome through an independent signal such as version, process state, telemetry, Pulse, logs, or a product health check.
Define success before you run anything #
For every fleet operation write down:
- target cohort;
- command/change;
- preconditions;
- expected resulting state;
- acceptable failure rate;
- stop/rollback condition;
- independent verification signal.
Example:
Restart
northstar-agenton 25 tagged staging devices. Success means all 25 reconnect, the process is active, and the expected binary checksum/version is visible within the verification window.
That is much stronger than “run systemctl restart on 25 devices.”
Start with the smallest meaningful cohort #
A normal rollout shape is:
- one representative device;
- a small staging tag/cohort;
- one site/customer/product generation;
- a wider production cohort;
- remainder of compatible fleet.
Do not make the first production test the entire fleet because the command appears simple.
Check preconditions #
Before executing, confirm facts that change the risk:
- devices are the intended product generation;
- required storage/memory is available;
- devices are sufficiently online;
- current software/configuration is known;
- dependencies are healthy;
- maintenance window/customer impact is acceptable;
- rollback path exists where required.
Pulse and inventory data can help reveal unexpected variation before action.
Execute through a bounded mechanism #
Use Fleet Jobs or another supported orchestration path that gives you explicit target scope and progress rather than running an uncontrolled shell loop from an engineer's laptop.
For scripts/actions:
- make them safe to retry where possible;
- fail clearly;
- avoid unbounded loops;
- avoid interactive assumptions;
- emit structured/useful output;
- respect local safety and product ownership boundaries.
Do not use job completion as the only verification #
A command can exit successfully while the intended outcome is wrong.
Examples:
- package command succeeds but old process is still running;
- restart succeeds but application immediately crashes;
- config file writes but application rejects it;
- container starts but reports unhealthy;
- device reboots but never reconnects.
Verify the outcome through a separate path.
Good independent verification signals #
Depending on the change:
- process/service active state;
- binary checksum;
- package/application version;
- container image/version;
- product telemetry value;
- reported configuration;
- health/dependency test;
- expected log event;
- Device Class Pulse conformity;
- reconnect after reboot.
Prefer a signal that represents the actual customer/product outcome rather than merely another view of the same command execution.
Treat non-reporting devices explicitly #
After a change, separate:
- succeeded and verified;
- command failed;
- command succeeded but verification failed;
- device went offline;
- device never participated because it was offline;
- result unknown.
“98% successful” can hide the two devices you most need to investigate if the unknowns are folded into the percentage.
Stop conditions matter #
Define when the rollout stops automatically or operationally.
Examples:
- more than 2% fail verification;
- any device fails to reconnect after reboot;
- a safety/product health signal degrades;
- unexpected product generation appears in the cohort;
- customer-facing error rate increases.
Do not continue merely because the job engine is capable of targeting more devices.
Rollback is another fleet change #
A rollback should have its own expected state and verification signal.
If the change cannot be rolled back, make that explicit before execution and use a smaller cohort accordingly.
Record enough evidence for incident review #
After the maintenance window retain:
- who initiated the change;
- target selection;
- command/version/configuration;
- timestamps;
- failures/unknowns;
- verification result;
- follow-up actions.
This turns “we pushed something last night” into an auditable operational event.