Make generated run lifecycle autonomous

This commit is contained in:
npc0-hue
2026-08-06 18:57:42 +08:00
parent acec5e4367
commit 4e78957a60
22 changed files with 583 additions and 236 deletions
@@ -2,14 +2,14 @@
Platform currently stores `serverInstances.state` as both desired state and observed runtime state. `CompleteRunJob` projects successful lifecycle jobs directly into that field, so a previous `process.start` success can leave a server as `running` even after the actual Run-managed process is gone. A manually started generated Run can register and heartbeat, but platform will not dispatch another start job because it trusts the stale stored state.
Run already has the safer primitive: plugin-declared `process.status` executes inside the generated Run workspace and returns a redacted `processState`. This change uses that existing channel as the observed runtime source.
Run already has the safer primitive: it owns the generated package startup path and can report redacted `processState` facts from inside the generated Run workspace. This change uses Run reports as the observed runtime source and avoids Platform registration-time probes.
## Goals / Non-Goals
**Goals:**
- Make generated Run startup reconcile stale platform lifecycle state through a platform-dispatched, Run-executed `process.status` job.
- Stop generated Run startup from relying on a platform-dispatched registration-time `process.status` job.
- Project server state from Run `processState` for status/start/stop lifecycle results.
- Preserve platform ownership of authorization, command dispatch, leases, and audit.
- Preserve Platform ownership of authorization, leases, and audit for explicit operator requests while treating generated Run bootstrap as Run-owned.
**Non-Goals:**
- Add a new live telemetry protocol or raw process list to heartbeat.
@@ -18,9 +18,8 @@ Run already has the safer primitive: plugin-declared `process.status` executes i
## Decisions
- Use `process.status` rather than adding heartbeat fields. This keeps state reconciliation inside the existing job lease, capability, audit, and plugin-declared action model.
- Queue status reconciliation on generated Run registration when stored state is `running` or `failed`. Those states are the ones most likely to be stale after a manually restarted Run or process crash.
- Skip reconciliation when an active lifecycle job already exists for the server. The active job is already the current control operation and should not be raced by a status probe.
- Do not queue status reconciliation on generated Run registration. Registration confirms identity and session only; Run-owned lifecycle/status reports correct stale Platform projections.
- Keep explicit `process.status` result projection for operator-requested or Run-reported status flows that are not registration bootstrap side effects.
- Project `process.status` into lifecycle state with conservative mapping: `running` => `running`, `stopped/not-started` => `stopped`, unexpected `exited` => `failed`, operator-stopped `exited` => `stopped`.
## Risks / Trade-offs