# Merged Plan: fix interactive `cloud login` session teardown ## Goal Fix the user-visible bug where interactive `cloud login` with a wrong password prints the retry message, then drops back to a fresh CraftOS shell with lost in-memory command history instead of prompting again. This should be handled as one end-to-end fix plan, split into small independently testable changes. ## Current Assessment The interactive `cloud login` branch itself is already covered and appears logically correct: - wrong password returns an unauthorized verdict; - the program queues `login-restore`; - it prints `invalid password, try again (ctrl+t to cancel)`; - it loops back to `readPassword()`. The observed shell/history loss points to the surrounding runtime being torn down, not to the login loop simply returning. There are two plausible and concrete runtime failure mechanisms in the current code. Either one can explain the same user-visible symptom, and both are worth fixing because both affect boot-session stability. ## Bugs / Fix Areas ### 1. Event-loop stop event collision `packages/trapos-core/apis/eventloop.lua` allocates loop IDs from a module-local counter and stops a running loop by queueing a synthetic global event named `@libeventloop/END_OF_LOOP/`. If `/apis/eventloop.lua` is loaded as independent module instances, two loops can both use ID `1`. The short-lived verifier loop created by `cloud login` can then queue an end event that also matches the boot event loop. Since `startup/boot.lua` runs shell + boot loop under `parallel.waitForAny`, a boot-loop exit kills the shell coroutine and loses shell history. Fix: - Allocate event loop IDs from a `_G`-backed process-global counter. - Avoid queueing a global synthetic stop event when `stopLoop()` is called from inside the same event loop's handler or timeout. Use an internal `stopRequested` flag and break after the current dispatch. - Keep queued synthetic stop events for external stop calls, where the loop may need to wake from `os.pullEventRaw()`. ### 2. Boot event-loop handler errors can tear down the session `eventloop.lua` currently dispatches handlers directly. A thrown handler escapes `runLoop()`. The boot event loop is run directly inside `parallel.waitForAny(shellFn, eventLoopFn)`, so an unhandled server/daemon error can terminate the TrapOS startup session and drop the user into a fresh shell. Fix: - Add protected handler dispatch as an explicit event-loop option, for example `createEventLoop({ onError = function(eventName, err) ... end })`. - Use protected dispatch for `_G.bootEventLoop` in `packages/trapos-boot/startup/boot.lua`. - Keep the default unprotected for short-lived program-local loops so development/test failures still fail loudly unless explicitly isolated. ### 3. Cloud daemon websocket open can throw during reconnect `packages/trapos-cloud/apis/libcloud.lua` calls `httpLike.websocketAsync(url)` without `pcall`. During wrong-password verification the daemon rapidly reconnects with the candidate secret, receives an unauthorized response, restores the persisted secret, and reconnects again. If `websocketAsync` throws during that churn, the error can currently escape through the boot event loop. Fix: - Forward-declare `scheduleReconnect` so `connect()` can call it. - Wrap `httpLike.websocketAsync(url)` in `pcall`. - On failure, log a warning and schedule a normal reconnect instead of throwing. ## Out Of Scope For This Fix ### Stale hello `messageId` matching Matching hello responses by the active hello `messageId` is useful cloud-login hardening. It can prevent stale hello responses from an older reconnect from being attributed to the current candidate. However, this does not directly explain shell/session teardown. Treat it as a follow-up unless the main fixes still leave wrong verdict attribution in manual or probe testing. ### Persisted shell history Do not add persisted shell history as a workaround. The correct fix is to stop tearing down the TrapOS session. ## Implementation Steps ### Step 1: Add regression coverage `packages/trapos-core/tests/eventloop.lua`: - Add a test that simulates independent module loads of `/apis/eventloop.lua`. - Create one boot-like loop and one verifier-like loop from separate module instances. - Stop the verifier-like loop from inside its handler. - Assert the boot-like loop still handles a later probe event. - Explicitly stop the boot-like loop at the end so the test cannot hang. `packages/trapos-core/tests/eventloop.lua`: - Add a protected-dispatch test where one handler throws but a sibling or later handler/event still runs. - Assert the configured error sink receives the error. `packages/trapos-cloud/tests/cloud.lua`: - Add a daemon test where `http.websocketAsync` throws during `startSession()` or a reconnect. - Assert the error does not propagate. - Assert a reconnect timer is scheduled. Keep existing `packages/trapos-cloud/tests/cloud-program.lua` retry tests. They already prove the program-level wrong-then-right interactive branch. ### Step 2: Fix `eventloop.lua` Change `packages/trapos-core/apis/eventloop.lua`: - Replace module-local event-loop ID allocation with a `_G` counter. - Add same-loop `stopRequested` handling. - Add optional protected dispatch support for handlers and timeout callbacks if practical. At minimum, protect regular event handlers used by boot servers. - Preserve existing `api.STOP` unregister behavior on successful handler return. - Preserve existing error contracts for duplicate registration and stopping an already stopped loop. ### Step 3: Use protected boot event loop Change `packages/trapos-boot/startup/boot.lua`: - Construct `_G.bootEventLoop` with an error handler. - Error handler should report the failure using `printError` if available, otherwise `print`. - The boot loop should continue after a server handler error. ### Step 4: Harden cloud reconnect Change `packages/trapos-cloud/apis/libcloud.lua`: - Forward-declare `scheduleReconnect` before `connect()`. - Wrap `httpLike.websocketAsync(url)` with `pcall`. - On failure, log `websocket open failed` and call `scheduleReconnect()`. - Avoid duplicate noisy logs while already in reconnect-loop mode. ### Step 5: Package versions Bump owning package versions and mirror them in `packages/index.json`: - `trapos-core` for `apis/eventloop.lua`. - `trapos-boot` for `startup/boot.lua`. - `trapos-cloud` for `apis/libcloud.lua`. Do not bump the full `trapos` meta-package unless repository release convention requires it for this change set. ## Verification Targeted iteration: ```sh just trapos-exec 'shell.run("/programs/runtest.lua", "--pretty", "/tests/eventloop.lua", "/tests/cloud.lua", "/tests/cloud-program.lua")' ``` Required after Lua/package edits: ```sh just check just test --pretty ``` Manual validation on a disposable TrapOS computer: 1. Deploy updated `trapos-core`, `trapos-boot`, and `trapos-cloud`. 2. Reboot so the cloud daemon and boot loop use the updated code. 3. Run `cloud login --force`. 4. Enter a known wrong password. 5. Confirm the retry message is printed and the password prompt returns. 6. Cancel or complete the login, then use up-arrow to confirm shell history still contains the command. 7. Enter the correct password and confirm `logged in`. 8. Run `cloud status` to confirm the daemon recovered normally. ## Recommended Split Keep this as one merged fix plan, but implement it in three reviewable chunks: 1. Event-loop stop collision and same-loop stop semantics. 2. Boot-loop protected dispatch. 3. Cloud reconnect crash-safety. This split makes each failure mode testable while still fixing the single user-facing bug end to end.