Skip to content

RTOP-302: Stop the PLC when a task is stuck in its scan - #212

Merged
thiagoralves merged 9 commits into
developmentfrom
bugfix/RTOP-302-stuck-task-watchdog
Oct 5, 2026
Merged

thiagoralves merged 9 commits into
developmentfrom
bugfix/RTOP-302-stuck-task-watchdog

Conversation

@thiagoralves

@thiagoralves thiagoralves commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Problem

RTOP-302: a PLC task stuck in its scan (for example an IEC WHILE loop that never exits) could not be stopped. The dispatcher only logged overruns. A STOP joined the worker with no timeout and wedged the runtime in TRANSITIONING_TO_STOP (every command answered COMMAND:BUSY) until a 120 s fallback forced ERROR with the stuck threads still alive. On a single core the spinning SCHED_FIFO task starved everything else. Outputs also kept their last value on every stop, and IEC task priorities were inverted (IEC 0, the highest, ran at FIFO 1).

What changed

  • Stuck-task watchdog (dispatcher). A task found still in the same scan on 10 consecutive due ticks (10 of its own periods) trips the dispatcher; the count resets whenever the task is found idle, and ticks replayed after a late dispatcher count as missed periods too. A task's first scan (initialisation, e.g. a Python FB starting) is not counted; the main watchdog trips it only past 10 s (PLC_FIRST_SCAN_TIMEOUT_MS). The dispatcher claims the stop, stops releasing scans, and lets each in-flight scan run until 10 periods after its release (10 s for a first scan). A task still inside IEC program code then gets SIGUSR2, whose handler siglongjmps to the existing per-thread recovery point. The unload runs afterwards on a SCHED_OTHER thread, and any stop that had to abort a task lands in ERROR.
  • Outputs off on every stop. After the drain, the journal is closed so plugin writes (HMI, s7comm) can't re-energise %Q, every %Q image slot is zeroed, synchronous plugins get one last cycle_start/cycle_end, and the zeroed image is held 500 ms (PLC_OUTPUTS_OFF_SETTLE_MS) for plugins that poll from their own threads.
  • Last resort: _exit(42). This happens when an aborted task does not exit within 2 s, a stop overruns its budget (max(10 × the longest task interval, 10 s) + 30.5 s), or the dispatcher stops ticking for max(10 base ticks, 1 s). The fatal path logs without taking a lock (log_emergency) and writes a marker (/run/runtime/watchdog_fault). The webserver restarts with --safe-mode --fault on exit 42, and the next boot also consumes the marker, so the runtime comes back in safe mode reporting ERROR even when the supervisor can't read the exit code. That boot, before the command socket opens, claims a stop, starts the configured plugins (including VPP board plugins), drives all outputs to 0 and stops them. If that step itself hangs, the next exit records it in the marker and the following boot skips it, relying on the hardware safe-state watchdog. Safe mode then survives later crashes until an upload clears it.
  • Priorities. Watchdog FIFO 99, dispatcher FIFO 98. IEC priority is clamped to 0..48 and mapped to FIFO 49 − priority, with a warning when clamped.
  • Priority inversion. Runtime mutexes are re-initialised with PTHREAD_PRIO_INHERIT, the retain store uses a PI RtMutex, and state reads are lock-free.
  • Teardown hardening. Program unload is serialised against a concurrent shutdown. The Python block loader blocks the abort signal around its fork/shm/stdio work. The unused per-task heartbeat and the dead plc_io_cycle.{cpp,h} are removed.
  • Docs. "Watchdog System" in docs/ARCHITECTURE.md, --fault in docs/DEVELOPMENT.md, and the lifecycle README coverage notes.

Release note

  • Task scheduling changes for existing projects. A task with IEC priority 1 moved from FIFO 1 to FIFO 48.
  • New trip condition. A lower-priority task delayed by higher-priority tasks for 10 of its own periods now trips the watchdog, where before it only ran at a reduced rate.

How it was verified

There is no PR-triggered C build or test in CI, so everything below was run locally.

Unit

  • tests/host/run.sh (all pass):
    • test_rt_mutex: the PI protocol is set, with a negative control.
    • test_task_policy: priority mapping and clamping.
    • test_image_outputs: every %Q slot is zeroed and inputs/memory are untouched. The test fails when the zeroing call is removed.
  • tests/pytest/runtimemanager/test_watchdog_exit.py: exit 42 → safe mode with --fault on the first exit; other exits keep the rapid-crash rule.
  • CI pytest set: 152 passed, 7 skipped. Plugin step: 26 passed.
  • pylint 10/10; black, isort and ruff clean on the changed Python.
  • clang-format 21.1.0 on the new or changed .c/.h lines. Files the hook would fully rewrite were left unformatted.

Lifecycle suite (tests/lifecycle, Docker, a real compiled program): 24/24 passed, 0 build warnings.

End to end. Real projects were uploaded with openplc-cli upload to a Docker runtime built from this branch, with each fixture's source and the compiled loop checked before reading the result:

Scenario Result
Original Opta Demo (while magicValue < 10 loop; spin confirmed in the uploaded .so) trips 10 periods (200 ms) after it locks, program unmapped, ERROR
Two tasks on a shared global, fast one spinning (IEC −5 and 100) ERROR every run; in one run the slow task was blocked on a global's mutex held by the aborted task, got its grace, was aborted, ERROR at 1.8 s; IEC 100 → FIFO 1 with a warning
Task taking ~8 periods per scan, 30 s no trips
STOP while a 2 s task is stuck ERROR at 16 s (10 periods after the stuck scan's release)
Remote Modbus TCP coil held TRUE device received False on both a normal stop and a fault stop; stayed True on the commit before the output change (negative control)
Abort swallowed (LD_PRELOAD shim) FATAL logged, marker written, exit 42, restarted as --safe-mode --fault, ERROR
Dispatcher hung in a plugin hook stall detected at ~1.05 s, exit 42, safe mode, ERROR
Marker present, plain restart (exit code not visible) boots in safe mode, ERROR
Healthy Opta Demo runs, stops, restarts, no trips
10 ms task with a ~1 s first scan before review fixes: ERROR at 100 ms; now runs, no trip
First scan that never ends trips at 10 s ("first scan still running after 10000 ms"), ERROR
Python FB on a 1 ms task runs; Python child started (the loader stayed under 10 ms on this host)
Modbus slave HMI writing the coil during the 500 ms settle (unit 1) before: device False then True again; now stays False
Exit 42 then safe-mode boot, coil held TRUE before: device stays True; now goes False ~1 s after the exit
Safe-mode boot outputs-off hung by a plugin second exit records the context; next boot skips it, ERROR, no loop
Single CPU (--cpuset-cpus=0): spin, shared globals, Modbus coil, STOP while stuck all pass; the webserver is starved while a task spins, but the dispatcher and watchdog still act

Manual (developer): tested with the Opta Demo project against a runtime built from this branch; reported working.

Not run / out of scope

  • Ceedling. It fails on development as well: project.yml excludes SOEM paths that are missing, and tests/support/ethercat_stubs.c needs soem/soem.h. Not changed here.
  • Mutexes that can't take priority inheritance (STruC++ GlobalVar std::mutex, the Python GIL, third-party plugin locks). A hang on one of them ends in _exit(42).
  • Boot outputs-off timing. It runs before the command socket opens, so the runtime is unreachable during it (about 1 s normally, up to the 40.5 s budget if a plugin hangs).
  • Overlap with open external PR fix(runtime): use progress counter for watchdog monitoring #208 (watchdog progress counter). It touches the same heartbeat code and will conflict.

No requirements document: this is a bug fix, and the Jira task holds the description and acceptance criteria. No cybersecurity risk assessment: the change touches local thread scheduling, the PLC stop path and the webserver's handling of the runtime exit code. No network exposure, authentication, cryptography, update mechanism or external file parsing is affected.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FWp9HZHNzjr7A2VLviZrg5

thiagoralves and others added 8 commits October 1, 2026 16:26
A SCHED_FIFO thread waiting on a lock held by a preempted SCHED_OTHER
thread can block forever on a single CPU. The dispatcher hits this on
every tick through plc_get_state().

- State reads are a lock-free atomic load; writers still serialise.
- state, task-array, completion, log and debug-write mutexes are
  upgraded to PTHREAD_PRIO_INHERIT before main.
- The retain store's std::mutex becomes RtMutex (PI, BasicLockable).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The IEC priority was used directly as the SCHED_FIFO priority, so IEC 0
(the highest) ran at the lowest real-time level. Tasks could also reach
FIFO 98, starving PREEMPT_RT IRQ threads at 50.

- IEC priority clamped to 0..48 and mapped to FIFO 49 - priority, with a
  warning when clamped.
- Dispatcher at FIFO 98, main watchdog at FIFO 99 so it can catch a
  stuck dispatcher. Both threads are named.
- pthread_setschedparam errors are reported from its return code.
- Host tests can link C sources compiled as C.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The dispatcher only logged overruns, and the stop path joined workers
with no way to make a spinning one leave its body, so an unbounded loop
in IEC code wedged the runtime in TRANSITIONING_TO_STOP.

- A task found still in one scan for 10 of its own periods trips the
  dispatcher: it claims the stop, drains every task, then completes the
  stop on a separate thread (the unload joins the dispatcher).
- The drain gives each in-flight scan until 10 periods after its
  release, then sends SIGUSR2; the handler siglongjmps to the task's
  recovery point when the task is inside its scan window.
- A stop that had to abort a task lands ERROR instead of STOPPED.
- The unused per-task heartbeat is removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Outputs kept their last value after a stop: the image was only cleared
after the plugins had already stopped, so remote and local I/O stayed
energised.

After the tasks are joined, the dispatcher zeroes every %Q image slot,
runs one last cycle_start/cycle_end so synchronous plugins write it,
and holds it for PLC_OUTPUTS_OFF_SETTLE_MS so plugins polling the image
from their own threads send it too. Program storage is not touched.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a task cannot be aborted, or the dispatcher itself stops ticking,
nothing in-process can recover: the runtime stayed wedged until a 120 s
fallback forced ERROR with the stuck threads still alive.

- The teardown gives an aborted (or woken idle) task 2 s to exit, then
  calls watchdog_fatal_exit().
- The watchdog (FIFO 99, 100 ms tick) exits when a stop outlasts
  10 x the longest task interval + settle + 30 s, or when the
  dispatcher has not ticked for max(10 base ticks, 1 s).
- watchdog_fatal_exit() logs through log_emergency(), which never waits
  on a lock, and calls _exit(42).
- The webserver restarts plc_main with --safe-mode --fault on exit 42;
  the runtime then reports ERROR without loading the program.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The abort window covers only the IEC program bodies, so a jump can
  never leave the image mutex owned by a dead thread.
- Transition workers are created with explicit SCHED_OTHER; a fault
  stop spawned by the dispatcher no longer runs the teardown at FIFO 98.
- unload_plc_program is serialised, so a shutdown racing a fault stop
  cannot unload or join twice.
- The fatal exit leaves a marker the next boot consumes, so safe mode
  also follows when the supervisor could not read the exit code.
- Teardown split into helpers, shared monotonic clock helper, named
  constants, escaped emergency log JSON, stale comments rewritten.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread core/src/plc_app/plc_state_manager.cpp Outdated
Comment thread core/src/plc_app/plc_state_manager.cpp
Comment thread webserver/runtimemanager.py Outdated
Comment thread core/src/plc_app/plc_state_manager.cpp
Comment thread core/src/plc_app/utils/watchdog.c
Comment thread core/src/plc_app/plc_main.c Outdated
Comment thread core/src/plc_app/utils/log.c Outdated
Comment thread core/src/plc_app/plc_io_cycle.cpp Outdated
- First scans are not counted as stuck; the main watchdog trips one
  only past PLC_FIRST_SCAN_TIMEOUT_MS (10 s), and the teardown and stop
  budget give a first scan the same allowance.
- Once outputs are zeroed on stop, the journal is closed so plugin
  writes (HMI, s7comm) cannot re-energise %Q during the settle.
- The safe-mode boot after a watchdog exit drives all outputs to 0
  through the configured plugins under a claimed stop; a hang there is
  recorded in the fault marker and skipped on the next boot.
- Safe mode survives later crashes until an upload clears it.
- --fault handling runs before the command socket exists.
- The Python block loader blocks the abort signal around fork/shm/stdio.
- Emergency log JSON escapes control characters.
- Dead plc_io_cycle.{cpp,h} removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@thiagoralves
thiagoralves merged commit 950d592 into development Oct 5, 2026
3 checks passed
@thiagoralves
thiagoralves deleted the bugfix/RTOP-302-stuck-task-watchdog branch October 5, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants