Repository navigation
RTOP-302: Stop the PLC when a task is stuck in its scan - #212
Merged
Merged
Conversation
A SCHED_FIFO thread waiting on a lock held by a preempted SCHED_OTHER thread can block forever on a single CPU. The dispatcher hits this on every tick through plc_get_state(). - State reads are a lock-free atomic load; writers still serialise. - state, task-array, completion, log and debug-write mutexes are upgraded to PTHREAD_PRIO_INHERIT before main. - The retain store's std::mutex becomes RtMutex (PI, BasicLockable). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The IEC priority was used directly as the SCHED_FIFO priority, so IEC 0 (the highest) ran at the lowest real-time level. Tasks could also reach FIFO 98, starving PREEMPT_RT IRQ threads at 50. - IEC priority clamped to 0..48 and mapped to FIFO 49 - priority, with a warning when clamped. - Dispatcher at FIFO 98, main watchdog at FIFO 99 so it can catch a stuck dispatcher. Both threads are named. - pthread_setschedparam errors are reported from its return code. - Host tests can link C sources compiled as C. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The dispatcher only logged overruns, and the stop path joined workers with no way to make a spinning one leave its body, so an unbounded loop in IEC code wedged the runtime in TRANSITIONING_TO_STOP. - A task found still in one scan for 10 of its own periods trips the dispatcher: it claims the stop, drains every task, then completes the stop on a separate thread (the unload joins the dispatcher). - The drain gives each in-flight scan until 10 periods after its release, then sends SIGUSR2; the handler siglongjmps to the task's recovery point when the task is inside its scan window. - A stop that had to abort a task lands ERROR instead of STOPPED. - The unused per-task heartbeat is removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Outputs kept their last value after a stop: the image was only cleared after the plugins had already stopped, so remote and local I/O stayed energised. After the tasks are joined, the dispatcher zeroes every %Q image slot, runs one last cycle_start/cycle_end so synchronous plugins write it, and holds it for PLC_OUTPUTS_OFF_SETTLE_MS so plugins polling the image from their own threads send it too. Program storage is not touched. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a task cannot be aborted, or the dispatcher itself stops ticking, nothing in-process can recover: the runtime stayed wedged until a 120 s fallback forced ERROR with the stuck threads still alive. - The teardown gives an aborted (or woken idle) task 2 s to exit, then calls watchdog_fatal_exit(). - The watchdog (FIFO 99, 100 ms tick) exits when a stop outlasts 10 x the longest task interval + settle + 30 s, or when the dispatcher has not ticked for max(10 base ticks, 1 s). - watchdog_fatal_exit() logs through log_emergency(), which never waits on a lock, and calls _exit(42). - The webserver restarts plc_main with --safe-mode --fault on exit 42; the runtime then reports ERROR without loading the program. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The abort window covers only the IEC program bodies, so a jump can never leave the image mutex owned by a dead thread. - Transition workers are created with explicit SCHED_OTHER; a fault stop spawned by the dispatcher no longer runs the teardown at FIFO 98. - unload_plc_program is serialised, so a shutdown racing a fault stop cannot unload or join twice. - The fatal exit leaves a marker the next boot consumes, so safe mode also follows when the supervisor could not read the exit code. - Teardown split into helpers, shared monotonic clock helper, named constants, escaped emergency log JSON, stale comments rewritten. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
marconetsf
reviewed
Oct 5, 2026
- First scans are not counted as stuck; the main watchdog trips one
only past PLC_FIRST_SCAN_TIMEOUT_MS (10 s), and the teardown and stop
budget give a first scan the same allowance.
- Once outputs are zeroed on stop, the journal is closed so plugin
writes (HMI, s7comm) cannot re-energise %Q during the settle.
- The safe-mode boot after a watchdog exit drives all outputs to 0
through the configured plugins under a claimed stop; a hang there is
recorded in the fault marker and skipped on the next boot.
- Safe mode survives later crashes until an upload clears it.
- --fault handling runs before the command socket exists.
- The Python block loader blocks the abort signal around fork/shm/stdio.
- Emergency log JSON escapes control characters.
- Dead plc_io_cycle.{cpp,h} removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
RTOP-302: a PLC task stuck in its scan (for example an IEC
WHILEloop that never exits) could not be stopped. The dispatcher only logged overruns. A STOP joined the worker with no timeout and wedged the runtime inTRANSITIONING_TO_STOP(every command answeredCOMMAND:BUSY) until a 120 s fallback forced ERROR with the stuck threads still alive. On a single core the spinning SCHED_FIFO task starved everything else. Outputs also kept their last value on every stop, and IEC task priorities were inverted (IEC 0, the highest, ran at FIFO 1).What changed
PLC_FIRST_SCAN_TIMEOUT_MS). The dispatcher claims the stop, stops releasing scans, and lets each in-flight scan run until 10 periods after its release (10 s for a first scan). A task still inside IEC program code then getsSIGUSR2, whose handlersiglongjmps to the existing per-thread recovery point. The unload runs afterwards on a SCHED_OTHER thread, and any stop that had to abort a task lands in ERROR.%Q, every%Qimage slot is zeroed, synchronous plugins get one lastcycle_start/cycle_end, and the zeroed image is held 500 ms (PLC_OUTPUTS_OFF_SETTLE_MS) for plugins that poll from their own threads._exit(42). This happens when an aborted task does not exit within 2 s, a stop overruns its budget (max(10 × the longest task interval, 10 s) + 30.5 s), or the dispatcher stops ticking for max(10 base ticks, 1 s). The fatal path logs without taking a lock (log_emergency) and writes a marker (/run/runtime/watchdog_fault). The webserver restarts with--safe-mode --faulton exit 42, and the next boot also consumes the marker, so the runtime comes back in safe mode reporting ERROR even when the supervisor can't read the exit code. That boot, before the command socket opens, claims a stop, starts the configured plugins (including VPP board plugins), drives all outputs to 0 and stops them. If that step itself hangs, the next exit records it in the marker and the following boot skips it, relying on the hardware safe-state watchdog. Safe mode then survives later crashes until an upload clears it.PTHREAD_PRIO_INHERIT, the retain store uses a PIRtMutex, and state reads are lock-free.plc_io_cycle.{cpp,h}are removed.docs/ARCHITECTURE.md,--faultindocs/DEVELOPMENT.md, and the lifecycle README coverage notes.Release note
How it was verified
There is no PR-triggered C build or test in CI, so everything below was run locally.
Unit
tests/host/run.sh(all pass):test_rt_mutex: the PI protocol is set, with a negative control.test_task_policy: priority mapping and clamping.test_image_outputs: every%Qslot is zeroed and inputs/memory are untouched. The test fails when the zeroing call is removed.tests/pytest/runtimemanager/test_watchdog_exit.py: exit 42 → safe mode with--faulton the first exit; other exits keep the rapid-crash rule..c/.hlines. Files the hook would fully rewrite were left unformatted.Lifecycle suite (
tests/lifecycle, Docker, a real compiled program): 24/24 passed, 0 build warnings.End to end. Real projects were uploaded with
openplc-cli uploadto a Docker runtime built from this branch, with each fixture's source and the compiled loop checked before reading the result:while magicValue < 10loop; spin confirmed in the uploaded.so)Falseon both a normal stop and a fault stop; stayedTrueon the commit before the output change (negative control)--safe-mode --fault, ERROR--cpuset-cpus=0): spin, shared globals, Modbus coil, STOP while stuckManual (developer): tested with the Opta Demo project against a runtime built from this branch; reported working.
Not run / out of scope
developmentas well:project.ymlexcludes SOEM paths that are missing, andtests/support/ethercat_stubs.cneedssoem/soem.h. Not changed here.GlobalVarstd::mutex, the Python GIL, third-party plugin locks). A hang on one of them ends in_exit(42).No requirements document: this is a bug fix, and the Jira task holds the description and acceptance criteria. No cybersecurity risk assessment: the change touches local thread scheduling, the PLC stop path and the webserver's handling of the runtime exit code. No network exposure, authentication, cryptography, update mechanism or external file parsing is affected.
🤖 Generated with Claude Code
https://claude.ai/code/session_01FWp9HZHNzjr7A2VLviZrg5