A C++ lifecycle orchestration layer for ROS 2 on resource-constrained embedded systems.
β‘ 82.1% faster boot time on real robot hardware πΎ 40.7% lower stable RAM vs a sequential per-node
ros2 run(CLI) baseline π§ Dependency-ordered, low-overhead ROS 2 lifecycle orchestration in C++
Validated on a commercial robot platform (IFA 2025 showcase).
β οΈ Scope of comparison: All results below are reported relative to a sequential per-noderos2 run(CLI) bring-up, i.e., each node started through a separateros2 runinvocation viastd::system(). We do not claim results relative to the parallel-capableros2 launchsystem.
Measured across three configurations (10 runs each) on the same hardware and node set:
- Boot time: 60.04 s β 10.74 s (β82.1%, sequential
ros2 runCLI β C++ DLM) - Stable RAM: 464.09 MiB β 275.07 MiB (β40.7%)
- Avg RAM during boot: 247.34 MiB β 189.58 MiB (β23.3%)
An ablation isolates the two effects (see Β§5): removing the CLI/interpreter front end (native fork/exec) accounts for essentially all of the memory saving, while dependency-aware parallel activation drives most of the remaining boot-time reduction.
π Tested on a low-end embedded SoC (LG DQ1, quad Cortex-A53 @ 1 GHz, 1 GB RAM).
π Boot time is measured from the manager's start log entry to the initialization-complete marker (emitted after all initialization worker threads join). Every run was verified to contain all 14 managed nodes in the ACTIVE state.
π Memory is whole-system used memory (from free), not per-process RSS. The stable value is sampled 90 s after all nodes reach ACTIVE; raw KiB counters are converted to MiB.
π These results suggest a practical deployment path for ROS 2 on highly constrained embedded hardware.
ROS 2 can remain challenging to deploy on low-cost embedded robots when memory budgets and startup constraints are tight.
This project demonstrates that ROS 2 can be:
- deployed on 1 GB-class hardware
- brought up with dependency-ordered lifecycle coordination
- used in real production environments
| Approach | Strengths | Limitations |
|---|---|---|
Per-node ros2 run (CLI, Python front end) |
Simple, no launch file needed | Interpreter overhead per node, sequential, high transient RAM |
ros2 launch (Python) |
Flexible, concurrent start, event handlers | Interpreter overhead; launch-based orchestration |
| systemd | Stable process management | No ROS 2 lifecycle awareness |
| ROS 2 Composition | Efficient intra-process execution | No system-level orchestration; shared address space |
| micro-ROS | Optimized for MCUs | Not for Linux-based systems |
| LifecycleManager | Dependency-ordered, low-overhead, lifecycle-aware | More integration effort than default launch-based workflows |
π§ [NOTICE] Project Status
This repository currently provides architecture documentation, YAML configuration examples, measurement methodology, and empirical validation results for the Deterministic Lifecycle Manager.
These materials are intended to make the design and observed system behavior technically reviewable before the full implementation is publicly released.
A partial or full source release under Apache License 2.0 is being prepared through internal compliance review.
- Overview - "What & Why?"
- Structural Pain Points in Production Systems - "Pain Point"
- Architecture & LifecycleManager - "Solution"
- Deterministic Boot Flow - "Deep-dive"
- Metrics & Validation - "Validation"
- Source, Build & License - "Open-Source Status"
This document introduces the "Deterministic Lifecycle Manager" (DLM), a C++βbased lifecycle orchestration service designed to run ROS 2 reliably on lowβend embedded robotic platforms. The target systems are cost-constrained commercial robotic platforms built on entry-level SoCs such as Rockchip PX30, LG DQ1, or other Cortex-A35/A53-class architectures, typically equipped with around 1 GB of RAM. On these platforms, boot-time behavior and peak resource usage critically impact system stability.
During system integration and productionβequivalent platform evaluation, we repeatedly encountered critical issues when bringing up ROS 2 nodes through an interpreter-based, per-node ros2 run (CLI) path on resourceβconstrained hardware.
The most common problems were:
- High baseline RAM usage before application logic starts
- Frequent OOM (Out Of Memory) kills during early boot
- Unstable and non-deterministic startup sequences in production images Similar issues were repeatedly observed during evaluation on multiple low-end platforms, with the detailed quantitative results in this repository collected on the LG DQ1-based test platform. In our evaluated configuration, these issues were not sufficiently resolved through parameter tuning, configuration changes, or partial optimization. Our working conclusion was that orchestration overhead in the evaluated interpreter-based launch path was a major contributor under tight memory constraints.
To address these limitations, we implemented the Deterministic Lifecycle Manager as a minimal, native C++ service that launches and supervises ROS 2 nodes directly as OS-level processes. Key characteristics include:
- No Python front-end dependency in the native launch path on the target system
- ROS 2 nodes built and executed as native C++ binaries
- Hardware-Aware Concurrent Spawning: Utilizes C++ worker threads alongside standard POSIX primitives (
fork(),execvp(),SIGCHLD) to balance parallel execution. By bounding the number of simultaneous process launches to the system hardware concurrency (CPU cores), it limits CPU contention and OS scheduler overhead. - Dual-Mode "Benchmark Mode": Supports toggling between native C++ spawning and legacy per-node CLI execution (
std::system("ros2 run β¦")) to allow direct A/B performance comparisons. This design preserves standard ROS 2 lifecycle semantics and DDS-based communication while reducing the runtime overhead associated with the evaluated interpreter-based launch path.
π‘ Note: The name "Deterministic" refers to dependency-constrained, repeatable completion (activation respects the dependency partial order and the same set of nodes reaches ACTIVE on every run) β not hard real-time or exact-timing guarantees. Under parallel activation, exact per-run timing varies with thread scheduling.
This section explains why an interpreter-based, per-node bring-up path becomes a reliability bottleneck on lowβend SoCs, based on issues repeatedly observed during productionβequivalent system integration.
Interpreter-based bring-up (e.g., per-node ros2 run, or ros2 launch) relies on Python processes that are loaded and initialized during system boot.
- On low-cost hardware with 1 GB or less of RAM, this introduces substantial overhead before any application logic starts.
- As the number of nodes increases, Python interpreter initialization and runtime management consume a significant portion of system resources, frequently leading to memory pressure and OOM (Out Of Memory) events during early boot.
- A per-node CLI bring-up does not enforce strict OSβlevel startup ordering or readiness guarantees.
- As a result, dependent nodes may start before prerequisite nodes are fully initialized, leading to race conditions and unstable boot behavior in production environments.
CLI/launch-based bring-up focuses on process creation and parameter loading.
- After startup, lifecycle state transitions are not centrally coordinated.
- In systems managing many nodes, this results in fragmented lifecycle handling, uncoordinated state transitions, and the absence of a single authoritative component responsible for global system stateβan important weakness on resourceβconstrained platforms.
ROS 2 lifecycle states are intentionally minimal and lowβlevel, while production robots operate in missionβlevel modes such as standby or navigation.
- Without centralized orchestration, mapping missionβlevel behavior to coordinated lifecycle transitions across multiple nodes becomes errorβprone and difficult to validate, especially on lowβend SoCs where deterministic behavior is critical.
π‘ Summary On lowβend embedded platforms, interpreter-based bring-up introduces overhead and nonβdeterminism during boot and state transitions. In our evaluated setup, these limitations were not sufficiently mitigated through configuration alone. This motivated a native, dependency-ordered lifecycle orchestration approach.
The Lifecycle Manager is a multi-threaded C++ ROS 2 node that acts as a centralized lifecycle coordinator. It separates communication from orchestration across distinct execution contexts:
- Spin Thread: Runs a ROS 2 executor dedicated to communications and service callbacks.
- Orchestration Thread: Manages the core orchestration loop, including package spawning and the
processQueue()mechanism. This non-blocking queue serializes state-transition requests so they are processed in a repeatable order. - Bounded Worker Concurrency: Concurrent process spawning is dispatched to a pool capped at the system hardware concurrency, bounding the number of simultaneous node transitions.
It operates as a ROS 2-native orchestration component within a standard ROS 2 system, while avoiding dependence on the Python front-end path in the evaluated deployment mode.
The system is structured around five core modules:
flowchart TD
App["<b>APPLICATION LAYER</b><br/>(Requests device state changes)"]
YAML["<b>Configuration YAML</b><br/>(Source of Truth)"]
subgraph Manager ["LIFECYCLE MANAGER (Native C++)"]
direction TB
SL["Service Layer<br/>(Queue Manager)"]
Conf["Configuration<br/>(YAML Parser)"]
Core["Orchestration Core<br/>(Coordinates Launcher & Engine)"]
NL["Node Launcher<br/>(fork/exec/SIGCHLD)"]
TE["Transition Engine<br/>(State Machine Logic)"]
LC["Lifecycle Client<br/>(Service Interface)"]
SL --> Core
Conf --> Core
Core --> NL
Core --> TE
TE --> LC
end
subgraph Nodes ["MANAGED ROS 2 NODES"]
direction LR
NA["Node A"]
NB["Node B"]
NN["Node N"]
end
App -- "ROS 2 Service" --> SL
YAML --> Conf
NL -- "Native Execution" --> Nodes
LC -- "Get/ChangeState" --> Nodes
style App fill:#f8f9fa,stroke:#343a40,stroke-width:2px
style Manager fill:#f8f9fa,stroke:#343a40,stroke-width:2px
style Nodes fill:#f8f9fa,stroke:#343a40,stroke-width:2px
style SL fill:#fff,stroke:#343a40
style Conf fill:#fff,stroke:#343a40
style Core fill:#fff,stroke:#343a40
style TE fill:#fff,stroke:#343a40
style NL fill:#fff,stroke:#343a40
style LC fill:#fff,stroke:#343a40
style YAML fill:#fff,stroke:#343a40,stroke-width:2px
style NA fill:#fff,stroke:#343a40
style NB fill:#fff,stroke:#343a40
style NN fill:#fff,stroke:#343a40
- Service Layer β Exposes the
/lifecycle_transition_deviceROS 2 service and manages a thread-safe work queue. It supports string-based state transition requests (e.g., "NORMAL", "SLEEP") for improved human readability and CLI usability. - Config Engine β A centralized YAML parser that serves as the single source of truth. It supports name-based device state definitions, removing the need for fragile numeric indexing.
- Transition Engine β The central orchestrator that coordinates lifecycle state machines across multiple packages. It serializes state transition requests to prevent race conditions during concurrent updates.
- Node Launcher β Handles native process spawning via POSIX
fork/execand detects child-process termination viaSIGCHLD(asynchronously reaping exited children to prevent zombie accumulation). It manages dynamic path resolution (ament_index_cpp) and per-process log redirection.- Benchmarking Mode: Includes a specialized execution path for legacy per-node CLI/script launches (via
std::system()), enabling A/B testing against the native path.
- Benchmarking Mode: Includes a specialized execution path for legacy per-node CLI/script launches (via
- Lifecycle Client β Interfaces with managed nodes using standard ROS 2
GetStateandChangeStateservices, with per-operation timeouts and bounded retries.
Architecture Independence β By relying exclusively on POSIX standard system calls (fork, execvp, sigaction) and standard ROS 2 APIs, the Manager is designed to be portable across common Linux targets such as ARM64 and x86_64, subject to standard ROS 2 and platform integration constraints.
The following diagram illustrates the shift from an interpreter-based, per-node CLI bring-up to a lightweight C++ native orchestrator.
flowchart TD
subgraph Old [β Before: sequential per-node ros2 run CLI]
P_N1["ros2 run (Python CLI wrapper) β Node A"]
P_N2["ros2 run (Python CLI wrapper) β Node B"]
P_N3["ros2 run (Python CLI wrapper) β Node C"]
P_N1 --> P_N2 --> P_N3
P_Note["π¨ one CLI wrapper per node Β· sequential Β· high transient RAM"] -.-> P_N1
end
subgraph New [β
After: C++ Lifecycle Manager]
C_Manager{"C++ LifecycleManager (Native OS Process, Low RAM)"}
C_Manager == "Concurrent fork() & execvp()" ==> C_N1[Node Process A]
C_Manager == "Concurrent fork() & execvp()" ==> C_N2[Node Process B]
C_Manager == "Concurrent fork() & execvp()" ==> C_N3[Node Process C]
C_N1 -. "Lifecycle State: Active" .-> C_Manager
C_N2 -. "Lifecycle State: Active" .-> C_Manager
C_N3 -. "Lifecycle State: Active" .-> C_Manager
C_Note["π‘ bounded concurrency & dependency ordering"] -.-> C_Manager
end
style P_N1 fill:#ffebee,stroke:#c62828,stroke-width:2px
style P_N2 fill:#ffcdd2,stroke:#c62828,stroke-width:2px
style P_N3 fill:#ffcdd2,stroke:#c62828,stroke-width:2px
style C_Manager fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
All orchestration behavior is defined declaratively in a single YAML file:
# Example: Configuration with device state names
LIFECYCLE_MANAGER_CONFIG:
DEVICE_STATE_NAMES: ["NORMAL", "SLEEP", "POWERSAVE"]
USE_LAUNCH_SCRIPT: false
NODE_TRANSITION_STRATEGY: "parallel"
PACKAGE_slam_package:
PACKAGE_ENABLE: true
NODE_slam_node:
EXECUTABLE: slam_node
ARGUMENT: ["--param_a", "value_a"]
DEPENDENCY: ["lidar_package,lidar_node"]
DEVICE_INIT: ACTIVE
DEVICE_STATE_NORMAL: ACTIVE
DEVICE_STATE_SLEEP: INACTIVE
DEVICE_STATE_POWERSAVE: FINALIZEDKey configuration capabilities:
- Launch mode (
USE_LAUNCH_SCRIPT): Selects between native binary spawning (fork/exec) or per-node CLI/script execution (benchmarking) - Transition strategy (
NODE_TRANSITION_STRATEGY): Parallel (multi-threaded) or sequential execution - Executable arguments (
ARGUMENT): Command-line arguments for the native process - Dependency declaration (
DEPENDENCY): Inter-node startup ordering and readiness polling - Initialization target (
DEVICE_INIT): Specific lifecycle state for initial boot - Device state mapping (
DEVICE_STATE_<NAME>): Per-node lifecycle targets for each robot mission state
The ROS 2 lifecycle standard does not allow direct transitions between certain primary states. The Lifecycle Manager resolves all intermediate steps automatically:
UNCONFIGURED β ACTIVE : Configure β Activate (two-step)
ACTIVE β UNCONFIGURED : Deactivate β Cleanup (two-step)
UNCONFIGURED β INACTIVE : Configure
INACTIVE β ACTIVE : Activate
ACTIVE β INACTIVE : Deactivate
Any state β FINALIZED : Appropriate shutdown transition
The application layer simply declares a target state; the Lifecycle Manager resolves and executes all intermediate transitions transparently, each wrapped with configurable retry and timeout policies.
To bridge the semantic gap between robot missions and low-level lifecycle states, the Lifecycle Manager introduces a "Device State" abstraction.
| Node \ Device State | NORMAL | SLEEP | POWERSAVE |
|---|---|---|---|
slam_node |
ACTIVE | INACTIVE | FINALIZED |
lidar_node |
ACTIVE | INACTIVE | FINALIZED |
navigation_node |
ACTIVE | INACTIVE | FINALIZED |
motor_node |
ACTIVE | INACTIVE | FINALIZED |
camera_node |
ACTIVE | ACTIVE | INACTIVE |
diagnostic_node |
ACTIVE | ACTIVE | ACTIVE |
To switch the robot from "NORMAL" to "SLEEP", just call:
ros2 service call /lifecycle_transition_device lifecycle_manager_msgs/srv/TransitionDevice "{request: 'SLEEP'}"This single call automatically transitions each node to its matching target state β navigation and motor stop, while camera and diagnostics stay running. This design decouples mission logic from lifecycle management.
The Lifecycle Manager follows a rigorous, dependency-ordered sequence to ensure all nodes are prepared and synchronized.
By identifying independent node groups at runtime from YAML dependency declarations (e.g., DEPENDENCY: ["pkg_name,node_name"]), the system initializes multiple packages concurrently. Rather than scaling with the total node count N, completion time is bounded by both the dependency critical path and the available parallelism β approximately max(L, W/c), where L is the longest dependency chain, W is the total node-initialization work, and c is the concurrency cap (hardware concurrency). In other words, boot time is governed by dependency depth and the concurrency cap rather than by node count alone.
flowchart TD
Start("[ SYSTEM STARTUP ]") --> YAML
YAML["<b>YAML Configuration</b><br/>(Source of Truth)"] --> Exec
Exec["<b>Execution Strategy</b><br/>(Parallel vs Sequential)"] --> Path
subgraph Path ["[ ORCHESTRATOR PATH ] - (Parallel / Seq Loop)"]
direction TB
Check["<b>Check if Enabled</b><br/>(f_packageLaunch flag)"] --> Launch
Launch["<b>Package Launch</b><br/>(Native fork/exec)"] --> Dep
Dep["<b>Dependency & State Check</b><br/>(GetState + Dep Polling)"] --> Trans
Trans["<b>State Transition</b><br/>(ChangeState Client)"]
end
Path --> Ready("[ SYSTEM READY ]")
style Path fill:#f8f9fa,stroke:#343a40,stroke-width:2px,stroke-dasharray: 5 5
style Start fill:#e7f3ff,stroke:#007bff,stroke-width:2px
style Ready fill:#d4edda,stroke:#28a745,stroke-width:2px
style YAML fill:#fff,stroke:#343a40,stroke-width:1px
style Exec fill:#fff,stroke:#343a40,stroke-width:1px
style Check fill:#fff,stroke:#343a40,stroke-width:1px
style Launch fill:#fff,stroke:#343a40,stroke-width:1px
style Dep fill:#fff,stroke:#343a40,stroke-width:1px
style Trans fill:#fff,stroke:#343a40,stroke-width:1px
- Load YAML Configuration β Reads all package/node definitions, dependencies, and device-state mappings from a single YAML file.
- Select Execution Strategy β Determines parallel (multi-threaded) or sequential mode per YAML configuration.
- Native Process Spawning β Each package binary is launched via
fork/execwith executable paths resolved dynamically byament_index_cpp. Per-process log files are created with timestamps. - Dependency Wait β For each node, polls the managed-state of declared dependency nodes until they reach
ACTIVE(30-second timeout). - Lifecycle State Polling β Calls
GetStateservice with retry until the node responds, confirming it is alive and ready for state transitions. - Initial State Transition β Applies the
DEVICE_INITtarget lifecycle state via the multi-step state machine (e.g., triggeringConfigure β Activatefor anACTIVEtarget).
After all nodes reach their initial states, the system enters the operational phase β waiting for Device State requests via /lifecycle_transition_device and applying per-node lifecycle targets through the TransitionEngine.
This section validates the impact of the Deterministic Lifecycle Manager using measurements collected on a pre-production commercial cleaning robot platform publicly showcased at IFA 2025.
For evaluation purposes, the product's software stack was ported to ROS 2, system interfaces were redesigned, and the C++ DLM was deployed directly on the target hardware. The goal was realistic system-level evaluation under severe production constraints (memory pressure and strict boot-time requirements). All measurements were performed on identical hardware using the same ROS 2 node set.
| ITEM | Specification |
|---|---|
| HW Platform (AP) | LG DQ1 (Cortex-A53Γ4 @ 1 GHz) |
| RAM | 1 GB |
| eMMC | 4 GB |
| ROS 2 Distribution | ROS 2 Humble |
| Managed ROS 2 Nodes | 14 lifecycle nodes + 1 non-lifecycle transform process |
| OS | Yocto-based Linux |
| Yocto Version | Kirkstone |
Three configurations were evaluated. All bring up the identical 14 nodes with the same lifecycle targets and parameters on the same hardware and software stack; only the orchestration mechanism differs:
- (1) Baseline β sequential per-node
ros2 run(CLI): Each node is started through a separateros2 runinvocation (which loads the ROS 2 Python command-line front end), one after another (USE_LAUNCH_SCRIPT: true,NODE_TRANSITION_STRATEGY: sequential). - (2) Native sequential:
Nodes are started as native
fork/execprocesses, sequentially, without a Python front end. - (3) Proposed β C++ DLM (native + parallel):
Native
fork/execcombined with dependency-aware parallel activation.
This staging isolates the contribution of removing the CLI/interpreter front end from that of dependency-aware scheduling.
No other system components or ROS 2 node implementations were changed between the configurations. Because our baseline is the sequential per-node CLI bring-up, we do not claim results relative to the parallel-capable
ros2 launchsystem.
| Metric | ros2 run CLI (seq.) |
Native exec (seq.) | C++ DLM (parallel) |
|---|---|---|---|
| Boot time (s) | 60.04 Β± 0.64 | 34.62 Β± 0.51 | 10.74 Β± 0.77 |
| Avg RAM during boot (MiB) | 247.34 Β± 1.32 | 140.16 Β± 1.17 | 189.58 Β± 8.44 |
| Stable RAM, 90 s after boot (MiB) | 464.09 Β± 1.05 | 273.37 Β± 1.88 | 275.07 Β± 1.49 |
Values are averaged over ten repeated runs (Mean Β± Std. Dev.). Memory is whole-system used memory (
free), converted from KiB to MiB.
- Boot completion time was reduced from 60.04 s to 10.74 s, a β82.1% reduction (baseline β DLM).
- Stable memory usage after startup decreased from 464.09 MiB to 275.07 MiB, a β40.7% reduction.
- Ablation: removing the CLI/interpreter front end alone (config 1 β 2) cut boot time to 34.62 s (β42.3%) and accounts for essentially all of the memory saving (464.09 β 273.37 MiB, β41.1%); adding dependency-aware parallel activation (config 2 β 3) cut boot time to 10.74 s while leaving stable memory essentially unchanged (+0.6%). In absolute terms the two effects contribute comparably to the boot-time reduction (25.4 s and 23.9 s).
Note on variance: the parallel configuration shows higher relative run-to-run timing variance than the sequential baselines (coefficient of variation β 7% vs β 1%). What is repeatable across runs is the outcome: activation respects the dependency partial order and the same set of nodes reaches ACTIVE on every run.
Legacy Sequential ros2 run CLI |
Proposed Parallel Native C++ DLM |
|---|---|
![]() |
![]() |
Left: Legacy Sequential per-node ros2 run CLI | Right: Proposed Parallel Native C++ DLM
-
β±οΈ Boot Time (baseline β DLM)
ros2 runCLI (seq.):60.04 s- C++ DLM (parallel):
10.74 s - Improvement: β 82.1%
(60.04 s β 10.74 s)
-
πΎ Stable RAM (baseline β DLM)
ros2 runCLI (seq.):464.09 MiB- C++ DLM (parallel):
275.07 MiB - Improvement: β 40.7%
(464.09 MiB β 275.07 MiB)
All experiments were conducted on identical hardware and software configurations.
- Same nodes and execution graph
- Same workload
- Only the orchestration mechanism was changed (sequential per-node
ros2 runCLI β nativefork/execβ native + dependency-aware parallel)
Detailed configuration and setup are available in this repository.
A side-by-side boot comparison (Legacy per-node ros2 run CLI vs. Native C++ DLM) is shown in the Performance Charts section above.
π‘ Benchmark Reproducibility: The comparison was performed within the same LifecycleManager instance by toggling the
USE_LAUNCH_SCRIPTandNODE_TRANSITION_STRATEGYflags, comparing identical node sets launched via per-noderos2 runCLI vs. native C++ spawning (sequential and parallel). This supports the interpretation that a large portion of the improvement is attributable to orchestration-path differences.
Conclusion: By removing the Python front end from the evaluated runtime path, startup memory pressure was reduced, and OOM events observed in the baseline configuration were not reproduced in the evaluated DLM configuration. These improvements were achieved without modifying the ROS 2 nodes themselves and resulted in a repeatable boot sequence (in outcome) on the target hardware.
- Architecture:
x86_64(Development/PC) andARM64/AArch64(Embedded Target) - OS: Ubuntu 22.04 (Jammy) or later / Linux (Yocto-based)
- ROS 2: Developed and evaluated on Humble; designed to be portable to Iron and Jazzy
- Compiler: C++17 or higher (Required for
<filesystem>and modern C++ features) - Dependencies:
rclcpp,lifecycle_msgs,lifecycle_manager_msgs,ament_index_cpp
The full implementation is not yet publicly available in this repository.
Build and run instructions will be added once the source release is completed.
Currently available in this repository:
- architecture documentation
- YAML configuration examples
- measurement methodology
- empirical validation results
This project is licensed under the Apache License 2.0, allowing internal commercial use as well as future open-source contributions.
As described at the top of this document, this repository currently provides architectural documentation, configuration examples, and empirical validation results that describe the execution model and lifecycle behavior of the system.
The complete C++ source code is planned to be released under the same license following completion of internal compliance procedures.

