Skip to content

RMW_EVENT_MESSAGE_LOST is plumbed but never raised #292

Description

@YuanYuYuan

RMW_EVENT_MESSAGE_LOST is fully plumbed in hiroz and never raised. A subscriber that asks for it gets a status that is always zero, so in-transit message loss is silently invisible to any ROS 2 application running on hiroz.

Found while comparing hiroz's subscriber queue behaviour against rmw_zenoh_cpp for #250.

What exists

piece state
ZenohEventType::MessageLost ✅ defined (crates/hiroz/src/event.rs)
rmw_event_type 3 → MessageLost mapping ✅ crates/rmw-zenoh-rs/src/rmw.rs
rmw_message_lost_status_t fill-in on rmw_take_event ✅ rmw.rs
the sequence number needed to detect loss ✅ Attachment::sequence_number is serialised on every sample
anything that calls update_event_status(MessageLost, ..) ❌ nothing, outside #[cfg(test)]

So the event is reachable from the public rmw API, always reports zero, and the data required to populate it is already on the wire.

What upstream does

rmw_zenoh_cpp's SubscriptionData::add_new_message compares each arriving message's sequence number against the last one seen from that publisher GID, and raises the event on a gap:

// Check for messages lost if the new sequence number is not monotonically increasing.
const int64_t seq_increment = std::abs(msg->attachment.sequence_number() - last_known_pub_it->second);
if (seq_increment > 1) {
    int32_t num_msg_lost = /* clamped seq_increment - 1 */;
    events_mgr_->update_event_status(ZENOH_EVENT_MESSAGE_LOST, std::move(num_msg_lost));
}
last_known_published_msg_[gid_hash] = msg->attachment.sequence_number();

It keeps a per-publisher last_known_published_msg_ map for exactly this.

Scope — what this is not

Queue-depth drops are out of scope, in both implementations. A subscriber dropping its own oldest queued message pops from the front; the sequence-gap check compares the arriving message against the previous arrival, so a depth-drop cannot produce a gap. Upstream logs those at debug level and does not raise MESSAGE_LOST for them either. hiroz logs them with an escalating warn!.

This issue is about loss in transit — samples that never arrived — which is what the ROS event means and what neither implementation infers from its own queue.

Acceptance

  • A per-publisher-GID last-sequence map on the subscriber
  • On a gap, update_shared_event_status(.., MessageLost, gap - 1) — via the shared entry point, not under a lock
  • Sequence numbers reset/absent are handled without a false positive on the first sample from a publisher
  • A test that drops a sample in transit and observes a non-zero rmw_message_lost_status_t, demonstrated to report zero without the fix

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions