Repository navigation
Update main branch to 3.7.6 - #14
Merged
Merged
Conversation
Improves an existing error log Log the remote provider's ip:port when a local offer is rejected as already offered remotely. while at it, adds a fake socket test to replicate the scenario.
Fix integer overflow in ASSIGN_CLIENT payload deserializer The ASSIGN_CLIENT payload deserializer used the addition form pos + name_len > _size, which wraps in uint32_t. A crafted name_len = 0xFFFFFFFF bypasses the guard and yields a ~4 GB string_view over a tiny buffer → heap over-read (crash / DoS). The has_address branch had the same issue.
We cannot use __android_log_print if we're not properly escaping the formatting.. typical (and very bad) printf like vulnerability! This was reported in COVESA, see: COVESA#1054 Use __android_log_write instead, which does no formatting, and is anyhow a much better fit
Append the registered application name + uid to routing-manager error logs (impl + stub) that print a client id Client-id-only errors made it hard to tell which application caused a routing failure. Each affected VSOMEIP_ERROR/_P log in the impl and stub now shows 0x1234 (app_name, uid=1000) via a get_client_info() helper. The uid is dropped for TCP clients and unknown mappings; remote-client and static send_local logs are unchanged, as the info can't be resolved there.
Promote vsomeip security logs to error level. This will enable the newly promoted logs to be easily searched in Brian ingestion system and to be contemplated on vsomeip post test.
It's outright dumb not to do it unconditionally, it is the same no matter whether uds/tcp/mixed-mode is used
All of them, hopefully for the last time
Improves error message for event registration conflict The error message about registering events after offering a service printed no identifiers, so it was impossible to tell which service/instance triggered the wrong behavior from the logs. This PR adds the affected service+instance to the message.
Same as 22bbb48, makes no sense to do this check only with security turned on
Add AI Tools to .gitignore
The Android build fails while compiling implementation/utility/src/service_instance_map.cpp:
external/bmw/boost/boost/container_hash/detail/hash_float.hpp:251:25:
error: no member named 'equal_to' in namespace 'std'
return std::equal_to<T>()(v, 0);
Root cause is an include-order bug:
service_instance_map.hpp includes <boost/functional/hash.hpp> before <unordered_map>.
Until now the header was only included by other translation units that had already pulled in <functional>,
so std::equal_to was defined by the time the Boost header was parsed. The bug was masked.
Commit 22b77f4 introduced a new service_instance_map.cpp that includes
service_instance_map.hpp as its first include, with nothing pulling in <functional> beforehand.
The protocol hard-assumes that specific messages come from the router. Ensure it always does so, regardless of security config Promote some logs to errors
Remove useless and fragile socat invocation. For some reason lost in the fog in history, the ssh connection to the network test slaves uses a socat command that does not seem to serve any purpose. After introducing test case isolation, that could actually lead to spurious breakdowns of the VXLAN setup, ultimately failing test cases due to the source port being fixed and back-to-back slave connections potentially running into a TIME_WAIT race. Remove the socat invocation, which should also remove the test framework instability we're seeing.
Add full-logging threshold to trace connector to reduce performance impact of large messages. Problem The Trace Connector logs every SOME/IP message with its full payload by default, flooding logs with volumous IPC (~100 kB camera frames, ~300 kB Android hexstreams). No way to cap payload logging by size, nor to exempt a single large message from such a cap. Changes full_logging_threshold — new tracing.full_logging_threshold config (bytes). Under the default filter, messages larger than this (16-byte header + payload) trace header-only; smaller ones trace in full. Default 2048 (from a ~303 k-message histogram:, truncates only the multi-kB tail). 0 disables it; sub-16 values clamp to 16; parser rejects negatives/garbage/overflow. trace_result_e — replaced matches()'s ambiguous pair<bool,bool> with a DROP_FILTER/DEFAULT enum plus a pure should_log_full() policy, cleanly splitting "positive filter → always full" from "default → subject to threshold." full-payload filter — mirror of header-only: forces matching messages bypass the full-logging threshold. Lets the size cap stay active for all traffic while exempting one large service/method. Tests and docs updated accordingly.
Test was hammering the send, which could overload the receive buffer of the socket, leading to some lost messages. The test was expected to receive the end_message, that never came Increase delay a bit.
SomeIP data from fake udp sockets is no longer being printed. This PR adds the logging of parsed someip message on gate.
Split the pending event registration set into producer vs. consumer sets routing_manager_client kept both provider-side and consumer-side pending event registrations in a single pending_event_registrations_ vector guarded by its own dedicated pending_event_registrations_mutex_. Every access already ran under the matching role lock (provider_mutex_ for offered events, consumer_mutex_ for consumed ones), so the extra mutex was redundant. This splits the vector by role — pending_provided_event_registrations_ under provider_mutex_ and pending_consumed_event_registrations_ under consumer_mutex_.
It is ALWAYS BAD for a build to touch the source directory. Every generated file - no matter whether it's a compiled binary, intermediate source files, or configuration, should go into the build directory Fix the few last bad examples, one of them was even hardcoding /build.. also a very bad practice While at it, very minor cleanups; use PROJECT_SOURCE_DIR/PROJECT_BINARY_DIR in top-level CMakeLists.txt
Removes some unused declarations Removes some unused declarations at routing_manager_client Notice it while looking into rmc code.
On a client-id reuse, only tear down the old client's provider connection if that endpoint is still bound to the old client, and only drop the shared sec_client mapping once the client has no endpoint left in either role. Commit ec47d64 decoupled consumer- and provider-side error handling, but in on_routing_info (a consumer-side signal) a new client id at a known address:port still tore down the old client's provider role unconditionally. Since client ids are recycled across peers, after an STR resume a new app can already hold a healthy provider connection there — and the teardown would close the new client's link instead of the departed one's. Additionally, the shared client→sec_client mapping was cleaned only in remove_local_provider. As it is shared by both roles, tearing down the provider dropped it even when a healthy consumer connection to the same client remained, leaving that good consumer without its mapping.
Enables the receive-side offer check and adds regression tests to cover this change.
Extends the protocol tests to cover additional client/router/server message sequences: Service offer/re-offer ordering relative to event registration: Service release; Event unsubscription Stopping individual events.
Remove service_requests_ map Commit d61bf2e made routing_manager_impl::requested_services_ be populated unconditionally (also for locally-offered services). As a result it now holds essentially the same data as routing_manager_stub::service_requests_. Given this, we can drop one of the maps.
Escalate client endpoint reconnect log to ERROR after a certain time. There were reported cases where a client endpoint can't reach its remote peer, it would retry forever, only logging a WARNING on every failed attempt. This is ok in some cases, but should be signaled when it takes too much time. This change makes the existing per-attempt "Restarting socket due to ..." log escalate from WARNING to ERROR once the outage has lasted continuously for more than VSOMEIP_RECONNECT_TIMEOUT.
This change fixes a segmentation fault in both overloads of the add_remote_service_info method.
Warn users when provided/consumed events or handlers are registered too late or registered more than once. These logs were made to catch ordering mistakes caused by API misuse, not races! This PR adds logs for the following cases: Late event-registration errors Late handler-registration errors Duplicate-registration The late/duplicate checks are best-effort and API-level only. The routing manager is queried outside the application's handler mutexes to avoid lock-ordering issues, so a concurrent offer/request/subscribe on another thread can still be missed.
Remove the fixed client ID configuration from the network-tests Commit 4062d4c added a log-only check in on_message that flags an incoming message whose client id belongs to the receiving host's own range (which mis-routes the response to a local client instead of the remote host). It was gated to skip hosts on the default diagnosis address — a test workaround — so it never ran exactly where client-id collisions are most likely. This removes that guard so the check runs on every host. It is safe: is_valid_client_id has a single caller and only emits a log line — no message-delivery behaviour changes. The accompanying test changes make the network-tests realistic (distinct diagnosis per boardnet side + dynamically-assigned client ids) so the warning is meaningful and not spurious. All fixed client IDs have been removed except for the following: client_id_test_same_client_ids* - The purpose of the tests is to ensure that the client IDs are the same on both the master and the slave. client_id_utility* - The purpose of the test is to have client IDs both within and outside the range defined by the diagnosis. configuration_test- Check whether the specified client ID matches the one configured in the application e2e_profile_04/07* - Use the client ID to check the full CRC subscribe_notify_test_same_client_ids* - the purpose of the tests is to ensure that the client IDs are the same on both the master and the slave. In addition, some tests have been updated so that they no longer rely on the fixed client ID but instead check the dynamic client ID.
Fix files not adhering to current clang-format expectations. Reformat codebase with current version of clang-format version and ruleset, in order to avoid unnecessary formatting in unrelated patchsets.
These changes allow to build all the tests with the CI clang compiler.
Make the connection-drop check reliable by detecting the drop event instead of the momentary
disconnected state.
The wait_for_connection_drop returned whether the socket pair is disconnected at the sampling instant.
After a partial-read drop the consumer reconnects within ~0.4 ms,
16:50:58.529619 [fake-socket] calling cancel on: {fd: 11, type: tcp, role: server, app: server}
16:50:58.530039 [fake-socket] added: client_to_server to the known connections
so a waiter woken under load re-evaluated only after the connection had healed
and timed-out — an intermittent failure of test_partial_read_leads_to_connection_drop
The predicate now also succeeds when socket_count_ has grown since the wait began (a
drop-and-reconnect occurred), so the result no longer depends on catching the transient window.
Make the logger more robust against log-after-exit races. Since we still have threads that may log after main() has exited, we observe crashes that happen when the logger tries to use libdlt, mainly due to data structures like the log context already having been destroyed. The safest, albeit slightly ugly way to deal with this is to make the logger an "immortal singleton": It is simply never deallocated, thus its memory remains valid throughout the lifetime of the process The system will reclaim the memory anyway when tearing down the process. In order for leak sanitizer not to treat it as a memory leak, the allocated memory is referenced by a static pointer, thus it is referenced for the duration of the process, and never marked as lost. Ensure deregistration of the log context through a static guard that will trigger this when it gets destroyed during thread cleanup. This will set the context's log_level_ptr attribute to null, and any attempt by another thread to log afterwards will trigger a nullptr check that leads libdlt to safely discard the message early on, without the risk of triggering undefined behavior. Be a good boyscout too, and add a set of unit tests for the logger.
uint/size_t alone is fine, and anyhow shorter. Note as well that: interface/ is changed, but it was already a mix of both technically true that does not have to define these global types, but in practice it does (and see 1.!) This commit does nothing else except those changes + applying formatter
API is somewhat unfortunate, passes an untyped uint16_t to the user callback (which.. has to match the CommonAPI CallStatus definition..!) While we do not do a major API break, we can at least make the codebase better; define an enum class, use it internally, comment where appropriate Minor cleanups (mostly removing useless helpers) while at it
A lot of these logs are definitely errors: unparseable files, nonsense values, duplicate-entry-for-same-application, ..
Removes warning log from routing which caused flood of logging The intent of the logging was to show the messages that were received remotely and not forward by the router and therefore dropped. Unfortunately, this causes big flood of messages on target. Hopefully, we'll work around with data provided from statistics logs
Install ccache in the build image so that compiler invocations are cached across builds. This gives an overall build time improvement of around 45-50%. Also install git and nginx, which are required by the build and by the report publishing steps.
Fix security_test in debug build Due to the combination of delays caused by valgrind-memcheck and debug build, the test fails because it does not receive the availability response within 10 seconds. Upon analyzing one of the logs, the service was offered 20 ms after the client failed to receive the availability. To resolve the issue, scaled_timeout was applied, similar to other tests with the same issue.
They were not formatted before, somehow, and the CI is now enforcing it, somehow
Gate the late registration log on consumer's register_message_handler_ext CAPI currently does request_service before register_message_handler, resulting in a spam of late registration logs on our ECUs. This PR gates the log so that only providers use it.
Add more timeout scaling to fake socket tests. Under sanitizers, we stilll see sporadic failures in fake socket tests caused by exceeding tight timeouts. Add timeout scaling to the major culprits to make them more robust under heavy load in CI.
It was never implemented, it has a default that makes no sense, and it was blindly copy-pasted all over the place, including example configuration files Therefore drop it
Disable the initial wait phase by default. The 0-to-3000ms default is not sane, and generally from a POSIX PoV it does not make sense to have it (there are enough delays in vsomeipd starting up, other applications starting up..) Reduce the FIND/OFFER debounce time to 20ms. The 500ms default is also far from sane when the cycle time is 1000ms. It makes sense to have a debounce time (we DO NOT want to send a multicast FIND/OFFER per service!), but a small window is sufficient, as applications, once they startup, offer/request everything in a small window of time as well Especially the FIND/OFFER debounce times have been the source of delays in the past
Add a connection timeout for connections with the router Add a connection timeout for clients with the router. After this timeout is reached the application logs an error message and continues to try to connect with the router
Under stress, the e2e_profile_0x_test_external occasionally fails because it receives a duplicate initial notification and the test is expecting different/incremental payloads. This happens because, when the initial event is lost or experiences a very long delay in being sent, it is detected as lost.
Fix usei unit tests Instead of posting the messages to an actual socket, call the mocked on_message_received_unlocked directly so no messages are dropped
Normal registration log, downgrade it to VSOMEIP_INFO_P
It is only used to remove boardnet subscriptions, but anyhow the router already removes those via subscription expiry, so it's unnecessary Besides the useless code, it is harmful - it guarantees a thundering herd on suspend, because very libvsomeip client is woken up
Fix duplicate registration warn to match major The warning log for a availibility handler duplicate registration depends only on the service/instance, even though the handler also depends on the major/minor version. Therefore, if two handlers are registered for the same service/instance but with different major versions, the lib will report a warning for a duplicate registration even though they are different handlers. Now, it only logs the handler if the major version also matches.
Adopts the Android 17 bring-up changes Changes Remove Android.mk Add host_supported to libvsomeip3. Add three alias stubs — libvsomeip3-cfg, libvsomeip3-e2e, libvsomeip3-sd
The test had a mechanism to re-send the offer if the connection wasn't established in time, however, it would do so with the same sessionID, causing a reboot detection and switching ports
Drop message if we are not interested in the service before doing SOME/IP-TP, not after While at it, springle some ANY_* defs in the codebase Found by Mayur Agnihotri, reported in private
Make it so configs are reused and not needlessly loaded every time an app starts on a single process. Keeps behavior if a config is defined by VSOMEIP_CONFIGURATION_<name>. Minor change to configuration_impl to avoid unnecessary copies. First app now loads all of the config files, being the router or not Policies are now unique per app and loaded by configuration_impl but copied for each application. Added tests for configuration behavior.
szb640
approved these changes
Sep 21, 2026
Stek210588
added a commit
that referenced
this pull request
Sep 28, 2026
This reverts commit 9cc8753.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Update the main branch of the fork to 3.7.6 to keep it up-to-date to the base repository