Skip to content

Update main branch to 3.7.6 - #14

Merged
Stek210588 merged 82 commits into
mainfrom
main_3.7.6
Sep 21, 2026
Merged

Stek210588 merged 82 commits into
mainfrom
main_3.7.6

Conversation

@Stek210588

Copy link
Copy Markdown
Collaborator

Update the main branch of the fork to 3.7.6 to keep it up-to-date to the base repository

Jorge Saraiva and others added 30 commits September 17, 2026 15:18
Improves an existing error log
Log the remote provider's ip:port when a local offer is rejected as already offered remotely.
while at it,
adds a fake socket test to replicate the scenario.
Fix integer overflow in ASSIGN_CLIENT payload deserializer
The ASSIGN_CLIENT payload deserializer used the addition form
pos + name_len > _size, which wraps in uint32_t.
A crafted name_len = 0xFFFFFFFF bypasses the guard and yields
a ~4 GB string_view over a tiny buffer → heap over-read (crash / DoS).
The has_address branch had the same issue.
We cannot use __android_log_print if we're not properly escaping the
formatting.. typical (and very bad) printf like vulnerability!
This was reported in COVESA, see:
COVESA#1054
Use __android_log_write instead, which does no formatting, and is
anyhow a much better fit
Append the registered application name + uid to routing-manager error logs (impl + stub)
that print a client id
Client-id-only errors made it hard to tell which application caused a routing failure. Each affected
VSOMEIP_ERROR/_P log in the impl and stub now shows 0x1234 (app_name, uid=1000) via a
get_client_info() helper. The uid is dropped for TCP clients and unknown mappings;
remote-client and static send_local logs are unchanged, as the info can't be resolved there.
Promote vsomeip security logs to error level.
This will enable the newly promoted logs to be easily searched in Brian ingestion system
and to be contemplated on vsomeip post test.
It's outright dumb not to do it unconditionally, it is the same no
matter whether uds/tcp/mixed-mode is used
All of them, hopefully for the last time
Improves error message for event registration conflict
The error message about registering events after offering a service printed no identifiers,
so it was impossible to tell which service/instance triggered the wrong behavior from the logs.
This PR adds the affected service+instance to the message.
Same as 22bbb48, makes no sense to do this check only with security turned on
Add AI Tools to .gitignore
The Android build fails while compiling implementation/utility/src/service_instance_map.cpp:
external/bmw/boost/boost/container_hash/detail/hash_float.hpp:251:25:
error: no member named 'equal_to' in namespace 'std'
            return std::equal_to<T>()(v, 0);

Root cause is an include-order bug:

service_instance_map.hpp includes <boost/functional/hash.hpp> before <unordered_map>.
Until now the header was only included by other translation units that had already pulled in <functional>,
so std::equal_to was defined by the time the Boost header was parsed. The bug was masked.
Commit 22b77f4 introduced a new service_instance_map.cpp that includes
service_instance_map.hpp as its first include, with nothing pulling in <functional> beforehand.
The protocol hard-assumes that specific messages come from the router. Ensure
it always does so, regardless of security config
Promote some logs to errors
Remove useless and fragile socat invocation.
For some reason lost in the fog in history, the ssh connection to the
network test slaves uses a socat command that does not seem to serve
any purpose. After introducing test case isolation, that could actually
lead to spurious breakdowns of the VXLAN setup, ultimately failing
test cases due to the source port being fixed and back-to-back
slave connections potentially running into a TIME_WAIT race.
Remove the socat invocation, which should also remove the test
framework instability we're seeing.
Add full-logging threshold to trace connector to reduce performance
impact of large messages.
Problem
The Trace Connector logs every SOME/IP message with its full payload by
default, flooding logs with volumous IPC (~100 kB camera frames, ~300 kB
Android hexstreams). No way to cap payload logging by size, nor to exempt a
single large message from such a cap.
Changes
full_logging_threshold — new tracing.full_logging_threshold config (bytes).
Under the default filter, messages larger than this (16-byte header + payload)
trace header-only; smaller ones trace in full. Default 2048 (from a ~303
k-message histogram:, truncates only the multi-kB tail). 0 disables it; sub-16
values clamp to 16; parser rejects negatives/garbage/overflow.
trace_result_e — replaced matches()'s ambiguous pair<bool,bool> with a
DROP_FILTER/DEFAULT enum plus a pure should_log_full() policy, cleanly
splitting "positive filter → always full" from "default → subject to threshold."
full-payload filter — mirror of header-only: forces matching messages
bypass the full-logging threshold. Lets the size cap stay active for all
traffic while exempting one large service/method.
Tests and docs updated accordingly.
Test was hammering the send, which could overload
the receive buffer of the socket, leading to some lost
messages.
The test was expected to receive the end_message, that
never came
Increase delay a bit.
SomeIP data from fake udp sockets is no longer being printed.
This PR adds the logging of parsed someip message on gate.
Split the pending event registration set into producer vs. consumer sets
routing_manager_client kept both provider-side and consumer-side pending event registrations
in a single pending_event_registrations_ vector guarded by its own dedicated
pending_event_registrations_mutex_. Every access already ran under the matching role lock
(provider_mutex_ for offered events, consumer_mutex_ for consumed ones), so the extra mutex
was redundant. This splits the vector by role — pending_provided_event_registrations_ under
provider_mutex_ and pending_consumed_event_registrations_ under consumer_mutex_.
It is ALWAYS BAD for a build to touch the source directory. Every generated
file - no matter whether it's a compiled binary, intermediate source files, or
configuration, should go into the build directory
Fix the few last bad examples, one of them was even hardcoding /build..
also a very bad practice
While at it, very minor cleanups; use PROJECT_SOURCE_DIR/PROJECT_BINARY_DIR in
top-level CMakeLists.txt
Removes some unused declarations
Removes some unused declarations at routing_manager_client
Notice it while looking into rmc code.
On a client-id reuse, only tear down the old client's provider connection if that endpoint is still bound to the old
client, and only drop the shared sec_client mapping once the client has no endpoint left in either role.
Commit ec47d64 decoupled consumer- and provider-side error handling, but in on_routing_info (a consumer-side
signal) a new client id at a known address:port still tore down the old client's provider role unconditionally.
Since client ids are recycled across peers, after an STR resume a new app can already hold a healthy provider
connection there — and the teardown would close the new client's link instead of the departed one's.
Additionally, the shared client→sec_client mapping was cleaned only in remove_local_provider. As it is shared by
both roles, tearing down the provider dropped it even when a healthy consumer connection to the same client
remained, leaving that good consumer without its mapping.
Enables the receive-side offer check and adds regression tests to cover this change.
Extends the protocol tests to cover additional client/router/server message sequences:

Service offer/re-offer ordering relative to event registration:
Service release;
Event unsubscription
Stopping individual events.
Remove service_requests_ map
Commit d61bf2e made routing_manager_impl::requested_services_ be populated unconditionally (also for locally-offered
services). As a result it now holds essentially the same data as routing_manager_stub::service_requests_.
Given this, we can drop one of the maps.
Escalate client endpoint reconnect log to ERROR after a certain time.
There were reported cases where a client endpoint can't reach its
remote peer, it would retry forever,
only logging a WARNING on every failed attempt.
This is ok in some cases, but should be signaled when it takes
too much time.
This change makes the existing per-attempt "Restarting socket due to ..."
log escalate from WARNING to ERROR once the outage has lasted
continuously for more than VSOMEIP_RECONNECT_TIMEOUT.
This change fixes a segmentation fault in both overloads of the
add_remote_service_info method.
Warn users when provided/consumed events or handlers are registered too late or registered more than once.
These logs were made to catch ordering mistakes caused by API misuse, not races!
This PR adds logs for the following cases:

Late event-registration errors
Late handler-registration errors
Duplicate-registration

The late/duplicate checks are best-effort and API-level only. The routing manager is queried outside the
application's handler mutexes to avoid lock-ordering issues, so a concurrent offer/request/subscribe on
another thread can still be missed.
Remove the fixed client ID configuration from the network-tests
Commit 4062d4c added a log-only check in on_message that flags an incoming message whose client id belongs to the
receiving host's own range (which mis-routes the response to a local client instead of the remote host). It was gated
to skip hosts on the default diagnosis address — a test workaround — so it never ran exactly where client-id
collisions are most likely.
This removes that guard so the check runs on every host. It is safe: is_valid_client_id has a single caller and only
emits a log line — no message-delivery behaviour changes. The accompanying test changes make the network-tests
realistic (distinct diagnosis per boardnet side + dynamically-assigned client ids) so the warning is meaningful and
not spurious.
All fixed client IDs have been removed except for the following:

client_id_test_same_client_ids* - The purpose of the tests is to ensure that the client IDs are the same on both
the master and the slave.
client_id_utility* - The purpose of the test is to have client IDs both within and outside the range defined by the
diagnosis.
configuration_test- Check whether the specified client ID matches the one configured in the application
e2e_profile_04/07* - Use the client ID to check the full CRC
subscribe_notify_test_same_client_ids* - the purpose of the tests is to ensure that the client IDs are the same on
both the master and the slave.

In addition, some tests have been updated so that they no longer rely on the fixed client ID but instead check the
dynamic client ID.
Fix files not adhering to current clang-format expectations.
Reformat codebase with current version of clang-format version
and ruleset, in order to avoid unnecessary formatting in unrelated
patchsets.
These changes allow to build all the tests with the
CI clang compiler.
Make the connection-drop check reliable by detecting the drop event instead of the momentary
disconnected state.
The wait_for_connection_drop returned whether the socket pair is disconnected at the sampling instant.
After a partial-read drop the consumer reconnects within ~0.4 ms,
16:50:58.529619  [fake-socket] calling cancel on: {fd: 11, type: tcp, role: server, app: server}
16:50:58.530039  [fake-socket] added: client_to_server to the known connections

so a waiter woken under load re-evaluated only after the connection had healed
and timed-out — an intermittent failure of test_partial_read_leads_to_connection_drop
The predicate now also succeeds when socket_count_ has grown since the wait began (a
drop-and-reconnect occurred), so the result no longer depends on catching the transient window.
manuelnickschas and others added 24 commits September 17, 2026 15:18
Make the logger more robust against log-after-exit races.
Since we still have threads that may log after main() has exited,
we observe crashes that happen when the logger tries to use libdlt,
mainly due to data structures like the log context already having
been destroyed.
The safest, albeit slightly ugly way to deal with this is to make
the logger an "immortal singleton": It is simply never deallocated,
thus its memory remains valid throughout the lifetime of the process
The system will reclaim the memory anyway when tearing down the
process. In order for leak sanitizer not to treat it as a memory leak,
the allocated memory is referenced by a static pointer, thus it is
referenced for the duration of the process, and never marked as lost.
Ensure deregistration of the log context through a static guard
that will trigger this when it gets destroyed during thread cleanup.
This will set the context's log_level_ptr attribute to null, and
any attempt by another thread to log afterwards will trigger
a nullptr check that leads libdlt to safely discard the message
early on, without the risk of triggering undefined behavior.
Be a good boyscout too, and add a set of unit tests for the logger.
uint/size_t alone is fine, and anyhow shorter. Note as well that:

interface/ is changed, but it was already a mix of both
technically true that  does not have to define these global
types, but in practice it does (and see 1.!)

This commit does nothing else except those changes + applying formatter
API is somewhat unfortunate, passes an untyped uint16_t to the user
callback (which.. has to match the CommonAPI CallStatus definition..!)
While we do not do a major API break, we can at least make the codebase
better; define an enum class, use it internally, comment where
appropriate
Minor cleanups (mostly removing useless helpers) while at it
A lot of these logs are definitely errors: unparseable files, nonsense
values, duplicate-entry-for-same-application, ..
Removes warning log from routing which caused flood of logging
The intent of the logging was to show the messages that were received remotely and
not forward by the router and therefore dropped.
Unfortunately, this causes big flood of messages on target.
Hopefully, we'll work around with data provided from statistics logs
Install ccache in the build image so that compiler invocations are cached
across builds. This gives an overall build time improvement of around 45-50%.
Also install git and nginx, which are required by the build and by the
report publishing steps.
Fix security_test in debug build
Due to the combination of delays caused by valgrind-memcheck and debug build, the test fails because it does
not receive the availability response within 10 seconds.
Upon analyzing one of the logs, the service was offered 20 ms after the client failed to receive the availability.
To resolve the issue, scaled_timeout was applied, similar to other tests with the same issue.
They were not formatted before, somehow, and the CI is now enforcing it,
somehow
Gate the late registration log on consumer's register_message_handler_ext
CAPI currently does request_service before register_message_handler, resulting in a spam of
late registration logs on our ECUs.
This PR gates the log so that only providers use it.
Add more timeout scaling to fake socket tests.
Under sanitizers, we stilll see sporadic failures in fake
socket tests caused by exceeding tight timeouts.
Add timeout scaling to the major culprits to make them more
robust under heavy load in CI.
It was never implemented, it has a default that makes no sense, and it
was blindly copy-pasted all over the place, including example
configuration files
Therefore drop it
Disable the initial wait phase by default. The 0-to-3000ms default is
not sane, and generally from a POSIX PoV it does not make sense to have
it
(there are enough delays in vsomeipd starting up, other applications
starting up..)
Reduce the FIND/OFFER debounce time to 20ms. The 500ms default is also
far from sane when the cycle time is 1000ms. It makes sense to have a
debounce time (we DO NOT want to send a multicast FIND/OFFER per
service!), but a small window is sufficient, as applications, once they
startup, offer/request everything in a small window of time as well
Especially the FIND/OFFER debounce times have been the source of delays
in the past
Add a connection timeout for connections with the router
Add a connection timeout for clients with the router.
After this timeout is reached the application logs an error message
and continues to try to connect with the router
Under stress, the e2e_profile_0x_test_external occasionally fails because it receives a duplicate initial notification
and the test is expecting different/incremental payloads.
This happens because, when the initial event is lost or experiences a very long delay in being sent, it is detected as
lost.
Fix usei unit tests
Instead of posting the messages to an actual socket,
call the mocked on_message_received_unlocked directly
so no messages are dropped
Normal registration log, downgrade it to VSOMEIP_INFO_P
It is only used to remove boardnet subscriptions, but anyhow the router
already removes those via subscription expiry, so it's unnecessary
Besides the useless code, it is harmful - it guarantees a thundering
herd on suspend, because very libvsomeip client is woken up
Fix duplicate registration warn to match major
The warning log for a availibility handler duplicate registration depends only on the service/instance, even though
the handler also depends on the major/minor version. Therefore, if two handlers are registered for the same
service/instance but with different major versions, the lib will report a warning for a duplicate
registration even though they are different handlers.
Now, it only logs the handler if the major version also matches.
Adopts the Android 17 bring-up changes
Changes

Remove Android.mk
Add host_supported to libvsomeip3.
Add three alias stubs — libvsomeip3-cfg, libvsomeip3-e2e, libvsomeip3-sd
The test had a mechanism to re-send the offer if the connection
wasn't established in time, however, it would do so with the same
sessionID, causing a reboot detection and switching ports
Drop message if we are not interested in the service before doing
SOME/IP-TP, not after
While at it, springle some ANY_* defs in the codebase
Found by Mayur Agnihotri, reported in private
Make it so configs are reused and not needlessly loaded
every time an app starts on a single process.
Keeps behavior if a config is defined by VSOMEIP_CONFIGURATION_<name>.
Minor change to configuration_impl to avoid unnecessary copies.
First app now loads all of the config files, being the router or not
Policies are now unique per app and loaded by configuration_impl
but copied for each application.
Added tests for configuration behavior.
@Stek210588
Stek210588 requested a review from a team September 18, 2026 10:11
Comment thread test/network_tests/client_id_tests/client_id_test_utility.cpp Dismissed
Comment thread test/network_tests/client_id_tests/client_id_test_utility.cpp Dismissed
@Stek210588
Stek210588 merged commit 9cc8753 into main Sep 21, 2026
5 checks passed
@Stek210588
Stek210588 deleted the main_3.7.6 branch September 21, 2026 08:37
Stek210588 added a commit that referenced this pull request Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants