Introduction
#112 has marked the end of the PoC phase of Quent by demonstrating we can define composable and reusable application event models from single & relatively straightforward definitions from which an enormous amount of boilerplatey code can be generated to tie everything from instrumentation library all the way to analysis and UI together.
The current PoC-grade implementation has quite a lot of tech debt that needs to be erased in order to stabilize the project and provide a solid foundation for both users and a next wave of features. To this end, #128 has dissected the entire project into somewhat atomic parts, and provides a proposal for a more solid foundation.
This is a tracking issue for migrating the project to that architecture. The end result will be a first beta release (v0.1, this may require a few follow-ups after this issue is completed, so consider this an umbrella issue).
Schema-based instrumentation architecture
Schema
The new architecture centers itself around a "schema" for event data, which at the core only describes those things required to emit, transport, and store event data. This only includes:
- Entity types, including their events and event attributes
- Record types (compile-time defined attributes)
Each element in the schema can carry metadata for arbitrary purposes, as well as "constraints"
Constraints
The idea is to map all other currently existing concepts (tree-forming references, FSMs, resources, query engine stuff) onto this schema as opaque "constraints". Constraints will be a special type of metadata that impose certain restrictions or rules on the entities, events, or data. These restrictions/rules are (should be) validated in any layer of Quent that does something with these constraints.
Whether some layer of Quent can/should do something with these constraints depends on the constraint. Therefore, the benefit of introducing these constraints as explicitly separated from the core concepts, is that certain layers can be made oblivious to these constraints such that they can continue to work even if these constrains change. In the long term, the idea is that third parties can also author domain- or application-specific constraints to impose certain semantics on the events.
Example
Quent's stack consists of several layers, roughly speaking:
- Model definition: e.g. a YAML-based DSL to capture a schema
- Instrumentation: code generation step to produce a strongly typed library + potential bindings for C++, Python, ...
- Application:
- Capture & store: event export & storage
- Analysis: semantic analysis of events & serving a front-end
- Presentation: e.g. a UI, CLI tool, MCP endpoint, etc.
Consider the concept of allowing entities to emit events according to the topology defined through Finite-State-Machine semantics. This restricts the sequence of events an entity can emit is a non-core constraint, as it does not influence how to emit, transport, or store event data.
A code generator for a Rust-based instrumentation library that is aware of this constraint may choose to apply a typestate pattern to make it impossible for entity handles to transition into an illegal state.
However, if a C++ code generator has yet to catch up with such a feature, it can still generate an instrumentation library that can emit events for these FSM entities without requiring any modifications (although of course the client code should produce events in the correct order).
Some other layers don't care about these constraints at all. E.g. in translating the Quent schema to an Apache Arrow schema to store events in e.g. Parquet files, the FSM constraint has no significant meaning.
Goal
The goal of this umbrella issue is to:
- define the schema
- introduce a YAML-based DSL to capture the schema rather than relying on Rust macros
- migrate the instrumentation layer to use the schema during code generation
The goal of this issue is not to:
4. migrate the analysis layer to use the schema during code generation. While the current quent-model-macros crate does take care of generating some boilerplate for analysis, this boilerplate will have to be hand-written until the analysis side migrates over as well, which is going to be another bigger kind of effort explicitly not included in this work).
5. Make layers 1, 3, 4, 5, 6 as flexible as the "constraints" mechanism described above, continue to read below.
After this issue
In addition to defining schema and constraints to be able to declare arbitrary event semantics without breaking the entire stack, a longer term goal is expanding the flexibility of schema constraints to other pieces of the stack. This will allow influencing how schemas are captured e.g. in the YAML-based DSL, in code generation steps, analysis, and UI components for visualations.
Such a flexibile system into which we can provide well-isolated sets of features, where only those layers in which the semantics can improve the stack need to be appended, will be called "semantic modules" a.k.a. mods, as described in the main README.
In a way these mods represent a thin vertical slice of the entire stack, so changes to these mods are minimally impactful for other parts of the whole stack (unless there are dependencies between mods). In the even longer term I think it would be nice if those modules can be authored by third parties.
I propose this because I strongly believe that:
- event semantics are strongly tied to instrumentation syntax and visualization (or other forms of human or agentic digested events). We need a system that is flexible enough to support introducing any kind of semantics at every layer where the user sees fit in order to serve a wide variety of applications, while retaining end-to-end type safety and performance.
- agentic coding benefits greatly from well-curated modules for which only the final bits of application-specific glue are to be provided.
Plan
The plan is to introduce (or refactor) the following crates in this order:
This tracking issue ends there, as this provides everything necessary to retain all currently existing functionality, and will stabilize everything on the data capture side, such that historical data produced by early adopters is less likely to become unusable by future versions (although I think it is too early to conclude we could perhaps never have breaking changes to the definition of the schema itself).
These changes will incorporate solutions for the following pre-existing issues:
#186 #144 #127 #79 #491 #492
Introduction
#112 has marked the end of the PoC phase of Quent by demonstrating we can define composable and reusable application event models from single & relatively straightforward definitions from which an enormous amount of boilerplatey code can be generated to tie everything from instrumentation library all the way to analysis and UI together.
The current PoC-grade implementation has quite a lot of tech debt that needs to be erased in order to stabilize the project and provide a solid foundation for both users and a next wave of features. To this end, #128 has dissected the entire project into somewhat atomic parts, and provides a proposal for a more solid foundation.
This is a tracking issue for migrating the project to that architecture. The end result will be a first beta release (v0.1, this may require a few follow-ups after this issue is completed, so consider this an umbrella issue).
Schema-based instrumentation architecture
Schema
The new architecture centers itself around a "schema" for event data, which at the core only describes those things required to emit, transport, and store event data. This only includes:
Each element in the schema can carry metadata for arbitrary purposes, as well as "constraints"
Constraints
The idea is to map all other currently existing concepts (tree-forming references, FSMs, resources, query engine stuff) onto this schema as opaque "constraints". Constraints will be a special type of metadata that impose certain restrictions or rules on the entities, events, or data. These restrictions/rules are (should be) validated in any layer of Quent that does something with these constraints.
Whether some layer of Quent can/should do something with these constraints depends on the constraint. Therefore, the benefit of introducing these constraints as explicitly separated from the core concepts, is that certain layers can be made oblivious to these constraints such that they can continue to work even if these constrains change. In the long term, the idea is that third parties can also author domain- or application-specific constraints to impose certain semantics on the events.
Example
Quent's stack consists of several layers, roughly speaking:
Consider the concept of allowing entities to emit events according to the topology defined through Finite-State-Machine semantics. This restricts the sequence of events an entity can emit is a non-core constraint, as it does not influence how to emit, transport, or store event data.
A code generator for a Rust-based instrumentation library that is aware of this constraint may choose to apply a typestate pattern to make it impossible for entity handles to transition into an illegal state.
However, if a C++ code generator has yet to catch up with such a feature, it can still generate an instrumentation library that can emit events for these FSM entities without requiring any modifications (although of course the client code should produce events in the correct order).
Some other layers don't care about these constraints at all. E.g. in translating the Quent schema to an Apache Arrow schema to store events in e.g. Parquet files, the FSM constraint has no significant meaning.
Goal
The goal of this umbrella issue is to:
The goal of this issue is not to:
4. migrate the analysis layer to use the schema during code generation. While the current
quent-model-macroscrate does take care of generating some boilerplate for analysis, this boilerplate will have to be hand-written until the analysis side migrates over as well, which is going to be another bigger kind of effort explicitly not included in this work).5. Make layers 1, 3, 4, 5, 6 as flexible as the "constraints" mechanism described above, continue to read below.
After this issue
In addition to defining schema and constraints to be able to declare arbitrary event semantics without breaking the entire stack, a longer term goal is expanding the flexibility of schema constraints to other pieces of the stack. This will allow influencing how schemas are captured e.g. in the YAML-based DSL, in code generation steps, analysis, and UI components for visualations.
Such a flexibile system into which we can provide well-isolated sets of features, where only those layers in which the semantics can improve the stack need to be appended, will be called "semantic modules" a.k.a. mods, as described in the main README.
In a way these mods represent a thin vertical slice of the entire stack, so changes to these mods are minimally impactful for other parts of the whole stack (unless there are dependencies between mods). In the even longer term I think it would be nice if those modules can be authored by third parties.
I propose this because I strongly believe that:
Plan
The plan is to introduce (or refactor) the following crates in this order:
1.quent-schema: captures the structure of application event model data and introduces the concept of opaque constraints that can be implemented by any built-in or third-party crates to constrain events to arbitrary conventions.2.quent-constraints: constraints trait and canonical validation path for schemas3.quent-fsm: built-in constraint where events represent FSM transitions and must be sequenced to follow a well-defined topology4.quent-ref-<x>: built-in constraint for entity refs, so they can be given semantically enriched data-carrying roles and constraintsquent-ref-target: feat: add ref-target constraint crate #199quent-ref-tree: feat: add ref-tree constraint crate #2005.quent-instrumentation-<X>quent-instrumentation-build: Generates instrumentation libs from schemaquent-instrumentation: Generic types for any instrumentation lib + migrating PoC crates to the new arch:6.quent-yaml: YAML application event model definition for schemas to use in build.rs7.quent-cpp: Codegen for C++ targets8.quent-python: Codegen for Python targets9.quent-resource: constraint + modeling syntax for resources and usages of those resources in FSM states10.Post-migration tasksThis tracking issue ends there, as this provides everything necessary to retain all currently existing functionality, and will stabilize everything on the data capture side, such that historical data produced by early adopters is less likely to become unusable by future versions (although I think it is too early to conclude we could perhaps never have breaking changes to the definition of the schema itself).
These changes will incorporate solutions for the following pre-existing issues:
#186 #144 #127 #79 #491 #492