Currently, the contract between the generator and the executor is completely arbitrary. It could be JSONL, binary, space delimited, or CSV as long as the executor knows how to parse it. Meanwhile, the contract between the executor and impa is strictly a CSV line: data_id,duration_nanos.
- The Problem: Arbitrary string parsing (like splitting by spaces or commas) introduces massive I/O and CPU overhead. In micro-benchmarking, the time it takes an executor to deserialize a massive string might skew the actual algorithmic performance you are trying to measure.
- The Streaming Solution: Elevate the framework by introducing native support for structured, high-performance streaming protocols.
- Providing official, zero-copy serialization wrappers (e.g., Apache Arrow IPC or Protocol Buffers) would allow generators to blast binary streams to executors with near-zero overhead.
- We could introduce an
impa standard library (e.g., impa-rs, impa-py, impa-zig) that handles the boilerplate of reading the standard stream and reporting the metrics back.
Currently, the contract between the generator and the executor is completely arbitrary. It could be JSONL, binary, space delimited, or CSV as long as the executor knows how to parse it. Meanwhile, the contract between the executor and
impais strictly a CSV line:data_id,duration_nanos.impastandard library (e.g.,impa-rs,impa-py,impa-zig) that handles the boilerplate of reading the standard stream and reporting the metrics back.