Every value in a message describes its own form, so a decoder needs no schema hint to read it. It's the format proof artifacts, traces and certificates travel on between the compiler, the prover and the checker.
A proof pipeline is a set of machines passing artifacts to each other. The compiler emits an obligation, the prover discharges it, and an independent checker re-derives the result rather than trusting either. Those artifacts are large, and they cross a wire on the way.
A checker that spends longer parsing an artifact than checking it has moved the bottleneck without removing it. That is the only reason a wire format appears on a silicon verification site, and it's the reason this one is measured to usable rather than to parsed.
Timed to usable, not to parsed
A zero copy format can report a decode time near zero and then charge for the work on first access, so every value in the message is touched before the clock stops. Cap'n Proto is read zero copy, which is its best case.
Encode is fastest on all six workloads as well, against bincode, postcard, MessagePack, CBOR, JSON, Apache Arrow, Protobuf with gRPC and Cap'n Proto.
| Workload | LOGOS | Best rival |
|---|---|---|
| Int array, random, n = 1000 | 58 nsZero copy | Cap'n Proto, 211 ns |
| Float array, random f64, n = 1000 | 84 nsVarint | bincode, 678 ns |
| Float time series, n = 1000 | 82 nsVarint | bincode, 520 ns |
| Point list, n = 1000 | 133 nsZero copy | Cap'n Proto, 373 ns |
| Record list, n = 200 | 566 nsFixed memcpy | Cap'n Proto, 1.26 µs |
| String list, n = 200 | 297 nsVarint | Apache Arrow, 960 ns |
Out of the box the codec auto compresses and the rivals ship raw, and on that comparison it produces the smallest message on 6 of 6 workloads, a geometric mean of 1.5x smaller. That is a real transport advantage, because it's what a link actually carries by default. It is not a claim about the encoding.
Give every rival the same compression and the result is 3 of 6. On random primitive arrays a memcpy is a memcpy, and no amount of framing changes that. Both numbers describe the same six measurements, and neither one belongs on the page without the other.
| Workload | LOGOS | Best rival | Result |
|---|---|---|---|
| Int array, random | 3 KB | Protobuf and gRPC, 3 KB | Tie |
| Float array, random f64 | 5 KB | postcard, 5 KB | Rival smaller |
| Float time series | 3 KB | JSON, 2 KB | Rival smaller |
| Point list | 4 KB | postcard, 5 KB | LOGOS smaller |
| Record list | 2 KB | postcard, 2 KB | LOGOS smaller |
| String list | 2 KB | postcard, 2 KB | Rival smaller |
This is the workload Cap'n Proto is built for, and the one that matters when a checker needs a single field out of a large artifact. Over 2000 messages the struct view dial reads a field in 15 ns against Cap'n Proto at 44 ns, 2.93x quicker.
The gap against the self describing formats is a different order, because they have no choice but to decode the whole message before they can index into it. bincode takes 44.54 µs, JSON 140.70 µs and CBOR 240.89 µs for the same single field read.
Where a column has structure the encoder can name, it ships the structure rather than the values. These are the standard columnar techniques Parquet and Arrow are built on, applied automatically when they win and skipped when they do not.
A boolean column encodes at one bit per value, 7.71x smaller than postcard against a byte per value, measured on a real boolean column. Low cardinality strings are dictionary encoded, 13.03x smaller than postcard, on categorical labels where the same values recur.
The figures below are not general results and are not counted in any claim on this page. Each one is a column built to have perfect structure, so the encoder ships the rule that generates the data instead of the data. They show where the ceiling sits when structure is total, and nothing about a real payload.
The last row is measured against Protobuf with gRPC. The rest are against postcard. Read them as a demonstration of the mechanism, not as a compression ratio anyone should expect.
| Contrived shape | What ships | Factor |
|---|---|---|
| Affine progression | Base, stride and count | 333.56x |
| Polynomial, 3i squared minus 5i plus 7 | Its finite difference seeds | 372.11x |
| Consecutive integer set | Base, stride and count | 242.25x |
| Integer keyed map | Two such columns | 279.00x |
| Repetitive integer column | The run, once | 613.91x |
There is no single encoding. A frame is written with the dial that suits the link it crosses, and the auto structure setting runs a bake off per column and keeps the winner.
Four further capabilities exist and are deliberately absent from every measurement above, so no number on this page depends on them: an FNV checksum tag, type identifier elision over a link with a known schema, Reed and Solomon forward error correction, and subtree deduplication with back references.
| Dial | What it's for |
|---|---|
| Varint, LEB128 | Smallest bytes. The default on a bandwidth bound link |
| Fixed memcpy | Raw 8 byte integers. Fastest decode on a datacenter or RDMA link, at about four times the size |
| Group varint | Varint class size, with several integers decoded at once under SIMD |
| Auto structure | A per column bake off across delta, delta of delta, frame of reference bit packing, run length, dictionary, affine and polynomial. Ships the smallest and never larger than varint |
| Xor delta floats | Gorilla style. Slowly varying float streams shrink, and high entropy data falls back to memcpy |
| Struct view | An offset table per field, so any one field reads in constant time |
| Compression | deflate, lz4 or zstd over the frame, kept only where it's smaller |
Payloads are seeded random rather than a tidy sequence from zero to n. A sequence would hand a structural codec a free win on every workload, which is precisely what the showcase section above demonstrates. The fair numbers only mean something because the input was not built for the encoder.
Every codec encodes the same logical data on the same machine. Size is exact bytes with no envelope, and encode and decode are nanoseconds per whole message operation, taken with a warm up. The size advantage holds on 6 of 6 workloads out of the box, and on 3 of 6 once every rival is given the same compression. Showcase figures come from generated columns where the codec ships the generating rule instead of the data, and they stay out of every summary figure.
Measured on Intel Core i9-14900K, Ubuntu 24.04.2 LTS x86_64. Compared against bincode, postcard, MessagePack, CBOR, JSON, Apache Arrow, Protobuf with gRPC and Cap'n Proto.
Source: Every workload, every dial, and the full methodology, logicaffeine.com(opens in a new tab)
What are the artifacts this format is built to carry?