Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,057 words · 4 segments analyzed
When you need to parse data for downstream processing, the enthusiasm to get the processing outcome quickly might induce you to overlook one interesting nuance: the cost of how it’s parsed. See this “parse me a u64” function for example: fn parse_id(bytes: &[u8]) -> Result<u64, ParserError> { if bytes.len() < 8 { return Err(ParserError::InputTooShortForU64); } let owned = bytes[..8].to_vec(); Ok(u64::from_le_bytes(owned.try_into().map_err(ParserError::InvalidU64)?)) } It checks the length, copies the right number of bytes from the slice into a Vec and parses those as a u64. Returns adequate Errs for production. All looks short and safe. Now let’s examine it more carefully. For a type as simple as a number, should the heap be used or not? What about syscalls? In Rust, a Vec always uses the heap. That allocation might or might not require a syscall. That would depend on whether the allocator already has a free block or has to ask the OS for more memory which can easily happen under load. A lot of applications will not care about this detail but some need to care. What happens when you care? Since eight bytes in a buffer are already a u64, and from_le_bytes knows how to read them into a u64, it is fair to say the Vec I allocated in that .to_vec() is accidental complexity on the way to the result. A kind of internal complexity to achieve an end. But what if we find ways to not need that? If “making room” for the final result is not required, then we save time, effort, runtime cost, and code maintenance of that “making room”. Let’s push design to read in place and see what happens. We’re already receiving a reference &[u8] of where to read, so the next thing is to know how much exactly to read into a u64. This “reading in place” technique is what people call zero-copy, because whatever you’re doing for the processing, you are already not requiring a copy to start doing it (working on &[u8] is that contract). These eight bytes still move to the stack for from_le_bytes, but we skipped doing anything with the heap. The strength that Rust has in its contracts (types) is used to produce in a program zero-cost abstractions. So when it reads the right types in data and has all error modes properly modeled and handled, what you get is all that strength in its runtime. Let’s use abstractions that won’t copy now.
Here is the previous example but in a zero-copy implementation instead: fn parse_id(bytes: &[u8]) -> Result<u64, ParserError> { let raw: [u8; 8] = bytes.get(..8) .ok_or(ParserError::InputTooShortForU64)?
.try_into().map_err(ParserError::InvalidU64)?; Ok(u64::from_le_bytes(raw)) } This time we read the bytes in whichever place they already are, validate we can read 8 valid bytes there and use them to infallibly produce a u64. And what if a type is more complex than a u64? A longer message is the same technique, repeated over a new contract shaped differently. After all, this reusability is what makes zero-copy a technique.
For example in the command codec I am writing for orderflow, a command is an enum. A limit order carries an account, a client order id and other details. A cancel by order id is two integers. A market order has no price field. The bytes in the buffer are a different width for each engine command. /// EngineCommand is a client request to change order state. /// When a matching engine receives work, it uses this type so the variant is the /// layout: NewLimit and NewMarket are separate, and New has no engine OrderId. #[derive(Debug, Copy, Clone, PartialEq, Eq, Hash)] pub enum EngineCommand { /// Request to open a limit order. NewLimit { account_id: AccountId, client_order_id: ClientOrderId, instrument_id: InstrumentId, side: Side, price: Price, quantity: Quantity, }, /// Request to open a market order. NewMarket { account_id: AccountId, client_order_id: ClientOrderId, instrument_id: InstrumentId, side: Side, quantity: Quantity, }, /// Request to cancel by engine OrderId. CancelByOrder { order_id: OrderId }, ... For these commands, a kind byte says which layout I am looking at. Each layout has one length. I check that length, then I read fields into new types. Take NewLimit for example, price is a count of ticks, an i64, quantity is a count of lots, and for the rest? All of them are intentionally Copy. This makes commands not only very compact but, together, also unlock something. Notice the #[derive(Debug, Copy... in its definition? Nothing in enum EngineCommand owns a buffer. Their Copyness unlocks the full enum’s “copyness”. It’s a very deliberate design choice to enable the engine to parse using the zero-copy technique so not only the parsing but maybe further downstream processing of commands can be done “in place”. const NEW_LIMIT_LEN: usize = 57; fn decode_new_limit(payload: &[u8]) -> Result<EngineCommand, DecodeError> { if payload.len() != NEW_LIMIT_LEN { return Err(DecodeError::Length); } Ok(EngineCommand::NewLimit { account_id: AccountId::new(read_u64(payload, 8)), client_order_id: ClientOrderId::new(read_u64(payload, 16)), instrument_id: InstrumentId::new(read_u64(payload, 24)), side: decode_side(payload[32])?, price: Price::new(read_i64(payload, 33)), quantity: Quantity::new(read_u128(payload, 41)), }) } As you can expect, these commands live in an envelope with a sequence number and a kind type that defines how they are read from the stream: match kind { Kind::NewLimit => decode_new_limit(payload), Kind::NewMarket => decode_new_market(payload), Kind::CancelByOrder => decode_cancel_by_order(payload), Kind::CancelByClient => decode_cancel_by_client(payload), Kind::Replace => decode_replace(payload), } NewLimit is 57 bytes. CancelByOrder is 16. Everything is determined how it is to be read, and that’s guaranteed by the compiler. What’s left is making invalid data impossible to ingest. Something you do by handling extensibly each possible failure mode. And for the writes? Well, we do the same design intent we used for the zero-copy, but for writing. That is, the encoder writes those same offsets into a buffer the caller already owns. Every variant is a fixed pile of integers that fits predictably in the data buffer. No intermediate allocations required while writing. The journal frame and the datagram both hand over the payload region and share one write. Neither path grows a private copy of the command. This enum is interesting because the variant changes the memory layout but it stays free of the heap. In the same way these eight bytes were already a number, a limit order is fifty-seven, in known places. The rest is getting them wrapped in an envelop with a sequence number and kind byte that tells me where to look for them safely, consistently, and in the right order.