Demystifying Binary Protocols
When I started working with binary protocols, the format felt like a mystery or some advanced low-level bit magic, but it's not as hard as it seems as long as you stick to 1 byte as the smallest unit.
We will start from parsing a simple comma-separated list to parsing a binary-encoded message. I am only including code for the decoder; writing the encoders can be done as homework.
Comma-separated list
// string.ts
function parseString(msg: string): string[] {
const result: string[] = [];
let segment = "";
for (let i = 0; i < msg.length; i++) {
if (msg[i] === ",") {
result.push(segment);
segment = "";
} else {
segment += msg[i];
}
}
result.push(segment);
return result;
}
console.log(parseString("hello,world"));

These algorithms should be simple enough; it's a greedy approach: collect characters until we hit a comma (,), then store the segment and clear the temporary variable for the next segment. It has an O(n) time complexity, where n is the size of the msg string.
Pascal string
Pascal string, or length-encoded string, is a string where we put the length of the string at the start of the string, so we don't have to read each character: "05hello05world"
If we reserve the first 2 characters for the length, we can read just the n characters and avoid O(n) time complexity.
// pascalString.ts
const LENGTH_SIZE = 2;
function processPascalString(msg: string): string[] {
const result: string[] = [];
let start = 0;
while (start < msg.length) {
const length = parseInt(msg.slice(start, start + LENGTH_SIZE));
if (
Number.isNaN(length) ||
!Number.isFinite(length) ||
!Number.isSafeInteger(length)
) {
throw new Error(`string is not length encoded`);
}
start += LENGTH_SIZE;
const string = msg.slice(start, start + length);
result.push(string);
start = start + length;
}
return result;
}
console.log(processPascalString("05hello05world"));

We are starting to develop our protocol format. if the first 2 characters are not a valid integer we can throw an error and stop execution. The time complexity is O(m), where m is the number of strings encoded in the msg
But we are limited to only passing around strings. This is not very useful; we can improve our protocol by adding type information to the msg, as we did with length info.
Pascal string with types
function isFloat(value: number) {
return typeof value === "number" && !Number.isNaN(value) && value % 1 !== 0;
}
const LENGTH_SIZE = 2;
const TYPE_SIZE = 2;
const BOOL_SIZE = 2;
const TYPE_MAP = {
"00": "string",
"01": "int",
"02": "float",
"03": "boolean",
};
const TYPE_PARSER = {
string: (string: string) => string,
int: (number: string) => {
const int = parseInt(number, 10);
if (
Number.isNaN(int) ||
!Number.isFinite(int) ||
!Number.isSafeInteger(int)
) {
throw new Error(`${number} is not an int`);
}
return int;
},
float: (number: string) => {
const float = parseFloat(number);
if (Number.isNaN(float) || !Number.isFinite(float) || !isFloat(float)) {
throw new Error(`${number} is not a float`);
}
return float;
},
boolean: (string: string) => {
switch (string) {
case "01":
return true;
case "00":
return false;
default:
throw new Error(`${string} not boolean`);
}
},
};
function pascalStringWithType(msg: string): (string | number | boolean)[] {
const result: (string | number)[] = [];
let start = 0;
while (start < msg.length) {
const type = TYPE_MAP[msg.slice(start, start + TYPE_SIZE)];
if (!type) {
throw new Error(`unknown type ${type}`);
}
const typeParser = TYPE_PARSER[type];
start += TYPE_SIZE;
const length =
type === "boolean"
? BOOL_SIZE
: TYPE_PARSER["int"](msg.slice(start, start + LENGTH_SIZE));
start += type === "boolean" ? 0 : LENGTH_SIZE;
const string = msg.slice(start, start + length);
result.push(typeParser(string));
start = start + length;
}
return result;
}
console.log(pascalStringWithType("0005hello0005world01021202063.141503010300"));

We are using the first 4 characters for type and length info; they are commonly referred to as the header. This gives us enough space to encode 99 different values for type info, which cover almost all primitive types and any special type your protocol might make up
We read the type and the length, then skip the header byte, read the data, and parse it based on type:
- a string is 00
- an int is 01
- a float is 02
- and a bool is 03, where 00 is false, and 01 is true
We can now encode much more data in our protocol. And the time complexity is still O(m)
But this representation is wasteful; we are using 4 bytes to represent a max length of 99
Each character is 2 bytes long, whereas one u8 (a 1-byte unsigned int) can represent a max length of 255. So we can go from using 4 bytes for length and type each to 1 byte for length and type and represent a length of 255 and 255 different types if we just switch to a binary protocol
Binary protocol
// bytes.ts
const LENGTH_SIZE = 1;
const TYPE_SIZE = 1;
const INT_FLOAT_SIZE = 4;
const BOOL_SIZE = 1;
const TYPE_MAP = {
0x00: "string",
0x01: "int",
0x02: "float",
0x03: "boolean",
};
function parseBytes(bytes: Uint8Array): (string | number | boolean)[] {
const result: (string | number | boolean)[] = [];
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
const textDecoder = new TextDecoder();
let start = 0;
while (start < bytes.length) {
const type = view.getUint8(start);
start += TYPE_SIZE; // Advance 1 byte for Type
let length: number;
if (TYPE_MAP[type] === "string") {
length = view.getUint8(start);
start += LENGTH_SIZE; // Advance 1 byte for String Length
} else if (TYPE_MAP[type] === "int" || TYPE_MAP[type] === "float") {
length = INT_FLOAT_SIZE; // 32-bit int & float are 4 bytes
} else if (TYPE_MAP[type] === "boolean") {
length = BOOL_SIZE; // 1 byte
} else {
throw new Error(`Unknown type tag: 0x${type.toString(16)}`);
}
const payload = bytes.subarray(start, start + length);
const payloadView = new DataView(
bytes.buffer,
bytes.byteOffset + start,
length,
);
switch (TYPE_MAP[type]) {
case "string":
result.push(textDecoder.decode(payload));
break;
case "int":
result.push(payloadView.getUint32(0));
break;
case "float":
// Rounded to match single-precision float representation
result.push(Math.round(payloadView.getFloat32(0) * 10000) / 10000);
break;
case "boolean":
result.push(payloadView.getUint8(0) === 1);
break;
}
start += length;
}
return result;
}
const testBytes = new Uint8Array([
// 1. "hello" (Type: 0x00, Len: 5, Value: "hello")
0x00, 0x05, 0x68, 0x65, 0x6c, 0x6c, 0x6f,
// 2. "world" (Type: 0x00, Len: 5, Value: "world")
0x00, 0x05, 0x77, 0x6f, 0x72, 0x6c, 0x64,
// 3. 12 (Type: 0x01, Value: 12 as 4-byte uint)
0x01, 0x00, 0x00, 0x00, 0x0c,
// 4. 3.1415 (Type: 0x02, Value: 3.1415 as 4-byte float)
0x02, 0x40, 0x49, 0x0e, 0x56,
// 5. true (Type: 0x03, Value: 1)
0x03, 0x01,
// 6. false (Type: 0x03, Value: 0)
0x03, 0x00,
]);
console.log(parseBytes(testBytes));

The code is a bit more complex, but the core is the same: we read the type, then the length, then the data, and parse it.
If you look at the code, we only encode length for strings because the rest- uint32, float32, and bool are of fixed size: 4 bytes for int and float, and 1 byte for boolean, so we don't need to encode size info for them. Hopefully this helps demystify the binary protocol for you. We can be more efficient by grouping all booleans and using 1 bit per boolean (1 for true, 0 for false), but that makes the code more complex because we need bitmasks to encode and decode data for a simple protocol; this is enough. You can use it to represent a KV protocol where even indices are keys and odd indices are values
