Most serialization libraries let you choose the format: you describe a type and the library decides the bytes. scodec is for the opposite situation. The format is already fixed by an RFC, a device datasheet, a file specification or a legacy system, and your job is to read and write it exactly, bit for bit. scodec gives you a small type, Codec[A], that both encodes a value to bits and decodes bits back to a value, plus combinators to build large codecs out of small ones so the Scala code reads like the protocol diagram.
This article explains the model from first principles, builds a realistic frame codec with a magic number, bit flags, a length prefix and a tagged union of message types, streams it from a file with fs2, and covers how to test it and how it fails. The code targets scodec 2 on Scala 3; the differences for scodec 1 on Scala 2 are noted where they matter.
Bits first: scodec-bits
scodec works on bits, not bytes, because many binary formats do: a TCP header packs 4-bit and 1-bit fields, and video and compression formats rarely align to bytes. The companion library scodec-bits provides two immutable types. ByteVector is a byte sequence and BitVector is a bit sequence that need not be a multiple of eight. Both are persistent structures with cheap slicing and concatenation, so a decoder can take drop(16) of a large buffer without copying it.
Literals are checked at compile time: hex"cafe" builds a ByteVector and bin"101" a three-bit BitVector, and a typo in either is a compile error rather than a runtime surprise. Convert with .bits and .bytes; the second pads a partial final byte, which matters when you compare lengths.
The Codec contract
A Codec[A] is an Encoder[A] and a Decoder[A]. Encoding returns Attempt[BitVector]. Decoding returns Attempt[DecodeResult[A]], where a DecodeResult holds the decoded value and the remainder: the bits that were not consumed. Attempt is either Successful or Failure(err), an explicit error value rather than an exception, so a malformed packet is an ordinary result you can log and skip.
The remainder is the key to composition. To decode a header followed by a body, run the header codec, take its remainder, and run the body codec on that. Every combinator in the library is this idea with a different policy about what the next codec is and how the values are combined.
// build.sbt: scodec-core 2.x for Scala 3, 1.11.x for Scala 2 (check Maven Central for the latest)
// libraryDependencies += "org.scodec" %% "scodec-core" % "<2.x version>"
import scodec.*
import scodec.bits.*
import scodec.codecs.*
val bytesIn: ByteVector = hex"0102ff" // compile-time checked literal
val bitsIn: BitVector = bin"101" // exactly 3 bits, not a byte
val three: Codec[(Int, Int, Int)] = uint8 :: uint8 :: uint8
three.decode(bytesIn.bits)
// Attempt.Successful(DecodeResult((1,2,255), BitVector(empty)))
three.encode((1, 2, 256))
// Attempt.Failure(...): 256 is out of range for uint8In scodec 2, :: builds a codec for a Scala 3 tuple, and .as[CaseClass] maps a tuple codec onto a case class with the same field types in the same order. In scodec 1 the same operator built a shapeless HList, so code that pattern-matches on HLists is the main thing to rewrite when migrating; the .as calls usually survive unchanged.
The primitive codecs and their traps
The primitives map directly to protocol vocabulary, and each one has a sharp edge worth knowing:
uint8,uint16,int32and friends are big-endian. Little-endian variants end inL, such asuint16L. Choosing the wrong one decodes without error and yields a wrong number, the worst kind of bug.uint32decodes to aLongbecause an unsigned 32-bit value does not fit in anInt. Encoding checks the range, as the 256 example above shows.boolis a single bit. Two booleans followed by a byte leave everything after them misaligned by six bits unless you account for the padding.utf8andbyteswith no size consume everything that remains. They are only correct as the last field or inside a size-bounding combinator.constant(bits)writes fixed bits and fails decoding if they do not match: the natural home for magic numbers and reserved-zero fields.variableSizeBytes(sizeCodec, valueCodec)writes the encoded length, then the value, and on decode gives the inner codec only that many bytes. This is how length-prefixed fields work.provide(value)consumes nothing and always yields the value: useful for case objects.
Worked example: a framed protocol
Consider a small device protocol. Every frame starts with the magic bytes CA FE. A header byte holds the protocol version; the next byte holds an ack flag, an urgent flag and six reserved bits that must be zero. Then comes a 16-bit length and a message: a one-byte type tag followed by a body that depends on the tag. Type 1 is a ping with an unsigned 32-bit id, type 2 carries data for a stream, and type 3 closes the connection.
import scodec.*
import scodec.bits.*
import scodec.codecs.*
final case class Header(version: Int, ack: Boolean, urgent: Boolean)
sealed trait Msg
final case class Ping(id: Long) extends Msg
final case class Data(stream: Int, payload: ByteVector) extends Msg
case object Close extends Msg
final case class Frame(header: Header, msg: Msg)
object Wire:
// 8 bits of version, 2 flag bits, 6 reserved bits that must be written as zero
val header: Codec[Header] =
((uint8 :: bool :: bool) <~ constant(bin"000000")).as[Header]
val ping: Codec[Ping] = uint32.xmap(Ping(_), _.id) // uint32 is a Long
val data: Codec[Data] = (uint16 :: variableSizeBytes(uint16, bytes)).as[Data]
val close: Codec[Close.type] = provide(Close)
val msg: Codec[Msg] =
discriminated[Msg].by(uint8)
.typecase(1, ping)
.typecase(2, data)
.typecase(3, close)
// magic, header, then the whole message length-prefixed so unknown types can be skipped
val frame: Codec[Frame] =
(constant(hex"CAFE") ~> (header :: variableSizeBytes(uint16, msg))).as[Frame]
.withContext("frame")Read it against the specification and each line matches a sentence. <~ keeps the value on its left and discards a Unit on its right; ~> does the reverse, which is how the magic number and reserved bits stay out of the case classes. xmap converts between the decoded Long and Ping for a single-field class.
discriminated[Msg].by(uint8) builds a codec for the sealed trait: on encode it finds the first typecase whose type matches the value and writes its tag, on decode it reads the tag and picks the codec. An unknown tag is a decode failure, not a crash. Wrapping the message in variableSizeBytes(uint16, msg) is a deliberate design choice: because the length is known before the tag is read, a newer peer can add type 4 and an older reader can skip the frame instead of losing its place in the stream.
When a later field's shape depends on an earlier value, use flatPrepend: it decodes the first value and then calls your function to choose the codec for the rest. That is the general form behind count-then-items and version-dependent layouts. Prefer the specific combinators when they fit, because a dependent codec is harder to read and to test.
val f = Frame(Header(1, ack = true, urgent = false), Data(7, hex"68656c6c6f"))
val wire: BitVector = Wire.frame.encode(f).require // throws on failure: tests only
// cafe 01 80 000a 02 0007 0005 68656c6c6f
Wire.frame.decode(wire) match
case Attempt.Successful(DecodeResult(frame, rest)) if rest.isEmpty => handle(frame)
case Attempt.Successful(DecodeResult(_, rest)) =>
log.warn(s"${rest.size} trailing bits after frame") // protocol bug or framing slip
case Attempt.Failure(err) =>
log.warn(s"bad frame: ${err.messageWithContext}") // e.g. frame/...: causeThe encoded bytes follow the specification directly: cafe, version 01, 80 for ack set and urgent clear, a length of ten bytes, tag 02, stream 0007, the payload length 0005 and the five bytes of "hello". .withContext("frame") prefixes failures with a path, so an error reads like "frame/...: cause" instead of a bare message, which matters when you are debugging a capture at 3 a.m. Always inspect the remainder; a successful decode that leaves bits over usually means the framing is wrong. decodeValue returns only the value, so use it only where the remainder genuinely does not matter.
Streaming with fs2-scodec
Real inputs arrive in chunks that do not respect frame boundaries: a network read might end half way through a length prefix. The fs2-scodec module, formerly the separate scodec-stream project, bridges codecs and fs2 streams. StreamDecoder.many(codec) emits each value as soon as it is decoded, StreamDecoder.once decodes a single value, and toPipeByte turns a decoder into a pipe from bytes to values. When the codec reports insufficient bits, the decoder waits for the next chunk instead of failing.
// libraryDependencies += "co.fs2" %% "fs2-scodec" % "<fs2 3.x version>"
import cats.effect.{IO, IOApp}
import fs2.Stream
import fs2.interop.scodec.*
import fs2.io.file.{Files, Path}
object Replay extends IOApp.Simple:
val frames: StreamDecoder[Frame] = StreamDecoder.many(Wire.frame)
def run: IO[Unit] =
Files[IO].readAll(Path("capture.bin"))
.through(frames.toPipeByte) // buffers across chunk boundaries
.evalMap(f => IO.println(f.msg))
.compile
.drainMemory stays bounded by the largest frame rather than the file size. Two cautions apply. A corrupted length prefix can claim 65,535 bytes and make the decoder wait for data that will never come, so put timeouts on network streams. And a hard decode failure ends the stream; if you must resynchronize after garbage, scan for the magic bytes yourself before handing data to the codec. For fs2 concepts such as pipes and chunking see fs2 streams, and for the effect type that runs it see Cats Effect.
Testing codecs
Codecs are pure functions, so they are among the easiest code to test well. Use two kinds of test. Round-trip properties check that decoding an encoded value returns it with an empty remainder, over generated values. Golden vectors check against bytes produced by the other side: a capture from the device, the example in the RFC, or the output of the reference implementation. Round trips alone can pass while both directions share the same bug, such as the wrong endianness; golden vectors catch that.
// Round trip: decode(encode(x)) == x, over generated values
forAll(genFrame) { f =>
val bits = Wire.frame.encode(f).require
assertEquals(Wire.frame.decode(bits).require, DecodeResult(f, BitVector.empty))
}
// Golden vector: bytes captured from the other implementation or the spec
test("data frame matches spec example") {
val golden = hex"cafe0180000a020007000568656c6c6f".bits
assertEquals(Wire.frame.decodeValue(golden).require.msg, Data(7, hex"68656c6c6f"))
}Generators must respect the field ranges, or the round-trip test fails on encode rather than finding real bugs. See testing in Scala for property-based testing setup.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Plausible but wrong numbers | Endianness mismatch (uint16 vs uint16L) | Golden vectors from the peer; name little-endian fields explicitly |
| Everything after a flag field is garbage | bool is one bit; padding not modelled | Model reserved bits with constant or ignore |
| String field swallows the rest of the frame | Unbounded utf8 or bytes | Wrap in variableSizeBytes or fixedSizeBytes |
| Decode succeeds with bits left over | Framing wrong or a field missing | Assert rest.isEmpty; fail loudly |
| Old readers break on new message types | No length before the tag | Length-prefix the message so unknown types can be skipped |
| Stream stalls | Corrupt length prefix waiting for data | Timeouts; validate lengths against a maximum |
| Encode fails in production | Value out of range for the field | Validate at construction; use refined or opaque types |
Trade-offs and when not to use scodec
scodec's strengths are fidelity and clarity: the codec reads like the specification, both directions come from one definition so they cannot drift, and errors are values with context. The cost is throughput. Bit vectors and small allocations per field are slower than a hand-written ByteBuffer parser, so for a hot path handling millions of messages per second, benchmark first and consider hand-coding the few hottest message types while keeping scodec for the rest and for tests.
If you control both ends and simply need to exchange data, use a schema-based format such as Protocol Buffers or Avro instead; they give you evolution rules and cross-language tooling. derives Codec on a case class is convenient for internal formats, but for a wire format written by someone else, spell out every field codec so the byte layout is visible and reviewable. Constrained field types, as discussed in Scala 3 opaque types, pair well with codecs because range checks move to construction time.
What to do next
- Pick one binary format you already parse by hand and write its header as a scodec codec.
- Collect three golden byte captures from the real peer and turn them into tests before refactoring.
- Add a round-trip property test with generators that respect every field's range.
- Length-prefix any tagged union you design, so old readers can skip unknown types.
- Add
withContextat each layer and logmessageWithContexton failures. - Stream large inputs with
StreamDecoder.manyandtoPipeByte, with timeouts on network sources. - Benchmark the hot path; hand-code only what the numbers say you must.