Skip to content

Work with large binaries — no LFS

In TOVIO, a 4 GB model checkpoint is just a file. You commit it the same way you commit a one-line text change — no git lfs track, no pointer files, no smudge or clean filters, no separate store to host and authenticate against. This is binary parity: every file type is treated equally.

Local commit and networked sync both work

Content-defined chunking and dedup are part of the Phase 0 offline core — you can commit huge binaries right now and get incremental storage locally. Sharing them with a team over the wire rides on networked sync, which is built (over TLS, with sparse path-scope). The chunk model is the same either way. Production binaries passed the distinct-Docker-host exercise under a real operator CA; public-release, broader HA/load, hosted-parity, and external evidence remain Phase 5 gates.

Why Git made this painful

Git stores each version of a file as a whole object, and its packfile delta compression rarely finds much to share between two versions of a compressed or opaque binary format. A one-byte change to a large asset effectively re-stores the whole thing, and history balloons — so the community bolted on Git LFS: a pointer file in the repo, the real bytes in an external store, and a whole new set of credentials, endpoints, and failure modes. TOVIO replaces that bolt-on with chunking built into the object store.

How it works: FastCDC chunking

Files at or above a threshold (default 1 MiB) are split with FastCDC, a content-defined chunking algorithm. Instead of fixed-size blocks, FastCDC places chunk boundaries based on the content, so inserting a byte near the front doesn't reshuffle every chunk after it.

  • Each chunk becomes a content-addressed chunk object.
  • A chunk-manifest lists the ordered chunk addresses and reconstructs the file. The file's identity is its manifest address, not a hash of the reassembled bytes.
  • Identical chunks dedupe automatically across files, versions, and lanes — the object store treats a put of an address it already holds as a no-op.

The practical payoff: edit a large asset, and TOVIO re-stores only the chunks that actually changed. Re-commit the same asset on another lane, and it costs nothing — the chunks already exist. (This holds for files stored in the clear; protected files behave differently.)

Just commit it

There's no special command and no tracking step. tovio commit handles a 4 GB binary and a one-line text edit through the same path, and its output is the same either way — a file count, not a chunk report:

$ tovio commit -m "Add trained model checkpoint"
✓ committed chg:9d1c44  blake3:1f4c…
  1 file(s) · new change chg:9d1c45 · undo with `tovio undo`

IDs are illustrative. The chunking is invisible at the command line; to see it, read the file's chunk-manifest.

A small edit re-stores only what changed

Because boundaries are content-defined, a localized edit touches a handful of chunks, not the whole file. The next commit stores just those — the rest of the file is shared with the previous version, and the new manifest simply re-references the unchanged chunk addresses.

That is a storage property, not a timing claim. Public macro-performance numbers remain gated on reviewed pinned-host calibration and accepted Phase 5 performance evidence.

The per-file ceiling

A single file's reconstructed (plaintext) size is capped at 16 GiB. Reads refuse to reassemble a chunk-manifest that declares more than that, before any reassembly runs — a manifest's declared total is chosen by whoever wrote it, so an untrusted one could otherwise claim an enormous size and force an out-of-memory read. The Forge's per-object replication ceiling is the same 16 GiB, so a file that reads locally is a file that can replicate. Split anything larger into parts.

Tuning chunk sizes (advanced)

The defaults suit most repositories. They are part of the repository's format identity, though — changing them changes every boundary and therefore every address, so it is a deliberate, repo-level decision (tovio policy chunking, an authenticated write that applies to future snapshots), never a per-commit flag. --reset returns to the defaults.

Parameter Flag Default Range
min --min 32 KiB 4 KiB – 256 KiB
avg --average 64 KiB 16 KiB – 1 MiB
max --max 256 KiB 32 KiB – 4 MiB
threshold --threshold 1 MiB at least 64 KiB

The three windows also have to stay strictly ordered — min < average < max. Equal or inverted windows are refused before the chunker sees them, because they leave no useful normalization region.

Smaller chunks dedupe more finely but cost more manifest overhead; larger chunks do the reverse. Leave these alone unless you have a measured reason to change them.

Encrypted large files chunk, but do not dedupe

For a policy-protected large file, chunking happens before encryption: boundaries are found on the plaintext, then each chunk is sealed independently as a policy-chunk. The chunk-manifest itself stays clear, because it lists addresses rather than content.

Be clear-eyed about what that buys you, though. Sealing is deliberately randomized — a fresh key per file and a distinct nonce per chunk — so two chunks with identical plaintext produce different ciphertext and different addresses, even under the same policy and even inside the same file. Protected chunks therefore do not dedupe in practice, and editing a protected binary re-seals the whole file rather than just the changed region. The incremental-storage payoff above is a property of clear files.

Randomization is the point, not an oversight: dedupe across encrypted content would need deterministic encryption, which leaks plaintext equality across access domains. TOVIO takes the storage cost instead.

An unchanged protected file is not re-sealed

The snapshotter reuses the prior sealed object for a protected file whose plaintext and resolved recipient set are both unchanged. Without that, every commit would give the file a fresh address and a three-way merge would read it as a false conflict. So committing repeatedly without touching a protected binary costs nothing; it's the edits that are expensive.

Sparse clones keep big assets off machines that don't need them

Over the wire, a sparse clone lets a teammate pull only the subtree they work in — the large assets under a path they never touch simply don't materialize on their machine. Binary parity doesn't mean everyone carries every gigabyte.

When to lock instead of merge

Chunking gives you efficient storage for binaries, but it doesn't make a binary mergeable — two divergent edits of a checkpoint still can't be reconciled. If two people might edit the same binary at once, reach for a lock to serialize that one path. Storage parity and lock-based coordination are complementary tools.

Recap

  • Commit binaries directly — no LFS, no pointers, no filters, no external store.
  • FastCDC chunks files ≥ 1 MiB by content; a small edit re-stores only changed chunks.
  • Identical chunks dedupe across files, versions, and lanes automatically.
  • Protected files chunk before encryption but do not dedupe — sealing is randomized per file and per chunk. Dedup and incremental re-storage are clear-file properties.
  • tovio policy chunking tunes the windows for future snapshots; they must satisfy min < average < max.
  • A single file reconstructs up to 16 GiB; larger files are refused on read and on replication.
  • For concurrent edits of an unmergeable binary, add a lock.

Where to next

Last reviewed September 9, 2026

Suggest an improvement to this page Not for security reports — see disclosure