Skip to content

Behavioral versioning

Here is a problem version control has never had a real answer for: when your software is an AI system, the source code is only part of what determines how it behaves. The same code, pointed at a different model, or given a different system prompt, or handed a different tool manifest, behaves like a different product. Git versions the code and is blind to the rest.

TOVIO adds a concept with no Git equivalent: behavioral versioning — versioning the full behavioral surface of a deployed AI system, diffable independently of the code it ships in.

Phase note

Phase 4 behavioral snapshot, list, diff, and rollback are implemented (tovio behavioral, Advanced tier). diff reports which of the six recorded fields changed and attaches an advisory risk note (none / low / medium / high). The shipped snapshot is the flattened v1 form — one hash per descriptor plus a policy-version label; the richer nested surface in the data model (per-prompt hashes, eval results) is the specified target, not what is recorded today.

Code is not the whole behavior

For an AI feature, the things that change how it behaves live in several places at once:

  • the model and its version (or hash),
  • the prompts and templates that steer it,
  • the tools it can call — its tool manifest,
  • its memory and retrieval configuration,
  • its guardrails, and the eval results that say whether it's still good enough to ship.

Change any one of these with no change to a single line of code, and the system's behavior can shift dramatically. A model provider silently updates a checkpoint; someone tweaks a system prompt; a tool is added to the manifest. In a Git world, your commit history shows nothing — yet the product your users experience is different. The thing that actually changed is invisible to the system that's supposed to track changes.

The behavioral snapshot

TOVIO's answer is a behavioral snapshot: a versioned object that captures that surface — a hash of the model identifier, of the prompt-template surface, of the tool manifest, of the memory configuration and of the retrieval configuration, plus the guardrail/policy version label — and links it to the commit that implements it. (Eval results are part of the specified target shape; today's snapshot does not record them.)

Mental model

A commit, but for behavior instead of files. The same way a code commit pins an exact snapshot of your source, a behavioral snapshot pins an exact snapshot of how the AI system behaves — and stores it as a first-class, content-addressed object alongside the code.

Because it's a real versioned object, behavior gets the things versioned things get:

  • tovio behavioral snapshot — hash the supplied descriptors into a behavioral object and link it to the current commit (every descriptor is optional; a partial snapshot is well-defined).
  • tovio behavioral list — list the snapshots recorded across reachable history.
  • tovio behavioral diff <a> <b> — compare two commits' behavioral surfaces and see which fields changed — model, prompts, tools, memory, retrieval, or policy — even when the code diff is empty.
  • tovio behavioral rollback <commit> — re-record an earlier commit's behavioral surface against the current commit, without touching the tree — the way you'd revert code.

Why diffing behavior matters

The diff is the payoff. "The code didn't change but the behavior did" stops being an invisible, un-investigable mystery and becomes a line in a diff:

$ tovio behavioral diff <old-commit> <new-commit>
Behavioral fields changed: model, tools
Risk: high — model or policy changed; re-run safety and acceptance evaluations before rollout

The risk note is deterministic and advisory: a model or policy change is high, a tools/memory/retrieval change is medium, a prompt-only change is low, and a side with no recorded snapshot is reported as high rather than silently optimistic. It never blocks an operation.

Now the questions you could never answer have answers — once you can name the two commits, because a snapshot records the behavioral surface and not how it scored. Why did quality regress last Tuesday? Because the model hash moved. Our evaluation run passed on this commit and failed on that one — what differs? The tool manifest, while the prompt hash is identical. Can we get the good behavior back? Roll back to the commit whose snapshot had it.

How it relates to agents and provenance

Behavioral versioning and agent provenance are complementary, and it's worth keeping them distinct:

  • Provenance records, on each commit, which model and prompt an agent used to produce that change — a property of an action that already happened.
  • Behavioral versioning versions the AI system's behavioral surface as a thing in its own right, so you can snapshot, diff, and roll it back over time — a property of the system, tracked longitudinally.

Together they give you both the per-action record (provenance) and the over-time history (behavioral versions) of the AI in your product.

Where to go next

Last reviewed September 9, 2026

Suggest an improvement to this page Not for security reports — see disclosure