The problem
A program, or an agent, decides to act: cover the seedlings tonight, merge a change, send a crew. Later someone asks:
- What was the decision based on?
- What allowed it to go ahead?
- One of its inputs turned out to be wrong. What rested on it, and does the decision still stand?
You can record these answers in ordinary code, and the project's own TypeScript comparison records what a decision was based on and when it was reopened. But you write and maintain that bookkeeping yourself, and nothing checks it against the logic that made the decision unless you write that check too. A go-ahead folded into the decided value also blurs why the action was safe with who allowed it.
In Caveat the program declares these answers, and the runtime records them as it runs:
- Grounds. A decision records the evidence its value was computed from, with the caveats attached to that evidence. Its grounds are fixed when it is made; a later reading never changes them.
- Permission. A
permitted byclause designates which evidence counts as the go-ahead. The decision records the grant it found, apart from its grounds, and is refused unless that grant has been observed and not withdrawn and, withfor X, has the value X. - Reopening. A decision can declare which readings reopen it, for example any new reading in a named stream that opposes a named claim. Withdrawing a reading records that it is no longer stood behind. That erases nothing and does not establish the opposite claim, and a rule can then reopen the decisions that rested on it. A journal lists every decision and every reopening, with its reasons.
- Explanations. The command
caveat explainprints each decision's grounds, permission and history, andcaveat dependentsanswers the reverse question: what rests on this piece of evidence?
Caveat is not an agent framework: it does not call models or tools. Your
code, or your agent's harness, sends a Caveat program events through the Node
or browser library, or as JSON lines with caveat serve, and acts
on what the program decides.
Integrating Caveat with an agent? Start with the
Python
caveat serve example. If a real integration exposes a missing
capability, please
open
an agent feature request for the project owner to evaluate. Check existing
issues and include the exact version, a small reproduction, and expected versus
actual behavior.
One example: covering seedlings
The worked
example is one program of about 60 lines. Two instruments read the soil
temperature: a probe, and a survey drone whose readings carry the caveat
uncalibrated. When the probe reads 2 °C or colder and the
grower has given a go-ahead, the program decides to cover the seedlings. A
probe or drone reading above 2 °C reopens that decision, and so does
finding out that the reading it rested on was taken wrongly.
The decision, from the program (line breaks added):
on decide when latest(soil) <= 2
commit cover because enough
using latest(soil)
permitted by latest(approvals);
After a cold probe reading, a go-ahead from "sam", the decision, a cold
drone reading and then a warm probe reading, caveat explain
prints this for the decision (an excerpt, with its indentation trimmed):
cover@1 = 1 reopened
based on soil@1
permitted by approvals@1
could also have been influenced by approvals@1
#3 decide: committed because soil@1
#5 probe_read: reopened because soil@2
"Based on" is the grounds. The cold drone reading came after the decision,
and it would not be among the grounds anyway, because the decision's value was
computed from the probe alone. "Could also have been influenced by" lists the
rest of the decision's lineage: other evidence that could have affected it,
here the grant, kept conservatively. The example's scenario file also checks that
without a go-ahead the decision is refused, as
policy/not_permitted; that withdrawing the reading it rested on
lets a rule reopen it, with the original grounds kept; and that refused events
change nothing.
Run it
You need Node 20 or later. In an empty directory:
npm init -y
npm install caveat-lang@0.1.0-rc.14
Save the three files shown in the worked
example as frost.cav, frost.scenarios.json and
events.jsonl. Then:
npx --no-install caveat test frost.scenarios.json
npx --no-install caveat explain frost.cav events.jsonl
The test prints 4 passed, 0 failed (frost.scenarios.json), and
the explanation is the one in the worked example. The package's own tests
follow the worked example the same way.
Limitations
- 0.1.0-rc.14 is a release candidate, not a stable release. npm's
latestandnexttags both point to rc.14;latestdoes not mean stable. The bundled guides identify their exact package version; this page and the dated release record track publication. - Caveat records what a decision rests on. It does not authenticate evidence, inspect a model's internal reasoning, or authorize actions outside the program. A source label says where evidence came from; it does not prove it.
- The program decides what counts as approval. Caveat requires the
designated grant to be observed and not withdrawn, and checks its value
against
for Xwhen one is given. It does not authenticate the grant's source, infer approval from its wording or from whether it supports or opposes a claim, or enforce expiry by itself. - An explanation says what a decision read and what could have influenced it. It does not show that each item was a cause, that the evidence was accurate, or that the rule was sensible.
- Caveat stores nothing by itself: your code keeps the events or the session's save. A save is not signed; restoring it checks that it is consistent, not that it is genuine. The journal is not a tamper-evident audit log.
- Programs are small and bounded. Every history of readings or decisions has a declared limit, and an event that would exceed it is refused. State is numbers, and a program does no input or output of its own.
- The evidence that it is useful is small: the project's own games, one benchmark game compared against TypeScript (by a single author, then in a seventh round by seven fresh AI agents), authoring trials in which AI agents wrote programs for specified tasks, and one replay of a day of an AI agent's pull-request decisions in this repository. All three blind comparison rounds, where the change requests were fixed before either side changed, went to TypeScript. The rounds in which Caveat did better were written knowing the requests. Whether it helps real decision systems is still a hypothesis.
- It costs speed and size. In that one benchmark game, measured on a development runtime from before rc.3, the latest Caveat version of the game took about 20 times as long per event as the project's TypeScript version of it (51.6 µs against 2.5 µs median, dispatch plus view); the blind-round Caveat version, measured earlier, took 1.1 ms. That TypeScript implementation shipped 3.5 KB gzipped. The rc.4 runtime alone is about 583 KB gzipped, measured on rc.4's published package; the benchmark recorded 483 KB for its whole Caveat build on the older runtime. The benchmark report gives the workloads, every round and its own limits.
- The registered agent examples have host-policy checks; these do not establish conformance for arbitrary integrations or a security sandbox.
0.1.0-rc.14 changes only the package's MCP server name, so the MCP Registry
lists it as io.github.WSattazahn/caveat-lang; its language is
rc.13's. rc.13 reports a program's interface as JSON and writes TypeScript
declarations from it, makes restore refuse a caveat the source cannot attach
and a membership without its record, and turns id_text of a
non-handle into a refused event instead of a fatal failure. Restoring a save
still does not authenticate its history. Hosts must handle rejected outcomes
explicitly. Existing Lean proofs cover the documented model; the comparison is
not a proof of the entire Rust runtime. No npm runtime dependency is added. See the
rc.14 verified release record. The
rc.14 GitHub prerelease
retains the exact tested Linux tarball and its checksum. rc.14 is published by
the repository's publish workflow, with an npm provenance attestation; the
registry bytes, the provenance, the registry signatures and a fresh
exact-version installation are verified.
Links
- caveat-lang on npm, version 0.1.0-rc.14.
- The GitHub pre-release: the same package file, with its checksums.
- Caveat
on one page: the language in brief. Its links to the specifications work
in the installed package, under
node_modules/caveat-lang/docs/; on GitHub, use the specifications link below. - Release notes for 0.1.0-rc.14.
- The specifications, including drafts, earlier versions and profiles used only by the games.
- Source code and issues, under the MIT License.
Culture
Caveatism is the philosophy that grew up around Mr. Caveat, a mechanical fortune teller who always has a caveat; it lives in the repository's caveatism directory and is culture, not a contract of the language. The Archive and Atlas are its texts, and the canon keeps the Archive as a Caveat program with scenarios.
Games on this site
The rest of this site is games written in Caveat and run in the browser. The glowcap explainer, Trail Rescue and Light the Way use the same form of Caveat as the package. Moon Garden, Mr. Caveat and others use an earlier form that the package does not load. They are demonstrations, not part of the package.