<project_root>/.mareforma/graph.db
(WAL mode, ACID). Schema version: 2.
Claims table
Row-level CHECK: a row cannot say a human validated it without the envelope proving one did. Setting
validated_by or validated_at requires a non-NULL validation_signature. The CHECK is the row-level belt to the trigger’s transition-level suspenders.
Indexes
State-machine triggers
Triggers enforce at the storage layer the rules a claim cannot be edited out of. Defense in depth: a tampered Python interpreter cannot relax these rules.claims_update_status_terminal: retracted is terminal. Any UPDATE that transitions a row out of retracted raises mareforma:state:retracted_is_terminal. To resurrect a withdrawn finding, assert a new claim citing the old via contradicts=[<old_claim_id>].
claims_validation_is_terminal: a recorded validation is terminal. BEFORE UPDATE of validation_signature, validator_keyid, validated_by or validated_at on a row that already carries a validation_signature, any change to one of those four raises mareforma:append_only:validation_is_terminal. The row holds one signed envelope, so overwriting it would erase the first attestation and leave the graph no record that it was ever there. To revise, retract the claim and assert a new one.
claims_signed_fields_no_laundering: append-only over the signed predicate. It lives in _SIGNED_FIELDS_TRIGGER_SQL rather than _SCHEMA_SQL, so every open_db() drops and re-creates it and an existing database gains the current watch list. Refuses any direct-SQL UPDATE that changes a signed-predicate value (text, classification, generated_by, supports_json, contradicts_json, source_name, artifact_hash, ev_*, evidence_json, statement_cid, prev_hash, created_at), the asserter_keyid denormalisation, the observed_grounding verdict, or the predicate_payload the audit path re-checks a GROUNDED verdict against, on a row whose signature_bundle IS NOT NULL. Value-comparison fires only when something actually changed, so multi-column UPDATEs that re-emit unchanged values (e.g. status-only edits via update_claim) pass through. It also refuses to set signature_bundle back to NULL on a signed row, because that one write would clear the guard on this trigger and on claims_signed_no_delete; the Rekor attachment’s non-NULL rewrite stays legal. Raises mareforma:append_only:signed_field_locked.
claims_signed_no_delete: the delete-side twin of claims_signed_fields_no_laundering. BEFORE DELETE on a row whose signature_bundle IS NOT NULL, raises mareforma:append_only:signed_claim_no_delete. A signed claim cannot be wiped from the local graph while its Rekor entry and chain hash persist. Unsigned claims (legacy / no-key mode) carry no cryptographic commitment and stay deletable.
contradiction_invalidates_older: AFTER INSERT on contradiction_verdicts. Sets claims.t_invalid = NEW.created_at on the older of the two referenced claims (lex-smaller claim_id as the deterministic tie-break when timestamps collide), idempotent via WHERE t_invalid IS NULL.
replication_verdicts_append_only + replication_verdicts_no_delete: UPDATE on the signed columns and any DELETE both raise mareforma:append_only:verdict_locked / verdict_delete_blocked. Same for contradiction_verdicts via the symmetric pair.
Valid values
mareforma.schema():
replication_verdicts table
Signed replication verdicts produced by enrolled validators. The OSS core accepts verdicts from any enrolled identity; the predicates that GENERATE verdicts (semantic-cluster, cross-method, hash-match, shared-resolved-upstream) live outside the OSS and callGraph.record_replication_verdict() to write here. Append-only at the SQL layer: UPDATE on signed columns and DELETE are both refused by triggers.
No side effect on the claims it names: the verdict is a signed record of one party corroborating another, and it lifts nothing.
contradiction_verdicts table
Signed contradiction verdicts. Same shape asreplication_verdicts but binds a refutation between two claims. INSERT fires the contradiction_invalidates_older trigger which sets t_invalid on the older referenced claim.
rekor_inclusions table
Sidecar recording every successful Sigstore-Rekor submission, written by_record_rekor_inclusion as step 3 of the Rekor saga. Step 4 (the claims-row UPDATE that attaches the Rekor coords to signature_bundle) reads from this table on retry instead of re-submitting, so a single Rekor submission produces exactly one log entry even when the local UPDATE crashes mid-saga.
Append-only at the SQL layer: both UPDATE and DELETE are refused by triggers, mirroring the verdict-table protections. The saga’s write uses INSERT ON CONFLICT(claim_id) DO NOTHING, so a legitimate retry on the same claim_id is a silent no-op (the original row is preserved) and a SQL-writer cannot launder forged Rekor coords through the recovery path in refresh_unsigned().
validators table
The per-project set of enrolled public keys. Onlyhuman-typed rows can sign off on a claim.
The first key opened against a fresh
graph.db auto-enrolls as the root with a self-signed envelope (BEGIN IMMEDIATE guards against two simultaneous opens both becoming roots). The chain walk enforces a singleton-root invariant: if two rows have keyid == enrolled_by_keyid, neither is trusted. Walk is capped at 64 hops.
Removal is intentionally unsupported currently; validator history is append-only.
project_policy table
A root-signed, single-row (id = 1) declaration of project-wide trust policy. It carries rekor_required (findings must be witnessed by the transparency log before they can converge, backing require_rekor_witnessing) and strict_promotion_required (a converging pair must carry data on both sides, which the open(strict_promotion=True) flag declared before that flag was removed). Both are one-way once declared and bind every writer, not just the handle that declared them. The signed envelope is the authority; the flat columns are a denormalized read cache. The envelope payload carries its own version, so a declaration signed before a flag existed keeps verifying under the field list it was signed with. Extending the policy re-signs the row, which moves created_at, so each flag also carries the time it was first declared and a check that grandfathers earlier claims reads that instead. Restore verifies the envelope against the enrolled root before enforcing.
supports_revision table
A single-row (id = 1) monotonic counter over the claim_supports cache, bumped by every claim insert and every supports-edge change. It lives in graph.db rather than in the cache file, so it commits in the same database as the row it describes. The cache stamps the revision it was built from; a mismatch means the cache missed a mutation (a crash between the two WAL commits, or a writer that did not maintain it) and the cache is rebuilt on the next open. The claim-count check alone cannot see an in-place supports edit, which moves no count. Additive: an existing graph gains the table and its row on the next open, and that open rebuilds the cache once.
claims_fts table
The FTS5 virtual table behind text search. It is independent ofclaims (not content=claims), so storage is the only price of the feature and the sync triggers stay readable. claim_id is UNINDEXED, stored for join-back but not tokenized; the unicode61 tokenizer folds diacritics (remove_diacritics 2), so “gene” matches “géné”.
Three triggers keep it in lockstep with claims, all AFTER the write, so an IntegrityError on the wrapping statement rolls back the claim row and the search index together:
claims_fts_ai: AFTER INSERT, inserts the newclaim_idandtextclaims_fts_ad: AFTER DELETE, deletes the row’s index entryclaims_fts_au: AFTER UPDATE OFtext, rewrites the indexed text.textis a signed field, so this fires only on the unsigned-edit path
Trust layer tables
The trust layer (see Findings) adds seven tables for structured findings. They are additive:CREATE TABLE IF NOT EXISTS, and
adding them moves no version. A finding is an evidence tree
(finding → evidence_lines → contrasts → effect_estimates), anchored to a
content-addressed proposition and a pre-registered prediction, and attested
by an existing signed claim.
propositions table
The content-addressed unit of sameness.content_id (PK) is the answer hash;
frame_id is the question hash.
Indexes:
idx_prop_frame (frame_id), idx_prop_frame_dir (frame_id, direction).
predictions table
The pre-registered plan, bound to one proposition.plan_id (PK) is content-addressed over (content_id, prediction fields), so registering the same plan twice is a no-op.
A registered plan is append-only:
predictions_append_only (BEFORE UPDATE of every immutable column) and predictions_no_delete (BEFORE DELETE) raise mareforma:append_only:prediction_locked / prediction_delete_blocked, so the gap between registration and evidence is a real pre-registration guarantee.
plan_retirements table
A plan written by a release with a wider alpha bound can state a rule no gate can run, and the row above can be neither corrected nor removed.graph.retire_plan(plan_id, alpha=..., reason=...) records the way out: the plan, the plan that supersedes it (the same rule at an alpha the gate can run), and why. The read path then gates that plan’s evidence under the replacement, so the lines count again instead of dropping.
A retirement is append-only like the plan it retires:
plan_retirements_append_only (BEFORE UPDATE) and plan_retirements_no_delete (BEFORE DELETE) raise mareforma:append_only:plan_retirement_locked / plan_retirement_delete_blocked.
None of these columns is signed, so both the read path and restore re-derive the row from the attestation: the claim’s text renders plan, replacement and reason, a row its claim does not render resolves nothing on read and fails the restore with kind='claim_unverified'.
Resolution is reached only from a plan whose own rule cannot be run, but that premise is suppliable: rewriting a live plan’s alpha to a value no gate can discriminate at sends its lines into resolution, and a planted replacement at a stricter alpha re-gates a refutation to NEUTRAL, which is counted rather than skipped. That is why the attestation is checked on the read and not only on restore.
findings table
One attestation plus its computed bearing on a proposition under a plan.evidence_lines table
One line of evidence; a finding may carry several. Independence is counted by pairwise-distinct model, dataset, and signer: a corroborating line counts only when it stands on a distinct model/method as well as a distinct signer and dataset, so a same-model rerun is one line, not two.contrasts table
The comparison a line quantifies (control type only, for now).effect_estimates table
The estimate the gate reads. Minimal metafor-named field set.
The two trust axes are derived on read, not stored.
Status (per content_id,
the answer) is computed from the independent supporting / refuting line counts,
counted by pairwise-distinct model, dataset, and signer (status_policy@v4), so
improving the rule later is a new policy over the same data, not a migration.
question_status (per frame_id, the question: consistent / divided) is
derived alongside it from the same computation. Both are what a reader should
read trust off. The stored ladder that used to sit beside them is gone.
schema_census and schema_guards_seen tables
Two tables that record the state of the schema’s own write guards, because a dropped guard is the one tamper a read cannot infer afterwards. Every trigger in the schema is reconciled againstsqlite_master on every open, and
_ADDITIVE_TABLES_SQL recreates its own triggers on every open as well, so a
guard that was gone is back before the open returns. Nothing in the file
afterwards says it had gone, while the rows it let someone delete stay deleted.
The census runs first, ahead of both repairs, and writes down what was absent.
Ordering is the whole mechanism: run after either repair and it sees a healed
schema and reports clean forever. A read surface consults the record rather than
re-deriving from sqlite_master, and reads the union of every observation, not
the latest: a guard that came back is not a guard that was never gone.
schema_guards_seen is what separates “this graph never had that table” from
“somebody took that table away”. A guard is expected while its table is present,
which keeps a graph written before a table existed off the tamper report, and a
guard this graph has carried stays expected however its table is treated after.
It is written at the end of an open, from the guards actually present, and it
only grows.
Both tables carry no-delete and append-only guards of their own. Once every
guard heals on every open they are the only record that anything happened, and a
store of tamper evidence the tamperer can empty is not a record of anything.
Those guards refuse a DELETE and an UPDATE. They cannot refuse a DROP TABLE.
Dropping schema_census alone is enough: its guards go with it, the next open
rebuilds it empty, and every observation ever recorded is gone while the seen
store and the guarded table sit untouched. No scheme confined to one file the
attacker can write to will change that. What
the census does buy is that any smaller version of the attack is on the record,
including dropping the guarded table by itself, and that the census travels in
the backup, so a graph emptied this way disagrees with a claims.toml that has
to be found and edited as well.
schema_census:
A row is written only when something is missing, and only when the set differs
from the last one recorded, so a long-lived process does not bury the
observation that matters.
observed_at carries no primary key: appending rather
than replacing is what lets the delete guard be absolute.
schema_guards_seen:
grounding_attestations table
What the observer computed, carried inclaims.toml so recovery can be held to
the standard the write path holds.
observed_grounding is the one signal on a claim meant not to be the producer’s
own word. The write path enforces that: a verdict the process’s observer minted
is stored as the observer’s snapshot, and anything else is marked DECLARED with
its GROUNDED claim neutralised to OPAQUE. Restore never passed through that
check. It writes the axis straight out of claims.toml, so a producer could
export a claim, edit GROUNDED into it, re-sign with their own key, restore, and
every read surface rendered the result exactly like an execution mareforma
watched.
Restore cannot re-run the check instead. The register it reads is in-process and
keyed on a receipt digest, so it dies with the process that built it, and a
fresh restore would strip the axis off every honest claim along with the forged
one. So the observer’s word travels in the file.
A row exists only where the observer’s own record was kept. A declared verdict
gets none, and that absence is the signal.
grounding_attestation_state() answers in one word. attested means a row is
present, binds this claim’s current statement, names the axis the claim stores,
and verifies under the asserting key. unattested means no row, which is the
ordinary state for a declared verdict and for any claim written before the table
existed. broken means a row is present and fails one of those, which is a
stronger signal than absence and is never folded into it.
What this buys is parity, not prevention. The observer runs inside the
producer’s process and the producer holds the key, so a producer determined
enough to re-sign a claim can build an attestation too. There is no
cryptographic asymmetry between the producer at write time and the same producer
later. What it ends is the ordinary act, editing the axis and nothing else, the
same lazy path the schema census closes for a dropped trigger. The trust map’s
grounding residual says which state a claim is in rather than implying more.
verdict_chain table
One row per verdict recorded from the version that introduced this table onwards, each carrying the tip of the chain behind it and a signature over that tip made by the verdict’s own issuer. A graph rebuilt byrestore can hold fewer. The backup writer stops at the
first link that does not follow the one before it, so a wrecked graph backs up
as the part of its chain that still holds rather than as the wreckage, and the
file records how many links it left out and why. restore refuses such a file
unless you pass trust_unaccounted_backup=True, and reports the count.
Every claim, validator and verdict in claims.toml carries its own signature,
so nobody can change what a row says. Nothing in the file signs which rows are
in it. Delete a verdict’s entry and restore rebuilds a graph that never had
it, reports clean, and nothing in the file disagrees. This table is what makes
that absence speak.
The issuer signs each link, not the project root. Nothing holds a private key
when the backup is written, and the two verdict paths are the only mutations
that require a signer, so that is where a real key is in hand. The issuer
attests only what an issuer is entitled to attest: that the verdict set behind
their verdict hashed to prev_tip when they issued it.
verify_verdict_chain() recomputes the chain and checks every signature. What a
clean result rules out, stated as narrowly as it holds: no verdict has been
taken out of the middle by anyone holding no enrolled key. Removing a verdict
means removing its link, and the next link then has to be re-signed over the
gap, its tip recomputed, and the verdict it covers made to verify under the key
the link names. An outside attacker with file access and the project operator
are held out by that.
An enrolled peer is not. A verdict’s signed payload carries no issuer, so a
peer can claim a surviving verdict, re-sign it under its own key, and re-sign
the link to match; every check then passes on a chain it just shortened. Closing
that needs the issuer inside the verdict’s signed bytes, which changes bytes an
already-released reader rebuilds. Nor is the issuer of the verdicts held out,
and it cannot be: a key can always restate its own view of its own verdicts.
Two things it does not say. A removed suffix leaves a shorter chain that
verifies, so length is reported by verdict_chain_coverage() rather than
checked. And verdicts recorded before this table existed carry no link, which
the same covered-versus-total pair is what makes visible.
The link binds to the verdict’s signature rather than its row, because the
signature is the one field only the issuer could have produced. There is no
foreign key on verdict_id: verdicts live in two tables, so the reference is
not expressible, and a link left pointing at a verdict that is gone is the
evidence rather than a violation to cascade away.
Schema versioning
The schema version is stored in SQLite’suser_version pragma.
Do not delete
graph.db. It holds the claim chain and every signature, and
claims.toml cannot reconstruct them: the backup carries what each row says, and
the chain is what says which rows were there. A graph this build will not open is
almost always a graph a newer build wrote, and the answer is to upgrade rather
than to start again.
The migration registry carries one route, 1 to 2, so a graph written by an
earlier release migrates on the open that meets it. The schema change and the
version bump commit in one transaction or neither does, so a crash leaves the
graph at the version it started from with every claim, signature and chain link
intact. Nothing on that path tells anyone to delete a file.
The upgrade is one way. A graph at 2 is refused by any release that expects
1, which is every release before this one, so upgrade every machine that
shares a project rather than a subset of them.
Storage
graph.db is stored at <project_root>/.mareforma/graph.db. Created automatically
on first mareforma.open(). The .mareforma/ directory is created if it does not exist.
claims.toml at the project root is a human-readable backup of all claims,
written after every mutation. It is not the source of truth (graph.db is)
but it survives graph.db deletion.
Runtime PRAGMAs
open_db() sets these connection-level PRAGMAs on every open: