VAiDR · Verified AI Decision Records

The evidence
layer for AI.

Your AI can only be as valuable as what you can prove about it. VAiDR is the proof: client-controlled records of what your agents actually did, built to be handed to someone who does not trust you.

  • Your policy. Your rules.
  • No trust in Stones AI required.
  • Your data never leaves your VPC.

A log is not a defensible record.

Enterprises are moving from small AI pilots to hundreds of agents taking real action at machine speed. When a regulator, court, or insurer asks what your AI did, most teams have a log: plain text rows, editable by anyone with access, rotated on a retention schedule.

Any system can show that an agent acted. Few can prove what it did, who authorized it, and that the record has not been altered since. Here is what is missing when the subpoena arrives:

  • No identity binding to the agent and human that acted
  • No independent proof the record existed at that moment
  • No tamper evidence if a row is changed later

With VAiDR each of those becomes a property of the record itself, fixed at the moment the decision is made rather than assembled afterward. A record assembled after the fact inherits the trustworthiness of whatever assembled it. One captured at the moment of decision does not.

Every agent action becomes a signed, witnessed record.

Step 01

Observe

VAiDR sits where your agents act, as an inline control plane over every model and tool call. It captures which agent acted, which human authorized it, which policy version was in force, and which model ran. Payload content is hashed, not shipped.

Step 02

Govern

You author the policy, not us. It runs in the request path, before the model is called. Block the action, hold it for a person, allow it with a flag, or allow it. Each rule declares whether it fails closed or fails open when a check cannot run, and that choice is part of the auditable policy.

Step 03

Attest

Every decision becomes a record signed with a hybrid Ed25519 and ML-DSA-65 pair, then witnessed in a public append only log. Years later, someone holding the record and your public key can check it without contacting you or us.

Control before the action, not an alert after.

Monitoring observes, it never intervenes. By the time an alert fires, the data has left, the claim is denied, the funds have moved. For agents acting at machine speed, the only control that matters happens before the action, and every enforcement decision is itself recorded and signed.

Block

The request is rejected before the model is contacted. A signed record of the block is still written.

Hold for a human

The request pauses for operator approval. A person approves or rejects before anything proceeds.

Allow with flag

The request proceeds and the record carries the rule that matched and what it found.

Allow

The request proceeds normally and a standard signed record is written.

You author the policy. We do not decide what is safe for you. Rules are plain, auditor readable conditions on prompts, tools, models, and agents. Fail closed is the default for high stakes rules: if the check cannot run, the request does not either.

Your data never leaves your VPC. What leaves is proof.

A record has two audiences and they are owed different things. You need to see what happened, the exchange included. Everyone else needs to be convinced it happened, without seeing any of it. One record does both because proof and disclosure are separable.

Inside your environment, where it never left.

  • The full decision record, including the exchange
  • Prompts, inputs, and model outputs
  • Documents and data the agent touched
  • Every enforcement decision and the rule behind it

Enough to prove it, nothing to read.

  • A canonical SHA-256 hash of the request and response
  • Ed25519 and ML-DSA-65 signatures
  • Your published key fingerprint and a timestamp
  • Agent identity, policy version, enforcement outcome
  • Token counts, model name, latency

If Stones AI were breached tomorrow, none of your confidential data would be in it, because none of it was ever there.

Hash only witnessing is the default. Full body ingest exists as an explicit opt in, for teams who want it and can say why. The signing key is generated and held by you. Stones AI operates one of the two witnesses, which puts us in the chain, and the mitigation is that nobody has to trust us: the other witness is public, and disagreement between the two is itself the tamper signal.

Anyone you hand a record to can check it.

Not us on their behalf. Them, with a public key and a command, without contacting you or Stones AI. That is a stronger claim than saying we verify it, and it is the whole reason the instrument works.

The verifier reduces the record to canonical bytes, recomputes the hash, checks both signatures, and confirms the entry in the public log. It refuses to read a public key out of the record it is checking, because that would be circular. Every trust anchor comes from outside the record.

$ vadr-verify --pubkey ./customer-acme.pem record.json

[✓] Canonical form derived
[✓] Hash chain consistency
[✓] Ed25519 signature
[✓] ML-DSA-65 signature
[✓] Rekor entry exists
[✓] Witnessed log inclusion proof

PASS — record verifies end-to-end

What a pass proves is narrow and worth stating precisely: this record is byte identical to what was signed at decision time, and it existed then. We make no claim that its contents are accurate. Identity and context fields are asserted by your systems, and the record proves they have not changed since.

Two ways in. Neither rewrites your app.

Inline proxy

Point your provider base URL at the proxy. The calling application needs no SDK, no code change, and no knowledge that VAiDR is there. It does add a network hop and a failure domain, which is the honest tradeoff against the SDK.

Benchmarked at 1.79ms p50 added per call and 805 requests per second on one worker, measured 2026-08-22 against a mock upstream. Your numbers will differ against a real provider. We publish the conditions because a number without them is an adjective.

Python SDK

In process, when you want the capture inside your own code path. No extra hop and no extra failure domain, with signing at roughly 0.4ms off the caller's thread.

Anthropic Messages, OpenAI Chat Completions, and Azure OpenAI. Bedrock and Vertex are not implemented yet, and we would rather tell you that here than in week two of a pilot.

Instrumented actions produce verifiable evidence. Uninstrumented actions produce a legible gap. We do not claim to capture every decision your AI makes, and any vendor who does is describing a deployment you do not have.

Built for businesses that take risk seriously.

Financial services, insurance, legal, healthcare. The exposure is sharpest where AI actions have to be answered for. What decides whether you need this is less your industry than what your AI is actually doing.

Agents that act on their own

Your agent calls tools, moves data, and finishes tasks without a person watching each step. Every one of those actions is yours. The record holds what it did, under which rule, and who set that rule.

Agents that talk to your customers

What your agent tells a customer is a promise your company made. When someone disputes what they were told, you can open the exchange rather than reconstruct it, and prove to the other side that what you opened is what was said.

Teams using models on regulated work

Your people put model output into filings, advice, and reports that carry your name. The record holds which model produced it, what went in, and who approved it before it went out.

AI that decides about people

Lending, claims, eligibility, hiring. When a decision goes against someone, you may have to explain it. The record holds the inputs, the policy version in force at the time, and the outcome, so the review judges the decision against the rule that actually applied.

Three stages, each crediting the next.

The pilot is a deposit, not a purchase. What you pay at each stage sits inside the next number rather than on top of it.

  1. 01

    Pilot

    Four weeks

    Two hours to scope, one week to integrate, one month of records accumulating against a policy you wrote. It runs in observe mode: your policy evaluates every decision and each record carries the rule that matched, without VAiDR standing in front of production.

  2. 02

    Design partner

    Six months

    You pick the workflows and shape the roadmap. The test is a reconciliation: calls we recorded against calls your provider billed for that workload, one to one, against a number we do not produce. Anything on their side that we cannot account for is traffic that went around the control, and we name it.

  3. 03

    Client

    Annual

    Nothing about the records changes. Same signed artifacts, same verifier, same independence. The relationship stops being about proving the thing works and starts being about running it.

We will also tell you when to stop. If a verification failure goes unexplained, or we cannot account for an in scope path, we did not deliver, and you should walk.

V1 alpha, in pre-production hardening.

Built and tested end to end, with 1,004 tests running in CI. GA target is Q1 2027, which is when SEC 17a-4 grade production retention lands. Design partner pilots run before that, and design partners shape what ships.

There has been no third-party security audit yet, signing keys live in a file rather than an HSM, and record bodies are not WORM stored by default. Those are on the trust page with the rest of the open register, because you would find out eventually and it is better if you find out from us.

Early enough that your input shapes what this becomes.

If you are deploying AI agents in a regulated environment and cannot yet answer for what they did, we want to talk. Bring one workflow and the policy you wish were running on it.