Skip to content

Worked examples

Three reports you can open, read and check yourself.

These are three examples: outcomes, obligations, and agentic commerce. Each report is a single self-contained file. The sealed records are inside it, and the checks run in your browser with no network.

The lead agentic-commerce example is paid inference on mesh-llm. It is an architecture page rather than a report file. So is the NANDA Town deal room. The agent-to-agent negotiation is an independent build you run yourself. For how a judged row is made, see how judging works.

Examples 1 and 2 share data: both are built over the same tau2 airline agent sessions. Example 3 does not. It is a different scenario, a consumer-credit mapping, and it has nothing to do with the airline sessions.

Each report ends with the same stamp, “Countersigned: none”. That stamp is empty on purpose, because no second party has signed these reports.

Example 1 · outcomes

Did the agent do the job, as a buyer would ask it?

What it showsOutcome judgments in a buyer's words, over real tau2 airline transcripts: N of M per outcome, with each row's evidence class shown (fact, judged, or not yet showable).

What it doesn'tIt says nothing about EU AI Act or other obligations; that is example 2. A “pass” is a judged verdict against a pinned rubric and nothing more.

Report dated 16 Sept 2026 · 3.9 MB file (about 0.5 MB to download, compressed) · judged against the Outcomes pack

Verify it yourself: the report's last section re-runs ten checks on the sealed records inside the file, in your browser, offline. Change one byte of the file and the result turns to “Bundle verification failed”.

Illustrative · simulated figures

What a month looks like, illustrative

What it showsThe format we deliver: outcomes, daily AI grading, weekly blind human check, drill-downs, laid out over one month for an airline support agent.

What it doesn'tThe volumes and the weekly human checks are simulated to show the format. It is not a customer's report, no human agreement figure has been measured, and it is not sealed. For the sealed report, open example 1.

Format sample, September 2026 · the days marked REAL link to real tau2 airline test conversations; everything else is simulated

Example 2 · obligations · EU AI Act

Which EU AI Act obligations can this record even speak to?

What it showsAll 37 rows of the EU AI Act Obligations pack, run over the same airline sessions as example 1. 34 rows read “not present” because a generic airline transcript does not carry the field the obligation needs. 3 judged rows read “not checked” because no judge was run.

What it doesn'tNo row reads established or failed. This is a scan of what the record can support, not a determination under the Act, and it is not legal advice.

Report dated 16 Sept 2026 · 0.2 MB file

Verify it yourself: the report's last section re-runs the same ten checks, in your browser, offline. One changed byte fails them.

Example 3 · agentic commerce · consumer credit

An agent that extends credit: what can a regulator check?

What it showsQuadX AI's FCA obligations mapping (22 July 2026) for a buy-now-pay-later credit agent. It has 11 rows, and each is graded by who can check it: only the operator, a party to the transaction, or an independent third party. Rows 4 and 5 are backed by a real digest that two parties recomputed independently.

What it doesn'tThe other 9 rows are mapping only; they have not been run against any data. This is different data from examples 1 and 2, and it is not legal advice.

QuadX, on its own rows: “'Good outcome' remains a supervisory quality judgment; the record proves the process, not the quality of the result.”

The FCA clause mapping is QuadX AI's FCA Obligations Mapping v1.0 (22 July 2026): QuadX AI reviewed and corrected the Action State Group strawman and verified the FCA clause references against the FCA Handbook as at July 2026.

Mapping by QuadX AI · report dated 16 Sept 2026 · 0.2 MB file

The A2A + AP2 example on the open standard's site shows the payment side of agentic commerce: the called agent seals a record of each payment action, bound to the mandate that authorised it.

Verify it yourself: the report's last section re-runs the same ten checks, in your browser, offline. One changed byte fails them.

Agentic commerce · Project NANDA

Two agents agree a deal: can a stranger check the terms?

What it showsA buyer's agent negotiates with two competing seller agents on NANDA Town. A person clears the acceptance, and the agreed terms are sealed into one record per side, over the same digest. The records are designed so that anyone can recompute that digest and check both, once the service is public.

What it doesn'tNo payment moves, because the example stops at the agreement. The room seals both records; each side does not seal its own. The deal-room service itself is not published, and the page says which parts are merged and which are not yet public.

Architecture page with a flow table and a claims table · built during Project NANDA's July 2026 hackathon

Verify it yourself: every claim on the page says what it rests on, and links the NANDA Town pull request or commit where the source is public.

Agentic commerce · agent to agent · independent build

One buyer, two suppliers, three rounds: who agreed to what?

What it showsA buyer agent haggles with two competing supplier agents over three rounds and narrates its reasoning before each move. It picks a winner against a private mandate, and turns down the cheaper final offer because that offer breaks a hard constraint. Each side keeps its own signed record, and a stranger who trusts neither can check the whole transaction offline.

What it doesn'tNo payment moves. The agents' reasoning is scripted, with no model call. It is demo code that runs on your machine, not a hosted service or a supported adapter.

Built by Jody (independent agentic engineer) · Apache-2.0 · installs capsule-emit from PyPI

Verify it yourself: bash RUN.sh plays the negotiation in a terminal, and bash RUN.sh --browser opens it as a page that needs no network. The repo also commits the evidence bundle from one run, which you can re-check as it is.

Method · behind every judged row

How is a judged row judged?

What it showsThe rules a judged row follows. The judge is pinned and sealed with each verdict, and the rubric is frozen before the period. Recomputed and judged rows are never blended, several judges can read the same row, and a blind human sample is reported as an agreement figure.

What it doesn'tIt gives no agreement figure for these examples, because none has been measured yet. The outcomes report's human check reads “no human-check data”.

Method page with a claims table · a method, not a report

Which rows a report checks comes from a pack. Read the two packs →