Example 1 · outcomes
Did the agent do the job, as a buyer would ask it?
What it showsOutcome judgments in a buyer's words, over real tau2 airline transcripts: N of M per outcome, with each row's evidence class shown (fact, judged, or not yet showable).
What it doesn'tIt says nothing about EU AI Act or other obligations; that is example 2. A “pass” is a judged verdict against a pinned rubric and nothing more.
Report dated 16 Sept 2026 · 3.9 MB file (about 0.5 MB to download, compressed) · judged against the Outcomes pack
Verify it yourself: the report's last section re-runs ten checks on the sealed records inside the file, in your browser, offline. Change one byte of the file and the result turns to “Bundle verification failed”.
Illustrative · simulated figures
What a month looks like, illustrative
What it showsThe format we deliver: outcomes, daily AI grading, weekly blind human check, drill-downs, laid out over one month for an airline support agent.
What it doesn'tThe volumes and the weekly human checks are simulated to show the format. It is not a customer's report, no human agreement figure has been measured, and it is not sealed. For the sealed report, open example 1.
Format sample, September 2026 · the days marked REAL link to real tau2 airline test conversations; everything else is simulated
Example 2 · obligations · EU AI Act
Which EU AI Act obligations can this record even speak to?
What it showsAll 37 rows of the EU AI Act Obligations pack, run over the same airline sessions as example 1. 34 rows read “not present” because a generic airline transcript does not carry the field the obligation needs. 3 judged rows read “not checked” because no judge was run.
What it doesn'tNo row reads established or failed. This is a scan of what the record can support, not a determination under the Act, and it is not legal advice.
Report dated 16 Sept 2026 · 0.2 MB file
Verify it yourself: the report's last section re-runs the same ten checks, in your browser, offline. One changed byte fails them.
Agentic commerce · lead example · mesh-llm
Paid inference between strangers: who paid, and who can show it?
What it showsHow two mesh-llm nodes each keep their own book of one paid exchange: the payer's payment events (merged upstream in #2035), the provider's matching half (proposed in issue #2017), and how the two join on the payment hash both wallets share.
What it doesn'tThat the inference was correct, which upstream does not check. It also doesn't show the swapped halves or the Evidence tab working today: those are not yet published, and the page marks each one.
Architecture page with a diagram, a lifecycle table and a claims table · a design, not a report
Verify it yourself: every claim on the page says what it rests on, and links the mesh-llm pull request, issue or merge commit where the source is public.
Example 3 · agentic commerce · consumer credit
An agent that extends credit: what can a regulator check?
What it showsQuadX AI's FCA obligations mapping (22 July 2026) for a buy-now-pay-later credit agent. It has 11 rows, and each is graded by who can check it: only the operator, a party to the transaction, or an independent third party. Rows 4 and 5 are backed by a real digest that two parties recomputed independently.
What it doesn'tThe other 9 rows are mapping only; they have not been run against any data. This is different data from examples 1 and 2, and it is not legal advice.
QuadX, on its own rows: “'Good outcome' remains a supervisory quality judgment; the record proves the process, not the quality of the result.”
The FCA clause mapping is QuadX AI's FCA Obligations Mapping v1.0 (22 July 2026): QuadX AI reviewed and corrected the Action State Group strawman and verified the FCA clause references against the FCA Handbook as at July 2026.
Mapping by QuadX AI · report dated 16 Sept 2026 · 0.2 MB file
The A2A + AP2 example on the open standard's site shows the payment side of agentic commerce: the called agent seals a record of each payment action, bound to the mandate that authorised it.
Verify it yourself: the report's last section re-runs the same ten checks, in your browser, offline. One changed byte fails them.
Agentic commerce · Project NANDA
Two agents agree a deal: can a stranger check the terms?
What it showsA buyer's agent negotiates with two competing seller agents on NANDA Town. A person clears the acceptance, and the agreed terms are sealed into one record per side, over the same digest. The records are designed so that anyone can recompute that digest and check both, once the service is public.
What it doesn'tNo payment moves, because the example stops at the agreement. The room seals both records; each side does not seal its own. The deal-room service itself is not published, and the page says which parts are merged and which are not yet public.
Architecture page with a flow table and a claims table · built during Project NANDA's July 2026 hackathon
Verify it yourself: every claim on the page says what it rests on, and links the NANDA Town pull request or commit where the source is public.
Agentic commerce · agent to agent · independent build
One buyer, two suppliers, three rounds: who agreed to what?
What it showsA buyer agent haggles with two competing supplier agents over three rounds and narrates its reasoning before each move. It picks a winner against a private mandate, and turns down the cheaper final offer because that offer breaks a hard constraint. Each side keeps its own signed record, and a stranger who trusts neither can check the whole transaction offline.
What it doesn'tNo payment moves. The agents' reasoning is scripted, with no model call. It is demo code that runs on your machine, not a hosted service or a supported adapter.
Built by Jody (independent agentic engineer) · Apache-2.0 · installs capsule-emit from PyPI
Verify it yourself: bash RUN.sh plays the negotiation in a terminal, and bash RUN.sh --browser opens it as a page that needs no network. The repo also commits the evidence bundle from one run, which you can re-check as it is.
Method · behind every judged row
How is a judged row judged?
What it showsThe rules a judged row follows. The judge is pinned and sealed with each verdict, and the rubric is frozen before the period. Recomputed and judged rows are never blended, several judges can read the same row, and a blind human sample is reported as an agreement figure.
What it doesn'tIt gives no agreement figure for these examples, because none has been measured yet. The outcomes report's human check reads “no human-check data”.
Method page with a claims table · a method, not a report