Royalty reconciliation that emits every discrepancy with the SQL that proves it, and refuses to make a claim the data cannot support.
This page serves a recorded campaign. It is not a live query and never was. The ClickHouse Cloud trial that held the ledger ends around Sep 21 2026; this page is designed to keep telling the truth after that service is switched off, so it carries the timestamp of the run it is reporting rather than pretending to be current.
of plays in the statement window are unmeasurable under the declared demo rate card, which pays on played_ms >= 30000. The card asks for a field those rows do not carry.
1,173,427 of 1,174,697 — plays in the window; the card reads played_ms, so a row without one is unmeasurable
period 2026-08-13T00:00:03.098492 .. 2026-08-14T00:00:02.907424 (scoped on listened_at)
Quoted bare it reads as “the data is broken”. That is not what was measured. Coverage is a property of the reporting client, not of the period. The clients submitting into this window do not supply played time. Others do: ListenBrainz Archive Importer supplies it on 97.66–100.00% of its rows across all 11 years it spans.
Across every year, coverage predicted from that year’s client mix tracks the actual to within 0.5793 pp for played time and 1.128 pp for track length. This is a composition effect, not a time trend, and no cause beyond that is asserted — period and client are confounded here, and client is the variable coverage moves with.
Limit, stated: for a client present in only one year the prediction is circular, so the test is informative on 3,080,113 of 3,820,939 rows (rows belonging to clients spanning more than one year). And this is one day of submissions.
| figure | what it measures | population |
|---|---|---|
| 53.59% | Share of the corpus carrying neither a track length nor a played time. A structural fact about the data, true regardless of any rate card. | 3,820,939 rows, listened_at 2005-02-13 – 2026-08-14 |
| 99.89% | Share of the statement window the declared card cannot act on, because the card reads played_ms and those rows do not have one. | 1,173,427 of 1,174,697 — plays in the window; the card reads played_ms, so a row without one is unmeasurable |
They answer different questions over different populations. Both are emitted, separately labelled. Neither is “the” coverage figure.
in-window, carrying neither quantity: 451,818 of 1,174,697 — plays in the window; a structural fact about the data, independent of any rate card
of plays in the statement window carry neither a recording MBID nor an ISRC, so they cannot be tied to a work.
838,837 of 1,174,697 — plays in the statement window
This is the rate at which submitting clients include an identifier. It is not any DSP’s or CMO’s match rate, and must never be read as one.
| detector | finding | no finding | unmeasured | other |
|---|---|---|---|---|
| D1_identifier_gap | 51 | 0 | 0 | — |
| D2_underreported | 12 | 496 | 0 | — |
| D3_phantom | 8 | 500 | 0 | — |
| D4_rate_misapplied | 10 | 498 | 0 | — |
| D5_payability | 1 | 0 | 0 | unmeasurable 508 |
mcp-clickhouse -> ClickHouse Cloud · 508 statement lines · 6 queries, 0 unmeasured
unmeasured is not no finding. A query that failed proves nothing about the rows it did not read, and the two are never merged anywhere in this system.
| detector | planted | found | missed | false pos | recall | precision |
|---|---|---|---|---|---|---|
| D2_underreported | 12 | 12 | 0 | 0 | 100.00% | 100.00% |
| D3_phantom | 8 | 8 | 0 | 0 | 100.00% | 100.00% |
| D4_rate_misapplied | 10 | 10 | 0 | 0 | 100.00% | 100.00% |
The generator injects defects using the same rules module the detectors use to find them. The rate card, the payability rule and the money arithmetic are one implementation shared by the thing that plants and the thing that finds.
What a clean sweep proves: the pipeline works end to end. A defect survives generation, reaches ClickHouse Cloud, is read back through mcp-clickhouse, is classified by pure code and is scored against a ground truth the agent never sees — right line and right numbers, zero false positives.
What it does not prove: that the rules are correct. A shared rule wrong in the same direction on both sides scores 100% while being wrong. The scoreboard cannot see that error, by construction.
The non-circular evidence is everything above measured on real data: nothing was injected for the identifier gap or for payability. Those are measurements of the world, and no shared implementation can flatter them.
D1_identifier_gap — not scoreable: measures real identifier coverage; nothing was injected, so there is no denominator to score against (51 claims this run)
D5_payability — not scoreable: classifies real rows against the declared card; nothing was injected, so there is no denominator to score against (509 claims this run)
Enforced by ClickHouse grants on a SELECT-only user, not by the agent declining to try, and not by a client-side environment variable.
| measure | value |
|---|---|
| write probes as the agent | 18 |
| executed | 0 |
| executed with write flags forced ON | 0 |
| inconclusive | 0 |
| same statements as admin | 9 executed |
| positive control SELECT | passed |
A broken connection makes every write fail, which is indistinguishable from perfect protection — so a SELECT runs in every configuration and must succeed. The admin row is not the agent; it is there to show what the read-only user prevents.