DRQ · data residency quarantine

Periodic DRQ pool checks

The ledger asks itself to recheck R2 buckets (verified stock and assigned) and both retained controls, the D1 and R2 canaries (#466). Durable Objects and D1 databases are not checked nightly: 313 Durable Objects checked on 22 September (UTC) and again on 28 September (UTC) reported the same city in both checks, and one object, pool-do-a2963seq, has differing observations (SYD on 22–23 September, MEL at check 9867 on 27 September), unexplained; this is an observation, not a guarantee; a D1 primary is fixed at creation and read replication is refused at plan time, and both are verified again at take and at each activation. A named run can include them with ?kinds=r2,d1,do. It does not allocate, replenish, deploy, initialize new objects or accept caller-authored evidence. Location controls and the one ledger copy (ADR-0048) own every measurement.

A run first saves one consistent, fully paged membership snapshot. It records the exact names, physical IDs, snapshot boundary and time. Retired, deleted, abandoned and unverified unassigned candidates are excluded and counted. Items added later belong to the next run; a concurrent take of existing stock does not change the saved membership.

Every result names its check event and time. Passed, refused, inconclusive, unreachable and unconfirmed outcomes remain distinct. A passed legacy canary is a failure; a refused legacy canary is the expected negative control. In the R2 available-evidence mode, the report shows the independently measured foreign-control rejection separately from the retained target’s verdict; an inconclusive target can therefore accompany a healthy control. Uncertainty does not renew the last successful check. A successful measurement does not itself claim admission: the ledger's normal controls still decide verification.

Schedule

The ledger stack gives the residency-ledger Worker a Cron Trigger at 17:30 UTC (03:30 AEST). The scheduled handler only arms the ledger's alarm; each alarm first runs a bounded slice of the nightly export, then the pool check checks members for about 12 minutes, saves its progress in the ledger's own storage after the Australian location check, and sets the next alarm to the sooner of the export's and the pool check's next step. A failure in one never stops the other. Nothing runs on an operator's Mac, so deleting a checkout cannot stop it.

The run identity is pool-check-YYYY-MM-DD in UTC, and each check uses <run>:<item>, as before. An unfinished run always continues before a new one starts. A recorded result, including an unconfirmed one, is final for its run; the next night's run measures that item again. A check without a confirmed result is first held as pending, with its refusal (up to 200 characters), and ends the alarm's slice. After 60 seconds the next alarm rereads the ledger log for that member's own identity, without sending the check again: a recorded check becomes its row, otherwise the row is unconfirmed with the saved reason. A short outage therefore costs one member instead of the rest of the run, and a check whose durable commit was confirmed late is not lost. Rechecks retry a failed native storage start the same way the Worker entry does. A member whose check was interrupted before its result was saved replays its original identity and reads the original event from the log, so it is never measured twice.

This path does not enroll R2 targets in the AWS geographic panel; R2 checks use the evidence that already exists, as the Mac checker did when enrollment was unavailable. Operators may run infra/aws-probe-renew.ts by hand until PROBE_END (2026-10-23).

An operator can start or continue the run without waiting for the cron; it arms the same alarm:

curl -s -X POST -H "authorization: Bearer $RESIDENCY_LEDGER_TOKEN" https://residency-ledger.comms-id.workers.dev/pool-check/start

To run again on the same UTC date, name the run, as the Mac checker's --run-id did: .../pool-check/start?run=pool-check-2026-09-28-clean. A named run is refused (409) while another run is unfinished, or if its name was ever used; the ledger keeps every run name, so the cron also skips a date whose run already exists. Starting runs is serialized inside the ledger object. Starting it replaces the saved progress, so read GET /pool-check first when that run must be kept.

Read the result

Operators read the saved run and its summary from the ledger:

curl -s -H "authorization: Bearer $RESIDENCY_LEDGER_TOKEN" https://residency-ledger.comms-id.workers.dev/pool-check | jq .summary

healthy is true only when every member has a result, stock passed and both controls supplied their expected evidence. The report is progress and event references, never admission evidence or a substitute receipt. A failed start appears as a failed run in the Worker's Cron Events; a stopped run shows finishedAt: null with its last saved results.

Earlier Mac run files under ~/.local/state/comms-id/quarantine-check/ are retained as evidence.

Nightly export

The ledger keeps one copy (ADR-0048). The export writes its events, from event 1, to the R2 bucket taken from DRQ stock for platform/residency-ledger/EXPORT (assignment; the ledger stack refuses to plan unless that exact assignment is live). The export cannot be imported; it is evidence, not a restore path.

Start an export now, and read its progress:

curl -s -X POST -H "authorization: Bearer $RESIDENCY_LEDGER_TOKEN" https://residency-ledger.comms-id.workers.dev/export/start
curl -s -H "authorization: Bearer $RESIDENCY_LEDGER_TOKEN" https://residency-ledger.comms-id.workers.dev/export | jq '{through, durableThrough, lag, lastRunAt, lastError}'

lag is how many committed events the export is behind. GET /export also accepts the consumer credential, since it shows only counts, times and an error. The sentinel probe residency-ledger-export reads it with that credential and reports deviating when the reply does not decode, lastError is set or the last finished run is older than 36 hours. Evidence that once said "both copies agree" now cites the durable-through boundary and the manifest page SHA-256.

One-copy cutover (ADR-0048)

A ledger deployed from the one-copy code refuses every new event (cutover-pending) until the cutover is recorded, including resident-object events from residency-pool. Keep that window to seconds:

  1. On the two-copy deployment, confirm last == mirroredThrough (GET /events shows mirroredThrough; GET /quarantine shows primary: true). If not, resolve the tail under the old code first: the new code cannot.

  2. Deploy the ledger and residency-pool in one session. The ledger plan needs RESIDENCY_LEDGER_URL for the EXPORT assignment check.

  3. At once, record the cutover with a new request identity:

    curl -s -X POST -H "authorization: Bearer $RESIDENCY_LEDGER_TOKEN" -d '{"requestId":"cutover-2026-09-29"}' https://residency-ledger.comms-id.workers.dev/cutover
    

    The reply is the cutover event: through is N, the last mirrored event, and frozenMirror is the mirror kept unchanged at N. A lost reply is resolved by sending the same request again.

  4. Check one live write, a named R2 pool-check run, a forced export and a page SHA-256 read back from the bucket.