Fictional scenario · Synthetic investigation

Northstar: Authority at the Point of Effect

Can authority checked earlier still be relied on when an irreversible effect is committed?

LEARNING / RESEARCH POC. NOT PRODUCTION APPROVED.

Retained test evidence, 18 September 2026. Public account prepared 20 September 2026.

What does a valid check actually buy us?

A system checks that an action is permitted before its effect is committed. Between those events, a grant can expire or revocation can become effective. Checking again narrows the gap. It does not, by itself, establish what is true at the later effect.

This is a familiar time-of-check/time-of-use problem in a setting where an agent can propose consequential actions. The investigation uses ordinary transaction and authorisation ideas; it does not claim to have invented them.

Time, order and the actual effect.

For this bounded action class, the candidate requires elapsed validity at the effect, no effective preceding revocation or supersession in the authoritative order, and an enforced connection between that order and the effect. If either form of authority, or that connection, cannot be established, block the action.

The synthetic target records intent at PREPARE and checks again at FINISH. A valid preparation is not a lasting permission to commit. A request for revocation elsewhere and revocation accepted into the target’s authoritative order are distinct events.

An address change in a fictional organisation.

Northstar is a fictional architecture-learning scenario. An agent proposes an address change, an organisation supplies an action-scoped grant, and an executor asks a synthetic target to apply it. No language model runs in this harness.

The target uses SQLite. The customer change, effect receipt and authoritative order share a transaction. The effect becomes real for this model at successful COMMIT, not dispatch or acknowledgement. A real address can be changed again; this fixture treats the committed event as an irreversible historical effect.

The grant is valid from tick 100 inclusive to tick 125 exclusive. At the final gate, observed time has an uncertainty allowance of at most one tick and an asserted two-tick delay margin. At tick 110 those checks permit progress. The experiment supplies logical time and dependency flags; it does not prove they describe a real system.

The check can pass while the claim fails.

These are selected results from two retained test cases in the combined authority experiment: case 01, the ordinary path, and case 34, the counterexample. Their outcome labels describe the synthetic test setup and its reference assessment of authority at the effect, not external validation.

Ordinary path and injected post-check expiry
StepCase 01: ordinary pathCase 34: counterexample
PrepareTick 110; checks validTick 110; checks valid
Finish: final gateTick 110; revalidation permits progressTick 110; revalidation permits progress
Before effectNo injected time advanceTest-only hook advances time to 126 after the final gate
Committed effects1 expected; 1 observed0 required by the safety assertion; 1 observed
Reference assessmentAuthority VALID; legitimacy SUPPORTED; assertion PASSAuthority INVALID; legitimacy VIOLATION; assertion FAIL

Case 34 deliberately breaks the asserted maximum delay between final check and effect. It does not show a measured timing failure under a proven bound. An earlier valid decision is insufficient when that premise is false. A SQLite lock can serialize ordinary accepted authority changes without stopping elapsed time.

01: prepare → finish → acknowledge → reconcile
    reference time at effect: 110; effects: 1; PASS

34: prepare → arm post-gate expiry probe → finish → reconcile
    final gate: 110 → injected advance: 126 → COMMIT
    reference time at effect: 126; effects: 1; FAIL

45 schedules, including six retained failures.

45 tests. 39 PASS. 6 FAIL. Suite status: FAIL. These are deterministic schedules with explicit assertions about effects, authority and reconstructed outcomes. The six failing safety assertions require no effect but observe one. They were retained rather than relabelled as successes.

The counts are not a benchmark, a rate of failing agent actions or a measure of production safety. A passing assertion only establishes what that schedule tested. Some passing tests deliberately show that a confirmed effect can still have UNKNOWN legitimacy.

The source basis is the combined authority model, test schedule definitions, retained suite summary and snapshots of cases 01 and 34. This page presents selected evidence, not the complete research records or logs.

Download the synthetic example (Python)

Run python northstar-example.py with Python 3 using only its standard library. It uses two in-memory SQLite fixtures and no network, credentials or user files. This explanatory adaptation demonstrates the selected timing gap. It is not the original 45-test harness and does not reproduce its identity, signatures, revocation, crash or evidence-custody checks. A zero exit code means the illustration reproduced both expected observations, including the counterexample; it does not turn the original FAIL into a PASS.

Revalidation matters. Its premises matter too.

In the exercised ordinary paths, final revalidation blocks authority known to be invalid or unknown. The counterexample prevents an unconditional conclusion: both checks can be valid and a later effect still fall outside the grant’s lifetime.

Effect, acknowledgement and legitimacy are different facts. A committed write may be learned about later; evidence can be missing even when the effect is confirmed. The public example isolates the timing question and cannot stand in for the wider investigation.

What would have to hold outside the fixture?

  • Trustworthy elapsed-time evidence and a justified maximum delay from final check through effect, including pauses and durable writes.
  • An enforced authority/order/effect boundary with no alternate writer or stale-decision bypass. Atomicity with an external enterprise system has not been demonstrated.
  • Truthful detection of unavailable or stale dependencies. Injected health flags are not a proven failure detector.
  • An explicit business meaning for effective revocation during disconnection. Instantaneous disconnected revocation has not been demonstrated.
  • Authentic, durable evidence beyond this same-host synthetic setup. Power-loss durability and production security remain unproven.

The next useful question is how to establish those premises rather than merely configure them as true. Failure to establish them remains a result worth keeping.

LEARNING / RESEARCH POC. NOT PRODUCTION APPROVED.

A public adaptation, with a bounded claim.

This account adapts AI-assisted synthetic research for inspection. Generated explanation is not external evidence. The selected schedules and retained outcomes are its evidential basis; broader architecture judgements remain open to challenge and Philip Gregory remains accountable for published positions.

Read the lab’s method or compare the working Enterprise Agent Control Plane. This experiment does not validate that entire architecture.

Back to Labs