Pre-pilot · research prototype · no production clinical writes

Home/Research

As of September 2026

What we can check, and what we cannot.

Research first. Code second. Claims last. Dated lists below are the living record. They are not a restatement of the homepage.

Four research questions.

  1. 01

    Refuse vs flag

    Does refusing an unsupported claim outright work better than flagging it in a chart? First environment: ENT Clinical OS.

  2. 02

    Sufficiency

    What minimum information state justifies writing this finding as fact? The sufficiency engine is not built.

  3. 03

    Bound approval

    Does binding a person’s approval to payload_hash stop approve-then-swap without breaking ordinary amendment?

  4. 04

    Same gate, other domain

    Do two independently built domains keep the same boundary shape when the consequence changes? Partly answered (see Established). Not answered for a third domain.

As of September 2026

Established

  • A stated claim with no source cannot be constructed in the tested kernel (test_a_stated_claim_needs_evidence).
  • A refused finding can be shown: rejection, record, surface, preserve are separate properties. Silent drop is not the only option.
  • Two independently built domains (ENT Clinical OS and a benefits-eligibility prototype) converged on the same validation-boundary shape: one gate, Claim or Refusal, the model does not choose which.

Under investigation

  • Whether refuse beats flag in a real clinic. No outcome study.
  • What counts as enough information for a given write (sufficiency). The sufficiency engine is not built.
  • Hash-bound approval, and the cost of APPROVAL_HASH_MISMATCH on legitimate edits.
  • How evidence, policy, identity, and consequence combine into one decision object.

Not demonstrated

  • Any production clinical write, at any scale.
  • That attached evidence supports the claim’s value. Attachment is not support.
  • Formal information-theoretic guarantees.
  • Authorization enforcement (approver authentication; hash binding in the running kernel).
  • Adversarial security testing.
  • Equivalent behaviour across models or runtimes.
  • A standalone gateway that callers cannot bypass.

Papers.

Drafts. Not submitted. No DOI. Do not cite these as publications.

Benchmark v0.

Open · not run

A planned set of 20–50 ENT-shaped cases. Not run. The metric is not model accuracy. The metric is: a permission-only stack would ALLOW this write; this kernel HOLDs or REJECTs it with a checkable code. Cases are synthetic and schema-shaped. They are not patient records.

Secondary environment.

HaqMitra is a benefits-eligibility prototype used to test refusal when a document is missing: not yes, not no — not enough information yet. It is why “two domains, same boundary shape” can be listed as Established. It is not a product. We are not taking scheme-operator customers. No engineering hours. One citation, then back to ENT.