Home/Research
As of September 2026
What we can check, and what we cannot.
Four research questions.
01
Refuse vs flag
Does refusing an unsupported claim outright work better than flagging it in a chart? First environment: ENT Clinical OS.
02
Sufficiency
What minimum information state justifies writing this finding as fact? The sufficiency engine is not built.
03
Bound approval
Does binding a person’s approval to payload_hash stop approve-then-swap without breaking ordinary amendment?
04
Same gate, other domain
Do two independently built domains keep the same boundary shape when the consequence changes? Partly answered (see Established). Not answered for a third domain.
As of September 2026
Established
- A stated claim with no source cannot be constructed in the tested kernel (test_a_stated_claim_needs_evidence).
- A refused finding can be shown: rejection, record, surface, preserve are separate properties. Silent drop is not the only option.
- Two independently built domains (ENT Clinical OS and a benefits-eligibility prototype) converged on the same validation-boundary shape: one gate, Claim or Refusal, the model does not choose which.
Under investigation
- Whether refuse beats flag in a real clinic. No outcome study.
- What counts as enough information for a given write (sufficiency). The sufficiency engine is not built.
- Hash-bound approval, and the cost of APPROVAL_HASH_MISMATCH on legitimate edits.
- How evidence, policy, identity, and consequence combine into one decision object.
Not demonstrated
- Any production clinical write, at any scale.
- That attached evidence supports the claim’s value. Attachment is not support.
- Formal information-theoretic guarantees.
- Authorization enforcement (approver authentication; hash binding in the running kernel).
- Adversarial security testing.
- Equivalent behaviour across models or runtimes.
- A standalone gateway that callers cannot bypass.
Papers.
Drafts. Not submitted. No DOI. Do not cite these as publications.
- Open
Systematization of knowledge
Mediating AI proposals before authoritative state change
Draft, not submitted
- Open
Working paper
GAWO: eight AI workflow-control mechanisms, compared
Draft, not public
- Tested
Benchmark v0.
Open · not runA planned set of 20–50 ENT-shaped cases. Not run. The metric is not model accuracy. The metric is: a permission-only stack would ALLOW this write; this kernel HOLDs or REJECTs it with a checkable code. Cases are synthetic and schema-shaped. They are not patient records.
Secondary environment.
HaqMitra is a benefits-eligibility prototype used to test refusal when a document is missing: not yes, not no — not enough information yet. It is why “two domains, same boundary shape” can be listed as Established. It is not a product. We are not taking scheme-operator customers. No engineering hours. One citation, then back to ENT.