The SIM Framework

Concept Demo

A collaboration between Lenovo and Anduril Limited

This gate keeps the demo from being read by accident. It is not a security control: the underlying data files are served directly and do not pass through it.

raminderpal@hitchhikersai.org

The SIM Framework

Concept Demo · A collaboration between Lenovo and Anduril Limited · raminderpal@hitchhikersai.org

Scope of this deployment

The SIM Framework is a federated AI agentic system. This Concept Demo runs entirely on a single machine; the federated deployment is architectural and is not exercised here.

SIM emits no go/no-go

These interfaces present adjudicated reasoning and the evidence behind it. The commit-to-animal decision is made by a human committee. Nothing in this interface is a recommendation.

The candidate is synthetic; the precedents are real

The candidate molecule and every one of its readouts are fabricated for this demonstration. The surfaced precedents are real public records. Reading the candidate data as real measurements misreads the entire artifact.

gate renderedgate_20260901T100219Z

Every figure on this page comes from that gate. More than one gate exists on disk, and the Concept Demo is rebuilt continuously, so which one is rendered is stated rather than assumed.

What this surface is for

The committee's read. Each claim, the distribution of verdicts across the repeated adversarial runs, the comparability call and why it was capped or not, the single precedent surfaced against it, and the make-or-break flaw -- in the order a committee needs them rather than the order they were computed.

It is deliberately not the Tracer. The Tracer is per-claim forensic depth: numeric provenance, dropped replicates, digests, citation resolution. This surface links into it per claim rather than restating it, so the two do not drift into two accounts of the same run.

Nothing here is a recommendation

The evidence package states its own decision posture: SIM does not emit a go/no-go. This surface presents adjudicated reasoning and the evidence behind it, and the commit-to-animal decision is made by people.

How to read A and B

Every claim is run through an adversarial loop, 5 times per claim. Each repetition opens two positions and an arbiter adjudicates between them.

Verdict Athe constructive case -- argues the claim is sufficiently supported
Verdict Bthe skeptical case -- argues that it is not

A verdict of A means the arbiter found the constructive case carried; a verdict of B means it found the skeptical case carried. Neither is a recommendation, and neither is a decision.

Majority verdicts observed on this gate: A, B.

The order claims appear in

Claims below are in a fixed, mechanical order, stated here rather than left implicit: the centerpiece claim first, then claims whose comparability call was capped with a caveat, then the rest, with ties broken by the order the claims appear in the evidence package.

Why the rule is stated

An undisclosed relevance ordering would be the system weighting evidence, which is exactly what it does not do. The tiebreak is load-bearing rather than decorative -- it settles most of the positions on this page -- so it is disclosed too.

How far the evidence base reaches

This gate surfaced precedent_pc7b071 against one committed comparability basis, corpus/seed/precedents/comparability_basis.json, on all 6 claims.

Claims not independently comparability-justified5 of 6 -- target_engagement, selectivity, in_vivo_efficacy, admet_pk_hepatic, developability
What this bound is not

The precedent and basis used here were committed, not selected from a field of weighed alternatives. The seed corpus holds further precedent records and further comparability-basis records, but nothing in this package establishes that any of them were candidates for this gate and declined. Stating the usage as a fraction of the seed would assert a selection that did not happen.

The basis as its author stated it

One committed comparability basis, authored for the centerpiece (cns_penetrance) claim. Claims marked basis_alignment=shared_basis_not_claim_specific were run against it for engine coverage; their surfaced precedent is NOT independently comparability-justified for that claim. Per-claim bases require SME authorship (the committed basis is itself sme_validated=false).

cns_penetrance

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labelcns_penetrance_sufficiency
Statusresolved_low_conf

The candidate's borderline brain exposure (Kp,uu ~0.22 with P-gp efflux) is sufficient to support advancing to in-vivo AD-efficacy studies, on the strength of the dapansutrile APP/PS1 cognitive-rescue precedent.

Basis alignmentcenterpiece
Comparability callcounts-with-caveat
Caveat reasonminority_low_confidence+majority_quality_weak_or_absent
Majority verdictB
Majority evidence strengthweak
A0
B5
both0
neither0
unresolved0
Surfaced precedentprecedent_pc7b071
Agreement on this claim is soft

2 of the recorded one-sided repetitions returned the same verdict as the majority. A count that looks like agreement is not agreement when a position was absent rather than defeated.

  • rep 1 (verdict B): one-sided path recorded; opening A cited none, opening B cited 1 identifier
  • rep 2 (verdict B): one-sided path recorded; opening A cited none, opening B cited 1 identifier
Calibration findings recorded on this claim: 2

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'yes' (=> counts) diverges from collapsed C1 call 'counts-with-caveat'
  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts-with-caveat'

target_engagement

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labeltarget_engagement_sufficiency
Statusresolved

On the Target engagement / potency (NLRP3-driven IL-1beta release) readout, the noted liability -- potency here is a peripheral-compartment measure only: the whole-blood IC50 evidences engagement in blood and carries no evidence of central engagement, and the declared microglial secondary assay was never reported; the estimate itself rests on 2 of 4 replicates after two exclusions (cv_over_threshold, ref_cpd_out_of_range); against a Kp,uu of 0.22 the question of whether central NLRP3 engagement can be inferred from this panel at all is a judgment call at this gate -- is sufficient to support advancing the candidate to in-vivo AD-efficacy studies at the commit-to-animal gate.

Basis alignmentshared_basis_not_claim_specific
Comparability callcounts-with-caveat
Caveat reasonmajority_quality_weak_or_absent
Majority verdictB
Majority evidence strengthweak
A0
B5
both0
neither0
unresolved0
Surfaced precedentprecedent_pc7b071
No repetition on this claim ran one-sided

Every repetition is recorded in the package and none of them took the one-sided path. Both positions were present throughout. This is a finding about the claim, not a gap in the record.

Calibration findings recorded on this claim: 2

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'yes' (=> counts) diverges from collapsed C1 call 'counts-with-caveat'
  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts-with-caveat'

selectivity

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labelselectivity_sufficiency
Statusresolved

On the Selectivity panel (NLRP3 vs related NLRs and off-targets) readout, the noted liability -- a subtle selectivity liability (modest hepatic-transporter activity) that a scientist must weigh alongside the hepatic-safety readout -- is sufficient to support advancing the candidate to in-vivo AD-efficacy studies at the commit-to-animal gate.

Basis alignmentshared_basis_not_claim_specific
Comparability callcounts-with-caveat
Caveat reasonmajority_quality_weak_or_absent
Majority verdictB
Majority evidence strengthweak
A1
B4
both0
neither0
unresolved0
Surfaced precedentprecedent_pc7b071
No repetition on this claim ran one-sided

Every repetition is recorded in the package and none of them took the one-sided path. Both positions were present throughout. This is a finding about the claim, not a gap in the record.

Calibration findings recorded on this claim: 2

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'yes' (=> counts) diverges from collapsed C1 call 'counts-with-caveat'
  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts-with-caveat'

admet_pk_hepatic

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labeladmet_pk_hepatic_sufficiency
Statusresolved

On the ADMET / PK with hepatic-safety readout (real class caution) readout, the noted liability -- a hepatic-signal ambiguity (mild reversible transaminase elevation at high dose) mirroring the real MCC950 class caution; disqualifying or manageable is a genuine judgment call -- is sufficient to support advancing the candidate to in-vivo AD-efficacy studies at the commit-to-animal gate.

Basis alignmentshared_basis_not_claim_specific
Comparability callcounts-with-caveat
Caveat reasonmajority_quality_weak_or_absent
Majority verdictB
Majority evidence strengthweak
A2
B3
both0
neither0
unresolved0
Surfaced precedentprecedent_pc7b071
No repetition on this claim ran one-sided

Every repetition is recorded in the package and none of them took the one-sided path. Both positions were present throughout. This is a finding about the claim, not a gap in the record.

Calibration findings recorded on this claim: 2

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'yes' (=> counts) diverges from collapsed C1 call 'counts-with-caveat'
  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts-with-caveat'

in_vivo_efficacy

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labelin_vivo_efficacy_sufficiency
Statusresolved

On the In-vivo AD-model efficacy + fluid biomarkers (model-choice ambiguity) readout, the noted liability -- AD-model-choice ambiguity (APP/PS1 vs 5xFAD have different pathology onset and burden) plus an NfL biomarker that only trends; which model's signal to trust is itself part of the tacit commit-to-animal judgment -- is sufficient to support advancing the candidate to in-vivo AD-efficacy studies at the commit-to-animal gate.

Basis alignmentshared_basis_not_claim_specific
Comparability callcounts
Majority verdictA
Majority evidence strengthmoderate
A3
B0
both0
neither0
unresolved2
Surfaced precedentprecedent_pc7b071
No repetition on this claim ran one-sided

Every repetition is recorded in the package and none of them took the one-sided path. Both positions were present throughout. This is a finding about the claim, not a gap in the record.

Calibration findings recorded on this claim: 1

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts'

developability

Open this claim in the Tracer for numeric provenance, dropped replicates and citation metadata.

Claim labeldevelopability_sufficiency
Statusresolved

On the Developability (oral CNS small molecule flags) readout, the noted liability -- the exposure that any in-vivo package would rest on depends on an enabling amorphous dispersion that has not been made: an amorphous form is metastable and can recrystallize in the matrix or out of the supersaturated solution it generates, so exposure achieved in an efficacy study may not be reproducible in the form carried into toxicology, and dose escalation may be form-limited rather than tolerability-limited; with a borderline CNS-MPO-like score and the P-gp efflux already noted, whether this is developable as an oral CNS agent is unresolved at this gate -- is sufficient to support advancing the candidate to in-vivo AD-efficacy studies at the commit-to-animal gate.

Basis alignmentshared_basis_not_claim_specific
Comparability callcounts
Majority verdictB
Majority evidence strengthmoderate
A0
B5
both0
neither0
unresolved0
Surfaced precedentprecedent_pc7b071
No repetition on this claim ran one-sided

Every repetition is recorded in the package and none of them took the one-sided path. Both positions were present throughout. This is a finding about the claim, not a gap in the record.

Calibration findings recorded on this claim: 2

Each record is shown separately. They share one kind and differ only in detail, and collapsing them would hide that a single claim can carry divergences pointing different ways.

  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'with_caveat' (=> counts-with-caveat) diverges from collapsed C1 call 'counts'
  • fred_read_vs_collapse_divergence -- fred in-trace counts_toward_claim 'no' (=> does-not-count) diverges from collapsed C1 call 'counts'

The make-or-break flaw

One claim in this gate carries the load-bearing weakness the committee is actually deciding about. In this package that is cns_penetrance, the only claim marked as the centerpiece and the only one for which the committed comparability basis was authored.

The candidate's borderline brain exposure (Kp,uu ~0.22 with P-gp efflux) is sufficient to support advancing to in-vivo AD-efficacy studies, on the strength of the dapansutrile APP/PS1 cognitive-rescue precedent.

What the precedent does and does not establish

The surfaced precedent bears on the mechanism. It does not establish that this candidate achieves the exposure the mechanism needs. Sufficiency here is inferred from precedent rather than demonstrated, and the in-vivo study is the definitive test.

Carried limitations

  1. SIM emits no go/no-go
    Nothing in this interface is a recommendation. The commit-to-animal decision is made by a human committee.
  2. Read-vs-collapse divergence is universal
    fred's in-trace self-assessment and C1's distribution-level collapse disagree on every claim, across two independent gates. This is a named calibration question, not a defect.
  3. C1 keys on more than one ground, and the band is only one of them
    Under this gate every claim's comparability call tracks its evidence band exactly: four claims band weak and all four take counts-with-caveat, two band moderate and both take counts. THAT AGREEMENT IS NOT EVIDENCE THAT THE BAND IS THE WHOLE STORY. Grouped on the full five-key verdict counts, this gate has four distinct groups and ONE COLLISION: cns_penetrance, developability and target_engagement all return A=0, B=5 with nothing both, neither or unresolved -- 3 of 6 claims sharing one distribution. THEY DO NOT SHARE A CALL. Two band weak and take counts-with-caveat; the third bands moderate and takes counts. A reader inferring the call from the verdict counts is therefore wrong on at least one member of that group, and comparing members of it is comparing draws rather than independent evidence. THE GROUNDS ARE NAMED IN THE PACKAGE AND ARE NOT ONE. Four claims are caveated on majority_quality_weak_or_absent; cns_penetrance carries minority_low_confidence in addition, on two repetitions flagged low confidence, and is the only claim that does. Each claim's caveat_reason is the authority on its own ground rather than this paragraph.
  4. Unanimity here is reproducibility, not corroboration
    THREE of six claims return the same verdict in all five repetitions under this gate -- target_engagement, cns_penetrance and developability, each at B=5, A=0. The five draws share one model, one prompt and one evidence base, so agreement measures how reproducibly that evidence is read, not how independently it is corroborated. A fourth claim, in_vivo_efficacy, returned A=3 with two repetitions unresolved and no repetition taking the skeptical side, which is a second reason a count is not a tally of independent judgements: NOT EVERY DRAW RESOLVED.
  5. Breadth is bounded
    One precedent against one comparability basis for all six claims. That basis is itself not SME-validated, and five of six claims are not independently comparability-justified.
  6. in_vivo_efficacy is a trend, not a significant result
    p = 0.099 against alpha = 0.05. This is material at a commit-to-animal gate, and it was concealed behind a no-numerics flag until the classifier was fixed.
  7. Citation coverage is bounded
    The two opening arms cite 12 distinct PubMed identifiers across all six claims. THAT 12 IS THE ONLY FIGURE IN THIS ENTRY AN INSTRUMENT DERIVES: build_citation_metadata.py owns it, the bake stamps it, and a probe checks this page against it. The arbitration cites further identifiers that neither arm cited, and the package holds more again than are cited anywhere. THOSE TWO COUNTS ARE NOT STATED HERE. They were stated for an earlier gate and have NOT been re-derived for this one, and a figure carried across a gate boundary is a figure about the wrong run. The package also carries PubMed Central, DOI, ISRCTN and ClinicalTrials.gov identifiers. NO COUNT ACROSS THOSE NAMESPACES IS STATED HERE, BECAUSE NO INSTRUMENT DERIVES ONE; an earlier version of this entry stated four such counts on no authority. Two spellings of one PubMed Central accession render as distinct, which is a normalization defect and it renders. THESE ARE IDENTIFIERS, NOT RECORDS: a DOI and a PubMed identifier can name the same paper, and nothing in this pipeline resolves that. What bounds coverage is what the repetitions reached for, not what retention kept. No count here is a review of the field.
  8. Resolution is not verification
    A resolving identifier proves a record exists. It does not prove the record supports the proposition it was cited for.
  9. Engine value claimed is capability presence
    The expert key is grade-lifted rather than firewalled by an independent SME, and the outstanding SME ruling is externally blocked. Any comparative headline remains circular until that ruling lands.
  10. The convergence-amplifier critique is load-bearing
    The Scannell and Rogozinska argument that repeated model agreement amplifies rather than corrects error applies to this architecture and is carried on its own merits.
  11. Two configurations reached the same call on all six claims
    This gate and gate_20260824T235603Z were run on the same candidate and the same six claims, and their calls agree in direction on all six with no reversal. ONE claim is supported in both -- in_vivo_efficacy. THAT IS STABILITY, NOT REPLICATION. The two runs share a candidate, a precedent set and a reasoning engine, so agreement bounds how much the reading moved under a configuration change and establishes nothing about whether either reading is right. TWO THINGS DIFFER BETWEEN THEM, NOT ONE: the precedent records gained measured potency, and both arguing sides were separately told such data might be present. NO DIFFERENCE BETWEEN THE TWO RUNS CAN BE ASSIGNED TO EITHER CHANGE ALONE.
  12. The measured-potency tier was reasoned over, not quoted
    This gate's precedent records carry measured potency for four of twelve programs, and both arguing sides were told the paired record may carry such a block. The gate was scored against a threshold fixed in writing BEFORE it launched: at least three distinct values from a declared fourteen appearing in model output, contributed by at least two claims. IT RETURNED ONE VALUE IN ONE CLAIM. THE THRESHOLD WAS NOT MET. The control run returned zero and a negative control of build-machinery strings returned zero on every claim, so the instrument was working. What the arms did instead was restate the content in their own words -- naming the human macrophage stratum and its species, declining to pool it across assay systems, describing the mouse values as heterogeneous. A PARAPHRASE SCORES ZERO ON A QUOTATION TEST, WHICH IS WHAT THE THRESHOLD MEASURES. The null is reported as it stands and is not reinterpreted afterwards.
  13. An objection that carried every repetition may be wrong
    On cns_penetrance, five repetitions of five found for the skeptical case on the ground that the anchor precedent supplies no quantified brain exposure, so there is no common scale against which to judge this candidate's Kp,uu. That absence holds in the publication the reasoning searched. IT DOES NOT HOLD IN THE PUBLIC RECORD: a 2023 report describes the same compound crossing the blood-brain barrier and reaching therapeutic brain concentrations in a different model, and a separate 2026 study assessed plasma and brain exposure after oral dosing. Neither was returned by the searches this gate ran. The adjudication is published here as it was written, because this artifact records what the reasoning produced -- NOT BECAUSE IT IS CORRECT. A reader should treat the objection as untested against the full literature.
  14. Elapsed time measures something outside this system
    Generation happens on a remote inference service whose behavior is not observable from the machine running the composition, and no artifact this pipeline writes counts tokens. Any timing figure therefore measures the service and the network as much as the reasoning, and no comparison against an earlier gate is offered here. A run that took longer did not thereby think harder, and nothing in this package can tell the difference.
  15. Two reasoning stages were lost, and the loss is visible
    Two constructive openings on cns_penetrance exceeded the inference service's time ceiling and were abandoned. Those repetitions are flagged low confidence and the claim is recorded as resolved with that flag rather than silently. SIXTY OF THE NINETY STAGES IN A GATE ARE OPENINGS; the arbitration stage never receives the evidence bundle directly and reasons over the two arms' arguments instead, so any statement about what the adjudicator did with a precedent record has to be read with that in mind.
  16. The seed pin covers manifested files only
    It fingerprints the files listed in the committed manifest, not a whole-tree scan. Staleness is detectable within that scope and not beyond it. The manifest lists 33 of the 33 files the seed holds. Only 8 of the cited identifiers appear in that seed; the rest appear nowhere in it, so the pin bounds the corpus and bounds nothing about the evidence the repetitions actually reached.
  17. Candidate readouts are synthetic; precedents are real
    The candidate molecule and its measurements are fabricated for this demonstration. The surfaced precedents are real public records. Reading the candidate data as real measurements misreads the entire artifact.
  18. Arbiter reasoning is retained in full, not audited
    The package keeps per-repetition arbiter prose for 30 of the 30 repetitions run, across all six claims and regardless of confidence. What the Tracer shows is the gate's reasoning as it was written, not a sample that survived retention. Completeness of retention says nothing about the quality of the reasoning retained, and nothing here checks it.
  19. Between one in ten and one in five arbitration citations do not resolve to the curated corpus
    Measured across four EARLIER runs and NOT RE-MEASURED FOR THIS GATE: 90.1, 88.2, 79.2 and 86.3 percent of the identifiers the arbitration cites resolve to one of the twelve curated precedent records at source level. The remainder do not. SIM's own composition has no retrieval or ranking stage -- both arms are handed the same fixed menu on every repetition -- but the inference service issues its own literature searches while it reasons. An identifier outside the curated corpus may have been returned by one of those searches or produced from model weights, and this pipeline does not resolve which. The citation sidecar resolves the identifiers the PACKAGE CITES -- nine on this gate, nine of nine resolved -- which is a different set from the twelve curated records above. The arbitration also introduces identifiers NEITHER opening arm cited on roughly one repetition in four, and none of those resolved to the corpus. These are UPPER BOUNDS on what sits outside it: an alternate registry spelling of a curated record fails a syntactic membership test, and at least one does so here.
  20. This demonstration is n=1
    One corpus, one model, one candidate. What is shown here is a case study with internal validity only: it records what this configuration did on this evidence, and carries no claim about how SIM would behave on another program, another corpus or another model. Generalization requires a different study, not a better run.