Experimental · Jev evaluation

Can a language model make the same call as the tree?

A second, independent approach to the same two questions the deterministic engine answers (is IV thrombolysis indicated, is EVT indicated), this time asking Jev (TypeSafe System One) to judge instead of walking the authored tree. Static snapshot of a completed evaluation; nothing here is wired into the live engine, CodeStrokeApp, or StrokeDecisionEngine.

Experimental · educational purposes only currently
The experiment

Does asking Jev the grade directly fail on knowledge, or on holding too much at once?

Two approaches were tried against the same 42-case set the tree itself is tested against (24 of TypeSafeStroke's own scenarios, plus 18 systematically generated boundary cases spanning time window, ASPECTS, occlusion site, mRS, age, and contraindication status). 66 total management questions, scored against the tree's own guideline-cited ground truth, using the same grade taxonomy already used throughout this site: recommended, reasonable, may_consider, individualize, not_recommended, harm, action (missing information, the tree's own "unknown is not no" property).

How decomposition works

Small questions in, the real engine synthesizes the verdict

Instead of one call asking for the final 7-way grade, each case gets a handful of small, narrow Choice questions: a contraindication-severity triage mirroring the tree's own worst-item-wins screen, a disabling-deficit call, and bucketed classifications of time window, core size, pre-stroke mRS, and NIHSS. Jev's answers are substituted into a copy of the case's true state at a representative value well inside the chosen band, and the same deterministic, parity-tested engine that runs the decision engine on this site computes the final verdict. Jev never computes a grade itself; only the tree does.

The one real gap

Three misses, all the same confirmed finding

All 3 remaining misses (of 66) are the identical, reproducible gap: for a large core (ASPECTS <6), Jev answered "clear" (nothing to flag) when it should have flagged a relative contraindication, despite the question's own criteria explicitly listing "ASPECTS <6" as a relative item. This was checked against a genuine guideline divergence already noted in the tree's own open questions (AHA's criterion is narrower than CSBPR's numeric cutoff) before being confirmed, on 2026-09-22, as a real miss rather than a defensible alternate reading. No other item (of roughly 26 explicit history and lab items in the severity question) produced a wrong answer anywhere in the 42-case set. This one item is the entire gap between 63/66 and a clean sweep.

Status

A promising early result, not a decision

This changes the picture from "not usable" to "one narrow, specific, confirmed gap", but nothing on this page has been wired into the live decision engine, CodeStrokeApp, or StrokeDecisionEngine. Static, illustrative, for educational purposes only currently, and reviewed the same way every open question on this site is: flagged plainly rather than smoothed over. Hemorrhagic (ICH) stroke and a live-query mode are both explicitly out of scope for this snapshot.