Discovery Loop · Honest Self-Test
Blinded Drug-Response Benchmark
A genuine closed-book test of the platform. For every HER2 mutation that has a published, cited drug-response outcome, we hide the answer, let the platform predict which drugs it will respond to or resist — reasoning only from the mutation's structure and mechanism — then reveal the real published result and score it honestly.
The predictor never sees the drug-response outcome, nor any text that names the drug or states the response. It reasons from structure alone.
Every prediction is written down and time-stamped before the published answer is attached and graded. The order is enforced by the data itself.
The score is compared to simply guessing the most common outcome. Reasoning only “works” if it beats that. Whatever the result, it is reported as-is.
Runs live against the cited mutation set. Takes about a minute while the reasoner thinks through each variant.
No benchmark has been run yet. Press Run benchmark to give the platform its first closed-book test.