Discovery Loop · Honest Self-Test

Blinded Drug-Response Benchmark

A genuine closed-book test of the platform. For every HER2 mutation that has a published, cited drug-response outcome, we hide the answer, let the platform predict which drugs it will respond to or resist — reasoning only from the mutation's structure and mechanism — then reveal the real published result and score it honestly.

Answer sealed first

The predictor never sees the drug-response outcome, nor any text that names the drug or states the response. It reasons from structure alone.

Commit before reveal

Every prediction is written down and time-stamped before the published answer is attached and graded. The order is enforced by the data itself.

Scored against a baseline

The score is compared to simply guessing the most common outcome. Reasoning only “works” if it beats that. Whatever the result, it is reported as-is.

Run the closed-book test

Runs live against the cited mutation set. Takes about a minute while the reasoner thinks through each variant.

No benchmark has been run yet. Press Run benchmark to give the platform its first closed-book test.