Digital Protein Foundry
Model Arena
Every candidate model faces the SAME blinded challenge set. Predictions are committed before any answer is revealed — then ground truth is unlocked and each prediction is graded by an impartial referee. Held-out challenges are scored separately; they are the real signal of whether a representation predicts reality.
Central law: predictions are locked in during the commit phase with ground truth hidden. No model can see an answer before committing, and held-out outcomes never feed back into model evolution.