Fast Protein-Ligand Scoring

Build a machine-learning surrogate for a computational docking pipeline. Given a protein target and a candidate ligand, your endpoint predicts the ligand's docking score. The goal is to rank promising molecules accurately while scoring them far faster than running the original docking simulation. Input: a target_id and a ligand SMILES string. Request contract: POST {"features": {"target_id": "CDK2", "smiles": "Cc1cc..."}} Response contract: 200 OK {"score": -8.74} Participants train a model from the provided public protein, ligand, and docking-score data, then deploy it as an inference endpoint. The private evaluation dataset covers five protein targets: CDK2, DPP4, MK14, PARP1, and SRC. The evaluator samples hidden examples from every target, sends one request per example to your endpoint, validates the returned score, and compares your predictions against the private reference docking scores. The private reference scores are never sent to your endpoint. The leaderboard score is a macro composite, computed separately per target and then averaged equally across all five: 0.60 times Spearman rank correlation, plus 0.20 times Pearson correlation, plus 0.20 times a term that rewards low RMSE relative to that target's training-data spread. MAE and RMSE are also reported individually for context but do not themselves drive the primary ranking beyond their role inside the composite. Before it is scored, every submission must pass a technical gate. Your endpoint must answer GET /health (beside your predict path, e.g. https://your-host/health for https://your-host/predict) with a 2xx status; at least 95% of prediction requests must succeed and return a finite numeric score; your average latency must be below 1 second per request; and any single request still running after 5 seconds is cut off and counts as failed. A submission that fails any part of the gate is marked failed and is not scored, so it does not appear on the leaderboard. Throughput is reported for context and does not affect rank.

Build a machine-learning surrogate for a computational docking pipeline. Given a protein target and a candidate ligand, your endpoint predicts the ligand's docking score. The goal is to rank promising molecules accurately while scoring them far faster than running the original docking simulation. Input: a target_id and a ligand SMILES string. Request contract: POST {"features": {"target_id": "CDK2", "smiles": "Cc1cc..."}} Response contract: 200 OK {"score": -8.74} Participants train a model from the provided public protein, ligand, and docking-score data, then deploy it as an inference endpoint. The private evaluation dataset covers five protein targets: CDK2, DPP4, MK14, PARP1, and SRC. The evaluator samples hidden examples from every target, sends one request per example to your endpoint, validates the returned score, and compares your predictions against the private reference docking scores. The private reference scores are never sent to your endpoint. The leaderboard score is a macro composite, computed separately per target and then averaged equally across all five: 0.60 times Spearman rank correlation, plus 0.20 times Pearson correlation, plus 0.20 times a term that rewards low RMSE relative to that target's training-data spread. MAE and RMSE are also reported individually for context but do not themselves drive the primary ranking beyond their role inside the composite. Before it is scored, every submission must pass a technical gate. Your endpoint must answer GET /health (beside your predict path, e.g. https://your-host/health for https://your-host/predict) with a 2xx status; at least 95% of prediction requests must succeed and return a finite numeric score; your average latency must be below 1 second per request; and any single request still running after 5 seconds is cut off and counts as failed. A submission that fails any part of the gate is marked failed and is not scored, so it does not appear on the leaderboard. Throughput is reported for context and does not affect rank.

Category
Life Sciences & Healthcare
Status
open

Success metrics

  • ▸Spearman Correlation
  • ▸MAE
  • ▸RMSE
  • ▸Latency
  • ▸Throughput
  • ▸Success Rate
  • ▸Pearson Correlation

· ranks on this challenge

Standings

RankSolverScoreSubmission
Loading…