Built as of 13 September 2026.
Measurements come from ChEMBL 37. The endpoints are Ki, IC50, Kd, EC50 and Kb.
Activity is derived from the reported value and unit rather than read from a precomputed column. Duplicate readings collapse by median, never by the most potent value, because the best value selects unit errors. Only exact readings are used; censored readings are dropped.
High-throughput screening is excluded. On the sibling GPCR model, comparisons drawn from withheld screening endpoints scored 0.538 where neither compound had been seen, against 0.773 on held-out medicinal chemistry.
| Stage | Count |
|---|---|
| Measurements after scoping to human single-protein targets carrying a family | 1,263,626 |
| Targets in that pool | 2,705 |
| Ligands in that pool | 794,638 |
| Cross-family comparisons formed | 89,888 |
| Comparisons fitted on | 81,199 |
| Comparisons held out | 8,689 |
| Ligands across those comparisons | 22,588 |
| Targets servable | 1,877 |
| Family pairings represented | 290 |
Four rules, and they are the substance of the method.
| Endpoint | Comparisons | Accuracy |
|---|---|---|
| IC50 | 6,577 | 0.776 |
| Ki | 1,440 | 0.697 |
| Kd | 558 | 0.588 |
| EC50 | 114 | 0.693 |
The Kd row is thin and weak and should not be leaned on.
The holdout is compound-disjoint: a ligand's InChIKey decides which side it lands on, so zero ligands appear on both sides. A held-out comparison therefore involves chemistry the model was never fitted on. It is not a temporal split.
A comparison is one row of three concatenated blocks, in a fixed order. Nothing is scored on its own and subtracted.
| Block | Dimensions | Contents |
|---|---|---|
| Sequence A | 480 | ESM2 esm2_t12_35M_UR50D, mean pooled over residues, full-length UniProt canonical sequence |
| Ligand | 1,038 | Morgan count fingerprint, radius 2, 1,024 bits, then 14 descriptors: molecular weight, heavy atom count, bond count, rotatable bonds, ring count, and per-element atom counts for C, N, O, S, F, Cl, Br, I and P |
| Sequence B | 480 | as sequence A |
Rows are sequence, ligand, sequence, so 2×480 + 1,038 = 1,998 columns. At prediction time the same symmetry is applied as an average over both argument orders, so comparing A with B and comparing B with A return probabilities summing to exactly 1.
A random forest. Every comparison is entered twice, once as given and once with the two sequence blocks exchanged and the label inverted, so no answer depends on which side a target was written on.
| Quantity | Value |
|---|---|
| Comparisons fitted on | 81,199 |
| Rows after the swap | 162,398 |
| Label balance | 0.5 |
| Held-out comparisons | 8,689 |
| Held-out accuracy | 0.750 |
Two curves and two axes, because the trade is the point. Accuracy climbs as the cutoff rises, and the number of comparisons you still get an answer for falls. Either curve on its own tells half the story.
Published beside the headline because a cross-family comparator can score well while knowing nothing about the compound.
| What is being asked | Accuracy |
|---|---|
| The model | 0.750 |
| Always pick whichever family usually wins that pairing | 0.654 |
| The same forest with the ligand removed | 0.710 |
The compound contributes about four points on top of target identity, reproduced at 3.9, 4.5 and 4.1 points across three independently built versions.
The serving image pins every version the forest was fitted under. RDKit computes 1,024 of every ligand block's 1,038 dimensions, so a different RDKit produces different features, scores worse, and nothing in the model detects it. The image refuses to build unless the joblib matches its manifest checksum, the 30 shipped reference predictions replay to 1e-6, and the both-orders average still returns exactly 0.5 for a target against itself.
| Package | Pinned |
|---|---|
| numpy | 2.2.6 |
| python | 3.10.15 |
| sklearn | 1.7.2 |