Why Indication and Predisposition were rescored

Answers to the 14–17 August 2026 review thread, with every figure recomputed from the atlas (207 associations, 177 drug × condition pairs, 21 drugs). Companion to the association explorer.

18 → 0rows with disease-alone P>0.05 still scoring Predisposition ≥ 70
66.9 → 10.5worst Indication spread between variants of one pair
1.1 → 79Indication, simvastatin × hypertension
96.6median Predisposition of the 9 genome-wide rows — unchanged

1. The objections

gosborcz: powinno być wysokie indication w tych przykładach … albo przynajmniej w okolicy 50-70
michkor: simva dostaja ludzie z nadcisnieniem … i raportowanie nadcisnienia, moze byc brakiem efektu prewencji, ale chyba nikt nie przyjmie ze moze to byc ADR
gosborcz: a p-value na disease alone jest mega słabe … czy to jest tylko po samych betach? … to skąd takie wysokie predisposition?
michkor: temporality na 91 jest dziwne

Both scoring objections were correct, and they turned out to have separate causes. Neither was a tuning problem.

2. Root cause A — Indication was a label flag, not a measurement

Only 10 of 207 rows matched a BNF licensed-indication code. Those scored 86.1–100; everything else had a median of 12.1 and was driven by a per-variant genetic channel. Two consequences:

The matcher also compared a three-character prefix as a substring, so M79.6 "pain in limb" counted as the labelled M79.7 "fibromyalgia". Parsing the code list properly (10 → 9 labelled rows) removes that.

3. The fix — a co-treatment statistic

The cohort already contains the missing measurement: the share of a drug's own users who carry the condition. It needs no curation, it is a property of the pair, and it says exactly what michkor said in words.

drugICD-10condition users with this condition
simvastatinI10Essential (primary) hypertension28.72%
atenololI251Atherosclerotic heart disease10.26%
simvastatinI251Atherosclerotic heart disease7.53%
simvastatinI48Atrial fibrillation and flutter7.30%
amlodipineE669Obesity, unspecified6.34%
atorvastatinZ955Presence of coronary angioplasty implant/g3.91%
simvastatinN390Urinary tract infection3.73%
simvastatinN40Enlarged prostate3.72%

Median across the atlas is 1.02%, maximum 28.72%. Two independent derivations exist — from genotype counts (A1_CASE_CT/(2·A1_CASE_FREQ)) and from EHR diagnosis counts (TEMP_n_total/N_regenie) — and they agree at Spearman 0.956 over the 150 rows where both are defined, so the EHR value is used only as a fallback to reach full coverage.

Indication is now the noisy-OR of three channels, with the label deliberately weakened from 0.85 to 0.4 so that it no longer saturates the scale on its own:

Indication  =  100 · [ 1 − ( 1 − 0.4·label ) · ( 1 − 0.8·cotx ) · ( 1 − ev(prescribed) ) ]
cotx  =  clip( log( prevalence / 0.5% ) / log( 30% / 0.5% ) , 0 , 1 )  0.5% counts for nothing, 30% saturates

The three channels are shown per association in the expanded row (Indication_label, Indication_cotreatment, Indication_genetics), so any score can be taken apart.

What this channel does not do: it measures co-occurrence in the treated population, which is what matters for confounding, but it cannot by itself separate "the drug is given for it" from "both are common in the same patients" — statin users are older, so urinary infection and enlarged prostate also rank highly above. The label channel and the expanded row are there to tell the two apart.

4. Root cause B — Predisposition ranked betas, not evidence

The old score correlated better with the raw effect size than with significance (Spearman 0.90 against |β| versus 0.85 against −log10P). Because the odds-ratio channel was a bare percentile rank of |β|, an estimate could be enormous and meaningless and still score at the top: 61 of 207 rows have |β| in the upper half while their 95% lower bound is zero.

drugICD-10gene −log10Pβ ± SE (disease-alone) Predisposition
ramiprilI64FDX11.177.93 ± 4.3496.60.0
simvastatinR600SPRTN0.973.69 ± 2.2991.80.0
simvastatinM513MEF2D1.261.76 ± 0.9291.70.0
ramiprilM109MATN41.110.73 ± 0.4184.70.0
bisoprololR252INTS70.711.54 ± 1.1883.80.0
simvastatinN951SPRTN0.592.78 ± 2.4483.40.0

The top row is the clearest case: an odds ratio of roughly 2800 with a confidence interval spanning three orders of magnitude either side, at P = 0.07, scored 96.6.

Evidence for an arm is now a conjunction of two factors that a noisy estimate cannot fake:

sig(X)  =  clip( −log10PX / 7.3 , 0 , 1 )   absolute: the genome-wide 5×10⁻⁸ bar, not this atlas's spread
eff(X)  =  rank( max( 0 , |βX| − 1.96·SEX ) )   only what survives its own error bar
ev(X)  =  dir(X) · sig(X) · eff(X)   significant AND defensible

All 18 rows that scored ≥ 70 without clearing P < 0.05 are now 0, while the 9 genuinely genome-wide rows keep a median of 96.6. The cost is a scale that reads 0 for most of the atlas — which is the honest answer to "does the disease GWAS show anything here", and why the simple toggle keeps the relative, rank-based view for ranking within the atlas.

This also settles the fibromyalgia question: at −log10P = 3.1 the amitriptyline locus is real but roughly four orders of magnitude short of a genome-wide peak, so it now scores in the thirties rather than 89 — matching the reviewer's "nie było widać piku na GWAS".

5. The disputed pairs, before and after

drugICD-10condition IndicationPredispositionco-treatment −log10P disease
simvastatinI10Essential (primary) hypertension1.179.18.20.028.72%0.33
simvastatinR030Elevated BP reading7.017.791.323.01.24%1.86
atenololR030Elevated BP reading, without diagn13.39.962.50.00.83%0.70
atenololI500Congestive heart failure4.87.976.414.90.75%1.36
amitriptylineM797Fibromyalgia86.351.989.339.11.38%3.11
tramadolR101Pain localized to upper abdomen11.820.429.40.01.42%0.45
ramiprilI64Stroke84.39.496.60.00.81%1.17
simvastatinG629Polyneuropathy, unspecified7.52.831.40.00.58%0.22
simvastatinI251Atherosclerotic heart disease100.0100.0100.094.77.53%39.98
citalopramM819Osteoporosis, unspecified4.213.680.626.31.00%2.23

Two intuitions the data did not support

R03.0 is not the beta-blocker's indication. The expectation that atenolol × R03.0 should score like a hypertension indication is understandable but wrong for that code: R03.0 is "elevated blood-pressure reading, without diagnosis of hypertension", so a patient already treated with atenolol is coded I10 instead. The data agrees — 0.83% of atenolol users carry R03.0, against 28.7% of simvastatin users carrying I10. The intuition is right for I10 and wrong for R03.0.

tramadol × R10.1 stays low, as expected. Upper abdominal pain is not a tramadol indication; if anything it is a recognised opioid gastrointestinal effect, which is the ADR reading rather than the indication reading.

7. Temporality — a real bug, then a residual limitation

The bug: undated events sat in the denominator

michkor: nie wiem czy temporability jest dobrze policzone przy missing data … np gaba - dizzines ma 66 temp, a jest before=1, after=72, missing=36

Correct, and it was a genuine defect. The source column TEMP_pct_after divides by TEMP_n_total, which counts events carrying no usable date, so every undated event was silently treated as "not after". The reported example computes as 72/(72 + 1 + undated) instead of 72/73:

gabapentin × dizzinessdated eventsundated Temporality
R42 Dizziness and giddiness7333.0%66.198.6

The scale now divides by the dated events only. 27 of 207 rows move by more than two points and 9 by more than 25. The direction of the error matters: an affected row was pushed down, i.e. towards "the condition precedes the drug", which is the indication signature — so the bug did not merely add noise, it manufactured false indication-like readings out of missing dates.

drugICD-10conditiondated events undatedTemporality
simvastatinR80Isolated proteinuria2586.2%13.396.0
amlodipineR80Isolated proteinuria2579.5%20.5100.0
lisinoprilR21Rash2467.1%32.9100.0
atorvastatinR21Rash6566.8%33.2100.0
amitriptylineR21Rash4857.5%42.5100.0

Every row still rests on at least 24 dated events, so nothing had to be dropped, but TEMP_n_dated and TEMP_pct_undated are now shown per association so the depth of evidence behind a temporal reading is visible.

Propagated to the clustering features

The same dilution reached the model through ADR_pct_after, ADR_pct_before and ADR_after_over_before, and correcting it there is not a substitution. On the dated basis after + before = 100 exactly, so those three features collapse onto one axis: the two shares become perfectly anti-correlated (r = −1.0) and the log-ratio is a rank-identical transform of either (Spearman 1.0). Left as-is, one axis would have been counted three times.

The block is therefore re-parameterised into the two dimensions that genuinely exist — direction, as the log-ratio over dated events (the better-conditioned encoding, since the dated after-share is piled against 100), and how much of the record is undated, which the old formulation smuggled in as the slack term 100 − pct_after − pct_before. The matrix goes from 44 to 43 features and the worst correlation inside the block falls from 0.89 to 0.22. The v30 rationale for adding ADR_pct_before survives: it was adopted as the strongest single separator of Indication from Effect, and on the dated basis that separator is the log-ratio.

Cluster assignments move modestly and only at higher k: k = 3 is essentially unchanged (ARI 0.98–1.00 for k-means), k = 4 gives ARI 0.77–0.82 with 11 of 207 rows changing camp, and k = 5 drops to 0.34–0.52 — a sign that the five-cluster solutions were partly resting on the triple-counted temporal axis. Notably the rows that move are not the ones with missing dates (median undated share 0% in both groups): the shift comes from removing the duplicate weighting, not from the handful of repaired values.

The residual limitation

Even corrected, a median of 94.7% of dated diagnoses fall after first exposure (it was 91.5% before the fix), and simvastatin × I10 shows only 2.1% before — clinically implausible, since statins are typically started in patients whose hypertension is already documented. The most likely explanation is that the prescription record reaches further back than the diagnosis record, so earlier diagnoses are invisible, and chronic conditions are re-coded at every later contact.

The signal is not empty: pairs where the condition is the drug's licensed indication still show more pre-exposure diagnoses (median 15.2% versus 5.2%). But the absolute level is compressed towards "after", so a high Temporality is weak evidence on its own and a low one is the informative case. That part is documented in the scale's own modal rather than silently corrected, because its fix belongs in the source records, not in the score.

8. Researched write-ups

Ten disputed associations were written up with web-searched sources; each appears in that association's expanded row in the explorer, with its sources and a verdict.

drugICD-10verdictheadline sources
amitriptylineM797licensed indicationGuideline-backed (weak) fibromyalgia use; p=8e-4 is no GWAS peak, so 89 is unsupportable5
atenololI500co-prescription populationHF codes track the cardiovascular disease atenolol treats; atenolol itself is not an HF drug5
atenololR030statistical artifactR03.0 is the no-hypertension-diagnosis code, so it is not atenolol's indication - I10 is5
citalopramM819plausible ADRReal SSRI class effect on bone (RR~1.6 fracture); DLEU1 signal weak, so keep genetics discounted6
ramiprilI64statistical artifactTextbook sparse-data bias: beta 7.93 (SE 4.34, p=0.067) is noise, not a 2800-fold risk5
simvastatinG629contestedGenuine ADR candidate, contested: positive case-control signal, null meta-analysis and RCT evidence5
simvastatinI10co-prescription populationHypertension marks who gets a statin, not what a statin does5
simvastatinI251licensed indicationTextbook positive control: 9p21 and LPA predispose to CAD, the reason simvastatin was prescribed5
simvastatinR030statistical artifactR03.0 is a residual coding category, not an ADR: weak, imprecise signal on 1.24% of users5
tramadolR101plausible ADRNot a tramadol indication; upper-abdominal pain is a recognised opioid GI adverse effect5

Drafted by web-searching research agents (Claude), one per association; every claim is backed by the listed sources, which were visited during drafting. Machine-drafted and pending expert review -- not a curated clinical statement.

6. Related work: the Mendelian-randomisation route

A genomic-led strategy to anticipate drug safety effects — Brian R. Ferolito, Andrea R. V. R. Horimoto, Kai Gravel-Pucillo, Daniel J. Golden, Hesam Dashti, Claudia Giambartolomei, Danielle Rasooly, Rachael Matty, Liam Gaziano, Yakov Tsepilov, Lauren Costa, Nicole Kosik, Harris Ioannidis, Mohd Karim, Giovanna Winicki, Fiona Hunter, Claudia Langenberg, John C. Whittaker, Tianxi Cai, Gina M. Peloso, Barbara Zdrazil, Maya Ghoussaini, Andrew R. Leach, Sumitra Muralidhar, Ines A. Smit, Juan P. Casas, J. Michael Gaziano, Kelly Cho, Alexandre C. Pereira (30 authors, incl. the VA Million Veteran Program), PLOS Genetics 2026, 10.1371/journal.pgen.1012211.

What they do

The authors build a genome-wide, target-level screen for predictable (on-target, Type A) adverse drug reactions using two-sample Mendelian randomisation. Drug targets are proxied not by drug exposure but by cis-acting molecular QTLs: expression QTLs and protein QTLs drawn from GTEx v8, eQTLGen, deCODE, Fenland and ARIC are used as instruments for the abundance of the product of each of 16,915 protein-coding genes, so that the 'exposure' in the MR is genetically predicted target level rather than any particular compound. These instruments are tested against 1,449 harmonised phenotypes assembled from the Million Veteran Program, FinnGen (R10) and UK Biobank, using the Wald ratio for single-variant instruments and inverse-variance weighting for multi-variant instruments, with MR-Egger as a sensitivity analysis and Bayesian colocalisation (posterior probability H4 > 0.5) as a confounding-by-LD check, at a Bonferroni threshold of p <= ~1.6e-9 over roughly 31.5 million gene-trait tests. The sign of the MR estimate is then read as a direction of pharmacological modulation (agonism versus inhibition), converting gene-trait hits into gene-mechanism pairs that represent liabilities expected if a drug moved the target in a given direction. Predictions are benchmarked against external safety evidence: FAERS spontaneous reports, LiverTox and DILIrank 2.0 for hepatotoxicity, clinical trials terminated early for safety, and orthogonal annotations from OMIM, ClinVar, rare-variant burden and mouse knockouts. The stated selling point is that, because the method needs only genetics and not a selective tool compound, it can be applied before a molecule exists.

What they report

The screen yields 58,276 significant gene-trait associations, which condense into 8,495 unique gene-mechanism pairs flagged as potential safety liabilities, and it recovers hundreds of adverse reactions already documented for approved drugs. Roughly 40% of predicted ADRs for targets of approved drugs were matched in FAERS, and pairs supported by colocalisation were more likely to be matched (OR ~1.62); the identified gene-mechanism pairs were ~1.7-fold enriched among 269 trials terminated early for safety reasons. Hepatotoxicity resolved into cholestatic (ALP/bilirubin) and hepatocellular (ALT/AST) axes, with targets hitting multiple liver modules showing ~2.55-fold higher odds of documented drug-induced liver injury. Immune-related pathways were prominent among the safety-associated mechanisms, suggesting particular liability for immune-modulating targets.

The assumptions their design rests on

MR is an instrumental-variable design resting on three assumptions: relevance (the variant is associated with the exposure), independence (the variant shares no common cause with the outcome, i.e. is not associated with confounders of the exposure-outcome relationship), and the exclusion restriction (the variant affects the outcome only through the exposure). Only relevance is directly testable; independence and the exclusion restriction cannot in general be verified empirically, and the exclusion restriction is the one most often violated, through horizontal pleiotropy or linkage disequilibrium. Restricting instruments to cis variants at the target locus strengthens rather than removes this: as Schmidt et al. argue, post-translational (vertical) pleiotropy is exactly what a drug acting on the protein would also produce and is therefore admissible, whereas pre-translational horizontal pleiotropy - alternative splicing, micro-RNA effects, other transcripts at the locus - still biases the estimate, which is why the present paper adds colocalisation and MR-Egger and explicitly notes that its results 'can be biased by horizontal pleiotropy due to the potential non-specific nature of genetic instruments'. Two further limitations are structural rather than statistical, and are conceded by the authors and by the methodological literature: a genetic instrument represents a lifelong, small, compensable perturbation present from conception, which need not recapitulate the magnitude, timing or duration of a drug given later in life for a limited period, so effect sizes and even the presence of an effect can differ. And a cis-MR estimate is an estimate for modulating the target, not for a marketed molecule: it cannot distinguish molecules within a class, cannot speak to dose, formulation or off-target chemistry, and by construction cannot capture idiosyncratic (Type B) reactions, reactive-metabolite toxicity, or compound-specific effects.

How this atlas differs

The two designs sit at opposite ends of the evidence chain. Their exposure is a genetically predicted, lifelong shift in the abundance of a target protein or transcript, evaluated in the general population; our exposure is an actually dispensed prescription, and our outcome is a clinical event observed within the users of that specific drug, in a national EHR-linked genotyped cohort. Consequently they can claim something about the target - that moving gene X in direction Y is likely to carry liability Z, before any molecule exists - but not about a particular product, its dose, or its short-term use; we can claim something about a marketed drug as actually prescribed, including dose-dependence and timing, but only for drugs with enough exposed users, and only for effects visible in routine care. Our design should not be read as Mendelian randomisation, and we do not present it as such. Conditioning the analysis on being treated makes treatment a collider: prescription is caused both by genotype-correlated liability (indication, severity, comorbidity, prior tolerability) and by the same clinical factors that drive the outcome, so restricting to users can induce associations between variants and outcomes that do not exist in the source population - the mechanism described as collider/selection bias by Munafo et al. and as index-event bias in genetic studies of subsequent events by Yaghootkar et al., where selection on an index event can produce both false positives and masked true effects. In addition, the exposure being contrasted is a prescription, which is not randomised at conception and is confounded by indication, so no instrumental-variable interpretation is available; our estimates are association estimates within a treated stratum, which is why we surround them with the Confidence, Indication, Predisposition, Frequency and Temporality scales and with the three comparison arms (the same condition drug-free, daily dose, and propensity to be prescribed) that are designed precisely to expose indication-driven and selection-driven signals. The designs are therefore complementary rather than competing: cis-MR generates target-level, pre-clinical hypotheses with a causal warrant but no compound specificity, while within-user GWAS provides post-marketing, compound- and dose-specific observation with no causal warrant, and a signal appearing in both is supported by two largely non-overlapping sets of assumptions.

Where the two could be compared

The most direct comparison is at the level of drug-ADR pairs: their gene-mechanism predictions can be mapped through the target-to-drug annotations (ChEMBL/Open Targets, which they already use) onto the drugs represented in our ADR arms, and the intersection counted - how many of their predicted liabilities for a target correspond to a condition we have power to test in users of a drug hitting that target. Within that intersection one can ask about direction agreement, checking whether the sign implied by their inferred mechanism of action (inhibition versus activation) matches the direction of the within-user effect for the corresponding variants, and whether the cis locus itself carries signal in our arm. A useful asymmetry test is to compare hit rates across our arms: a genuinely on-target, mechanism-mediated liability should be visible in the ADR arm and in the dose arm but ought to be weaker in the drug-free population arm and in the prescription-propensity arm, whereas an indication- or selection-driven signal should behave the opposite way. Finally, the two resources can be used as mutual enrichment benchmarks - whether our within-user hits are enriched among their predicted safety pairs relative to matched non-predicted pairs, and conversely - which is a more robust statement than any individual pair given the different assumption sets.

Publicly reported Journal Impact Factors for PLOS Genetics cluster around 3.7-3.9 in recent Journal Citation Reports editions (Wikipedia's infobox gives 3.7 for 2024, while several indexing aggregators report 3.9), but the figure could not be verified against a primary Clarivate JCR record, so it should be checked before publication.

Sources

Generated by scale_diagnostics.py from atlas_enriched.parquet and compute_scales, so the figures cannot drift from the scoring they describe. Rebuild with python -m abm_atlas.classifier_clustering.scale_diagnostics.