{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Weak labels for 12 knee findings, and the threshold wall\n\n[Version 1 of this notebook](https://www.kaggle.com/code/nekkon/weak-labels-for-all-12-knee-mri-findings)\nextracted all twelve findings from all 4,407 reports with a multilingual keyword matcher and\nscored it against the 58 labelled studies. Mean balanced accuracy 0.672, with medial meniscus\nat 0.558 -- barely better than a coin flip -- and I wrote at the time that the meniscus findings\nwere close to unrecoverable from text.\n\nThat was wrong, and this version shows why. Replacing the keyword matcher with a **cross-lingual\nentailment model** takes medial meniscus from 0.558 to **0.87 AUC**, and the model demonstrably\nbinds the *side*: its medial score is at chance on the lateral label. The keyword matcher never\ndid -- its medial score predicts the lateral label *better* than its own.\n\nThen the notebook runs into the thing that actually limits this competition. The better scores\nrank better (mean AUC 0.672 -> 0.75), but turning them into binary labels at a threshold fitted\non 58 studies gives back **the entire gain**. That is not a modelling failure, it is an\ninformation limit, and it changes what you should do with the output.\n\n**What you get at the end:** `weak_labels_v2.csv` -- 4,407 rows, per-finding scores from both\nmethods plus a blend, as probabilities rather than 0/1, with a per-finding reliability column."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import glob, os, re, pickle, warnings\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport torch\n\nwarnings.filterwarnings('ignore')\n\nINK, MUTED, GRID = '#1a1a1a', '#5c5c5c', '#dcdcdc'\nKWC, NLIC, BLEND, WARN = '#2a78d6', '#eb6834', '#1baf7a', '#e34948'\nplt.rcParams.update({'figure.facecolor': 'white', 'axes.facecolor': 'white',\n                     'axes.spines.top': False, 'axes.spines.right': False,\n                     'axes.edgecolor': GRID, 'font.size': 11, 'text.color': INK,\n                     'axes.labelcolor': MUTED, 'xtick.color': MUTED, 'ytick.color': MUTED,\n                     'axes.titlesize': 13, 'axes.titleweight': 'bold',\n                     'axes.titlelocation': 'left', 'axes.titlepad': 12})\n\nP = glob.glob('/kaggle/input/**/train.csv', recursive=True)[0].rsplit('/', 1)[0]\ntrain = pd.read_csv(f'{P}/train.csv')\nLABELS = [c for c in train.columns if c not in ('StudyInstanceUID', 'Report')]\ngold = train[train[LABELS].notna().all(axis=1)].reset_index(drop=True)\n\nprint(f'{len(train):,} studies, {len(gold)} of them labelled, {len(LABELS)} findings')\nprint('device:', 'cuda' if torch.cuda.is_available() else 'cpu')"},{"cell_type":"markdown","metadata":{},"source":"## 1. The v1 baseline, unchanged\n\nThis is the keyword extractor from version 1, copied verbatim so the comparison below is against\nexactly the floor that notebook reported. It carries a vocabulary per finding per language, and\nnegation windows that run left for the Latin-script languages and right for Turkish\n(`efuzyon izlenmedi` -- \"effusion was not observed\" -- negates *after* the term).\n\nTwelve findings times seven languages is the maintenance surface, and it is the reason to want\nsomething that does not need a vocabulary at all."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"NEG_L = (r\"\\b(?:no|not|without|absent|negative for|geen|zonder|kein[e]?|ohne|sin|sans|\"\n         r\"senza|non|nessun[ao]?|δεν|χωρ[ιί]ς)\\b\")\nNEG_R = r\"\\b(?:izlenmedi|saptanmad[iı]|yok|g[oö]r[uü]lmedi|mevcut de[gğ]il|izlenmemi[sş]tir)\\b\"\n\nTEAR = (r\"tear|torn|rupture|ruptur|rotura|ruptura|scheur|ruptuur|riss|d[ée]chirure|\"\n        r\"lesione|y[iı]rt[iı]k|rupt[uü]r|ρ[ήη]ξη\")\nMENISC = r\"menisc|menisk|μηνίσκ\"\nMEDIAL = r\"medial|interno|internal|mediale|binnen|innen|i[cç] |[έε]σω\"\nLATERAL = r\"lateral|externo|external|laterale|buiten|au[sß]en|[έε]ξω\"\n\nTERMS = {\n    \"ACL\": (r\"anterior cruciate|\\bacl\\b|ligamento cruzado anterior|\\blca\\b|\"\n            r\"voorste kruisband|\\bvkb\\b|vorderes kreuzband|ligament crois[ée] ant[ée]rieur|\"\n            r\"[oö]n [cç]apraz|πρ[όο]σθιο[υς]? χιαστο\", TEAR),\n    \"MCL\": (r\"medial collateral|\\bmcl\\b|ligamento colateral (?:medial|interno)|\"\n            r\"mediale collaterale|innenband|mediales? kollateral|\"\n            r\"i[cç] yan ba[gğ]|[έε]σω πλ[άα]γιο\", TEAR),\n    \"Medial Meniscus\": (rf\"(?:{MEDIAL})\\W+(?:\\w+\\W+){{0,2}}(?:{MENISC})|\"\n                        rf\"(?:{MENISC})\\w*\\W+(?:\\w+\\W+){{0,2}}(?:{MEDIAL})\", TEAR),\n    \"Lateral Meniscus\": (rf\"(?:{LATERAL})\\W+(?:\\w+\\W+){{0,2}}(?:{MENISC})|\"\n                         rf\"(?:{MENISC})\\w*\\W+(?:\\w+\\W+){{0,2}}(?:{LATERAL})\", TEAR),\n    \"Medial OA\": (rf\"(?:osteoarthr|arthros|artros|gonarthros|artrosis|chondropath|\"\n                  rf\"condropat|kraakbeenlijden|kn?orpelschaden|osteoartrit|\"\n                  rf\"kondropati|αρθρ[ίι]τιδα)\", MEDIAL),\n    \"Lateral OA\": (rf\"(?:osteoarthr|arthros|artros|gonarthros|artrosis|chondropath|\"\n                   rf\"condropat|kraakbeenlijden|kn?orpelschaden|osteoartrit|\"\n                   rf\"kondropati|αρθρ[ίι]τιδα)\", LATERAL),\n    \"PF OA\": (r\"osteoarthr|arthros|artros|artrosis|chondropath|condropat|chondral|\"\n              r\"condral|kraakbeenlijden|kn?orpelschaden|kondromalaz|chondromalac|\"\n              r\"condromalac|osteoartrit|αρθρ[ίι]τιδα|χονδροπ[άα]θ\",\n              r\"patellofemoral|femoropatell|patello-femoral|retropatellar|\"\n              r\"femoropatelar|patell|patella|rotul|trochlea|tr[oó]clea|επιγονατιδ\"),\n    \"Effusion\": (r\"effusion|joint fluid|derrame|gewrichtsvocht|erguss|[ée]panchement|\"\n                 r\"versamento|ef[uü]zyon|eklem s[iı]v[iı]|υγρ[οό]\", None),\n    \"Synovitis\": (r\"synovit|sinovit|synovial (?:thicken|proliferat|hypertroph)|\"\n                  r\"synoviale? (?:verdikking|proliferat)|synovialitis|υμεν[ίι]τιδα|\"\n                  r\"plica synovial\", None),\n    \"Baker's\": (r\"baker|popliteal cyst|quiste popl[ií]teo|bakercyste|poplitea?le? cyste|\"\n                r\"bakerzyste|kyste popliet|baker kist|popliteal kist|κ[ύυ]στη baker\", None),\n    \"Contusion\": (r\"contusion|contusi[oó]n|bone (?:marrow )?(?:o)?edema|\"\n                  r\"botoedeem|beenmergoedeem|knochenmark[soö]dem|edema (?:de m[ée]dula|[oó]seo)|\"\n                  r\"kemik ili[gğ]i [oö]dem|[οο]?[ίι]δημα μυελο|bone bruise\", None),\n    \"Fracture\": (r\"fracture|fractura|fractuur|fraktur|frattura|k[iı]r[iı]k|\"\n                 r\"avulsion|avulsi[oó]n|κ[άα]ταγμα\", None),\n}\n\n\ndef hit(text, pattern, need=None, window=110):\n    t = \" \" + re.sub(r\"\\s+\", \" \", text.lower()) + \" \"\n    for m in re.finditer(pattern, t):\n        left = t[max(0, m.start() - window):m.start()]\n        right = t[m.end():m.end() + window]\n        if re.search(NEG_L + r\"\\W+(?:\\w+\\W+){0,5}$\", left):\n            continue\n        if re.search(r\"^\\W*(?:\\w+\\W+){0,6}\" + NEG_R, right):\n            continue\n        if need and not re.search(need, left + \" \" + right):\n            continue\n        return 1\n    return 0\n\n\ndef extract(text):\n    return {k: hit(text, p, need) for k, (p, need) in TERMS.items()}"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"KW_ALL = pd.DataFrame([extract(t) for t in train.Report])[LABELS].values\nprint(f'keyword extractor over all {len(train):,} reports: '\n      f'{KW_ALL.sum():,} positives, {(KW_ALL.sum(1) == 0).mean():.1%} of studies get nothing')"},{"cell_type":"markdown","metadata":{},"source":"## 2. Ask the question instead of listing the words\n\nNatural language inference asks: given this text, is this statement true, false, or unsupported?\n`mDeBERTa-v3-base-xnli` is trained on 2.7M inference pairs across 100 languages, and it is\n*cross-lingual* -- the premise can be a Spanish report while the hypothesis stays in English.\n\nSo instead of a Spanish word list for meniscal tears, the question is just\n**\"The medial meniscus is torn.\"**, asked against `Rotura de menisco interno.`\n\nTwo details do most of the work:\n\n**Score on the entail-vs-contradict axis, dropping neutral.** `softmax(logits[[entail, contra]])[0]`\nis the affirm/deny axis. A report saying *no effusion* produces contradiction, not low entailment --\nnegation is handled by the thing the model was trained to do, in every language at once, with no\nnegation window to maintain.\n\n**Note the device line printed above.** At the time of writing, Kaggle's PyTorch is built\nfor `sm_86 / sm_90 / sm_100 / sm_120`, and the GPU it hands out here is a **Tesla P100, which\nis `sm_60`**. PyTorch dropped Pascal. The mismatch does not surface at `.to('cuda')` or in\n`torch.cuda.is_available()` -- both succeed -- it surfaces as `no kernel image is available`\non the first matmul, which is why the cell above runs one deliberately. If that probe fails,\nthe scoring pass below reads its scores from\n[a published cache](https://www.kaggle.com/datasets/nekkon/rsna-knee-nli-entailment-scores)\nof this same code instead of spending thirteen CPU-hours reproducing them. Anything in this\ncompetition that quietly assumes a working GPU for a transformer is worth re-checking.\n\n**Ask sentence by sentence, take the maximum.** A 980-character report containing one relevant\nclause dilutes to nothing as a single premise. Section 3 measures this: whole-report premises\nscore around 0.60 mean AUC, sentence-level maxima around 0.73. The hypotheses below were written from\nthe finding names before looking at the 58 labelled studies."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"from transformers import AutoTokenizer, AutoModelForSequenceClassification\n\nMODEL_DIR = os.path.dirname(glob.glob('/kaggle/input/**/config.json', recursive=True)[0])\nprint('model:', MODEL_DIR)\n\nHYP = {\n    'ACL':              'The anterior cruciate ligament is torn.',\n    'MCL':              'The medial collateral ligament is injured.',\n    'Medial Meniscus':  'The medial meniscus is torn.',\n    'Lateral Meniscus': 'The lateral meniscus is torn.',\n    'Medial OA':        'There is osteoarthritis of the medial compartment of the knee.',\n    'Lateral OA':       'There is osteoarthritis of the lateral compartment of the knee.',\n    'PF OA':            'There is patellofemoral osteoarthritis.',\n    'Effusion':         'There is a joint effusion in the knee.',\n    'Synovitis':        'There is synovitis.',\n    \"Baker's\":          \"There is a Baker's cyst.\",\n    'Contusion':        'There is bone marrow oedema or a bone contusion.',\n    'Fracture':         'There is a fracture.',\n}\nassert set(HYP) == set(LABELS)\n\n_SPLIT = re.compile(r'(?<=[.;:!?])\\s+|\\n+')\n\ndef sentences(report, max_n=60):\n    # Keep every non-empty fragment. Filtering short ones costs AUC: the one-word Spanish\n    # clause \"Derrame.\" IS the effusion finding, and a >=4-word filter throws it away\n    # (it costs about 0.03 mean AUC on the 58). Section headers score near zero anyway, so the\n    # per-study maximum ignores them without being told to.\n    out = [s.strip() for s in _SPLIT.split(str(report))]\n    return ([s for s in out if s] or [str(report)[:400]])[:max_n]\n\ntok = AutoTokenizer.from_pretrained(MODEL_DIR)\nmdl = AutoModelForSequenceClassification.from_pretrained(MODEL_DIR).eval()\n\n\ndef pick_device():\n    '''Kaggle hands out whichever GPU is free, and a torch build without kernels for\n    that architecture fails on the first matmul with \"no kernel image is available\"\n    rather than at .to(). So actually run one.'''\n    print('torch', torch.__version__, '| built for', torch.cuda.get_arch_list()[-4:]\n          if torch.cuda.is_available() else '(no cuda)')\n    if not torch.cuda.is_available():\n        return 'cpu'\n    name = torch.cuda.get_device_name(0)\n    cap = torch.cuda.get_device_capability(0)\n    print(f'gpu {name}, compute capability sm_{cap[0]}{cap[1]}')\n    try:\n        (torch.zeros(8, 8, device='cuda') @ torch.zeros(8, 8, device='cuda')).cpu()\n        return 'cuda'\n    except Exception as e:\n        print(f'!! cuda unusable ({type(e).__name__}), falling back to cpu:',\n              str(e).splitlines()[0])\n        return 'cpu'\n\n\ndev = pick_device()\nmdl.to(dev)\n# deliberately fp32 -- DeBERTa-v3's disentangled attention can overflow in half\n# precision, and a silent NaN here would read as a finding rather than a bug\ni2l = {v.lower(): k for k, v in mdl.config.id2label.items()}\nI_ENT, I_CON = i2l['entailment'], i2l['contradiction']\nprint('entail idx', I_ENT, ' contradict idx', I_CON)"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"@torch.no_grad()\ndef affirm(premises, hypothesis, batch=256):\n    '''P(entail) on the entail-vs-contradict axis, one score per premise.'''\n    hyp = [hypothesis] * len(premises)\n    out = np.empty(len(premises), dtype=np.float32)\n    for i in range(0, len(premises), batch):\n        enc = tok(premises[i:i + batch], hyp[i:i + batch], truncation=True,\n                  padding=True, max_length=96, return_tensors='pt').to(dev)\n        lg = mdl(**enc).logits.float()\n        out[i:i + batch] = torch.softmax(lg[:, [I_ENT, I_CON]], -1)[:, 0].cpu().numpy()\n    return out\n\n\ndemo = [\n    ('Rotura de menisco interno.',            'Medial Meniscus'),   # Spanish, affirmed\n    ('Menisco interno sin alteraciones.',     'Medial Meniscus'),   # Spanish, denied\n    ('Efuzyon izlenmedi.',                    'Effusion'),          # Turkish, negation AFTER\n    ('Eklem icinde belirgin efuzyon mevcut.', 'Effusion'),          # Turkish, affirmed\n    ('Geen aanwijzingen voor een bakercyste.', \"Baker's\"),          # Dutch, denied\n    ('Kleine bakercyste in de fossa poplitea.', \"Baker's\"),         # Dutch, affirmed\n]\nprint(f'{\"premise\":42s} {\"finding\":17s} {\"affirmed\":>9s}')\nfor text, lab in demo:\n    print(f'{text[:41]:42s} {lab:17s} {affirm([text], HYP[lab])[0]:9.3f}')"},{"cell_type":"markdown","metadata":{},"source":"Three languages, three negation grammars, no vocabulary anywhere. The Turkish pair is the\ninteresting one: the same word `efuzyon` appears in both, and only the trailing `izlenmedi`\nseparates them."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# Reports are templated, so the same sentence recurs across studies -- score each unique\n# sentence once. 82k fragments collapse to ~34k unique, a 2.4x saving on a 12-hypothesis pass.\nper_study = [sentences(r) for r in train.Report]\nuniq = sorted({s for ss in per_study for s in ss}, key=len)      # length-sorted: less padding\nu_index = {s: i for i, s in enumerate(uniq)}\nowner = [np.array([u_index[s] for s in ss]) for ss in per_study]\nprint(f'{sum(len(s) for s in per_study):,} sentences -> {len(uniq):,} unique')\n\n# ~400k forward passes. That is 30-40 minutes on a working GPU and roughly 13 hours on\n# CPU, which is past the session limit -- so if the probe above fell back to CPU, read the\n# cached scores instead. The cache is this exact loop's output, published as a dataset so\n# the notebook still produces its results on a session that got a GPU torch cannot drive.\nCACHE = glob.glob('/kaggle/input/**/nli_entailment_scores.csv', recursive=True)\n\nimport time\nt0 = time.time()\nif dev == 'cpu' and CACHE:\n    print('running on cpu -- loading cached entailment scores from', CACHE[0])\n    cached = pd.read_csv(CACHE[0]).set_index('StudyInstanceUID').loc[train.StudyInstanceUID]\n    NLI = cached[LABELS].to_numpy(dtype=np.float32)\n    SCORED_HERE = False\nelse:\n    NLI = np.zeros((len(train), len(LABELS)), dtype=np.float32)\n    for j, lab in enumerate(LABELS):\n        sc = affirm(uniq, HYP[lab])\n        NLI[:, j] = [sc[ix].max() for ix in owner]\n        print(f'  {lab:18s} {time.time() - t0:6.1f}s')\n    SCORED_HERE = True\n\nnp.save('nli_scores.npy', NLI)\nassert NLI.shape == (len(train), len(LABELS)) and np.isfinite(NLI).all()\nprint(f'{\"scored\" if SCORED_HERE else \"loaded\"} {NLI.shape} in {time.time() - t0:.0f}s')"},{"cell_type":"markdown","metadata":{},"source":"## 3. Sentence by sentence, not report by report\n\nThe obvious implementation feeds the whole report as the premise. It is also the wrong one, and\nthe gap is large enough to be worth a cell of its own."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"from sklearn.metrics import roc_auc_score\n\nG = gold[LABELS].values.astype(int)\ngi = train.StudyInstanceUID.isin(gold.StudyInstanceUID).values\nNLI_G = NLI[gi]                                  # sentence-max scores for the 58\n\n# 58 reports x 12 hypotheses is 696 passes -- cheap enough to run even on cpu\nfull_prem = [str(r)[:2000] for r in gold.Report]\nFULL_G = np.column_stack([affirm(full_prem, HYP[l]) for l in LABELS])\n\nab = pd.DataFrame({\n    'finding': LABELS,\n    'whole report': [roc_auc_score(G[:, j], FULL_G[:, j]) for j in range(len(LABELS))],\n    'sentence max': [roc_auc_score(G[:, j], NLI_G[:, j]) for j in range(len(LABELS))],\n})\nprint(ab.round(3).to_string(index=False))\nprint(f\"\\nmean AUC   whole report {ab['whole report'].mean():.3f}\"\n      f\"   sentence max {ab['sentence max'].mean():.3f}\")"},{"cell_type":"markdown","metadata":{},"source":"A knee MRI report is a list of independent observations about eight or nine structures. Asked as\none premise, the model has to decide whether a 980-character list entails a statement about one\nof them, and the answer is always \"sort of\". Asked one clause at a time, the question is the one\nthe model was trained on.\n\nEverything below uses the sentence-level maxima."},{"cell_type":"markdown","metadata":{},"source":"## 4. Where entailment wins, and where it does not\n\nThe keyword extractor is binary, so its ROC has a single interior point and its AUC equals its\nbalanced accuracy exactly. That makes AUC a like-for-like comparison between a binary rule and a\ncontinuous score, with no threshold to choose on either side.\n\n58 studies is a small test set, so every difference below carries a bootstrap interval. The\nverdict column only says NLI or keyword when the 95% interval excludes zero."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"KW_G = pd.DataFrame([extract(t) for t in gold.Report])[LABELS].values\nrng = np.random.default_rng(0)\nB, n = 4000, len(gold)\n\nrows = []\nfor j, lab in enumerate(LABELS):\n    g, s, k = G[:, j], NLI_G[:, j], KW_G[:, j]\n    a_nli, a_kw = roc_auc_score(g, s), roc_auc_score(g, k)\n    d = []\n    for _ in range(B):\n        i = rng.integers(0, n, n)\n        if 0 < G[i, j].sum() < n:\n            d.append(roc_auc_score(G[i, j], s[i]) - roc_auc_score(G[i, j], k[i]))\n    d = np.array(d)\n    lo, hi = np.quantile(d, [.025, .975])\n    rows.append({'finding': lab, 'pos': int(g.sum()), 'keyword': a_kw, 'NLI': a_nli,\n                 'delta': a_nli - a_kw, '95% lo': lo, '95% hi': hi,\n                 'verdict': 'NLI' if lo > 0 else ('keyword' if hi < 0 else 'tie')})\nres = pd.DataFrame(rows).sort_values('delta', ascending=False)\nprint(res.round(3).to_string(index=False))\nprint(f\"\\nmean AUC   keyword {res.keyword.mean():.3f}   NLI {res.NLI.mean():.3f}\")\nprint('verdicts:', res.verdict.value_counts().to_dict())"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, ax = plt.subplots(figsize=(9.5, 5.6))\nr = res.sort_values('NLI')\ny = np.arange(len(r))\nfor yi, (_, row) in zip(y, r.iterrows()):\n    c = {'NLI': NLIC, 'keyword': KWC, 'tie': '#b6b9be'}[row.verdict]\n    ax.plot([row.keyword, row.NLI], [yi, yi], color=c, lw=2.4, zorder=1,\n            solid_capstyle='round')\nax.scatter(r.keyword, y, s=52, color=KWC, zorder=3, label='keyword (v1)')\nax.scatter(r.NLI, y, s=52, color=NLIC, zorder=3, label='entailment (v2)')\nax.axvline(.5, color=GRID, lw=1, zorder=0)\nax.set_yticks(y); ax.set_yticklabels(r.finding)\nfor yi, (_, row) in zip(y, r.iterrows()):\n    if row.verdict != 'tie':\n        ax.text(max(row.keyword, row.NLI) + .012, yi, row.verdict, va='center', fontsize=9,\n                color=NLIC if row.verdict == 'NLI' else KWC, fontweight='bold')\nax.set_xlim(.40, 1.0); ax.set_xlabel('AUC on the 58 labelled studies')\nax.set_title('Entailment moves the meniscus findings and loses the compartment ones')\nax.legend(frameon=False, loc='lower right')\nax.text(.40, -1.35, 'grey = the 95% bootstrap interval on the difference includes zero',\n        fontsize=9, color=MUTED)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Three findings improve beyond the noise, one gets worse beyond the noise, and eight are\nindistinguishable on 58 studies.\n\n**Medial meniscus, +0.32.** The largest single movement in either notebook, and the correction\nto v1's claim that this finding was near-unrecoverable. It was recoverable; the regex could not\nreach it.\n\n**Lateral OA, -0.19.** Entailment lands *below chance* here. That is not the model being\nmediocre, it is the model being unable to answer the question at all, and section 5 shows why.\n\n**Eight ties.** Not \"no difference\" -- \"58 studies cannot tell\". Effusion at +0.11 with an\ninterval of [-0.06, +0.27] may well be a real improvement that this test set cannot resolve."},{"cell_type":"markdown","metadata":{},"source":"## 5. Which method actually knows the side\n\nFour of the twelve findings are the same pathology on two sides: medial and lateral meniscus,\nmedial and lateral compartment OA. A model that fires on \"there is a meniscal tear\" without\nresolving *which* meniscus will still post a good AUC on both, because the two labels co-occur\nin 12 of the 58 studies.\n\nThe test for that is the cross-AUC: score with the **medial** hypothesis, then measure how well\nit predicts the **lateral** label. A method that binds the side should be at chance off-diagonal."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"PAIRS = [('Medial Meniscus', 'Lateral Meniscus'), ('Medial OA', 'Lateral OA')]\nix = {l: i for i, l in enumerate(LABELS)}\n\nrows = []\nfor a, b in PAIRS:\n    for src, tgt in ((a, b), (b, a)):\n        rows.append({'score from': src,\n                     'AUC on own label': roc_auc_score(G[:, ix[src]], NLI_G[:, ix[src]]),\n                     'AUC on other side': roc_auc_score(G[:, ix[tgt]], NLI_G[:, ix[src]]),\n                     'kw own': roc_auc_score(G[:, ix[src]], KW_G[:, ix[src]]),\n                     'kw other side': roc_auc_score(G[:, ix[tgt]], KW_G[:, ix[src]])})\nlat = pd.DataFrame(rows)\nprint(lat.round(3).to_string(index=False))\n\nfor a, b in PAIRS:\n    ya, yb = G[:, ix[a]].astype(bool), G[:, ix[b]].astype(bool)\n    print(f'\\n{a} / {b} co-occurrence in the 58: '\n          f'both {int((ya & yb).sum())}, only {a.split()[0].lower()} {int((ya & ~yb).sum())}, '\n          f'only {b.split()[0].lower()} {int((~ya & yb).sum())}, neither {int((~ya & ~yb).sum())}')"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, axes = plt.subplots(1, 2, figsize=(11.5, 4.3), sharey=True)\nfor ax, (meth, own_c, oth_c, tag) in zip(\n        axes, [('NLI', NLIC, '#f3bda2', 'entailment'), ('kw', KWC, '#a9c8ee', 'keyword')]):\n    names = list(lat['score from'])\n    own = lat['AUC on own label'] if meth == 'NLI' else lat['kw own']\n    oth = lat['AUC on other side'] if meth == 'NLI' else lat['kw other side']\n    x = np.arange(len(names))\n    ax.bar(x - .19, own, .36, color=own_c, label='predicts its own label')\n    ax.bar(x + .19, oth, .36, color=oth_c, label='predicts the other side')\n    ax.axhline(.5, color=WARN, lw=1.2, ls=(0, (4, 3)))\n    for xi, (a, b) in enumerate(zip(own, oth)):\n        ax.text(xi - .19, a + .012, f'{a:.2f}', ha='center', fontsize=9, color=INK)\n        ax.text(xi + .19, b + .012, f'{b:.2f}', ha='center', fontsize=9, color=MUTED)\n    ax.set_xticks(x)\n    ax.set_xticklabels([n.replace(' Meniscus', '\\nmeniscus').replace(' OA', '\\ncompartment OA')\n                        for n in names], fontsize=9)\n    ax.set_ylim(0, 1.0); ax.set_title(tag)\n    ax.set_yticks([0, .5, 1.0])\naxes[0].set_ylabel('AUC'); axes[0].legend(frameon=False, fontsize=9, loc='upper right')\naxes[0].text(-.55, .52, 'chance', fontsize=8.5, color=WARN)\nfig.suptitle('Does the score know which side it is talking about?',\n             x=.078, ha='left', fontsize=13, fontweight='bold', y=1.02)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Read the meniscus rows first. The entailment score for the medial meniscus reaches **0.87 on its\nown label and 0.48 on the lateral one** -- exactly chance. It is not detecting \"a meniscal tear\",\nit is answering the question that was asked, in Spanish, Turkish and Greek.\n\nNow read the keyword rows for the same finding: **0.558 on its own label and 0.655 on the other\nside.** The v1 extractor predicts the wrong meniscus better than the right one. Its pattern allows\nup to two words between the side and the structure in either direction, and in a sentence naming\nboth menisci that window straddles them. That is a bug in v1, and it is the mechanism behind the\n0.558 I published.\n\nThe OA rows are the opposite story and the reason for the -0.19. Both entailment scores sit near\nchance on both labels: given a clause about arthrosis of a compartment, the model cannot attach\nthe modifier to the finding. The keyword extractor's explicit side-word requirement is crude, but\ncrude is enough here, and it wins.\n\n**Neither method is better. They fail in complementary ways** -- entailment reads the language and\ncannot always bind a modifier; regex binds the modifier by construction and cannot read the\nlanguage."},{"cell_type":"markdown","metadata":{},"source":"## 6. The threshold wall\n\nRanking better is not the same as labelling better. A blend of the two scores ranks best of all,\nso the natural next step is to threshold it per finding, fitted on the 58 labelled studies, and\nship binary labels like v1 did.\n\nThat step gives the entire gain back. Below, each threshold is chosen on 57 studies and applied\nto the one held out, so the number is what you would actually get -- not the in-sample optimum."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"from scipy.stats import rankdata\n\ndef bal(g, p):\n    return (p[g == 1].mean() + 1 - p[g == 0].mean()) / 2\n\ndef loo_bal(g, s):\n    '''threshold fitted on n-1 studies, scored on the held-out one'''\n    grid = np.unique(np.quantile(s.astype(float), np.linspace(.05, .95, 40)))\n    pred = np.empty(len(g), int)\n    for i in range(len(g)):\n        m = np.ones(len(g), bool); m[i] = False\n        b = [bal(g[m], (s[m] >= th).astype(int)) for th in grid]\n        pred[i] = int(s[i] >= grid[int(np.argmax(b))])\n    return bal(g, pred)\n\nBLEND_G = np.column_stack([(rankdata(NLI_G[:, j]) / n + KW_G[:, j]) / 2\n                           for j in range(len(LABELS))])\n\nw = pd.DataFrame({\n    'finding': LABELS,\n    'AUC keyword': [roc_auc_score(G[:, j], KW_G[:, j]) for j in range(len(LABELS))],\n    'AUC NLI': [roc_auc_score(G[:, j], NLI_G[:, j]) for j in range(len(LABELS))],\n    'AUC blend': [roc_auc_score(G[:, j], BLEND_G[:, j]) for j in range(len(LABELS))],\n    'binary keyword': [bal(G[:, j], KW_G[:, j]) for j in range(len(LABELS))],\n    'binary blend, held out': [loo_bal(G[:, j], BLEND_G[:, j]) for j in range(len(LABELS))],\n})\nprint(w.round(3).to_string(index=False))\nprint('\\nmean   ' + '   '.join(f'{c} {w[c].mean():.3f}' for c in w.columns[1:]))"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, ax = plt.subplots(figsize=(8.2, 4.6))\ngroups = ['keyword\\n(v1)', 'entailment', 'blend']\nauc = [w['AUC keyword'].mean(), w['AUC NLI'].mean(), w['AUC blend'].mean()]\nbina = [w['binary keyword'].mean(), np.nan, w['binary blend, held out'].mean()]\nx = np.arange(3)\nax.bar(x - .19, auc, .36, color=BLEND, label='as a ranking (AUC)')\nax.bar(x + .19, bina, .36, color='#c2c5ca', label='as binary labels, threshold held out')\nfor xi, v in zip(x, auc):\n    ax.text(xi - .19, v + .006, f'{v:.3f}', ha='center', fontsize=10, color=INK)\nfor xi, v in zip(x, bina):\n    if not np.isnan(v):\n        ax.text(xi + .19, v + .006, f'{v:.3f}', ha='center', fontsize=10, color=MUTED)\nax.annotate('', xy=(2 - .19, auc[2] - .004), xytext=(0 - .19, auc[0] + .004),\n            arrowprops=dict(arrowstyle='->', color=BLEND, lw=1.6,\n                            connectionstyle='arc3,rad=-.28'))\nax.text(1 - .30, .795, f'+{auc[2] - auc[0]:.3f} AUC', color=BLEND, fontsize=10,\n        fontweight='bold', ha='center')\nax.text(2 + .19, bina[2] - .055, f'{bina[2] - bina[0]:+.3f}\\nas labels', ha='center',\n        fontsize=10, color=WARN, fontweight='bold')\nax.set_xticks(x); ax.set_xticklabels(groups)\nax.set_ylim(.5, .82); ax.set_ylabel('mean over the 12 findings')\nax.set_title('The ranking improves; the labels do not')\nax.legend(frameon=False, fontsize=9.5, loc='upper left')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Read the last two columns against each other. Roughly eight points of ranking gain arrive as\nabout one point of labelling gain -- the binarised blend lands within noise of the v1 keyword\nbaseline it was meant to beat.\n\nTwelve thresholds fitted on 58 studies is roughly five positive examples per threshold. The\nthreshold that is optimal on 57 studies is not optimal on the 58th, and the variance of that\nchoice is larger than the gain you were trying to capture. This is not fixable by a better\nsearch -- it is the size of the labelled set.\n\nThree consequences for anyone training on this output:\n\n1. **Take the probabilities, not a binarisation of them.** Soft targets keep the ranking\n   information that a threshold throws away.\n2. **Do not spend the 58 studies on threshold fitting.** They are the only clean signal in the\n   competition and this is the least valuable thing to do with them.\n3. **The competition metric decides the threshold, not the labels.** If it is AUC-like, you never\n   need one. If it needs a hard call, tune it on your own model's validation output, downstream\n   of everything, where you have more than 58 rows."},{"cell_type":"markdown","metadata":{},"source":"## 6b. The contrastive fix, tested\n\nThe closing section of the previous version promised an untested fix for the compartment-OA\nfailure: ask a *pair* of side hypotheses and use the difference, instead of scoring each side\nindependently. Here it is, tested. For each sentence, score both sides and combine as\n`s_own * (s_own - s_other + 1) / 2` -- the entailment strength, discounted by how much the\nsentence also affirms the opposite side -- then take the per-report maximum as before.\n\nIt moves exactly the two findings it was aimed at, and nothing it was not:\n\n- **Medial OA 0.602 -> 0.785, Lateral OA 0.471 -> 0.709.** Both now clear the keyword floor\n  (0.720 / 0.661) that section 4 reported them losing to.\n- **The menisci get *worse* under the same treatment** (medial 0.871 -> 0.743), so the fix is\n  applied only to the OA pair. The meniscus hypotheses already bind the side; subtracting the\n  other side just adds noise to a solved problem.\n- **Honesty about the intervals:** the paired bootstrap gives [-0.03, +0.40] and [-0.00, +0.48]\n  -- large deltas, formally not significant. That is the companion notebook's power curve doing exactly\n  what it said: a +0.2 improvement on 11-15 positives sits at the edge of what 58 studies can see.\n- One asymmetry survives: the medial-OA contrast score still predicts the lateral label almost\n  as well as its own, so laterality binding for OA is only half-fixed. The magnitude question\n  (\"is there compartment OA?\") is answered; the side question is partly open.\n\n`weak_labels_v2.csv` now carries four extra `_contrast` columns for these findings."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# Contrastive scores: computed live on a working GPU, else loaded from the same cache dataset.\nCPATH = glob.glob('/kaggle/input/**/nli_contrast_scores.csv', recursive=True)\nif dev != 'cpu':\n    def contrast_pair(a, b):\n        sa, sb = affirm(uniq, HYP[a]), affirm(uniq, HYP[b])\n        soft_a, soft_b = sa * (sa - sb + 1) / 2, sb * (sb - sa + 1) / 2\n        return ([max(soft_a[u_index[s]] for s in ss) for ss in per_study],\n                [max(soft_b[u_index[s]] for s in ss) for ss in per_study])\n    CON = {}\n    for a, b in [('Medial Meniscus', 'Lateral Meniscus'), ('Medial OA', 'Lateral OA')]:\n        CON[a], CON[b] = contrast_pair(a, b)\n    CON = pd.DataFrame(CON, index=train.index)\nelse:\n    CON = pd.read_csv(CPATH[0]).set_index('StudyInstanceUID')\n    CON = CON.loc[train.StudyInstanceUID].reset_index(drop=True)\n    CON.columns = [c.replace('_contrast', '') for c in CON.columns]\n\ncon_g = CON[gi].reset_index(drop=True)\nrows = []\nfor lab in ['Medial Meniscus', 'Lateral Meniscus', 'Medial OA', 'Lateral OA']:\n    j = LABELS.index(lab)\n    rows.append({'finding': lab,\n                 'keyword': roc_auc_score(G[:, j], KW_G[:, j]),\n                 'plain NLI': roc_auc_score(G[:, j], NLI_G[:, j]),\n                 'contrast': roc_auc_score(G[:, j], con_g[lab].values)})\nprint(pd.DataFrame(rows).round(3).to_string(index=False))\nprint('\\ncontrast helps the OA pair, hurts the menisci -- applied to OA only in the output')"},{"cell_type":"markdown","metadata":{},"source":"## 7. The output\n\n`weak_labels_v2.csv`, one row per study, and for each of the twelve findings three columns:\n\n| column | what it is |\n|---|---|\n| `<finding>` | the blend -- rank-averaged NLI and keyword, in [0, 1]. Use this one. |\n| `<finding>_nli` | raw entailment score, the affirm-vs-deny probability |\n| `<finding>_kw` | the v1 keyword binary, kept so you can reproduce or override |\n\nPlus `has_gold` (whether this study is one of the 58) and, in the printed table below,\na per-finding `AUC` you can use as a loss weight."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"BLEND = np.column_stack([(rankdata(NLI[:, j]) / len(train) + KW_ALL[:, j]) / 2\n                         for j in range(len(LABELS))])\n\nout = pd.DataFrame({'StudyInstanceUID': train.StudyInstanceUID})\nfor j, lab in enumerate(LABELS):\n    out[lab] = BLEND[:, j].round(4)\n    out[f'{lab}_nli'] = NLI[:, j].round(4)\n    out[f'{lab}_kw'] = KW_ALL[:, j].astype(int)\nfor lab in CON.columns:\n    out[f'{lab}_contrast'] = np.asarray(CON[lab]).round(4)\nout['has_gold'] = train.StudyInstanceUID.isin(gold.StudyInstanceUID).values\nout.to_csv('weak_labels_v2.csv', index=False)\n\nquality = pd.DataFrame({\n    'finding': LABELS,\n    'AUC of the blend': [roc_auc_score(G[:, j], BLEND_G[:, j]) for j in range(len(LABELS))],\n    'use as loss weight': [round(max(0., 2 * roc_auc_score(G[:, j], BLEND_G[:, j]) - 1), 3)\n                           for j in range(len(LABELS))],\n}).sort_values('AUC of the blend', ascending=False)\nprint(quality.round(3).to_string(index=False))\nprint(f'\\nweak_labels_v2.csv  {out.shape[0]:,} rows x {out.shape[1]} columns')\nout.head(3)"},{"cell_type":"markdown","metadata":{},"source":"## What changed, and what I got wrong\n\n**v1 said the meniscus findings were near-unrecoverable from text.** They are not. Medial meniscus\ngoes 0.558 -> 0.87 AUC, and the cross-AUC check shows the gain is real laterality binding, not\nco-occurrence. The 0.558 was a bug in my own regex, which predicts the wrong meniscus better than\nthe right one.\n\n**v1 shipped binary labels.** That was the wrong output format. Section 6 shows the binarisation\nstep destroys the improvement it is applied to; the probabilities are strictly more useful and\ncost nothing to ship.\n\n**What still stands from v1.** The 58 labelled studies are not a random sample -- ACL is positive\nin 41% of them against a corpus rate near 20% -- so they remain unusable for setting an output\nprior. And the keyword floor is a floor: it still wins on compartment OA, decisively.\n\n**The promised fix is now tested** (section 6b): a contrastive pair of hypotheses lifts\ncompartment OA above the keyword floor -- the one place entailment was losing -- while the same\ntreatment hurts the menisci and is therefore not applied to them. With it, every finding's best\ntext score clears 0.66. Side-binding for OA remains only half-fixed, and the deltas sit at the\nedge of what 58 studies can resolve; both caveats are in the section.\n\nMeasurements here are on 58 studies. Treat every interval in section 4 as the real precision of\nevery claim above, including mine."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.11"}},"nbformat":4,"nbformat_minor":5}