{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"3504ba22-c3cb-44e7-bccb-6ac70f0bb591","cell_type":"markdown","source":"# RSNA Knee MRI — Multimodal Abnormality Detection: EDA\n\n**Goal:** for each *study*, predict 12 binary probabilities. Metric = **macro-averaged AUC-ROC** over the 12 labels.\n\nThis notebook answers, in order:\n\n1. What is in each file, and how do they join together?\n2. **Labels** — how many studies are actually labeled? prevalence, imbalance, co-occurrence.\n3. **Reports** — languages, length, and *can we recover labels from text?* (this is the crux of the competition)\n4. **Series metadata** — protocol heterogeneity, which planes/contrasts exist per study.\n5. **DICOM pixel level** — transfer syntaxes, geometry, scanners, intensity ranges, decode cost.\n6. **Train vs test** — domain shift checks.\n7. **Modelling implications** — validation design, what the metric rewards.\n\nEvery section ends with a short *Takeaway* cell so the analysis is interpretable, not just plotted.","metadata":{}},{"id":"f5aef201-3eb4-4fc0-8c28-de41552c5841","cell_type":"code","source":"import os, re, glob, json, math, random, warnings, time\nfrom collections import Counter, defaultdict\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\n\nwarnings.filterwarnings(\"ignore\")\npd.set_option(\"display.max_columns\", 100)\npd.set_option(\"display.width\", 200)\npd.set_option(\"display.max_colwidth\", 300)\n\nmpl.rcParams.update({\n    \"figure.dpi\": 110, \"font.size\": 9, \"axes.grid\": True,\n    \"grid.alpha\": .25, \"axes.spines.top\": False, \"axes.spines.right\": False,\n})\n\ntry:\n    import seaborn as sns\n    sns.set_palette(\"deep\")\n    HAS_SNS = True\nexcept ImportError:\n    HAS_SNS = False\n\nSEED = 42\nrandom.seed(SEED); np.random.seed(SEED)\nprint(\"pandas\", pd.__version__, \"| numpy\", np.__version__, \"| seaborn\", HAS_SNS)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:16:17.147091Z","iopub.execute_input":"2026-08-05T17:16:17.147279Z","iopub.status.idle":"2026-08-05T17:16:20.213853Z","shell.execute_reply.started":"2026-08-05T17:16:17.147258Z","shell.execute_reply":"2026-08-05T17:16:20.213034Z"}},"outputs":[],"execution_count":null},{"id":"91db596c-4636-46e6-a4d0-55fb55e8fe60","cell_type":"markdown","source":"## 1. Locate the competition directory\n","metadata":{}},{"id":"00641ef7-4fe7-47fa-a121-a5456d2f55ce","cell_type":"code","source":"DATA = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\nprint(\"DATA =\", DATA)\nfor p in sorted(DATA.glob(\"*\")):\n    size = f\"{p.stat().st_size/1e6:8.2f} MB\" if p.is_file() else \"     <dir>\"\n    print(f\"  {size}  {p.name}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:26:44.670239Z","iopub.execute_input":"2026-08-05T17:26:44.670489Z","iopub.status.idle":"2026-08-05T17:26:44.680476Z","shell.execute_reply.started":"2026-08-05T17:26:44.67047Z","shell.execute_reply":"2026-08-05T17:26:44.679764Z"}},"outputs":[],"execution_count":null},{"id":"bb4c25d4-ef16-4f48-82bd-ab43c6ab2d52","cell_type":"code","source":"LABELS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n          \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n          \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"]\n\nGROUPS = {                      # a clinically meaningful grouping used throughout\n    \"Ligament\":  [\"ACL\", \"MCL\"],\n    \"Meniscus\":  [\"Medial Meniscus\", \"Lateral Meniscus\"],\n    \"OA\":        [\"Medial OA\", \"Lateral OA\", \"PF OA\"],\n    \"Inflam\":    [\"Effusion\", \"Synovitis\", \"Baker's\"],\n    \"Trauma\":    [\"Contusion\", \"Fracture\"],\n}\nLABEL2GROUP = {l: g for g, ls in GROUPS.items() for l in ls}\n\ntrain        = pd.read_csv(DATA / \"train.csv\")\ntrain_series = pd.read_csv(DATA / \"train_series.csv\")\ntest         = pd.read_csv(DATA / \"test.csv\")\ntest_series  = pd.read_csv(DATA / \"test_series.csv\")\nsample_sub   = pd.read_csv(DATA / \"sample_submission.csv\")\n\nfor name, df in [(\"train\", train), (\"train_series\", train_series),\n                 (\"test\", test), (\"test_series\", test_series), (\"sample_submission\", sample_sub)]:\n    print(f\"{name:<20} {df.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:27:01.025211Z","iopub.execute_input":"2026-08-05T17:27:01.02543Z","iopub.status.idle":"2026-08-05T17:27:01.293716Z","shell.execute_reply.started":"2026-08-05T17:27:01.025412Z","shell.execute_reply":"2026-08-05T17:27:01.293021Z"}},"outputs":[],"execution_count":null},{"id":"2bbd73c9-27cc-4832-8eb3-c0f299b4259a","cell_type":"code","source":"train.head(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:27:25.178471Z","iopub.execute_input":"2026-08-05T17:27:25.178734Z","iopub.status.idle":"2026-08-05T17:27:25.211508Z","shell.execute_reply.started":"2026-08-05T17:27:25.178688Z","shell.execute_reply":"2026-08-05T17:27:25.210784Z"}},"outputs":[],"execution_count":null},{"id":"2d8a6e98-ff97-4af3-b6e1-fffc12312151","cell_type":"code","source":"# dtypes + missingness for train.csv\ninfo = pd.DataFrame({\n    \"dtype\":     train.dtypes.astype(str),\n    \"n_missing\": train.isna().sum(),\n    \"pct_missing\": (train.isna().mean() * 100).round(2),\n    \"n_unique\":  train.nunique(dropna=True),\n})\ninfo","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:28:23.543499Z","iopub.execute_input":"2026-08-05T17:28:23.543785Z","iopub.status.idle":"2026-08-05T17:28:23.577633Z","shell.execute_reply.started":"2026-08-05T17:28:23.543756Z","shell.execute_reply":"2026-08-05T17:28:23.577094Z"}},"outputs":[],"execution_count":null},{"id":"90f3e417-44c7-4d72-be56-6c46c698a33a","cell_type":"markdown","source":"### The single most important structural fact\n\nThe description says *\"Only a small subset of training studies carry per-condition labels\"*. That shows up as **NaN** in the 12 label columns. Everything about strategy follows from this ratio, so quantify it first.","metadata":{}},{"id":"d3bc1377-c0dd-4e7a-a5dc-a71e56c45c2e","cell_type":"code","source":"lab_present = train[LABELS].notna().all(axis=1)\nlab_partial  = train[LABELS].notna().any(axis=1) & ~lab_present\n\nn_all, n_lab, n_part = len(train), lab_present.sum(), lab_partial.sum()\nprint(f\"studies total            : {n_all:,}\")\nprint(f\"fully labeled            : {n_lab:,}  ({n_lab/n_all:.1%})\")\nprint(f\"partially labeled        : {n_part:,}\")\nprint(f\"unlabeled (report only)  : {n_all - n_lab - n_part:,}  ({(n_all-n_lab-n_part)/n_all:.1%})\")\nprint(f\"has a report             : {train['Report'].notna().sum():,}  ({train['Report'].notna().mean():.1%})\")\n\nlabeled   = train[lab_present].copy()\nunlabeled = train[~lab_present].copy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:29:24.595635Z","iopub.execute_input":"2026-08-05T17:29:24.595925Z","iopub.status.idle":"2026-08-05T17:29:24.611823Z","shell.execute_reply.started":"2026-08-05T17:29:24.595905Z","shell.execute_reply":"2026-08-05T17:29:24.611134Z"}},"outputs":[],"execution_count":null},{"id":"4638c6a8-c0e4-4290-856f-45f5d0fb9d50","cell_type":"markdown","source":"---\n## 2. Labels: prevalence, imbalance, co-occurrence","metadata":{}},{"id":"14c546d7-92a4-4e23-b6b9-873099f8aedd","cell_type":"code","source":"prev = (labeled[LABELS].mean().rename(\"prevalence\").to_frame()\n        .assign(n_pos=labeled[LABELS].sum().astype(int),\n                n_neg=(len(labeled) - labeled[LABELS].sum()).astype(int),\n                group=[LABEL2GROUP[l] for l in LABELS])\n        .sort_values(\"prevalence\", ascending=False))\nprev[\"prevalence\"] = prev[\"prevalence\"].round(4)\nprev","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:29:53.365673Z","iopub.execute_input":"2026-08-05T17:29:53.366037Z","iopub.status.idle":"2026-08-05T17:29:53.382481Z","shell.execute_reply.started":"2026-08-05T17:29:53.366019Z","shell.execute_reply":"2026-08-05T17:29:53.381717Z"}},"outputs":[],"execution_count":null},{"id":"ad07fb20-c430-4328-b887-a06a97b9c3d5","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(7, 4))\np = prev.sort_values(\"prevalence\")\ncolors = {\"Ligament\": \"#4C72B0\", \"Meniscus\": \"#DD8452\", \"OA\": \"#55A868\",\n          \"Inflam\": \"#C44E52\", \"Trauma\": \"#8172B3\"}\nax.barh(p.index, p[\"prevalence\"], color=[colors[g] for g in p[\"group\"]])\nfor y, (v, n) in enumerate(zip(p[\"prevalence\"], p[\"n_pos\"])):\n    ax.text(v + .005, y, f\"{v:.1%} (n={n})\", va=\"center\", fontsize=8)\nax.set_xlim(0, p[\"prevalence\"].max() * 1.35)\nax.set_xlabel(\"positive rate among fully-labeled studies\")\nax.set_title(\"Label prevalence — colour = clinical group\")\nhandles = [plt.Rectangle((0,0),1,1, color=c) for c in colors.values()]\nax.legend(handles, colors.keys(), fontsize=8, loc=\"lower right\", frameon=False)\nplt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:30:25.795345Z","iopub.execute_input":"2026-08-05T17:30:25.795588Z","iopub.status.idle":"2026-08-05T17:30:26.085833Z","shell.execute_reply.started":"2026-08-05T17:30:25.795568Z","shell.execute_reply":"2026-08-05T17:30:26.085253Z"}},"outputs":[],"execution_count":null},{"id":"143d31e9-3678-4fd3-aa94-d767c4cff603","cell_type":"markdown","source":"**How to read this:** macro-AUC weights every label equally, so a label with only a few dozen positives contributes 1/12 of the score *and* carries most of the variance. Any label under a few hundred positives is where the leaderboard will be won or lost — and also where your local CV will be noisiest.","metadata":{}},{"id":"09853b37-d982-4793-bf8b-e807b0901e2d","cell_type":"code","source":"# how many findings per study? (comorbidity is the norm in knee MRI)\nnpos = labeled[LABELS].sum(axis=1)\nfig, ax = plt.subplots(1, 2, figsize=(10, 3.4))\nax[0].hist(npos, bins=np.arange(-.5, npos.max() + 1.5), color=\"#4C72B0\", edgecolor=\"w\")\nax[0].set_xlabel(\"# positive labels in a study\"); ax[0].set_ylabel(\"studies\")\nax[0].set_title(f\"Findings per study (mean {npos.mean():.2f})\")\n\ngrp = pd.DataFrame({g: labeled[ls].max(axis=1) for g, ls in GROUPS.items()})\nax[1].bar(grp.columns, grp.mean(), color=\"#55A868\")\nax[1].set_ylabel(\"share of studies with >=1 finding\"); ax[1].set_title(\"Prevalence by clinical group\")\nax[1].tick_params(axis=\"x\", rotation=20)\nplt.tight_layout(); plt.show()\n\nprint(\"completely normal studies (all 12 = 0):\", f\"{(npos == 0).mean():.1%}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:30:50.605033Z","iopub.execute_input":"2026-08-05T17:30:50.605291Z","iopub.status.idle":"2026-08-05T17:30:50.822468Z","shell.execute_reply.started":"2026-08-05T17:30:50.605268Z","shell.execute_reply":"2026-08-05T17:30:50.821244Z"}},"outputs":[],"execution_count":null},{"id":"fb34fd33-242a-4166-a14b-c071a0fe85d3","cell_type":"code","source":"# co-occurrence: P(col | row) — asymmetric and far more informative than raw correlation\nM = labeled[LABELS].values.astype(float)\njoint = M.T @ M\ncond = joint / np.clip(M.sum(0)[:, None], 1, None)   # rows = conditioning label\ncond_df = pd.DataFrame(cond, index=LABELS, columns=LABELS)\n\nfig, ax = plt.subplots(1, 2, figsize=(13.5, 5.2))\nim = ax[0].imshow(cond_df.values, cmap=\"viridis\", vmin=0, vmax=1)\nax[0].set_xticks(range(12)); ax[0].set_xticklabels(LABELS, rotation=90, fontsize=7)\nax[0].set_yticks(range(12)); ax[0].set_yticklabels(LABELS, fontsize=7)\nax[0].set_title(\"P(column | row)\"); ax[0].grid(False)\nfor i in range(12):\n    for j in range(12):\n        v = cond_df.values[i, j]\n        ax[0].text(j, i, f\"{v:.2f}\"[1:], ha=\"center\", va=\"center\", fontsize=6,\n                   color=\"w\" if v < .6 else \"k\")\nplt.colorbar(im, ax=ax[0], fraction=.046)\n\n# phi (= Pearson on binaries) for symmetric association strength\nphi = labeled[LABELS].corr()\nim2 = ax[1].imshow(phi.values, cmap=\"RdBu_r\", vmin=-.6, vmax=.6)\nax[1].set_xticks(range(12)); ax[1].set_xticklabels(LABELS, rotation=90, fontsize=7)\nax[1].set_yticks(range(12)); ax[1].set_yticklabels(LABELS, fontsize=7)\nax[1].set_title(\"phi correlation\"); ax[1].grid(False)\nplt.colorbar(im2, ax=ax[1], fraction=.046)\nplt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:30:59.531672Z","iopub.execute_input":"2026-08-05T17:30:59.531998Z","iopub.status.idle":"2026-08-05T17:31:00.104489Z","shell.execute_reply.started":"2026-08-05T17:30:59.531974Z","shell.execute_reply":"2026-08-05T17:31:00.103736Z"}},"outputs":[],"execution_count":null},{"id":"865a1936-8aa0-4ec0-aafb-f9de1fbdd1f9","cell_type":"code","source":"# strongest pairs, ranked\npairs = (phi.where(np.triu(np.ones(phi.shape), 1).astype(bool))\n           .stack().sort_values(ascending=False))\nprint(\"STRONGEST POSITIVE ASSOCIATIONS\"); print(pairs.head(8).round(3))\nprint(\"\\nSTRONGEST NEGATIVE ASSOCIATIONS\"); print(pairs.tail(5).round(3))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:31:05.975956Z","iopub.execute_input":"2026-08-05T17:31:05.976212Z","iopub.status.idle":"2026-08-05T17:31:05.995565Z","shell.execute_reply.started":"2026-08-05T17:31:05.976192Z","shell.execute_reply":"2026-08-05T17:31:05.99491Z"}},"outputs":[],"execution_count":null},{"id":"94a570c4-f5b2-4af5-becb-dfc99dfa4252","cell_type":"markdown","source":"**Takeaway (labels).** Expect the OA compartments to travel together, effusion to co-occur with almost everything acute, and ACL to pull contusion along with it (the classic pivot-shift bone-bruise pattern). Two practical consequences:\n\n* Train **one multi-label head**, not 12 independent models — the shared representation gets label correlation for free.\n* Because labels are correlated, a *stacking* layer on the 12 raw logits (or simply averaging within a group) often gains a little AUC on the rare labels.","metadata":{}},{"id":"a266ef57-849b-4a22-8dcc-0e1349f77236","cell_type":"markdown","source":"---\n## 3. Reports: the second modality and the label factory","metadata":{}},{"id":"b521f894-3e70-4f68-91ea-18b5b3b322c3","cell_type":"code","source":"rep = train[\"Report\"].fillna(\"\")\ntrain[\"rep_chars\"] = rep.str.len()\ntrain[\"rep_words\"] = rep.str.split().str.len()\n\nfig, ax = plt.subplots(1, 2, figsize=(10, 3.2))\nax[0].hist(train.loc[train.rep_chars > 0, \"rep_chars\"], bins=80, color=\"#4C72B0\")\nax[0].set_xlabel(\"characters\"); ax[0].set_title(\"Report length\"); ax[0].set_yscale(\"log\")\nax[1].hist(train.loc[train.rep_words > 0, \"rep_words\"], bins=80, color=\"#DD8452\")\nax[1].set_xlabel(\"words\"); ax[1].set_title(\"Report word count\"); ax[1].set_yscale(\"log\")\nplt.tight_layout(); plt.show()\n\nprint(train[[\"rep_chars\", \"rep_words\"]].describe().round(1))\nprint(\"\\nempty reports:\", (train.rep_chars == 0).sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:33:17.825265Z","iopub.execute_input":"2026-08-05T17:33:17.82556Z","iopub.status.idle":"2026-08-05T17:33:18.447269Z","shell.execute_reply.started":"2026-08-05T17:33:17.825535Z","shell.execute_reply":"2026-08-05T17:33:18.446521Z"}},"outputs":[],"execution_count":null},{"id":"35a63772-fce7-4441-960c-26a2d1b7a4c3","cell_type":"code","source":"# --- crude but effective language fingerprinting (no external deps) ---\nSCRIPT_RANGES = {\n    \"Cyrillic\": (0x0400, 0x04FF), \"Greek\": (0x0370, 0x03FF), \"Arabic\": (0x0600, 0x06FF),\n    \"Hebrew\": (0x0590, 0x05FF), \"CJK\": (0x4E00, 0x9FFF), \"Hangul\": (0xAC00, 0xD7AF),\n    \"Kana\": (0x3040, 0x30FF), \"Thai\": (0x0E00, 0x0E7F), \"Devanagari\": (0x0900, 0x097F),\n}\nSTOPWORDS = {   # high-signal function words / radiology terms per language\n    \"English\":    [\"the \", \" and \", \" with \", \" no \", \"knee\", \"tear\", \"normal\"],\n    \"Spanish\":    [\" de \", \" del \", \" con \", \" sin \", \"rodilla\", \"menisco\", \"rotura\"],\n    \"Portuguese\": [\" de \", \" com \", \" sem \", \"joelho\", \"menisco\", \"sinais\"],\n    \"French\":     [\" de \", \" du \", \" avec \", \" sans \", \"genou\", \"menisque\", \"rupture\"],\n    \"German\":     [\" der \", \" und \", \" mit \", \" ohne \", \"kniegelenk\", \"riss\", \"kein\"],\n    \"Italian\":    [\" di \", \" con \", \" senza \", \"ginocchio\", \"menisco\", \"lesione\"],\n    \"Dutch\":      [\" van \", \" met \", \" geen \", \"knie\", \"meniscus\"],\n    \"Turkish\":    [\" ve \", \"diz \", \"menisk\", \"yirtik\"],\n}\ndef script_of(t):\n    counts = Counter()\n    for ch in t[:1500]:\n        o = ord(ch)\n        for name, (lo, hi) in SCRIPT_RANGES.items():\n            if lo <= o <= hi:\n                counts[name] += 1\n    return counts.most_common(1)[0][0] if counts and counts.most_common(1)[0][1] > 8 else \"Latin\"\n\ndef latin_lang(t):\n    t = \" \" + t.lower() + \" \"\n    scores = {L: sum(t.count(w) for w in ws) for L, ws in STOPWORDS.items()}\n    best = max(scores, key=scores.get)\n    return best if scores[best] >= 2 else \"Latin-unknown\"\n\ntrain[\"script\"] = rep.map(script_of)\ntrain[\"lang_guess\"] = np.where(train.script == \"Latin\", rep.map(latin_lang), train.script)\n\nfig, ax = plt.subplots(figsize=(7, 3.4))\nvc = train.lang_guess.value_counts()\nax.bar(vc.index, vc.values, color=\"#55A868\")\nax.set_yscale(\"log\"); ax.set_ylabel(\"studies (log)\"); ax.tick_params(axis=\"x\", rotation=45)\nax.set_title(\"Heuristic language / script of report\")\nplt.tight_layout(); plt.show()\ndisplay(vc.to_frame(\"n\").assign(pct=(vc / vc.sum() * 100).round(1)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:33:38.224651Z","iopub.execute_input":"2026-08-05T17:33:38.22495Z","iopub.status.idle":"2026-08-05T17:33:40.508372Z","shell.execute_reply.started":"2026-08-05T17:33:38.22493Z","shell.execute_reply":"2026-08-05T17:33:40.507786Z"}},"outputs":[],"execution_count":null},{"id":"d72eeb0d-db35-4a12-8b9b-8aa8957454fa","cell_type":"code","source":"# eyeball a few reports across languages\nfor lang in train.lang_guess.value_counts().head(4).index:\n    s = train.loc[train.lang_guess == lang, \"Report\"].dropna()\n    if not len(s): continue\n    print(\"=\" * 90); print(f\"### {lang}\"); print(\"=\" * 90)\n    print(s.iloc[0][:900].strip(), \"\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:33:58.215805Z","iopub.execute_input":"2026-08-05T17:33:58.216144Z","iopub.status.idle":"2026-08-05T17:33:58.226633Z","shell.execute_reply.started":"2026-08-05T17:33:58.216126Z","shell.execute_reply":"2026-08-05T17:33:58.226042Z"}},"outputs":[],"execution_count":null},{"id":"b4d8692d-4ba7-481b-9d97-5080d0e9c8c2","cell_type":"markdown","source":"### Can we recover labels from the report text?\n\nThis is the highest-leverage question in the whole EDA. If a keyword rule reproduces the human labels well on the labeled subset, you can **pseudo-label the entire unlabeled training set** and multiply your effective training data.\n\nThe rule below is deliberately simple — English + Spanish/Portuguese/French stems, with a **negation window** — and we score it against ground truth so you know exactly how much to trust it.","metadata":{}},{"id":"fe61f544-3859-4640-ad90-9ada62e6fac5","cell_type":"code","source":"# --- multilingual keyword patterns (extend these; this is a starting point) ---\nPAT = {\n \"ACL\":              r\"\\b(acl|lca|anterior cruciate|cruzado anterior|croise anterieur|kreuzband)\\b\",\n \"MCL\":              r\"\\b(mcl|lcm|medial collateral|colateral medial|collateral medial|innenband)\\b\",\n \"Medial Meniscus\":  r\"(medial meniscus|menisco (interno|medial)|menisque (interne|medial)|innenmeniskus)\",\n \"Lateral Meniscus\": r\"(lateral meniscus|menisco (externo|lateral)|menisque (externe|lateral)|aussenmeniskus)\",\n \"Medial OA\":        r\"(medial (compartment )?(osteoarthr|arthros|chondral loss|degenerative)|(gonartrosis|artrosis).{0,30}medial)\",\n \"Lateral OA\":       r\"(lateral (compartment )?(osteoarthr|arthros|chondral loss|degenerative)|(gonartrosis|artrosis).{0,30}lateral)\",\n \"PF OA\":            r\"(patellofemoral (osteoarthr|arthros|chondro|degenerativ)|femoropatelar|chondromalacia)\",\n \"Effusion\":         r\"(effusion|derrame|epanchement|erguss|joint fluid|liquido articular)\",\n \"Synovitis\":        r\"(synovitis|sinovitis|synovite|synovial (thicken|proliferat))\",\n \"Baker's\":          r\"(baker|popliteal cyst|quiste de baker|kyste poplite|poplitealzyste)\",\n \"Contusion\":        r\"(contusion|bone (bruise|marrow edema|marrow oedema)|edema (oseo|de la medula)|knochenmarkodem)\",\n \"Fracture\":         r\"(fracture|fractura|fratura|fraktur|avulsion)\",\n}\nNEG = r\"(no |not |without |absence of |negative for |sin |sem |pas de |kein |intact|normal|unremarkable|preserv)\"\nTEAR = r\"(tear|torn|rupture|rotura|ruptura|desgarro|lesion|riss|signal.{0,15}grade (2|3|iii|ii))\"\n\ndef negated(text, span_start, window=45):\n    return re.search(NEG, text[max(0, span_start - window):span_start]) is not None\n\ndef rule_predict(text):\n    t = re.sub(r\"\\s+\", \" \", str(text).lower())\n    out = {}\n    for lab, pat in PAT.items():\n        hit = 0\n        for m in re.finditer(pat, t):\n            if negated(t, m.start()):\n                continue\n            # ligaments & menisci need an injury word nearby, otherwise it is just anatomy description\n            if lab in (\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\"):\n                ctx = t[m.start(): m.end() + 90]\n                if not re.search(TEAR, ctx):\n                    continue\n            hit = 1; break\n        out[lab] = hit\n    return out\n\nrule_df = pd.DataFrame(list(labeled[\"Report\"].fillna(\"\").map(rule_predict)), index=labeled.index)\nprint(\"rule-based predictions computed for\", len(rule_df), \"labeled studies\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:34:13.156079Z","iopub.execute_input":"2026-08-05T17:34:13.156309Z","iopub.status.idle":"2026-08-05T17:34:13.184675Z","shell.execute_reply.started":"2026-08-05T17:34:13.156291Z","shell.execute_reply":"2026-08-05T17:34:13.184043Z"}},"outputs":[],"execution_count":null},{"id":"2683b252-6a79-40ab-a9f1-9886c2b44752","cell_type":"code","source":"def prf(y, p):\n    tp = int(((y == 1) & (p == 1)).sum()); fp = int(((y == 0) & (p == 1)).sum())\n    fn = int(((y == 1) & (p == 0)).sum())\n    prec = tp / (tp + fp) if tp + fp else np.nan\n    rec  = tp / (tp + fn) if tp + fn else np.nan\n    f1   = 2 * prec * rec / (prec + rec) if prec and rec else np.nan\n    return prec, rec, f1, tp, fp, fn\n\nrows = []\nfor lab in LABELS:\n    y = labeled[lab].astype(int).values\n    p = rule_df[lab].values\n    prec, rec, f1, tp, fp, fn = prf(y, p)\n    rows.append(dict(label=lab, prevalence=y.mean(), precision=prec, recall=rec,\n                     f1=f1, TP=tp, FP=fp, FN=fn))\nrule_eval = pd.DataFrame(rows).set_index(\"label\").round(3)\ndisplay(rule_eval.sort_values(\"f1\", ascending=False))\nprint(f\"macro F1 = {rule_eval.f1.mean():.3f} | macro precision = {rule_eval.precision.mean():.3f} \"\n      f\"| macro recall = {rule_eval.recall.mean():.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:34:59.043099Z","iopub.execute_input":"2026-08-05T17:34:59.043351Z","iopub.status.idle":"2026-08-05T17:34:59.06034Z","shell.execute_reply.started":"2026-08-05T17:34:59.043332Z","shell.execute_reply":"2026-08-05T17:34:59.059308Z"}},"outputs":[],"execution_count":null},{"id":"246ba5a5-672b-4c51-ad33-5e3b0054bea3","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(7, 4))\nr = rule_eval.sort_values(\"f1\")\nax.barh(r.index, r[\"precision\"], height=.38, label=\"precision\", color=\"#4C72B0\")\nax.barh([i + .4 for i in range(len(r))], r[\"recall\"], height=.38, label=\"recall\", color=\"#DD8452\")\nax.set_yticks([i + .2 for i in range(len(r))]); ax.set_yticklabels(r.index, fontsize=8)\nax.set_xlim(0, 1); ax.legend(frameon=False); ax.set_title(\"Keyword rule vs. ground-truth labels\")\nplt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:11.408277Z","iopub.execute_input":"2026-08-05T17:35:11.408521Z","iopub.status.idle":"2026-08-05T17:35:11.574663Z","shell.execute_reply.started":"2026-08-05T17:35:11.408501Z","shell.execute_reply":"2026-08-05T17:35:11.574121Z"}},"outputs":[],"execution_count":null},{"id":"dc827893-fb6b-4f8a-a788-b5058a8d1679","cell_type":"code","source":"# inspect the failures — this is how you improve the patterns\ndef show_errors(lab, kind=\"FN\", k=3):\n    y = labeled[lab].astype(int).values; p = rule_df[lab].values\n    mask = (y == 1) & (p == 0) if kind == \"FN\" else (y == 0) & (p == 1)\n    idx = np.where(mask)[0][:k]\n    print(\"#\" * 100); print(f\"### {lab} — {kind} examples ({mask.sum()} total)\"); print(\"#\" * 100)\n    for i in idx:\n        txt = re.sub(r\"\\s+\", \" \", str(labeled.iloc[i][\"Report\"]))\n        print(\"-\" * 90); print(txt[:700])\n\nworst = rule_eval.f1.idxmin()\nshow_errors(worst, \"FN\", 3)\nshow_errors(worst, \"FP\", 2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:18.6956Z","iopub.execute_input":"2026-08-05T17:35:18.696355Z","iopub.status.idle":"2026-08-05T17:35:18.703731Z","shell.execute_reply.started":"2026-08-05T17:35:18.696324Z","shell.execute_reply":"2026-08-05T17:35:18.703144Z"}},"outputs":[],"execution_count":null},{"id":"d8c8dbed-65fb-45ce-a22b-7ecdfdb958a3","cell_type":"markdown","source":"**Takeaway (text).** Read the F1 table as a *ceiling estimate for pseudo-labels*. Labels where precision is high but recall mediocre are safe to use as **positive-only** pseudo-labels; labels where precision is poor should be pseudo-labeled with soft targets or excluded.\n\nA stronger production path, in increasing order of effort:\n\n1. Extend these regexes per language using the FN/FP inspector above (cheap, gets you a long way).\n2. Machine-translate everything to English, then run one English rule set.\n3. Fine-tune a multilingual encoder (XLM-R / mDeBERTa) on the *labeled* subset using the report as input, then run it over the unlabeled reports to get **soft** pseudo-labels for the image model.\n\nCritically: **reports exist only for training.** At test time you get pixels alone. Text is a teacher, never an input to the final model.","metadata":{}},{"id":"d36fe5e0-0570-43fc-bcc2-3bb56dfcbca2","cell_type":"code","source":"# quick corpus view: most frequent tokens (after stripping numerics)\ntok = Counter()\nfor t in rep.sample(min(4000, (rep.str.len() > 0).sum()), random_state=SEED):\n    tok.update(w for w in re.findall(r\"[a-zA-Zaaaaeeeiiooouuunc]{4,}\", str(t).lower()))\npd.Series(dict(tok.most_common(40))).to_frame(\"count\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:29.343245Z","iopub.execute_input":"2026-08-05T17:35:29.343484Z","iopub.status.idle":"2026-08-05T17:35:29.506736Z","shell.execute_reply.started":"2026-08-05T17:35:29.343466Z","shell.execute_reply":"2026-08-05T17:35:29.506071Z"}},"outputs":[],"execution_count":null},{"id":"f673b898-a788-4738-8d84-b0ddeabae2e2","cell_type":"markdown","source":"---\n## 4. Series metadata: what protocol does each study actually contain?","metadata":{}},{"id":"4069efbb-82e0-4ac1-b438-c5ff1cd827fb","cell_type":"code","source":"display(train_series.head())\nprint(train_series.dtypes)\nprint(\"\\nseries rows:\", len(train_series), \"| unique studies:\", train_series.StudyInstanceUID.nunique())\nprint(\"studies in train.csv missing from train_series.csv:\",\n      len(set(train.StudyInstanceUID) - set(train_series.StudyInstanceUID)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:35.47818Z","iopub.execute_input":"2026-08-05T17:35:35.478418Z","iopub.status.idle":"2026-08-05T17:35:35.493293Z","shell.execute_reply.started":"2026-08-05T17:35:35.478401Z","shell.execute_reply":"2026-08-05T17:35:35.492556Z"}},"outputs":[],"execution_count":null},{"id":"f00e4b96-3dd1-42cd-9161-31bf0257278d","cell_type":"code","source":"spc = train_series.groupby(\"StudyInstanceUID\").size()\nfig, ax = plt.subplots(1, 3, figsize=(13, 3.2))\nax[0].hist(spc, bins=np.arange(.5, spc.max() + 1.5), color=\"#4C72B0\")\nax[0].set_title(f\"Series per study (median {spc.median():.0f})\"); ax[0].set_xlabel(\"# series\")\n\nvc = train_series.Anatomical_Plane.value_counts(dropna=False)\nax[1].bar(vc.index.astype(str), vc.values, color=\"#DD8452\"); ax[1].set_title(\"Plane\")\n\ncombo = (train_series.assign(k=train_series.Fluid_Sensitive.astype(str) + \"/\" +\n                              train_series.Fat_Suppression.astype(str))\n         .k.value_counts())\nax[2].bar(combo.index, combo.values, color=\"#55A868\")\nax[2].set_title(\"Fluid_Sensitive / Fat_Suppression\"); ax[2].tick_params(axis=\"x\", rotation=20)\nplt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:37.43462Z","iopub.execute_input":"2026-08-05T17:35:37.435363Z","iopub.status.idle":"2026-08-05T17:35:37.708217Z","shell.execute_reply.started":"2026-08-05T17:35:37.435344Z","shell.execute_reply":"2026-08-05T17:35:37.707512Z"}},"outputs":[],"execution_count":null},{"id":"348e62e2-a296-45d8-9ca5-6af95992af91","cell_type":"code","source":"# full protocol cross-tab: plane x contrast type\nct = pd.crosstab(train_series.Anatomical_Plane,\n                 train_series.Fluid_Sensitive.astype(str) + \"|\" + train_series.Fat_Suppression.astype(str))\ndisplay(ct)\nprint(\"\\nrow-normalised:\"); display((ct.div(ct.sum(1), axis=0) * 100).round(1))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:43.685284Z","iopub.execute_input":"2026-08-05T17:35:43.685556Z","iopub.status.idle":"2026-08-05T17:35:43.720437Z","shell.execute_reply.started":"2026-08-05T17:35:43.685536Z","shell.execute_reply":"2026-08-05T17:35:43.719978Z"}},"outputs":[],"execution_count":null},{"id":"acbf2049-2567-4bd8-93da-fa4c31e638d6","cell_type":"code","source":"# per-study protocol fingerprint: which (plane, fluid, fatsat) combos are present?\ntrain_series[\"combo\"] = (train_series.Anatomical_Plane.astype(str) + \"_F\" +\n                         train_series.Fluid_Sensitive.astype(str) + \"S\" +\n                         train_series.Fat_Suppression.astype(str))\npivot = (train_series.assign(v=1)\n         .pivot_table(index=\"StudyInstanceUID\", columns=\"combo\", values=\"v\",\n                      aggfunc=\"max\", fill_value=0))\ncov = pivot.mean().sort_values(ascending=False)\n\nfig, ax = plt.subplots(figsize=(8, 4))\nax.barh(cov.index[::-1], cov.values[::-1], color=\"#8172B3\")\nax.set_xlabel(\"fraction of studies containing this sequence type\")\nax.set_title(\"Sequence-type coverage across studies\")\nplt.tight_layout(); plt.show()\ndisplay(cov.to_frame(\"coverage\").round(3))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:49.104768Z","iopub.execute_input":"2026-08-05T17:35:49.105015Z","iopub.status.idle":"2026-08-05T17:35:49.255816Z","shell.execute_reply.started":"2026-08-05T17:35:49.104997Z","shell.execute_reply":"2026-08-05T17:35:49.255211Z"}},"outputs":[],"execution_count":null},{"id":"4fe51ad5-8a5f-4a9f-a641-9834e28fde25","cell_type":"code","source":"# how many studies are missing each plane entirely?\nplanes = (train_series.pivot_table(index=\"StudyInstanceUID\", columns=\"Anatomical_Plane\",\n                                   values=\"SeriesInstanceUID\", aggfunc=\"count\", fill_value=0) > 0)\nprint(\"studies missing each plane:\")\nprint((~planes).mean().round(3).to_frame(\"fraction_missing\"))\nprint(\"\\nmost common plane-sets:\")\nprint(planes.apply(lambda r: \"+\".join(sorted(planes.columns[r.values])), axis=1)\n            .value_counts().head(8))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:53.392812Z","iopub.execute_input":"2026-08-05T17:35:53.393063Z","iopub.status.idle":"2026-08-05T17:35:53.446044Z","shell.execute_reply.started":"2026-08-05T17:35:53.393045Z","shell.execute_reply":"2026-08-05T17:35:53.445408Z"}},"outputs":[],"execution_count":null},{"id":"407865dc-4e11-44ca-b517-041523b90ac1","cell_type":"markdown","source":"**Takeaway (protocol).** You cannot assume a fixed input tensor. Some studies lack an axial series, some lack fat-suppressed sequences, some have duplicates. Two robust designs:\n\n* **Per-sequence-type slots with dropout.** Define canonical slots (Sag-fluid-fatsat, Cor-fluid-fatsat, Ax-fluid-fatsat, Sag-T1…), fill what exists, zero-fill and mask what does not, and randomly drop slots during training so the model never depends on one.\n* **Series-agnostic MIL.** Encode every series independently, then attention-pool over all series in the study. Handles variable counts natively and is what usually wins these study-level tasks.\n\nAlso note which findings live in which plane — coronal fluid-sensitive for MCL and meniscal roots, sagittal for ACL, axial for patellofemoral cartilage. If you must subsample, subsample with anatomy in mind, not at random.","metadata":{}},{"id":"6283e774-9fee-4d1b-a232-c9aff920f466","cell_type":"markdown","source":"---\n## 5. DICOM level: geometry, scanners, encoding, intensities","metadata":{}},{"id":"4f42d410-0896-4ea1-96e7-196711c021d1","cell_type":"code","source":"try:\n    import pydicom\n    from pydicom.uid import UID\n    HAS_DCM = True\nexcept ImportError:\n    HAS_DCM = False\n    print(\"pydicom unavailable — skip section 5\")\n\nSERIES_DIR = DATA / \"train_series\"\nprint(\"series dir exists:\", SERIES_DIR.exists())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:35:59.124051Z","iopub.execute_input":"2026-08-05T17:35:59.12428Z","iopub.status.idle":"2026-08-05T17:35:59.656034Z","shell.execute_reply.started":"2026-08-05T17:35:59.124263Z","shell.execute_reply":"2026-08-05T17:35:59.655289Z"}},"outputs":[],"execution_count":null},{"id":"0de7b9a4-bc39-44ac-8853-66f1ca1caa44","cell_type":"code","source":"N_STUDIES_SCAN = 60        # raise for a more reliable picture; each study is a few file reads\n\ndef list_slices(study, series):\n    return sorted((SERIES_DIR / study / series).glob(\"*.dcm\"))\n\nif HAS_DCM and SERIES_DIR.exists():\n    study_ids = [s for s in train_series.StudyInstanceUID.unique() if (SERIES_DIR / s).exists()]\n    sample_studies = random.sample(study_ids, min(N_STUDIES_SCAN, len(study_ids)))\n    recs, t0 = [], time.time()\n    for st in sample_studies:\n        for se in train_series.loc[train_series.StudyInstanceUID == st, \"SeriesInstanceUID\"]:\n            files = list_slices(st, se)\n            if not files: continue\n            ds = pydicom.dcmread(files[0], stop_before_pixels=True, force=True)\n            g = lambda k, d=None: getattr(ds, k, d)\n            ps = g(\"PixelSpacing\", [np.nan, np.nan])\n            recs.append(dict(\n                StudyInstanceUID=st, SeriesInstanceUID=se, n_slices=len(files),\n                Rows=g(\"Rows\"), Columns=g(\"Columns\"),\n                px_y=float(ps[0]) if ps is not None else np.nan,\n                px_x=float(ps[1]) if ps is not None else np.nan,\n                SliceThickness=float(g(\"SliceThickness\") or np.nan),\n                Manufacturer=str(g(\"Manufacturer\", \"?\")).strip(),\n                Model=str(g(\"ManufacturerModelName\", \"?\")).strip(),\n                FieldStrength=g(\"MagneticFieldStrength\"),\n                TR=g(\"RepetitionTime\"), TE=g(\"EchoTime\"),\n                SeriesDescription=str(g(\"SeriesDescription\", \"\")).strip(),\n                BitsStored=g(\"BitsStored\"), PhotometricInterpretation=g(\"PhotometricInterpretation\"),\n                TransferSyntax=str(getattr(ds.file_meta, \"TransferSyntaxUID\", \"?\").name\n                                   if hasattr(ds, \"file_meta\") else \"?\"),\n                n_tags=len(ds),\n            ))\n    meta = pd.DataFrame(recs)\n    print(f\"scanned {len(sample_studies)} studies / {len(meta)} series in {time.time()-t0:.1f}s\")\n    display(meta.head())\nelse:\n    meta = pd.DataFrame()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:36:06.615687Z","iopub.execute_input":"2026-08-05T17:36:06.615978Z","iopub.status.idle":"2026-08-05T17:36:20.055205Z","shell.execute_reply.started":"2026-08-05T17:36:06.615959Z","shell.execute_reply":"2026-08-05T17:36:20.054372Z"}},"outputs":[],"execution_count":null},{"id":"3b687a47-e6f6-4ebb-8183-3d30ac164a69","cell_type":"code","source":"if len(meta):\n    fig, ax = plt.subplots(2, 3, figsize=(13, 6.5))\n    ax = ax.ravel()\n    ax[0].hist(meta.n_slices, bins=50, color=\"#4C72B0\"); ax[0].set_title(\"slices per series\"); ax[0].set_yscale(\"log\")\n    ax[1].hist(meta.Rows.dropna(), bins=40, color=\"#DD8452\"); ax[1].set_title(\"matrix Rows\")\n    ax[2].hist(meta.px_x.dropna(), bins=40, color=\"#55A868\"); ax[2].set_title(\"in-plane pixel spacing (mm)\")\n    ax[3].hist(meta.SliceThickness.dropna(), bins=40, color=\"#C44E52\"); ax[3].set_title(\"slice thickness (mm)\")\n    fs = meta.FieldStrength.dropna().astype(float)\n    ax[4].hist(fs, bins=30, color=\"#8172B3\"); ax[4].set_title(\"field strength (T)\")\n    mf = meta.Manufacturer.replace(\"\", \"?\").value_counts().head(8)\n    ax[5].barh(mf.index[::-1], mf.values[::-1], color=\"#937860\"); ax[5].set_title(\"manufacturer\")\n    plt.tight_layout(); plt.show()\n\n    print(meta[[\"n_slices\", \"Rows\", \"Columns\", \"px_x\", \"SliceThickness\"]].describe().round(2))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:36:34.464544Z","iopub.execute_input":"2026-08-05T17:36:34.464889Z","iopub.status.idle":"2026-08-05T17:36:35.406532Z","shell.execute_reply.started":"2026-08-05T17:36:34.464869Z","shell.execute_reply":"2026-08-05T17:36:35.405905Z"}},"outputs":[],"execution_count":null},{"id":"c962e561-be54-4268-92f3-e7786e223a24","cell_type":"code","source":"if len(meta):\n    print(\"TRANSFER SYNTAX MIX (drives decode speed — see benchmark below)\")\n    display(meta.TransferSyntax.value_counts().to_frame(\"n_series\"))\n    print(\"\\nPHOTOMETRIC / BIT DEPTH\")\n    display(pd.crosstab(meta.PhotometricInterpretation, meta.BitsStored))\n    print(\"\\nTOP SERIES DESCRIPTIONS (free text, useful as a weak protocol feature)\")\n    display(meta.SeriesDescription.str.lower().value_counts().head(25).to_frame(\"n\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:36:41.117758Z","iopub.execute_input":"2026-08-05T17:36:41.118053Z","iopub.status.idle":"2026-08-05T17:36:41.142266Z","shell.execute_reply.started":"2026-08-05T17:36:41.118033Z","shell.execute_reply":"2026-08-05T17:36:41.1416Z"}},"outputs":[],"execution_count":null},{"id":"7214ca66-9395-4328-bdd0-05df76f6e40a","cell_type":"code","source":"# decode-cost benchmark by transfer syntax — plan your preprocessing budget with this\nif HAS_DCM and len(meta):\n    times = defaultdict(list)\n    for _, r in meta.sample(min(40, len(meta)), random_state=SEED).iterrows():\n        f = list_slices(r.StudyInstanceUID, r.SeriesInstanceUID)\n        if not f: continue\n        t0 = time.time()\n        try:\n            _ = pydicom.dcmread(f[len(f)//2], force=True).pixel_array\n            times[r.TransferSyntax].append((time.time() - t0) * 1000)\n        except Exception as e:\n            times[f\"FAILED::{type(e).__name__}\"].append(np.nan)\n    for k, v in times.items():\n        print(f\"{k:<45} n={len(v):3d}  median {np.nanmedian(v):7.1f} ms/slice\")\n    print(\"\\nIf a syntax FAILS, install pylibjpeg / pylibjpeg-openjpeg / gdcm.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:36:48.604653Z","iopub.execute_input":"2026-08-05T17:36:48.604968Z","iopub.status.idle":"2026-08-05T17:36:49.286897Z","shell.execute_reply.started":"2026-08-05T17:36:48.604949Z","shell.execute_reply":"2026-08-05T17:36:49.286208Z"}},"outputs":[],"execution_count":null},{"id":"d53ac899-c158-4aa7-8272-648cbd27fbf8","cell_type":"code","source":"def load_series(study, series, max_slices=None):\n    files = list_slices(study, series)\n    if not files: return None\n    ds_list = []\n    for f in files:\n        try:\n            d = pydicom.dcmread(f, force=True)\n            _ = d.pixel_array\n            ds_list.append(d)\n        except Exception:\n            continue\n    if not ds_list: return None\n    ds_list.sort(key=lambda d: float(getattr(d, \"InstanceNumber\", 0) or 0))\n    if max_slices and len(ds_list) > max_slices:\n        idx = np.linspace(0, len(ds_list) - 1, max_slices).astype(int)\n        ds_list = [ds_list[i] for i in idx]\n    vol = []\n    for d in ds_list:\n        a = d.pixel_array.astype(np.float32)\n        a = a * float(getattr(d, \"RescaleSlope\", 1) or 1) + float(getattr(d, \"RescaleIntercept\", 0) or 0)\n        if getattr(d, \"PhotometricInterpretation\", \"\") == \"MONOCHROME1\":\n            a = a.max() - a\n        vol.append(a)\n    shapes = {v.shape for v in vol}\n    if len(shapes) > 1:                       # rare, but it happens\n        h = min(v.shape[0] for v in vol); w = min(v.shape[1] for v in vol)\n        vol = [v[:h, :w] for v in vol]\n    return np.stack(vol)\n\ndef window(x, lo=1, hi=99):\n    a, b = np.percentile(x, [lo, hi])\n    return np.clip((x - a) / max(b - a, 1e-6), 0, 1)\n\ndef montage(study, series, n=8, title=\"\"):\n    v = load_series(study, series, max_slices=n)\n    if v is None:\n        print(\"could not load\", series); return\n    v = window(v)\n    fig, ax = plt.subplots(1, len(v), figsize=(2 * len(v), 2.3))\n    ax = np.atleast_1d(ax)\n    for i, a in enumerate(ax):\n        a.imshow(v[i], cmap=\"gray\"); a.axis(\"off\"); a.set_title(f\"{i}\", fontsize=7)\n    fig.suptitle(title, fontsize=9); plt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:37:04.465842Z","iopub.execute_input":"2026-08-05T17:37:04.466557Z","iopub.status.idle":"2026-08-05T17:37:04.476821Z","shell.execute_reply.started":"2026-08-05T17:37:04.466535Z","shell.execute_reply":"2026-08-05T17:37:04.476127Z"}},"outputs":[],"execution_count":null},{"id":"8c93e0e3-669a-4a88-97bd-e4d3df96af1f","cell_type":"code","source":"# visualise one full study: every series, middle slices\nif HAS_DCM and SERIES_DIR.exists():\n    demo = sample_studies[0]\n    print(\"STUDY:\", demo)\n    row = train.loc[train.StudyInstanceUID == demo]\n    if len(row) and row[LABELS].notna().all(axis=1).iloc[0]:\n        pos = [l for l in LABELS if row.iloc[0][l] == 1]\n        print(\"labels:\", pos if pos else \"NORMAL\")\n    for _, s in train_series[train_series.StudyInstanceUID == demo].iterrows():\n        montage(demo, s.SeriesInstanceUID, n=6,\n                title=f\"{s.Anatomical_Plane} | fluid={s.Fluid_Sensitive} fatsat={s.Fat_Suppression}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:37:11.033339Z","iopub.execute_input":"2026-08-05T17:37:11.03362Z","iopub.status.idle":"2026-08-05T17:37:15.56736Z","shell.execute_reply.started":"2026-08-05T17:37:11.033598Z","shell.execute_reply":"2026-08-05T17:37:15.566693Z"}},"outputs":[],"execution_count":null},{"id":"598ecd4c-7512-45e5-9854-3a29a769bd41","cell_type":"code","source":"# side-by-side: a positive vs a negative case for one finding\ndef first_study_with(lab, val=1, plane=\"Sagittal\", fluid=1):\n    cand = labeled.loc[labeled[lab] == val, \"StudyInstanceUID\"]\n    for st in cand:\n        if not (SERIES_DIR / st).exists(): continue\n        ss = train_series[(train_series.StudyInstanceUID == st) &\n                          (train_series.Anatomical_Plane == plane) &\n                          (train_series.Fluid_Sensitive == fluid)]\n        if len(ss): return st, ss.iloc[0].SeriesInstanceUID\n    return None, None\n\nif HAS_DCM and SERIES_DIR.exists():\n    for lab, plane in [(\"ACL\", \"Sagittal\"), (\"Effusion\", \"Axial\"), (\"Medial Meniscus\", \"Coronal\")]:\n        for val in (1, 0):\n            st, se = first_study_with(lab, val, plane)\n            if st:\n                montage(st, se, n=6, title=f\"{lab} = {val}  ({plane}, fluid-sensitive)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:37:23.411516Z","iopub.execute_input":"2026-08-05T17:37:23.411806Z","iopub.status.idle":"2026-08-05T17:37:27.749915Z","shell.execute_reply.started":"2026-08-05T17:37:23.411788Z","shell.execute_reply":"2026-08-05T17:37:27.749012Z"}},"outputs":[],"execution_count":null},{"id":"e193c869-b3c5-4cdd-aee1-1358d30aba72","cell_type":"code","source":"# intensity statistics: why per-series normalisation is mandatory\nif HAS_DCM and len(meta):\n    stats = []\n    for _, r in meta.sample(min(25, len(meta)), random_state=SEED).iterrows():\n        v = load_series(r.StudyInstanceUID, r.SeriesInstanceUID, max_slices=5)\n        if v is None: continue\n        stats.append(dict(series=r.SeriesInstanceUID[-8:], mn=v.min(), p1=np.percentile(v, 1),\n                          med=np.median(v), p99=np.percentile(v, 99), mx=v.max(),\n                          fluid=r.SeriesInstanceUID))\n    S = pd.DataFrame(stats)\n    display(S[[\"series\", \"mn\", \"p1\", \"med\", \"p99\", \"mx\"]].round(1))\n    fig, ax = plt.subplots(figsize=(8, 3))\n    ax.boxplot([[s] for s in S.p99], vert=False, widths=.5)\n    ax.set_xlabel(\"99th-percentile intensity per series\"); ax.set_xscale(\"log\")\n    ax.set_title(\"Raw intensity scale varies by orders of magnitude across series\")\n    plt.tight_layout(); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:37:32.572945Z","iopub.execute_input":"2026-08-05T17:37:32.573194Z","iopub.status.idle":"2026-08-05T17:37:46.558622Z","shell.execute_reply.started":"2026-08-05T17:37:32.573172Z","shell.execute_reply":"2026-08-05T17:37:46.558054Z"}},"outputs":[],"execution_count":null},{"id":"cd825675-080c-4884-b196-149e47a199d2","cell_type":"markdown","source":"**Takeaway (pixels).** MRI intensities carry no absolute meaning — the same tissue can be 200 on one scanner and 4000 on another. Always normalise **per series** (percentile clip 0.5–99.5, then z-score or 0–1 scale). Resample to a common in-plane spacing if you crop, and remember slice thickness varies, so 3D kernels see different physical depths across studies.\n\nDecoding is likely your real bottleneck. Convert everything once to compressed `.npy`/HDF5 at fixed resolution before training rather than reading DICOM in the dataloader.","metadata":{}},{"id":"31862473-d3df-4dfe-9b7b-9eb05c70d04e","cell_type":"markdown","source":"---\n## 6. Train vs test: is the test set the same animal?","metadata":{}},{"id":"0730d9d5-ef8c-4146-a572-30c3a79c3cb1","cell_type":"code","source":"print(\"test studies:\", len(test), \"| test series rows:\", len(test_series))\ndisplay(test.head())\ndisplay(sample_sub.head(3))\nprint(\"\\nsubmission columns match spec:\",\n      list(sample_sub.columns) == [\"StudyInstanceUID\"] + LABELS)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:38:01.403526Z","iopub.execute_input":"2026-08-05T17:38:01.403831Z","iopub.status.idle":"2026-08-05T17:38:01.422492Z","shell.execute_reply.started":"2026-08-05T17:38:01.403811Z","shell.execute_reply":"2026-08-05T17:38:01.42191Z"}},"outputs":[],"execution_count":null},{"id":"5925f3f4-6751-47a8-9ab0-bffb9ae2b56c","cell_type":"code","source":"cmp = pd.DataFrame({\n    \"train\": train_series.Anatomical_Plane.value_counts(normalize=True),\n    \"test\":  test_series.Anatomical_Plane.value_counts(normalize=True),\n}).fillna(0).round(3)\ndisplay(cmp)\n\ncmp2 = pd.DataFrame({\n    \"train\": train_series.groupby(\"StudyInstanceUID\").size().describe(),\n    \"test\":  test_series.groupby(\"StudyInstanceUID\").size().describe(),\n}).round(2)\ndisplay(cmp2)\nprint(\"NOTE: the public example test files contain only 3 studies — these comparisons are\")\nprint(\"indicative only. The real test set is ~1300 studies and is swapped in at scoring time.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T17:38:03.365532Z","iopub.execute_input":"2026-08-05T17:38:03.365799Z","iopub.status.idle":"2026-08-05T17:38:03.388296Z","shell.execute_reply.started":"2026-08-05T17:38:03.365781Z","shell.execute_reply":"2026-08-05T17:38:03.387692Z"}},"outputs":[],"execution_count":null},{"id":"ecf048c3-8e5e-457c-a02d-d942767991a0","cell_type":"markdown","source":"**Takeaway (shift).** The organisers state explicitly that **prevalence is not guaranteed to match across train / public LB / private LB**. Two consequences:\n\n* AUC is prevalence-insensitive in expectation, which protects you — but its *variance* explodes when a label has few positives in the hidden set. Do not chase small LB moves on rare labels.\n* Never calibrate thresholds to training prevalence. Submit raw ranking scores; the metric only cares about order.","metadata":{}},{"id":"33611235-7532-4587-bbaa-0a4399c60731","cell_type":"markdown","source":"---\n## 7. Summary — what the EDA implies for the solution\n\n**Data structure**\n- Prediction unit = **study**; a study holds several series; a series holds 20–45 slices (long tail to hundreds).\n- Join key: `StudyInstanceUID` for labels, `StudyInstanceUID` + `SeriesInstanceUID` for the folder path.\n\n**Labels**\n- Only a minority of training studies are labeled; the rest ship with a report instead. Multi-label, heavily imbalanced, strongly correlated within clinical groups.\n- Metric is macro AUC, so **rare labels count as much as common ones** — and dominate the variance.","metadata":{}}]}