{"cells":[{"cell_type":"markdown","metadata":{},"source":"# The scan knows two things the report does not\n\nEvery public solution in this competition trains an image model on labels read out of the\nradiology reports, because 4,349 of the 4,407 studies have no other labels. That raises a\nquestion nobody has answered with a number: **once the image model is trained, does it know\nanything the report did not already say?** If it does not, the images are decoration and the\nwhole competition is text extraction.\n\nI trained one -- EfficientNet-B3 over the three planes, five folds, text-derived soft targets --\non a V100 and scored it, out of fold, on the 58 studies a radiologist actually annotated. Those\n58 never touched training. Then I put it next to the text scores from\n[my earlier notebook](https://www.kaggle.com/code/nekkon/weak-labels-for-all-12-knee-mri-findings)\nand asked, finding by finding, which one the radiologist agrees with.\n\nThe short answer: the images know two things the report does not -- **which compartment the\nosteoarthritis is in, and how much fluid there is** -- and on everything else they are a\nnoisier copy of the text they were trained on. The full answer, with the intervals that\n58 studies allow, is below.\n\nNothing here runs on a GPU. The training code is attached and documented; what this notebook\nexecutes is the evaluation, from the saved out-of-fold predictions."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import glob, json, warnings\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import roc_auc_score\nfrom scipy.stats import rankdata, spearmanr\nwarnings.filterwarnings('ignore')\n\nINK, MUTED, GRID = '#1a1a1a', '#5c5c5c', '#dcdcdc'\nTEXT, IMAGE, BLEND, WARN = '#2a78d6', '#eb6834', '#1baf7a', '#e34948'\nplt.rcParams.update({'figure.facecolor': 'white', 'axes.facecolor': 'white',\n                     'axes.spines.top': False, 'axes.spines.right': False,\n                     'axes.edgecolor': GRID, 'font.size': 11, 'text.color': INK,\n                     'axes.labelcolor': MUTED, 'xtick.color': MUTED, 'ytick.color': MUTED,\n                     'axes.titlesize': 13, 'axes.titleweight': 'bold',\n                     'axes.titlelocation': 'left', 'axes.titlepad': 12})\n\nP = glob.glob('/kaggle/input/**/train.csv', recursive=True)[0].rsplit('/', 1)[0]\ntrain = pd.read_csv(f'{P}/train.csv')\nLABELS = [c for c in train.columns if c not in ('StudyInstanceUID', 'Report')]\n\nOOF = glob.glob('/kaggle/input/**/pred_b3_gold_5folds.npy', recursive=True)[0].rsplit('/', 1)[0]\ngold_ids = pd.read_csv(f'{OOF}/gold_order.csv').StudyInstanceUID.tolist()\nG = train.set_index('StudyInstanceUID').loc[gold_ids][LABELS].values.astype(int)\nP5 = np.load(f'{OOF}/pred_b3_gold_5folds.npy')            # (5 folds, 58 studies, 12 findings)\nlogs = json.load(open(f'{OOF}/fold_logs.json'))\nscale = json.load(open(f'{OOF}/scale_logs.json'))\n\nWL = glob.glob('/kaggle/input/**/weak_labels_v2.csv', recursive=True)[0]\nT = pd.read_csv(WL).set_index('StudyInstanceUID').loc[gold_ids][LABELS].values\n\nprint(f'{len(gold_ids)} annotated studies, {P5.shape[0]} folds of image predictions, {len(LABELS)} findings')\nprint(f'positives per finding: {dict(zip(LABELS, G.sum(0)))}')"},{"cell_type":"markdown","metadata":{},"source":"## 1. What was trained, in one paragraph\n\nEach study becomes up to six *slots* -- sagittal, coronal and axial, each fluid-sensitive or\nnot -- and each slot is 24 slices spread over the middle 70% of the stack, cropped to a fixed\n140 mm and resampled to 336 px. Slice order comes from `ImagePositionPatient` projected on the\nslice normal, because file names are SOP UIDs and carry no order. Laterality comes from the\nsign of the patient-x coordinate. A shared EfficientNet-B3 encodes three adjacent slices as\none image; a per-finding attention head pools the slots, masking absent ones. Targets are the\nblended text probabilities from the weak-labels notebook -- soft, not binarised. Five folds by\nreport hash, so byte-identical reports stay together. Ten epochs, ~25 s each on a V100.\n\nTwo limits that matter for reading everything below. **Only 929 of the 4,407 studies were\nused** -- the DICOM corpus is 569 GB uncompressed and that is what fit on the disk -- so every\nfold trained on roughly 700 studies. And **the 58 annotated studies were never in training**;\nthey are the only place the number means what the competition scores."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"rows = []\nfor k in range(5):\n    best = max(logs[f'fold{k}']['log'], key=lambda r: r['weak_holdout_auc'])\n    rows.append({'fold': k, 'best epoch': best['ep'], 'AUC vs text labels (held-out fold)': best['weak_holdout_auc'],\n                 'AUC vs radiologist (58 gold)': best['gold_auc']})\nfolds = pd.DataFrame(rows)\nprint(folds.round(3).to_string(index=False))\nprint(f\"\\nmean   text-holdout {folds.iloc[:, 2].mean():.3f}   gold {folds.iloc[:, 3].mean():.3f}\"\n      f\"   gap {folds.iloc[:, 2].mean() - folds.iloc[:, 3].mean():+.3f}\")"},{"cell_type":"markdown","metadata":{},"source":"The gap between the two columns is the first result. Against the held-out *text* labels the\nmodel scores 0.73; against the *radiologist* it scores 0.62. A tenth of an AUC is what the\nmodel learned that was in the report and not in the knee -- template phrasing, the extractor's\nown biases, whatever the text carries that the image does not. It is a floor on label noise,\nmeasured rather than assumed."},{"cell_type":"markdown","metadata":{},"source":"## 2. Text, image, and the radiologist, finding by finding\n\nRank-average the five folds (the metric is AUC, so ranks are the only thing it reads), put the\nimage score next to the text score, and ask which one the radiologist agrees with. The\ninterval on the image column is a bootstrap over the 58 studies."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"def ranks(M):\n    return np.column_stack([rankdata(M[:, j]) for j in range(M.shape[1])])\n\nIMG = np.mean([ranks(P5[k]) for k in range(5)], 0)      # rank-averaged over folds\nTXT = ranks(T)\nauc = lambda S: np.array([roc_auc_score(G[:, j], S[:, j]) for j in range(len(LABELS))])\na_txt, a_img = auc(TXT), auc(IMG)\n\nrng = np.random.default_rng(0)\nlo, hi = [], []\nfor j in range(len(LABELS)):\n    b = []\n    for _ in range(2000):\n        i = rng.integers(0, len(G), len(G))\n        if 0 < G[i, j].sum() < len(i):\n            b.append(roc_auc_score(G[i, j], IMG[i, j]))\n    lo.append(np.percentile(b, 2.5)); hi.append(np.percentile(b, 97.5))\n\nres = pd.DataFrame({'finding': LABELS, 'positives': G.sum(0), 'text AUC': a_txt, 'image AUC': a_img,\n                    'image 95% lo': lo, 'image 95% hi': hi, 'image - text': a_img - a_txt})\nres = res.sort_values('image - text', ascending=False)\nprint(res.round(3).to_string(index=False))\nprint(f\"\\nmean AUC   text {a_txt.mean():.3f}   image {a_img.mean():.3f}\")"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, ax = plt.subplots(figsize=(9.5, 5.6))\nr = res.sort_values('image AUC'); y = np.arange(len(r))\nfor yi, (_, row) in zip(y, r.iterrows()):\n    c = IMAGE if row['image 95% lo'] > row['text AUC'] else (TEXT if row['image 95% hi'] < row['text AUC'] else '#b6b9be')\n    ax.plot([row['text AUC'], row['image AUC']], [yi, yi], color=c, lw=2.4, zorder=1, solid_capstyle='round')\n    ax.plot([row['image 95% lo'], row['image 95% hi']], [yi - .22, yi - .22], color=IMAGE, lw=1, alpha=.5)\nax.scatter(r['text AUC'], y, s=52, color=TEXT, zorder=3, label='text (NLI + keywords)')\nax.scatter(r['image AUC'], y, s=52, color=IMAGE, zorder=3, label='image (B3, 5-fold OOF)')\nax.axvline(.5, color=GRID, lw=1, zorder=0)\nax.set_yticks(y); ax.set_yticklabels(r.finding)\nfor yi, (_, row) in zip(y, r.iterrows()):\n    if row['image 95% lo'] > row['text AUC']:\n        ax.text(row['image AUC'] + .012, yi, 'image wins', va='center', fontsize=9, color=IMAGE, fontweight='bold')\nax.set_xlim(.40, 1.0); ax.set_xlabel('AUC against the radiologist, 58 studies')\nax.set_title('The image beats the report on two findings and trails it on the rest')\nax.legend(frameon=False, loc='lower right')\nax.text(.40, -1.4, 'thin bar = 95% bootstrap interval on the image AUC; grey = interval covers the text score',\n        fontsize=9, color=MUTED)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Two findings where the image wins outright, with the interval clear of the text score:\n\n**Lateral compartment OA, 0.77 against 0.56.** This is the finding the text side could never\nget right: the entailment model cannot attach *lateral* to *osteoarthritis*, the keyword\nextractor manages only a crude side-word rule, and the weak-labels notebook spent a whole\nsection on it. The scan does not have that problem. Which compartment the cartilage is worn in\nis a question about where things are in the image, and an encoder that sees the coronal plane\nanswers it directly. The text-derived label for this finding is barely above chance; the model\ntrained on it learned to see past it.\n\n**Effusion, 0.72 against 0.67.** Effusion is graded in reports -- *trace*, *small*, *moderate*\n-- and a binary text label throws the grade away. The image keeps it: how much fluid there is\nis visible on any fluid-sensitive sequence.\n\nEverywhere else the text wins, sometimes by a lot -- ACL 0.86 against 0.55, Baker's 0.88\nagainst 0.65, fracture 0.81 against 0.56. Those are findings that radiologists *name* in a\nreport, so the text label is nearly clean and the image, trained on 700 studies, cannot yet\nmatch a reader who saw the same knee and wrote down what was there."},{"cell_type":"markdown","metadata":{},"source":"## 3. Does adding the image to the text help at all?\n\nIf the image only knew what the text knew, a blend would add nothing. If it knew something\ndifferent, a blend would gain. Weighted rank blend, weight on the image side, bootstrap on\nthe difference against text alone."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"ws = [0.0, 0.1, 0.2, 0.3, 0.4, 0.5, 1.0]\nblend = {w: auc((1 - w) * TXT + w * IMG) for w in ws}\ntbl = pd.DataFrame({f'w={w}': blend[w] for w in ws}, index=LABELS)\ntbl.loc['MEAN'] = tbl.mean()\nprint(tbl.round(3).to_string())\n\nd = []\nfor _ in range(3000):\n    i = rng.integers(0, len(G), len(G)); ok = [j for j in range(len(LABELS)) if 0 < G[i, j].sum() < len(i)]\n    d.append(np.mean([roc_auc_score(G[i, j], (0.7 * TXT + 0.3 * IMG)[i, j]) - roc_auc_score(G[i, j], TXT[i, j]) for j in ok]))\nprint(f'\\nblend (w=0.3) minus text alone, mean over findings: {np.mean(d):+.3f}   95% [{np.percentile(d, 2.5):+.3f}, {np.percentile(d, 97.5):+.3f}]')"},{"cell_type":"markdown","metadata":{},"source":"Overall, the blend is a wash: +0.003 with an interval that straddles zero. Read that with the\ncompanion notebook's power curve in mind -- 58 studies detect a +0.05 improvement about 7% of\nthe time -- and it is the expected result whether or not the images help on average.\n\nBut the mean hides the structure. On the two findings above, weighting the image *up* helps:\nLateral OA goes from 0.56 to 0.69 at w=0.5, Effusion from 0.67 to 0.72. On ACL and Baker's,\nthe same weight *costs* six to seven points. There is no single weight that is right, because\nthe two sources are not two noisy views of the same signal -- they are good at different\nfindings. A per-finding weight, chosen on a holdout larger than 58 studies, is the obvious\nnext thing, and it is exactly what this test set cannot choose."},{"cell_type":"markdown","metadata":{},"source":"## 4. Is the image reading the knee, or the report?\n\nThe direct test. Correlate the image score with the text score, and with the radiologist,\nacross the 58 studies. Then look only at the studies where the text label is *wrong* --\ntext says one thing, radiologist says the other -- and ask which side the image takes."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"r_txt = np.mean([spearmanr(IMG[:, j], T[:, j]).correlation for j in range(len(LABELS))])\nr_gld = np.mean([spearmanr(IMG[:, j], G[:, j]).correlation for j in range(len(LABELS))])\nprint(f'mean Spearman   image ~ text {r_txt:.3f}    image ~ radiologist {r_gld:.3f}')\n\nside = []\nfor j, l in enumerate(LABELS):\n    tb = (T[:, j] >= np.median(T[:, j])).astype(int); wrong = tb != G[:, j]\n    if wrong.sum() >= 4 and 0 < G[wrong, j].sum() < wrong.sum():\n        side.append({'finding': l, 'studies where text is wrong': int(wrong.sum()),\n                     'image AUC on those': roc_auc_score(G[wrong, j], IMG[wrong, j])})\nside = pd.DataFrame(side)\nprint('\\nwhere the text label disagrees with the radiologist:')\nprint(side.round(3).to_string(index=False))\nprint(f'\\nmean image AUC on text-wrong studies: {side.iloc[:, 2].mean():.3f}   (0.5 = no opinion; <0.5 = sides with the text)')"},{"cell_type":"markdown","metadata":{},"source":"The image correlates with the text it was trained on twice as strongly as with the\nradiologist -- 0.40 against 0.20 -- and on the studies where the text is wrong, it sides with\nthe text more often than not (mean AUC 0.44, below chance). That is what training on\ntext-derived labels does: the model inherits the extractor's mistakes along with its\nknowledge. It is not a model of the knee; it is a model of the knee *filtered through how\nradiologists write about it*, with the two findings of section 2 as the exceptions where the\nimage signal was strong enough to override the label."},{"cell_type":"markdown","metadata":{},"source":"## 5. What more data buys\n\nEverything above was trained on 929 studies because that is what fit on one disk. The\nobvious question is whether the image side is data-limited. Same recipe, fold 0, on a\nquarter, a half and all of the studies:"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"sc = pd.DataFrame({'fraction of studies': [0.25, 0.5, 1.0], 'studies in training': [173, 347, 695],\n                   'AUC vs text labels': [scale[k]['weak'] for k in ('0.25', '0.5', '1.0')],\n                   'AUC vs radiologist': [scale[k]['gold'] for k in ('0.25', '0.5', '1.0')]})\nprint(sc.round(3).to_string(index=False))\n\nfig, ax = plt.subplots(figsize=(7.8, 4.4))\nx = sc['studies in training']\nax.plot(x, sc['AUC vs text labels'], color=TEXT, lw=2, marker='o', ms=7, label='vs held-out text labels')\nax.plot(x, sc['AUC vs radiologist'], color=IMAGE, lw=2, marker='o', ms=7, label='vs radiologist (58 gold)')\nfor xi, a, b in zip(x, sc['AUC vs text labels'], sc['AUC vs radiologist']):\n    ax.text(xi, a + .008, f'{a:.3f}', ha='center', fontsize=9.5, color=TEXT)\n    ax.text(xi, b - .018, f'{b:.3f}', ha='center', fontsize=9.5, color=IMAGE)\nax.axvline(4349, color=GRID, lw=1.2, ls=(0, (4, 3)))\nax.text(4349 - 80, .60, 'all 4,349\\nunannotated\\nstudies', ha='right', fontsize=9, color=MUTED)\nax.set_xscale('log'); ax.set_xticks([173, 347, 695, 1390, 2780, 4349]); ax.set_xticklabels(['173', '347', '695', '1.4k', '2.8k', '4.3k'])\nax.set_xlim(140, 5200); ax.set_ylim(.55, .80)\nax.set_xlabel('studies in training (log)'); ax.set_ylabel('mean AUC over 12 findings')\nax.set_title('Both curves are still rising at the edge of the disk')\nax.legend(frameon=False, loc='upper left')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"From a quarter to all of the 929, the radiologist AUC climbs 0.593 → 0.609 → 0.646 -- about\nfive points per doubling, with no sign of flattening. The full corpus is another two and a\nhalf doublings away. Extrapolating a straight line through three points is not evidence, but\nthe direction is: the image side is data-limited, not recipe-limited, and the leaderboard's\ntop scores come from teams who trained on all of it.\n\n## What this adds up to\n\n1. **Trained on text labels, the image model scores 0.73 against text and 0.62 against the\n   radiologist.** The gap is the label noise the model inherited, measured out of fold.\n2. **On two findings the scan beats the report: lateral compartment OA (0.77 vs 0.56) and\n   effusion (0.72 vs 0.67).** Both are things a report states badly -- which side, how much --\n   and an image states directly.\n3. **On the findings a radiologist names, the report still wins,** by up to thirty points on\n   ACL. Seven hundred studies do not teach a network what a reader who saw the knee wrote down.\n4. **A single blend weight is the wrong tool.** The sources are good at different findings;\n   the right weight is per-finding, and it needs a holdout bigger than 58 to choose.\n5. **The image side is data-limited.** Five points of radiologist AUC per doubling, still\n   rising at 929 studies, with 4,407 available.\n\nPredictions, fold logs and the training code are in the attached dataset\n(`rsna-knee-image-oof-gold`). If you have trained on the full corpus and have an out-of-fold\nnumber on the 58 annotated studies, I would like to put it on the section 5 chart."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.11"}},"nbformat":4,"nbformat_minor":5}