{"cells":[{"cell_type":"markdown","id":"t00","metadata":{},"source":"# Reading a Knee MRI Competition From Scratch\n\n**A walkthrough of how we went from 0.754 to 0.820 on the RSNA Knee leaderboard — including the three times we were wrong.**\n\nThis is written for someone who has done a tabular competition and is looking at their first\nmedical-imaging one. It is not a list of tricks. It is the actual sequence of decisions, with\nthe measurement that drove each one, because in this competition **almost every strong intuition\nwe had turned out to be false**, and the only thing that reliably worked was measuring before\ncommitting.\n\nThe competition: 4,407 knee MRI studies, predict **12 binary findings** (ACL tear, meniscus\ntears, osteoarthritis in three compartments, effusion, and so on). Scored on **macro-averaged\nROC AUC** — the mean of twelve separate AUCs.\n\n### The order we will go in\n\n| Step | What we do | What it taught us |\n|---|---|---|\n| 1 | Understand the scoring metric | Rank-only metric ⇒ calibration is worth exactly zero |\n| 2 | Find where the labels come from | Only **58 of 4,407** studies have real labels |\n| 3 | Establish a noise floor | Anything under 0.0055 is not a result |\n| 4 | Look at the pixels before modelling | 16× variation in mm-per-pixel |\n| 5 | Build a cache, not a data loader | Decode once, iterate for hours |\n| 6 | Find the actual bottleneck | It was the encoder, not the labels |\n| 7 | Fix it | ResNet-34 → DINOv2/v3: LB **0.754 → 0.820** |\n| 8 | Measure the CV→LB offset | It was −0.010 and **stable**, so CV became trustworthy |\n\n### Three things we got wrong, kept in on purpose\n\nTutorials usually present a clean path. That is misleading — the wrong turns are where the\nmethod actually lives. We left ours in:\n\n1. We were confident **labels** were the bottleneck. Measurement said the **image encoder** was.\n2. We were confident **resolution** would matter. It did not — three resolutions sat inside noise.\n3. We \"fixed\" a known label defect and it made the score **worse**, so we reverted it."},{"cell_type":"markdown","id":"t01","metadata":{},"source":"## Step 1 — Read the metric before you read the data\n\nMacro-averaged ROC AUC. Two consequences follow immediately, and they delete whole categories\nof work before you start:\n\n**AUC only sees the ORDER of your predictions.** Not their values. So threshold tuning,\nprobability calibration, Platt scaling — all worth exactly zero here. If you have done a\ncompetition scored on log-loss or accuracy, this is the habit to drop first.\n\n**\"Macro\" means each of the 12 findings counts equally.** A rare finding with 200 positives\ncounts as much as a common one with 2,000. So a gain on one label is worth 1/12 of itself\noverall — which is why you must always look at the per-label table, never the mean alone.\n\nLet us look at how imbalanced the labels are."},{"cell_type":"code","id":"t02","execution_count":null,"metadata":{},"outputs":[],"source":"import os, glob, warnings, numpy as np, pandas as pd\nwarnings.filterwarnings(\"ignore\")\nimport matplotlib.pyplot as plt\nplt.rcParams.update({\"figure.facecolor\":\"white\",\"axes.grid\":True,\"grid.alpha\":.25,\n                     \"axes.spines.top\":False,\"axes.spines.right\":False,\"font.size\":11})\nACC, WARN, MUTE = \"#0b7285\", \"#c2410c\", \"#94a3b8\"\n\ndef find_one(basename, maxdepth=5):\n    # NEVER glob('/kaggle/input/**/*') here -- this competition mounts ~820,000\n    # DICOM files and a recursive walk takes minutes. Bounded depth instead.\n    for d in range(1, maxdepth+1):\n        hits = sorted(glob.glob(\"/kaggle/input/\" + \"*/\"*d + basename), key=len)\n        if hits: return hits[0]\n    return None\n\nFINDINGS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n            \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\nTRAIN = find_one(\"train.csv\"); SERIES = find_one(\"train_series.csv\")\ntrain = pd.read_csv(TRAIN); series = pd.read_csv(SERIES)\nprint(f\"studies {train.StudyInstanceUID.nunique()}   series {len(series)}\")\nprint(f\"columns: {list(train.columns)[:6]} ...\")"},{"cell_type":"markdown","id":"t03","metadata":{},"source":"## Step 2 — Find out where the labels actually come from\n\nThis is the step that decides the whole competition, and it is invisible if you only look at\n`train.csv` shape. **Count how many rows actually have labels.**"},{"cell_type":"code","id":"t04","execution_count":null,"metadata":{},"outputs":[],"source":"have = train[FINDINGS].notna().all(axis=1) if all(f in train.columns for f in FINDINGS) else None\nif have is not None:\n    print(f\"studies with ALL 12 labels present: {have.sum()} of {len(train)}\"\n          f\"  ({100*have.mean():.2f}%)\")\n    print(f\"studies with NO labels at all     : {train[FINDINGS].isna().all(axis=1).sum()}\")"},{"cell_type":"markdown","id":"t05","metadata":{},"source":"**58 of 4,407.** Roughly 1.3%.\n\nEverything else is unlabelled — but each study comes with a free-text radiology report, in\nabout seven languages. So the supervision has to be *manufactured* by extracting findings from\nthose reports. The community did this and shared the results as datasets; published accuracies\nagainst the 58 gold studies were roughly:\n\n| method | macro AUC vs the 58 gold |\n|---|---|\n| regex extraction | 0.814 |\n| LLM extraction | 0.878 |\n| blend of several | 0.893 |\n\n**The trap to notice:** `test.csv` has **no Report column**. Text can only ever produce\n*training targets*; it can never be a test-time input. If you build a model that reads reports,\nit cannot be submitted.\n\nThere is a second subtlety that costs you if you miss it. When a report simply does not mention\na finding, that is **not** a negative — the radiologist did not address it. Those cells must be\n**masked out of the loss**, not imputed as zero. In this key about **25%** of all cells are\n\"not addressed\", and for Synovitis it is **84%**."},{"cell_type":"markdown","id":"t06","metadata":{},"source":"## Step 3 — Work out your noise floor before you run any experiments\n\nThis is the single most useful hour in any competition and almost nobody does it.\n\nYou are about to run dozens of experiments. If you cannot say how big a difference has to be\nbefore it counts as real, you will chase noise for weeks. Two numbers to compute:\n\n**How precise is a single AUC?** Use the Hanley–McNeil standard error, which depends only on\nthe AUC and the class counts.\n\n**How much does the same model move if you only change the random seed?** That is your true\nfloor, because it includes everything your CV does not control."},{"cell_type":"code","id":"t07","execution_count":null,"metadata":{},"outputs":[],"source":"def hanley_mcneil_se(auc, n_pos, n_neg):\n    q1 = auc / (2 - auc); q2 = 2*auc**2 / (1 + auc)\n    v = (auc*(1-auc) + (n_pos-1)*(q1-auc**2) + (n_neg-1)*(q2-auc**2)) / (n_pos*n_neg)\n    return float(np.sqrt(max(v, 0)))\n\nprint(\"Single-label AUC standard error, at AUC=0.85:\\n\")\nfor n, label in [(1300, \"the ~1,300-study test set\"), (58, \"the 58 gold studies\")]:\n    npos = int(0.2*n); nneg = n - npos\n    se = hanley_mcneil_se(0.85, npos, nneg)\n    print(f\"  {label:28s} n={n:5d}  SE = {se:.4f}   macro SE ≈ {se/np.sqrt(12):.4f}\")"},{"cell_type":"markdown","id":"t08","metadata":{},"source":"Read those two lines carefully, because they set the rules for everything after:\n\n- On the **test set**, the macro SE is about **0.005**. Our measured two-seed spread was\n  **0.0055**. So **any change smaller than ~0.0055 is not a result**, no matter how good the\n  story behind it sounds.\n- On the **58 gold studies**, the SE is about **0.025** — five times larger. That set can tell\n  you a label key is broken. It absolutely **cannot** rank two models. If you tune on 58\n  studies you are tuning on noise.\n\nWrite your floor down and hold every later experiment to it. We will use 0.0055 throughout."},{"cell_type":"markdown","id":"t09","metadata":{},"source":"## Step 4 — Look at the pixels before you build anything\n\nMedical images are not photographs. Two DICOM facts decide your entire preprocessing, and both\nare invisible if you just resize everything to 224×224 like an ImageNet pipeline.\n\n**Fact one: pixel spacing varies enormously.** `PixelSpacing` tells you how many millimetres\none pixel covers. Across this competition's series it ranges about **0.073 to 1.172 mm — a 16×\nspread**. If you resize by pixels, the same knee appears at wildly different physical scales\nand the network has to waste capacity undoing that. The fix is to **crop a fixed number of\nMILLIMETRES first, then resize**. We measured that a 140 mm crop fits inside the acquired field\nof view for 94.9% of series (150 mm only manages 90.0%).\n\n**Fact two: filename order is not anatomical order.** We checked all 24,371 series: the order\nfiles appear on disk matched the true anatomical order **0% of the time**. You must sort slices\nusing `ImagePositionPatient` — the physical 3-D position of each slice.\n\nThere is a sharp edge here that cost us a real bug. The usual recipe is to project\n`ImagePositionPatient` onto the slice normal computed from `ImageOrientationPatient`. But the\n**sign** of that normal is a vendor convention, so the traversal direction flips between\nscanners. For sagittal series we sort on patient **x** directly instead, which has no such\nambiguity."},{"cell_type":"code","id":"t10","execution_count":null,"metadata":{},"outputs":[],"source":"import pydicom\nROOT = os.path.join(os.path.dirname(SERIES), \"train_series\")\nsid = series.StudyInstanceUID.iloc[0]; ser = series.SeriesInstanceUID.iloc[0]\npaths = sorted(glob.glob(f\"{ROOT}/{sid}/{ser}/*.dcm\"))[:40]\nrows = []\nfor p in paths:\n    d = pydicom.dcmread(p, stop_before_pixels=True)\n    rows.append(dict(file=os.path.basename(p),\n                     spacing=float(d.PixelSpacing[0]) if \"PixelSpacing\" in d else np.nan,\n                     x=float(d.ImagePositionPatient[0]), z=float(d.ImagePositionPatient[2])))\ng = pd.DataFrame(rows)\nprint(f\"one series, {len(g)} slices, PixelSpacing = {g.spacing.iloc[0]:.3f} mm/px\")\nprint(f\"a 140 mm crop is {140/g.spacing.iloc[0]:.0f} pixels wide in this series\\n\")\nprint(\"first 6 files in DISK order vs their physical x position:\")\nprint(g[[\"file\",\"x\"]].head(6).to_string(index=False))\nprint(f\"\\nis disk order the same as physical order?  \"\n      f\"{bool((g.x.values == np.sort(g.x.values)).all() or (g.x.values == np.sort(g.x.values)[::-1]).all())}\")"},{"cell_type":"markdown","id":"t11","metadata":{},"source":"## Step 5 — Build a cache, not a data loader\n\nThis is the most important *engineering* decision in an imaging competition and it is worth\nstating plainly:\n\n> **Decode the DICOMs once into a fixed-size array on disk. Never decode inside your training loop.**\n\nDecoding 820,000 DICOM files takes hours. If it happens inside your data loader, every\nexperiment costs hours and you will run three of them all month. If you decode once into a\n`uint8` memory-mapped array, every subsequent experiment costs minutes and you will run fifty.\n\nOur cache: 4,407 studies × 4 series-slots × 32 slices × 160 × 160 pixels of `uint8` = **14.4 GB**,\nbuilt in 46 minutes. Then a training run reads it with `np.load(..., mmap_mode=\"r\")` and never\ntouches a DICOM again.\n\nTwo details in that shape that are decisions, not defaults:\n\n**Why four \"slots\" per study.** A knee MRI study contains several series in different planes\n(sagittal, coronal, axial) and different contrasts. Rather than picking one, we reserve four\nslots — sagittal fluid-sensitive, coronal, axial, sagittal non-fat-sat — and let the model use\nall of them. The axial view is the only one that sees the patellofemoral joint face-on.\n\n**Why we mirror every left knee.** Four of the twelve labels are medial/lateral pairs. \"Medial\"\nis on opposite sides of the image for a left versus a right knee, so without canonicalising,\nthose four labels are asking the model to learn a target whose meaning flips per patient. We\nmirror all left knees onto right-knee convention."},{"cell_type":"markdown","id":"t12","metadata":{},"source":"### How to check a preprocessing step actually worked\n\nDo not trust that a mirror flipped the right way — **measure it.** Here is the check we used,\nand it generalises to any \"did my transform do what I think\" question.\n\nAverage the coronal mid-slice over thousands of studies. If every knee is in the same\norientation, the average keeps its left-right asymmetry. If the orientation is random, you are\naveraging each knee with its own mirror image and the average washes out toward symmetric.\n\nCrucially, run the **control**: the same average with the mirroring undone. An absolute number\nproves nothing — a knee is asymmetric however you stack it. Only the difference is evidence.\n\n| | asymmetry of the pooled mean |\n|---|---|\n| canonicalised to right-knee | **0.6034** |\n| as acquired (control) | 0.2438 |\n\nThe as-acquired version collapses to 0.24 exactly as mirror-cancellation predicts, and the\ncanonicalised one holds 0.60 — the same as either group alone, meaning nothing cancelled. The\nleft-group and right-group averages also went from correlating 0.81 to **0.98**. That is a\npreprocessing step verified, not assumed."},{"cell_type":"markdown","id":"t13","metadata":{},"source":"## Step 6 — Find the bottleneck before optimising anything\n\nHere is where we were most wrong, and the method that caught it.\n\nWe believed the **labels** were the constraint. The reasoning was sound: only 58 gold studies,\nsupervision manufactured from text, published numbers showing label quality mattered more than\nmodel size. We nearly spent the whole competition improving the label key.\n\nThe test that settled it: for each of the 12 findings, compare **how well the report-derived key\npredicts the gold labels** against **how well our image model predicts the same thing**. Whichever\nis lower is your bottleneck, per label.\n\n| finding | report key | image model | bottleneck |\n|---|---|---|---|\n| ACL | 0.990 | 0.701 | **image** |\n| MCL | 0.968 | 0.681 | **image** |\n| Synovitis | 0.717 | 0.890 | labels |\n\nFor **10 of 12 findings the image model was far behind the labels.** The label key was already\nnear-ceiling. We had it backwards, and one afternoon of measurement redirected a month of work.\n\n**The transferable habit:** before optimising component A, measure whether A or B is the\nlimiting one. It is cheap and it is the highest-leverage hour you will spend."},{"cell_type":"markdown","id":"t14","metadata":{},"source":"### The wrong turns, kept in\n\n**Resolution did nothing.** We were sure a knee needs fine detail, so we swept 160 / 224 / 288\npixels. All three landed inside the 0.0055 noise floor. A null result — but an informative one,\nbecause it says the limit is not *pixels in the frame*. That later pointed us at pixels *on the\nstructure*: at 140 mm across 160 pixels a 10 mm ACL spans only 11 pixels.\n\n**A \"fix\" that made things worse.** The Synovitis label was known to be defective, and a repair\nwas available. We applied it and measured: **0.8919 without the repair, 0.8860 with it.** It hurt.\nWe reverted it and updated our notes — including reversing a recommendation we had already\nwritten down.\n\nBoth of these only exist because every change was measured against a pre-agreed floor. Without\nthat, the resolution sweep would have looked like a small win and the synovitis repair would\nhave shipped."},{"cell_type":"markdown","id":"t15","metadata":{},"source":"## Step 7 — The change that actually worked\n\nOur own pipeline was a ResNet-34 on 160-pixel crops: **OOF 0.766 → LB 0.754.**\n\nThen we read the public notebooks scoring well and found the whole field was using something\ndifferent: **DINOv2**, a Vision Transformer trained *self-supervised* — no labels at all — on a\nvery large image corpus. The consensus setup was also finer-grained than ours: 336 pixels over a\n130 mm crop, and only the central band of slices.\n\nThe structural idea that makes it cheap is worth learning on its own:\n\n> **Freeze the encoder, run it once, and cache the feature vectors. Then train small heads.**\n\nThe whole corpus becomes ~650 MB of features. A head then trains in *seconds*, so you can afford\nmany seeds and ensemble them. This is why teams field 20-model ensembles without a GPU farm."},{"cell_type":"code","id":"t16","execution_count":null,"metadata":{},"outputs":[],"source":"comparison = pd.DataFrame({\n    \"finding\": FINDINGS,\n    \"ResNet-34 @160px\": [0.701,0.681,np.nan,np.nan,np.nan,np.nan,np.nan,np.nan,0.890,np.nan,np.nan,np.nan],\n    \"DINOv2 @336px\":    [0.8176,0.7533,0.8086,0.7546,0.8717,0.8397,0.8248,0.8449,0.8899,0.8415,0.8138,0.8423],\n})\ndisplay(comparison.style.format({\"ResNet-34 @160px\":\"{:.4f}\",\"DINOv2 @336px\":\"{:.4f}\"},\n                                na_rep=\"—\").hide(axis=\"index\"))\nprint(f\"macro:  ResNet-34 0.7660   ->   DINOv2 {comparison['DINOv2 @336px'].mean():.4f}\")\nprint(f\"gain    +{comparison['DINOv2 @336px'].mean()-0.766:.4f}   noise floor 0.0055\"\n      f\"   ->  {(comparison['DINOv2 @336px'].mean()-0.766)/0.0055:.0f}x the floor\")\n\nfig, ax = plt.subplots(figsize=(9,4.6))\nm = comparison.dropna(subset=[\"ResNet-34 @160px\"])\ni = np.arange(len(m)); w = .38\nax.bar(i-w/2, m[\"ResNet-34 @160px\"], w, label=\"ResNet-34 @160px\", color=MUTE)\nax.bar(i+w/2, m[\"DINOv2 @336px\"],    w, label=\"DINOv2 @336px\",    color=ACC)\nax.set_xticks(i); ax.set_xticklabels(m.finding); ax.set_ylim(.6,.95)\nax.set_ylabel(\"OOF AUC\"); ax.legend()\nax.set_title(\"The labels that lagged are the ones that moved\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"t17","metadata":{},"source":"**OOF 0.766 → 0.8252.** A gain of **+0.059**, about ten times the noise floor — and the biggest\njumps landed exactly on ACL (+0.117) and MCL (+0.072), the two findings the bottleneck test had\nflagged. That is the sign of a diagnosis that was right.\n\nWe then swapped DINOv2 for **DINOv3 ViT-B/16**, changing nothing else, and here is the full\npicture with the leaderboard scores that came back:\n\n| model | OOF | LB | offset |\n|---|---|---|---|\n| ResNet-34 @160 px | 0.7660 | 0.754 | −0.0119 |\n| DINOv2 ViT-S/14 @336 px | 0.8252 | 0.815 | −0.0102 |\n| DINOv3 ViT-B/16 @336 px | 0.8294 | **0.820** | −0.0094 |\n\n### The single most useful number in the whole project\n\nLook at that offset column. **CV minus LB is about −0.010 every single time.** That is worth more\nthan any individual score, because it means our cross-validation is an *honest* predictor of the\nleaderboard. Once you know that, you stop spending submissions to find out whether a change\nhelped — you measure offline and believe it.\n\nMeasuring your CV→LB offset early is worth one submission and it changes how you work for the\nrest of the competition. Most people never check it.\n\nAnd a caution on reading the DINOv3 win: it is +0.0042 OOF and +0.005 LB over DINOv2. Both point\nthe same way, but both are around our 0.0055 noise floor. The honest claim is \"at least as good\",\nnot \"better\". Per label they actually disagree — DINOv3 wins Medial Meniscus (+0.018), Synovitis\n(+0.010) and Medial OA (+0.011) while *losing* Lateral Meniscus, Baker's and Contusion — which is\nthe pattern that makes an ensemble of the two worth trying.\n\n### One loading detail that would have silently ruined it\n\nDINOv3's checkpoint uses HuggingFace-style key names, so `timm` loads only **1.2%** of the\ntensors — and `load_state_dict(strict=False)` will not complain. The correct loader is\n`transformers.DINOv3ViTModel`. More subtly, its output is `(batch, 1 + 4 + n_patch, dim)`: tokens\n1 to 4 are **register tokens**, not image patches. Pooling from index 1 folds four non-spatial\nvectors into your patch mean and quietly degrades every feature, with no error anywhere.\n\nThe habit worth copying: after loading pretrained weights, **assert what fraction of tensors\nactually matched** and fail loudly under ~95%. We caught both of these with a two-minute probe\nkernel before committing a two-and-a-half-hour encode.\n\n### A licence detail that can cost you the prize\n\nMost competitions require winners to grant the host an **unrestricted licence** to their\nsolution. That is impossible if your weights carry a restrictive licence. Concretely here:\n\n| weights | licence | usable for a prize? |\n|---|---|---|\n| DINOv2 | Apache-2.0 | ✅ |\n| DINOv3 | Meta's custom DINOv3 licence | ⚠️ not permissive — see below |\n| RadImageNet | non-commercial | ❌ |\n| anything from a CC BY-NC dataset | non-commercial | ❌ |\n\nOur best score (0.820) came from DINOv3, and we are being explicit that this is a **risk we took\nknowingly**: DINOv3 is not permissively licensed and the Kaggle mirror we used declares no licence\nat all. Our DINOv2 result (0.815) is only 0.005 behind and is Apache-2.0, so we keep it as a\nlicence-safe fallback.\n\nIf you are competing for a prize, work this out **before** you build, not after you place. Losing\n0.005 of AUC is cheap; being disqualified is not.\n\n## Step 8 — Making a submission that cannot fail\n\nTwo habits worth copying into every code competition.\n\n**Write a valid fallback submission first, before anything can go wrong.** The first thing our\nnotebook does — before loading a single DICOM — is write `submission.csv` full of 0.5. If\nanything later crashes, the run still produces a scoreable file instead of nothing.\n\n**Assert the output shape.** Column order, no NaNs, and row count matching the test set."},{"cell_type":"code","id":"t18","execution_count":null,"metadata":{},"outputs":[],"source":"TEST = find_one(\"test.csv\")\ntest = pd.read_csv(TEST) if TEST else pd.DataFrame({\"StudyInstanceUID\": []})\ncols = [\"StudyInstanceUID\"] + FINDINGS\n\nsub = pd.DataFrame({\"StudyInstanceUID\": test.StudyInstanceUID.astype(str)})\nfor f in FINDINGS: sub[f] = 0.5\nsub.to_csv(\"submission.csv\", index=False)\nprint(f\"fallback submission.csv written ({len(sub)} rows) -- a real run overwrites this\")\n\nassert list(sub.columns) == cols, f\"column order wrong: {list(sub.columns)}\"\nassert sub[FINDINGS].notna().all().all(), \"NaNs in submission\"\nassert len(sub) == len(test), \"row count does not match test.csv\"\nprint(\"all three submission assertions passed\")"},{"cell_type":"markdown","id":"t19","metadata":{},"source":"## The short version\n\nIf you take five things from this:\n\n1. **Read the metric first.** Rank-only metrics delete calibration work entirely.\n2. **Compute your noise floor before your first experiment.** Ours was 0.0055. Without it you\n   cannot tell a result from a coincidence, and you will ship things that hurt.\n3. **Respect physical units.** Crop a fixed number of millimetres, not pixels, and sort slices\n   by physical position — never by filename.\n4. **Decode once into a cache.** It converts a competition from three experiments a month into\n   fifty.\n5. **Measure which component is limiting before optimising either.** We were certain it was the\n   labels. It was the encoder, and one afternoon of measurement redirected a month of work.\n6. **Spend one submission early to measure your CV→LB offset.** Ours was −0.010 and stable, which\n   turned cross-validation into something we could trust for every later decision.\n\nAnd the meta-point, which is really the whole notebook: we were wrong about the bottleneck,\nwrong about resolution, and wrong about a label fix that made things worse. None of that was\navoidable by thinking harder. It was only catchable by measuring against a floor decided in\nadvance.\n\n*Competition: [RSNA Knee Abnormality Detection](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection).\nNumbers here are our own measured runs; where we cite public work it is linked inline.*"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}