{"nbformat":4,"nbformat_minor":5,"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","pygments_lexer":"ipython3"}},"cells":[{"cell_type":"markdown","metadata":{},"source":"# RSNA Knee · 0.662 → 0.860 (process log, not a submission kernel)\n\nPublic best is **0.860**. The backbone never changed: DINOv2-S. Four steps did the real work:\n\n1. Stop treating every slice as if it carried the full study label. Train **study-bag mean BCE** instead (0.792 → 0.821).\n2. Stop sampling slices by filename. Sort by **ImagePositionPatient** and take **8 consecutive slices at stack center** (0.821 → 0.854).\n3. Mix the two public teacher CSVs **1:1** on that same geom8 cache (0.854 → 0.859).\n4. Keep that recipe; stack each selected slice as gray neighbors `[-1, 0, +1]` (0.859 → **0.860**). Thin public gain. Local 58-gold had said stop.\n\nThis notebook is a writeup. It does not write `submission.csv` and it does not attach scoring weights.\n\n**Do not use this kernel to submit.** CPU, Internet Off, no weights, no `submission.csv`.\n\nKernel: [this page](https://www.kaggle.com/code/moushunchen/rsna-knee-0-859-ipp-bag-after-study-level-bce) (same notebook; slug followed the title from 0.854 → 0.859).\n\nSnapshot: 2026-08-20 19:05 CST. The latest **scoring** kernel is separate and stays private. Two selected slots were **0.859 + 0.854** at this snapshot (not auto-changed). Best two public guns are **0.860 + 0.859**. **0.662 is never selected.**\n\n"},{"cell_type":"markdown","metadata":{},"source":"## Thanks — the label teachers (public CSVs, not other people's weights)\n\nSoft labels are why this run could climb off 0.662. We trained our own DINOv2-S on two **CC0 report-probability tables**. We did **not** load public high-scoring checkpoints, and we did **not** submit a Bend-the-Knee-style 0.91x ensemble as our own answer.\n\n| Teacher | Kaggle | File we actually trained on |\n|---|---|---|\n| **[pilkwang](https://www.kaggle.com/pilkwang)** | [rsna-knee-llm-labels](https://www.kaggle.com/datasets/pilkwang/rsna-knee-llm-labels) | `report_labels_v2.csv` — 0.786 / 0.792 / 0.821 / 0.854 |\n| **[stevenleehans](https://www.kaggle.com/stevenleehans)** | [rsna-knee-llm-report-labels](https://www.kaggle.com/datasets/stevenleehans/rsna-knee-llm-report-labels) | `llm_labels_v4_blend.csv` — mixed **1:1** with pilkwang for **0.859** (and the later knives) |\n\nAlso useful: pilkwang's [baseline writeup](https://www.kaggle.com/code/pilkwang/rsna-knee-baseline-v1) for the idea “train on LLM soft probabilities; do not skip 0.5”.\n\nThank you both for turning reports into trainable probabilities. Gold labels exist for only 58 studies. Without these tables, supervision collapses to the “said loudly in the report” subset, and public score stalls near 0.66.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 1. TL;DR — public ladder (the only ruler that counts)\n\n| Submit | What it was | Public |\n|---|---|---|\n| 55605980 / 55606376 | v12, selected on report-val 0.8424 | **0.662** — do not select |\n| 55607389 | A1: pilkwang soft labels, gold-58 pick, 1 fold | **0.786** |\n| 55608313 | same family, 3-fold logits mean | **0.792** |\n| 55610093 | study-bag, 3-fold logits mean, filename cache | **0.821** |\n| 55625266 | same bag + IPP **center-consecutive 8** | **0.854** |\n| 55629409 | same geom8 images + pilkwang⊗steven 1:1 teachers | **0.859** |\n| 55633558 | geom8-spread 5–95 eight-point v1 | COMPLETE, **no score** (hidden rerun crash) |\n| 55634369 | geom8-spread 5–95 eight-point v2 (decode fallback) | **0.844** — lost to center-8 |\n| **55643517** | **geom8-25d: center-8 RGB = neighbors `[-1,0,+1]`** | **0.860** current best |\n\nSilver (~0.921) is still the gap. After 0.859 we tried **one variable at a time**. Replacing center-8 with a 5–95 eight-point spread **lost**. Neighbor-channel 2.5D **won on public by +0.001** even though local 58-gold said stop (0.814 vs 0.837). Public is the ruler. The gain is thin — not a new platform.\n\n"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\nladder = pd.DataFrame(\n    [\n        [\"v12 report-val lock\", \"55606376\", 0.662, \"wrong ruler\"],\n        [\"A1 soft fold0\", \"55607389\", 0.786, \"gold-58 pick\"],\n        [\"slice 3-fold mean\", \"55608313\", 0.792, \"same family +0.006\"],\n        [\"study-bag filename\", \"55610093\", 0.821, \"study BCE +0.029\"],\n        [\"study-bag IPP geom8\", \"55625266\", 0.854, \"center-8 +0.033\"],\n        [\"geom8 + teachers 1:1\", \"55629409\", 0.859, \"same images +0.005\"],\n        [\"geom8 5-95 spread\", \"55634369\", 0.844, \"sampling lose -0.015\"],\n        [\"geom8 2.5D neighbors\", \"55643517\", 0.860, \"RGB [-1,0,+1] +0.001\"],\n    ],\n    columns=[\"run\", \"ref\", \"public_auc\", \"note\"],\n)\ndisplay(ladder)\n\nfig, ax = plt.subplots(figsize=(8.2, 3.6))\nxs = range(len(ladder))\nax.plot(xs, ladder[\"public_auc\"], marker=\"o\")\nax.set_xticks(list(xs), ladder[\"run\"], rotation=22, ha=\"right\")\nax.set_ylabel(\"public macro AUC\")\nax.set_ylim(0.64, 0.90)\nax.set_title(\"RSNA Knee public ladder — scoring kernels, not this log\")\nax.axhline(0.860, linestyle=\"--\", linewidth=1, label=\"current best 0.860\")\nax.legend(frameon=False)\nax.grid(True, axis=\"y\", alpha=0.3)\nplt.tight_layout()\nplt.show()\nprint(\"process log only; not a submission kernel\")\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 2. Data bugs (still true)\n\n1. Official DICOMs live under `train_series/` / `test_series/`, not `train_images/`.\n2. `train.csv` has 4407 studies; **all-12 non-null gold is 58 studies**.\n3. Extra labels come from report parsing. `0.5` / NaN / not-addressed is **not** negative.\n4. Split by **StudyInstanceUID**. Never leak slices of the same study.\n5. Submission row order is `sample_submission.csv`. Never hard-code 3 rows.\n6. This is a **code competition**. Upload a 3-row smoke CSV and you score garbage. Use CreateCodeSubmission / hidden rerun.\n7. `Fluid_Sensitive` equals `Fat_Suppression` on every `train_series.csv` row we checked. Do not run an ablation on a missing field.\n8. The **old** JPEG cache sorted files by **name**, then took 4 slices in the 20%–80% band. Filename order is not anatomy. **geom8** (the 0.854 / 0.859 cache) sorts by ImagePositionPatient and takes **8 consecutive slices at stack center**: `start = (n-8)//2`. Do not overwrite that cache; do not dump 570GB of DICOM to a laptop.\n9. A later cache that kept IPP order but took **eight linspace points in 5%–95%** is a different sample. On an 8-slice budget it lost to center-8. Do not reuse it as a drop-in replacement.\n\nMetric: per-label AUC, then macro over 12 columns.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 3. Why 0.662 happened\n\nWe locked a Kaggle val of **0.8424** on v12 (DINOv2-S, 336, 6 slots × 4 mid-band slices, no TTA). That val was **report labels**, not the 58-study gold, and not public.\n\nPublic came back **0.662**. The checkpoint was real; the **selection ruler** was not. Naive horizontal-flip TTA also lost on that same val (0.8424 → 0.838). Laterality-aware v13 lost macro even though Lateral OA moved. Those recipes stay dead.\n\nLesson: if validation is easy OA/effusion columns and the 58 gold studies are ignored, you can “lock 0.84” and still ship 0.66.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 4. Climb 1 — pilkwang soft labels + gold selection → 0.786, then 0.792\n\nSame DINOv2-S, same 336 JPEG cache, **no laterality flip**.\n\n- Labels: pilkwang LLM soft probabilities (`report_labels_v2`). Every finite value is trained (no 0.5 skip).\n- Model pick: **58-study hard-gold AUC**, never report-val.\n- Train: 1 epoch head-only, then unfreeze the last 6 blocks, 6 epochs, local RTX 5070 Ti.\n- Infer: Internet Off, Tesla T4, no TTA, no horizontal flip.\n\nA1 (fold 0, best epoch 5, gold 0.7797 on 27 studies) public **0.786**.\n\nThen fold 1 and fold 2 on the same recipe. Inference averages the three models’ **logits**, then study-mean, then sigmoid. Public **0.792** (+0.006). That is the usual isomorphic-ensemble bump.\n\nThose three `model.pt` files stay locked.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 5. Climb 2 — study-bag → 0.821\n\nThe remaining bug was **not** the backbone.\n\nOld train: every slice independently received the **study** vector. A slice that never shows ACL was still trained as ACL-positive.\n\nInference already did: encode slices → mean logits → sigmoid. Training optimized the wrong object.\n\n**bag-v1a**:\n\n- Group cached JPEGs per study (p50=20, max=24).\n- Batch 1–2 studies, pad + slice mask.\n- Mean-pool **logits**, one masked BCE at study level.\n- Same S backbone, BF16, 6 epochs, head then last-6, pilkwang soft labels, no flip.\n- Compare on **fixed epoch 6**, not per-fold gold early-stopping.\n- Infer: three fold weights, logits mean, no TTA.\n\nA paired 57-gold check vs the epoch-6 slice baseline was noisy (Δ +0.0088, P=0.70). Public still jumped **0.792 → 0.821** (+0.029). The hidden public set is larger, and it uses the exact quantity the bag loss trains.\n\nOn those 57 gold UIDs the bag helped Fracture / Effusion / Contusion; Lateral OA and Lateral Meniscus got worse. Macro public can still rise when the supervision object matches the metric.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 6. Climb 3 — IPP geometry, center-8 → **0.854**\n\nFilename order is not anatomy. The hidden rerun can read DICOM headers, so we rebuilt a **second** JPEG cache:\n\n- Sort slices by ImagePositionPatient along the stack.\n- Take **8 consecutive** slices at the **center** of the stack (`start = (n-8)//2`), not 4 name-sorted slices, and not a 5–95 spread.\n- New image directory only. The old filename cache stays frozen.\n- Same bag loss, same DINOv2-S, same 6-epoch lock, same 3-fold logits mean, still **no TTA / no horizontal flip**.\n- Selection: `fixed_epoch6`, not gold early-stopping.\n\nPublic **0.854** (ref **55625266**). Hidden rerun is slower (~40–55 min) because it reads IPP. bag-v1a was ~10 min on the old cache.\n\nThis cache is still the image recipe under 0.859 and under the 2.5D knife.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 7. Climb 4 — teacher mix 1:1 → **0.859**\n\nOne variable: labels. Images stayed geom8 center-8. Backbone, bag, 6-epoch lock, no flip, 3-fold logits mean: unchanged.\n\n- Mix: arithmetic mean of pilkwang `report_labels_v2` and steven `llm_labels_v4_blend` (1:1).\n- Train: same 3-fold bag recipe, local 5070 Ti, ~2 hours wall.\n- Local 58-gold vs the 0.854 run was noisy (+0.013 macro, keep-gate did not fire). We still submitted because public is the ruler.\n- Public **0.859** (ref **55629409**, +0.005). Was the best until 2.5D. Still one of the two selected slots at this snapshot.\n\nHonest band after this climb is still ~0.85–0.87, not 0.90. Silver is ~0.921.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 8. After 0.859 — two single-variable knives\n\n### 8a. 5–95 eight-point spread → **0.844** (failed)\n\nReplaced center-consecutive 8 with IPP order + **8 linspace points in 5%–95%**. Same 1:1 teachers, same bag, same 6 epochs.\n\n- Local 58-gold vs 0.859: 0.800 vs 0.837 (Δ −0.037). ACL was the hole (~−0.13). Gate: stop.\n- Hidden v1 (ref 55633558): editor smoke passed, hidden rerun crashed (likely edge-slice decode). No public score.\n- Hidden v2 (ref **55634369**): decode fallback to zeros. Public **0.844**.\n\nOn an **8-slice budget**, spreading across 5–95 is not a better geom8. Center-consecutive 8 stays. We do not retry this replacement. A *wider consecutive* window, or a *larger slice budget*, would be a new knife, not a redo of this one.\n\n### 8b. 2.5D neighbor RGB → **0.860** (thin public win)\n\nOne variable: representation. Same center-8 JPEGs, same 1:1 teachers. Each selected slice is stacked as gray neighbors `[-1, 0, +1]` (edges clamped). Not mixed with 5–95. Not a teacher change. Not more slices.\n\n- Train: 3-fold, 6 epochs, local 5070 Ti. All folds locked at epoch 6.\n- Local 58-gold vs 0.859: 0.814 vs 0.837 (Δ −0.023). Only PF OA was not worse. Gate: stop.\n- Public **0.860** (ref **55643517**, +0.001). Hidden rerun ~1 hour. Gold and public disagreed: gold pessimistic, public a hair up.\n\nThe 2.5D line stays as the current best scoring recipe. The gain is **one public thousandth**, not a new platform. Do not restack 5–95 on top of it.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 9. Locked recipe now (scoring kernel, not this log)\n\n- Backbone: timm DINOv2-S, pool = CLS ∥ patch-mean, 336\n- Cache: IPP-sorted, **8 consecutive center slices** (geom8). Filename cache is historical. 5–95 spread cache is not reused for this 8-slice recipe.\n- Train: study bag, flat mean logits, masked BCE, **pilkwang⊗steven 1:1**, no laterality. Scoring run adds neighbor RGB `[-1,0,+1]` on the same picks.\n- Select: fixed epoch 6 for the bag family\n- Infer: 3-fold logits mean, **no TTA**, **no horizontal flip**, Internet Off, T4. Scoring 2.5D kernel stacks `[-1,0,+1]` on the same center-8 picks.\n- Submit: CreateCodeSubmission hidden rerun, never the 3-row smoke CSV\n\nCurrent best public: **0.860** (ref 55643517). Previous best: **0.859** (ref 55629409). At this snapshot the two selected slots were still **0.859 + 0.854** (not auto-changed). **0.662 is never selected.** Do not overwrite geom8 / blend11 / 25d / bag-v1a / slice-0.792 weights.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 10. Dead ends (keep these killed)\n\n**Naive horizontal flip at test.** Lost validation. Do not add it because “TTA is free”.\n\n**Laterality-aware train (v13).** Lateral OA up, macro down.\n\n**Trusting report-val 0.8424.** Public 0.662. Never select that run.\n\n**Same-family fusion as the old 0.83 plan.** 0.786 → 0.792 was the whole isomorphic bump. Do not blend 0.792 back in.\n\n**Per-column `1-MAE` loss weights / gold×2 override.** That is not the same as blending two teacher CSVs. Parked.\n\n**RadImageNet DenseNet121 as a second family.** The official PyTorch zip has no weights license. Legal gate fail.\n\n**570GB DICOM dump to the PC.** JPEG cache only.\n\n**Second account / teaming.** One account: `moushunchen`.\n\n**Copying a 0.91x public ensemble as “our” submission.** We train our own S. Teachers = CSVs.\n\n**Replacing center-8 with 5–95 eight-point spread on an 8-slice budget.** Local gold and public both lost (0.844).\n\n**This process log as a submit kernel.** It does not write `submission.csv` on purpose.\n\n"},{"cell_type":"markdown","metadata":{},"source":"## 11. Next\n\n1. **2.5D is the current best public recipe** (0.860). Next knife: at most **one** new variable on that stack (wider consecutive window, more slices, or DINOv2-B). Do not restack 5–95. Do not mix teacher + sampling + 2.5D in one shot.\n2. If selecting two slots: **0.860 + 0.859** are the top public pair. Do not select 0.662. We do not auto-select.\n3. Rank-blend a second family later. 0.821 at most a small rank mix; 0.792 stays out; 0.844 stays out.\n4. Focal loss stays parked.\n\nExcluded from this upload: the latest infer/scoring kernel, all `model.pt` files, the gold-58 checkpoint, tokens, and local paths.\n\n"}]}