{"cells":[{"cell_type":"markdown","id":"e37309b5f6a2","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:10px solid #2a78d6;border-radius:8px;padding:26px 30px;margin:6px 0 20px 0;\"><div style=\"color:#c04b1c;font-size:12px;font-weight:700;letter-spacing:.16em;text-transform:uppercase;margin-bottom:10px;\">Imaging model</div><h1 style=\"color:#0b0b0b;font-size:31px;line-height:1.22;margin:0;font-weight:700;letter-spacing:-.01em;\">RSNA Knee Abnormality Detection: a patch token wider than the meniscus</h1><div style=\"color:#6a6862;font-size:13px;margin-top:12px;\">Frozen DINOv2 ViT-S/14 &middot; 336 px over a 150 mm field of view &middot; 576 patch tokens &middot; slices ordered by physical position</div></div>\n\n> **How this was made.** This notebook was written automatically by Claude Code running a\n> custom multi-agent LLM harness: one agent drafted the pipeline while separate\n> adversarial verifier agents attacked it, executed every cell that can run without a GPU\n> against a local mirror of the Kaggle input mount, and sent each defect back to this\n> notebook's builder script rather than patching the notebook. It is published to Kaggle\n> and submitted to the competition with the Kaggle CLI, the same path taken by the\n> companion EDA and metadata-baseline notebooks. **Code cells are collapsed by default**\n> so this reads as a report; click any *Show code* toggle to expand one.\n\nEvery public imaging kernel on this competition so far runs DINOv2 at **126 pixels**.\nThat number is not a hyperparameter. It is a measurement, and the measurement says the\nmodel cannot see the disease.\n\nEvery series in this corpus is acquired over a **150 mm field of view**. Four different\nacquisition matrices appear in the files (960, 896, 800 and 640 pixels square, at 0.1562,\n0.1674, 0.1875 and 0.2344 mm per pixel) and all four multiply out to 150.0 mm. So a\n126-pixel input is 1.19 mm per pixel, and one 14-pixel DINOv2 patch token covers\n**16.7 mm** of knee.\n\nA medial meniscus is 9 to 12 mm wide. A meniscal tear is 1 to 3 mm. One token holds the\nentire meniscus plus three times its area of surrounding bone and fat, and a 1 mm tear is\n**below the sampling limit** of the image the network is shown: at 1.19 mm per pixel it\noccupies less than one pixel, so it is gone before the patch embedding ever sees it.\n\nThat is an argument from sampling theory, so section 2 stops arguing and measures it on\nthe frozen backbone itself. Painting a lesion of known physical width onto a field and\nresizing that field to each candidate input size, a 1 to 3 mm feature moves the pooled\nembedding **two to five times further at 336 px than at 126 px**, while a 10 mm feature\ngains nothing. Resolution buys nothing for a meniscus and everything for a tear, and\n448 px beats 336 on no lesion width tested, which is where the ladder stops.\n\nThis notebook acts on that, and on three other things the frontier leaves on the table:\n\n| Axis | Public frontier | Here |\n|---|---|---|\n| Input size | 126 px, 81 patch tokens, 16.7 mm per token | **336 px, 576 patch tokens, 6.25 mm per token** |\n| Slices encoded | ~18 of ~166 per study, ~11% | **~43 of ~166, ~26%** |\n| Slice ordering | filename sort, which is SOP UID order, arbitrary | **projected physical position from the DICOM geometry** |\n| Intensity | per-slice percentile window | **per-series window, so a bright effusion stays bright** |\n| Series slots | 6 slots, 3 of which are \"the second series in this plane, whatever it is\" | **5 slots defined by contrast, measured fill 93%** |\n\nEverything above is arithmetic, and every number in it is recomputed in the cells below\nfrom the competition data rather than quoted from a previous run."},{"cell_type":"markdown","id":"fb554c46f4dd","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">1</span><span style=\"vertical-align:middle;\">Setup</span></h2></div>\n\nThread counts are pinned before NumPy and Torch are imported, because setting them\nafterwards is a no-op. The input root is detected rather than hardcoded: Kaggle mounts\nthis competition at `/kaggle/input/competitions/<slug>`, one level deeper than the usual\n`/kaggle/input/<slug>`, and a shallow scan of `/kaggle/input` finds the `competitions`\ndirectory, sees no `train.csv` in it and gives up. Version 1 of the companion baseline\nnotebook died exactly that way in public. Both shapes are named explicitly here and the\nfallback walks two levels. The resolved root is printed.\n\nEverything the pipeline is allowed to spend is declared in one cell: input size,\nper-slot slice budgets, header-read caps, wall-clock budgets. Nothing further down\ninvents a bound of its own.\n\nThe frozen backbone is loaded here too rather than later, because the next section\nmeasures against it. Section 7 covers what it is and why it stays frozen."},{"cell_type":"code","execution_count":null,"id":"8246589a88fe","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"import os\n\n# Pin thread counts BEFORE numpy / torch import. Setting them afterwards is a no-op:\n# the libraries read these at import time and never look again.\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\",\n           \"NUMEXPR_NUM_THREADS\", \"VECLIB_MAXIMUM_THREADS\"):\n    os.environ.setdefault(_v, \"4\")\n\nimport gc\nimport json\nimport math\nimport random\nimport re\nimport time\nimport unicodedata\nimport warnings\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nwarnings.filterwarnings(\"ignore\")\n\nSEED = 20260805\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\nLABELS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n          \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n          \"Contusion\", \"Fracture\"]\n\n# ---- The lever this notebook is built around --------------------------------\n# Field of view is 150 mm on every series in this corpus: four acquisition matrices\n# appear (960/896/800/640 px at 0.1562/0.1674/0.1875/0.2344 mm) and all four give\n# 150.0 mm. So input size alone fixes the physical size of a patch token.\nFOV_MM = 150.0\nPATCH_PX = 14                 # DINOv2 ViT-S/14\nEMBED_DIM = 384               # ViT-S hidden size\nN_BLOCKS = 12                 # ViT-S depth, used only for the FLOP arithmetic\n\nIMG_SIZE = 336                # 24 x 24 = 576 patch tokens, 6.25 mm per token\nSIZE_LADDER = [336, 280, 224]  # downshift order if the GPU phase blows the budget\n\n# ---- Series slots ------------------------------------------------------------\n# (plane, fluid_sensitive_flag, slices_to_encode). Five slots, not six: an Axial\n# non-fluid series exists for only 19.4% of studies, so that budget is spent on\n# slices in the slots that always fill. Section 4 shows the measured fill rates.\nSLOT_SPEC = [\n    (\"Sagittal\", 1, 12),\n    (\"Sagittal\", 0, 12),\n    (\"Coronal\", 1, 8),\n    (\"Coronal\", 0, 6),\n    (\"Axial\", 1, 8),\n]\nSLOT_NAMES = [\"{}_{}\".format(p, \"Fluid\" if f else \"Structural\")\n              for p, f, _ in SLOT_SPEC]\nN_SLOTS = len(SLOT_SPEC)\nMAX_SLICES = max(k for _, _, k in SLOT_SPEC)\n\n# Fractional window of the ordered stack each plane is sampled across. Sagittal and\n# Axial need the full extent (menisci sit at the medial and lateral edges of the\n# sagittal stack; Baker's cysts sit posteriorly and low on the axial stack), coronal\n# is tighter because its useful anterior-posterior range is narrower.\nPLANE_WINDOW = {\"Sagittal\": (0.12, 0.88), \"Coronal\": (0.20, 0.80),\n                \"Axial\": (0.12, 0.88)}\n\n# ---- Per-slice embedding ------------------------------------------------------\n# CLS + patch mean + patch max. The max is the channel a single anomalous token\n# reaches the head through; mean pooling over 576 tokens dilutes it 576-fold.\nSLICE_DIM = 3 * EMBED_DIM     # 1152\n\n# ---- Budgets ------------------------------------------------------------------\n# A Kaggle GPU kernel gets 9 hours. Everything below is inside 30,000 s, leaving\n# ~2,400 s of slack for commit, install and the head.\nWALL_BUDGET_S = 30000.0\nTEST_ENCODE_BUDGET_S = 9000.0      # test is encoded FIRST and is never subsampled\nTRAIN_ENCODE_BUDGET_S = 14000.0\nHEAD_BUDGET_S = 2400.0\nMAX_HEADER_READS_PER_SERIES = 64   # median series holds 30 slices\nPROBE_STUDIES = 6\nRERUN_TEST_STUDIES = 1300          # the host's stated hidden test size, projection only\nMIN_TRAIN_STUDIES = 1200\nGPU_BATCH = 32\nIO_WORKERS = max(2, min(8, (os.cpu_count() or 4)))\n\n# ---- Head ---------------------------------------------------------------------\nN_FOLDS = 5\n# The pooled matrix is 5 slots x 2 statistics x 1,152 dims wide, so the ridge probe\n# runs on a truncated SVD refitted inside each fold. 512 components rather than 256:\n# a fold costs 3.0 s instead of 2.3 s at the real width, and RidgeCV picks its own\n# alpha by leave-one-out GCV on the training split, so the extra components are\n# regularised rather than fitted blind.\nSVD_COMPONENTS = 512\nHEAD_HIDDEN = 256\nHEAD_EPOCHS = 18\nHEAD_LR = 2.5e-3\nHEAD_WD = 1e-2\nHEAD_DROPOUT = 0.2\nHEAD_SEEDS = (20260805, 20260806, 20260807)\nGOLD_WEIGHT = 8.0                  # gold studies count this much more in the loss\nBLEND_W = 0.5                      # fixed, not tuned; the curve is printed anyway\n\n# ---- Fine-tuning the backbone -------------------------------------------------\n# Section 7b. The frozen path encodes 46 slices per study and can recombine the\n# features DINOv2 already has, but it cannot change what it looks WITH. This path\n# trades coverage for adaptation: one 2.5D image per slot, taken at the slot's\n# centre, with the trailing transformer blocks opened for training. Five images\n# per study instead of 46, and an encoder that moves. The two are rank-blended, so\n# the shallow-wide model and the deep-narrow one both contribute.\nFINETUNE = True\nFT_UNFREEZE = 4               # trailing blocks opened for training; the rest fixed\nFT_ENCODER_LR = 1e-5          # two orders below the head: the encoder starts good\nFT_HEAD_LR = 2.5e-3           # the head is random and has everything to learn\nFT_WD = 1e-2\nFT_WIDTH = 256\nFT_DROPOUT = 0.2\nFT_EPOCHS = 8                 # ceiling only; the probe lowers it to fit the clock\nFT_STUDY_BATCH = 4            # studies per step, so 4 x N_SLOTS images per forward\nFT_BUDGET_S = 6000.0          # wall clock for the whole fine-tune stage\nFT_VAL_FOLD = 0               # one grouped fold held out; five folds do not fit\nFT_BLEND_W = 0.45             # rank weight on the fine-tuned model, fixed not tuned\nFT_SEED = 20260806\nFT_DISK_BUDGET_GB = 12.0      # /kaggle/working is 20 GB and holds the embeddings too\n\n# ---- Labels -------------------------------------------------------------------\n# ONE constant selects the training target. \"auto\" prefers LLM soft labels from an\n# attached dataset, then a rule-label file, then the inline rule extractor.\nLABEL_SOURCE = \"auto\"              # \"auto\" | \"llm\" | \"rule_file\" | \"rule_inline\"\nLLM_LABEL_PATTERNS = [\"labels_llm_*.parquet\", \"labels_llm_*.csv\"]\nRULE_LABEL_PATTERNS = [\"labels_rule_*.parquet\", \"labels_rule_*.csv\"]\n# Mirrors the builder's LABEL_DATASET_SOURCES, which writes kernel-metadata.json. It has\n# to exist INSIDE the notebook too, because the fail-closed guard below reads it at run\n# time; version 3 raised NameError on exactly this. Non-empty means a label dataset was\n# declared, so a fallback to rule labels is an error rather than a default.\nLABEL_DATASET_SOURCES = ['wguesdon/rsna-knee-llm-report-labels-opus']\n\n# ---- Cache --------------------------------------------------------------------\nSAVE_TRAIN_CACHE = False           # True to keep the train embeddings in the output\nLOAD_TRAIN_CACHE_FROM_DATASET = False\n\nprint(\"torch/GPU checked in the next cell; configuration first\")\nprint(\"input size        : {} px  ->  {:.3f} mm/px, {:.2f} mm per {}px token, {} tokens\".format(\n    IMG_SIZE, FOV_MM / IMG_SIZE, FOV_MM * PATCH_PX / IMG_SIZE, PATCH_PX,\n    (IMG_SIZE // PATCH_PX) ** 2))\nprint(\"slots             : {}\".format(\", \".join(\n    \"{}({})\".format(n, k) for n, (_, _, k) in zip(SLOT_NAMES, SLOT_SPEC))))\nprint(\"slices per study  : {} when every slot fills\".format(\n    sum(k for _, _, k in SLOT_SPEC)))\nprint(\"per-slice embedding: {} dims (CLS + patch mean + patch max)\".format(SLICE_DIM))\nprint(\"io workers        : {}\".format(IO_WORKERS))"},{"cell_type":"code","execution_count":null,"id":"2b2fd6a8af3d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"import matplotlib as mpl\nimport matplotlib.pyplot as plt\n\n# ---- Ink and surface ---------------------------------------------------------\n# The same design system as the companion EDA and baseline notebooks, so the three\n# read as one series. Three text tiers, every one at or above the WCAG AA 4.5:1\n# floor against SURFACE. Ratios are computed below, not eyeballed. GRID and AXIS\n# carry no text and are deliberately left light.\nSURFACE = \"#fcfcfb\"\nINK = \"#0b0b0b\"       # titles, tick labels                     19.2:1\nINK_2 = \"#52514e\"     # value annotations, detail line           7.7:1\nMUTED = \"#6a6862\"     # axis labels, captions                    5.4:1\nGRID = \"#e1e0d9\"\nAXIS = \"#c3c2b7\"\n\n# A bar does not have to clear a contrast floor; the sentence next to it does.\nPRIMARY = \"#2a78d6\"\nACCENT = \"#eb6834\"\nACCENT_TEXT = \"#c04b1c\"   # accent annotations                   4.8:1\nPRIMARY_DEEP = \"#256abf\"\nNEUTRAL = \"#c3c2b7\"\n\nmpl.rcParams.update({\n    \"figure.facecolor\": SURFACE,\n    \"figure.dpi\": 118,\n    \"savefig.facecolor\": SURFACE,\n    \"axes.facecolor\": SURFACE,\n    \"axes.edgecolor\": AXIS,\n    \"axes.linewidth\": 0.9,\n    \"axes.spines.top\": False,\n    \"axes.spines.right\": False,\n    \"axes.labelcolor\": MUTED,\n    \"axes.labelsize\": 12.5,\n    \"axes.titlesize\": 13.5,\n    \"axes.titlecolor\": INK,\n    \"axes.grid\": False,\n    \"grid.color\": GRID,\n    \"grid.linewidth\": 0.8,\n    \"grid.linestyle\": \"-\",\n    \"xtick.color\": MUTED,\n    \"ytick.color\": MUTED,\n    \"xtick.labelsize\": 12,\n    \"ytick.labelsize\": 12,\n    \"xtick.major.size\": 0,\n    \"ytick.major.size\": 0,\n    \"legend.frameon\": False,\n    \"legend.fontsize\": 12,\n    \"lines.linewidth\": 2.4,\n    \"lines.markersize\": 9,\n    \"font.family\": \"sans-serif\",\n    \"font.sans-serif\": [\"DejaVu Sans\", \"Segoe UI\", \"Helvetica\", \"Arial\"],\n    \"font.size\": 12.5,\n    \"text.color\": INK,\n})\n\n\ndef _rel_lum(colour):\n    \"\"\"WCAG relative luminance of a colour.\n\n    Args:\n        colour: Anything matplotlib can parse as a colour.\n\n    Returns:\n        Relative luminance in [0, 1].\n    \"\"\"\n    total = 0.0\n    for weight, chan in zip((0.2126, 0.7152, 0.0722), mpl.colors.to_rgb(colour)):\n        lin = chan / 12.92 if chan <= 0.03928 else ((chan + 0.055) / 1.055) ** 2.4\n        total += weight * lin\n    return total\n\n\ndef contrast(fg, bg):\n    \"\"\"WCAG 2.x contrast ratio between two colours.\n\n    Args:\n        fg: Foreground colour.\n        bg: Background colour.\n\n    Returns:\n        A ratio between 1.0 and 21.0. AA wants 4.5 for normal-size text.\n    \"\"\"\n    lo, hi = sorted((_rel_lum(fg), _rel_lum(bg)))\n    return (hi + 0.05) / (lo + 0.05)\n\n\ndef style(ax, grid_axis=\"x\"):\n    \"\"\"Apply the recessive grid and spine treatment to one axes.\n\n    Args:\n        ax: A matplotlib axes.\n        grid_axis: \"x\", \"y\" or \"none\". Gridlines run along this axis only,\n            behind the marks.\n\n    Returns:\n        The same axes, for chaining.\n    \"\"\"\n    for side in (\"top\", \"right\"):\n        ax.spines[side].set_visible(False)\n    for side in (\"left\", \"bottom\"):\n        ax.spines[side].set_color(AXIS)\n    ax.grid(False)\n    if grid_axis == \"x\":\n        ax.xaxis.grid(True, zorder=0)\n    elif grid_axis == \"y\":\n        ax.yaxis.grid(True, zorder=0)\n    ax.set_axisbelow(True)\n    ax.tick_params(length=0)\n    return ax\n\n\ndef _wrap(text, width_in, fontsize):\n    \"\"\"Wrap a header or caption string to the figure's own width.\n\n    The inline backend saves with a tight bounding box, so a text run wider than\n    the figure silently widens the exported PNG and strands the plot in one\n    corner.\n\n    Args:\n        text: The string to wrap.\n        width_in: Figure width in inches.\n        fontsize: Point size the string will be drawn at.\n\n    Returns:\n        The string with newlines inserted.\n    \"\"\"\n    import textwrap\n    chars = max(24, int((width_in - 0.15) * 72.0 / (fontsize * 0.53)))\n    return \"\\n\".join(textwrap.wrap(text, chars)) if text else \"\"\n\n\ndef finish(fig, finding, detail=\"\", caption=\"\", headroom=0.0, footroom=0.0):\n    \"\"\"Add the finding title, optional detail line and caption, then show.\n\n    Titles state the finding, not the mechanism. Reserved space is computed in\n    inches from the wrapped line count, so a tall figure and a short one get the\n    same visual header band.\n\n    Args:\n        fig: The figure to close out.\n        finding: The headline. One sentence, stating what the chart shows.\n        detail: Optional second line in secondary ink.\n        caption: Optional footer in muted ink, for provenance and caveats.\n        headroom: Extra inches reserved above the axes.\n        footroom: Extra inches reserved below the axes.\n\n    Returns:\n        None.\n    \"\"\"\n    w, h = fig.get_size_inches()\n    finding = _wrap(finding, w, 17.5)\n    detail = _wrap(detail, w, 13.0)\n    caption = _wrap(caption, w, 11.25)\n    n_find = finding.count(\"\\n\") + 1\n    n_det = (detail.count(\"\\n\") + 1) if detail else 0\n    n_cap = (caption.count(\"\\n\") + 1) if caption else 0\n\n    head_in = 0.32 + 0.31 * n_find + (0.10 + 0.25 * n_det) + headroom\n    foot_in = ((0.18 + 0.21 * n_cap) if caption else 0.13) + footroom\n\n    # tight_layout silently declines on any figure whose gridspec carries an explicit\n    # hspace, which is every stacked pair below, and only emits a UserWarning when it\n    # does. So the header band cannot be left to it: run it for the tick-label\n    # margins, then clamp top and bottom directly, which always applies.\n    import warnings as _warnings\n    try:\n        with _warnings.catch_warnings():\n            _warnings.simplefilter(\"ignore\", UserWarning)\n            fig.tight_layout(rect=(0.0, foot_in / h, 1.0, 1.0 - head_in / h))\n    except Exception:\n        pass\n    sp = fig.subplotpars\n    top = min(sp.top, 1.0 - head_in / h)\n    bottom = max(sp.bottom, foot_in / h)\n    if top > bottom + 0.05:\n        fig.subplots_adjust(top=top, bottom=bottom)\n\n    if caption:\n        try:\n            fig.canvas.draw()\n            rend = fig.canvas.get_renderer()\n            low = min(a.get_tightbbox(rend).y0 for a in fig.axes) / fig.bbox.height\n            shift = min(foot_in / h - low, 0.25)\n            sp = fig.subplotpars\n            if shift > 0.002 and sp.top > sp.bottom + shift + 0.05:\n                fig.subplots_adjust(bottom=sp.bottom + shift)\n        except Exception:\n            pass\n\n    fig.text(0.008, 1.0 - 0.20 / h, finding, ha=\"left\", va=\"top\",\n             fontsize=17.5, fontweight=\"bold\", color=INK, linespacing=1.25)\n    if detail:\n        fig.text(0.008, 1.0 - (0.25 + 0.31 * n_find) / h, detail, ha=\"left\",\n                 va=\"top\", fontsize=13, color=INK_2, linespacing=1.3)\n    if caption:\n        fig.text(0.008, 0.09 / h, caption, ha=\"left\", va=\"bottom\",\n                 fontsize=11.25, color=MUTED, linespacing=1.3)\n    plt.show()\n\n\ndef label_bars(ax, values, ypos, fmt=\"{:.2f}\", pad=0.012, color=None):\n    \"\"\"Write a value label just past the end of each horizontal bar.\n\n    Args:\n        ax: Axes holding the bars.\n        values: Sequence of bar lengths.\n        ypos: Sequence of bar centre positions.\n        fmt: Format string applied to each value.\n        pad: Horizontal offset in data units.\n        color: Text colour, defaulting to the secondary ink tier.\n\n    Returns:\n        None.\n    \"\"\"\n    for v, y in zip(values, ypos):\n        if v is None or (isinstance(v, float) and not np.isfinite(v)):\n            continue\n        ax.text(v + pad, y, fmt.format(v), va=\"center\", ha=\"left\", fontsize=12,\n                color=INK_2 if color is None else color)\n\n\nprint(\"design system ready | text contrast on the chart surface \"\n      \"(WCAG AA floor 4.5):\",\n      \", \".join(\"{} {:.1f}\".format(n, contrast(c, SURFACE)) for n, c in\n                [(\"INK\", INK), (\"INK_2\", INK_2), (\"MUTED\", MUTED),\n                 (\"ACCENT_TEXT\", ACCENT_TEXT)]))"},{"cell_type":"code","execution_count":null,"id":"75cb46e20760","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def detect_input_root():\n    \"\"\"Find the directory holding train.csv, without hardcoding a mount path.\n\n    Returns:\n        Path to the competition input root.\n\n    Raises:\n        FileNotFoundError: If no candidate directory contains train.csv.\n    \"\"\"\n    # Kaggle does not always mount a competition at /kaggle/input/<slug>. For this one\n    # it is /kaggle/input/competitions/<slug>, one level deeper, which a shallow scan\n    # of /kaggle/input misses: it finds the `competitions` directory, sees no train.csv\n    # inside it, and gives up. Version 1 of the companion baseline notebook died exactly\n    # that way in public. So both shapes are named explicitly and the fallback walks two\n    # levels rather than one.\n    candidates = [\n        Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n        Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n        Path(\"/mnt/data/Github/Kaggle/Competitions/rsna_knee_abnormality_detection\"\n             \"/data/kaggle_mirror/rsna-knee-abnormality-detection\"),\n        Path(\"/mnt/data/Github/Kaggle/Competitions/rsna_knee_abnormality_detection/data/raw\"),\n        Path(\"data/raw\"),\n        Path(\".\"),\n    ]\n    for c in candidates:\n        try:\n            if (c / \"train.csv\").is_file():\n                return c\n        except OSError:\n            continue\n\n    kin = Path(\"/kaggle/input\")\n    seen = []\n    if kin.is_dir():\n        for depth1 in sorted(p for p in kin.iterdir() if p.is_dir()):\n            seen.append(str(depth1))\n            try:\n                if (depth1 / \"train.csv\").is_file():\n                    return depth1\n                for depth2 in sorted(p for p in depth1.iterdir() if p.is_dir()):\n                    seen.append(str(depth2))\n                    if (depth2 / \"train.csv\").is_file():\n                        return depth2\n            except OSError:\n                continue\n\n    # Name what was checked AND what is actually mounted. A rerun failure is only\n    # debuggable from the log, and \"not found\" on its own says nothing.\n    raise FileNotFoundError(\n        \"no input root containing train.csv was found. Checked: \"\n        + \", \".join(str(c) for c in candidates)\n        + \". Directories present under /kaggle/input: \"\n        + (\", \".join(seen) if seen else \"none\"))\n\n\ndef find_subdir(root, names):\n    \"\"\"Return the first existing subdirectory of root among names, else None.\n\n    Args:\n        root: Directory to look inside.\n        names: Candidate directory names, tried in order.\n\n    Returns:\n        A Path, or None when none of the names is a directory.\n    \"\"\"\n    for n in names:\n        p = root / n\n        try:\n            if p.is_dir():\n                return p\n        except OSError:\n            continue\n    return None\n\n\ndef find_dinov2_dir():\n    \"\"\"Locate the mounted DINOv2 checkpoint directory.\n\n    The Kaggle model mount for metaresearch/dinov2/PyTorch/small/1 lands under\n    /kaggle/input/dinov2/..., but the exact casing and the trailing version\n    directory have moved before. So the explicit paths are tried first and then\n    /kaggle/input is walked for any directory holding a config.json whose\n    model_type is dinov2 next to a weights file. train_series and test_series are\n    pruned from the walk; they hold hundreds of thousands of files.\n\n    Returns:\n        A Path to the checkpoint directory, or None when nothing matches.\n    \"\"\"\n    explicit = [\n        Path(\"/kaggle/input/dinov2/pytorch/small/1\"),\n        Path(\"/kaggle/input/dinov2/PyTorch/small/1\"),\n        Path(\"/kaggle/input/dinov2/transformers/small/1\"),\n    ]\n    for p in explicit:\n        try:\n            if (p / \"config.json\").is_file():\n                return p\n        except OSError:\n            continue\n\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    skip = {\"train_series\", \"test_series\", \"train_images\", \"test_images\",\n            \".git\", \"__pycache__\"}\n    best = None\n    base_depth = len(base.parts)\n    for root, dirs, files in os.walk(base):\n        depth = len(Path(root).parts) - base_depth\n        dirs[:] = [d for d in dirs if d not in skip and depth < 6]\n        if \"config.json\" not in files:\n            continue\n        rp = Path(root)\n        if not any((rp / w).is_file() for w in\n                   (\"model.safetensors\", \"pytorch_model.bin\")):\n            continue\n        try:\n            cfg = json.loads((rp / \"config.json\").read_text())\n        except Exception:\n            continue\n        if \"dinov2\" not in str(cfg.get(\"model_type\", \"\")).lower() \\\n                and \"dinov2\" not in str(rp).lower():\n            continue\n        score = abs(int(cfg.get(\"hidden_size\", 10000)) - EMBED_DIM)\n        if best is None or score < best[0]:\n            best = (score, rp)\n    return best[1] if best else None\n\n\nROOT = detect_input_root()\nON_KAGGLE = str(ROOT).startswith(\"/kaggle/\")\nTRAIN_DCM = find_subdir(ROOT, [\"train_series\", \"train_images\"])\nTEST_DCM = find_subdir(ROOT, [\"test_series\", \"test_images\"])\nWORK = Path(\"/kaggle/working\") if Path(\"/kaggle/working\").is_dir() else Path(\".\")\n\ntrain = pd.read_csv(ROOT / \"train.csv\")\ntest = pd.read_csv(ROOT / \"test.csv\")\n\n\ndef read_optional(path, columns):\n    \"\"\"Read a CSV that may be absent, returning an empty typed frame instead.\n\n    Args:\n        path: CSV path.\n        columns: Column names for the empty fallback frame.\n\n    Returns:\n        A DataFrame, empty with the given columns when the file is missing.\n    \"\"\"\n    try:\n        if path.is_file():\n            return pd.read_csv(path)\n    except Exception as exc:\n        print(\"could not read\", path.name, \":\", exc)\n    return pd.DataFrame(columns=columns)\n\n\nSERIES_COLS = [\"StudyInstanceUID\", \"SeriesInstanceUID\", \"Fluid_Sensitive\",\n               \"Fat_Suppression\", \"Anatomical_Plane\"]\ntrain_series = read_optional(ROOT / \"train_series.csv\", SERIES_COLS)\ntest_series = read_optional(ROOT / \"test_series.csv\", SERIES_COLS)\nsample_sub = read_optional(ROOT / \"sample_submission.csv\",\n                           [\"StudyInstanceUID\"] + LABELS)\n\n# ---- Torch and the backbone --------------------------------------------------\nHAVE_TORCH = False\nDEVICE = \"cpu\"\nN_GPU = 0\ntry:\n    import torch\n    import torch.nn as nn\n    import torch.nn.functional as F\n    HAVE_TORCH = True\n    N_GPU = torch.cuda.device_count()\n    DEVICE = \"cuda\" if N_GPU > 0 else \"cpu\"\n    torch.manual_seed(SEED)\n    if N_GPU > 0:\n        torch.cuda.manual_seed_all(SEED)\n        torch.backends.cudnn.benchmark = True\nexcept Exception as exc:\n    print(\"torch unavailable:\", exc)\n\nHAVE_PYDICOM = False\ntry:\n    import pydicom\n    HAVE_PYDICOM = True\nexcept Exception as exc:\n    print(\"pydicom unavailable:\", exc)\n\n# The resize backend is environment detection, not preprocessing: the resolution\n# probe in section 2 needs it before section 6 exists. cv2 is present on Kaggle and\n# absent in some local environments, so the helper degrades rather than failing.\ntry:\n    import cv2\nexcept Exception:\n    cv2 = None\ntry:\n    from PIL import Image\nexcept Exception:\n    Image = None\n\n\ndef resize_2d(arr, size):\n    \"\"\"Resize a 2D float array to size x size, with three fallbacks.\n\n    cv2 is present on Kaggle and absent in some local environments, so the\n    function degrades rather than failing: cv2, then Pillow's float mode, then a\n    nearest-neighbour gather that needs nothing but NumPy.\n\n    Args:\n        arr: 2D float32 array.\n        size: Target side length in pixels.\n\n    Returns:\n        A float32 array of shape (size, size).\n    \"\"\"\n    h, w = arr.shape\n    if (h, w) == (size, size):\n        return arr.astype(np.float32)\n    if cv2 is not None:\n        interp = cv2.INTER_AREA if max(h, w) > size else cv2.INTER_LINEAR\n        return cv2.resize(arr.astype(np.float32), (size, size),\n                          interpolation=interp).astype(np.float32)\n    if Image is not None:\n        resample = Image.BILINEAR if max(h, w) <= size else Image.BOX\n        return np.asarray(\n            Image.fromarray(arr.astype(np.float32), mode=\"F\")\n            .resize((size, size), resample), dtype=np.float32)\n    yi = np.clip((np.arange(size) * h / size).astype(int), 0, h - 1)\n    xi = np.clip((np.arange(size) * w / size).astype(int), 0, w - 1)\n    return arr[np.ix_(yi, xi)].astype(np.float32)\n\n\nDINO_DIR = find_dinov2_dir()\nCAN_ENCODE = bool(HAVE_TORCH and HAVE_PYDICOM and DINO_DIR is not None\n                  and TRAIN_DCM is not None)\n\nprint(\"input root        :\", ROOT)\nprint(\"running on        :\", \"Kaggle\" if ON_KAGGLE else \"local mirror\")\nprint(\"train dicom dir   :\", TRAIN_DCM)\nprint(\"test dicom dir    :\", TEST_DCM)\nprint(\"writing to        :\", WORK)\nprint(\"torch             :\", \"yes\" if HAVE_TORCH else \"NO\",\n      \"| gpus:\", N_GPU, \"| device:\", DEVICE)\nprint(\"pydicom           :\", \"yes\" if HAVE_PYDICOM else \"NO\")\nprint(\"resize backend    :\", \"cv2\" if cv2 is not None else\n      (\"PIL\" if Image is not None else \"numpy\"))\nprint(\"dinov2 checkpoint :\", DINO_DIR if DINO_DIR else \"NOT FOUND\")\nprint(\"encoding possible :\", CAN_ENCODE)\nprint()\nprint(\"train studies     :\", len(train), \"| train series :\", len(train_series))\nprint(\"test studies      :\", len(test), \"| test series  :\", len(test_series))\n\ngold_mask = (train[LABELS].notna().all(axis=1)\n             if set(LABELS).issubset(train.columns)\n             else pd.Series(False, index=train.index))\nN_GOLD = int(gold_mask.sum())\ngold_ids = train.loc[gold_mask, \"StudyInstanceUID\"].astype(str).tolist()\nprint(\"gold-labelled     :\", N_GOLD, \"of\", len(train),\n      \"({:.2f}%)\".format(100.0 * N_GOLD / max(1, len(train))))\nif not CAN_ENCODE:\n    print()\n    print(\"NOTE: the imaging path cannot run here. Every later cell degrades to the\")\n    print(\"      series-manifest fallback and a valid submission is still written.\")"},{"cell_type":"code","execution_count":null,"id":"bcbc31807d56","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"DINO = None\nDINO_KW = {}\nDINO_INFO = \"not loaded\"\n\nif CAN_ENCODE:\n    # This try covers the load and nothing else. An earlier version wrapped the\n    # diagnostic prints below in the same block, and a NameError in one of those\n    # prints set CAN_ENCODE=False and sent the whole notebook down the\n    # metadata-only fallback while reporting \"could not load DINOv2\". A broad\n    # except around code that is not the thing being guarded turns a typo into a\n    # silent capability loss.\n    try:\n        import inspect\n        from transformers import AutoModel\n        DINO = AutoModel.from_pretrained(str(DINO_DIR), local_files_only=True)\n        DINO.eval()\n        for p in DINO.parameters():\n            p.requires_grad_(False)\n        DINO.to(DEVICE)\n        # HF's Dinov2 embeddings interpolate the positional grid unconditionally,\n        # but some releases also accept the flag explicitly. Decide from the\n        # signature rather than from a try/except around the forward pass, which\n        # would swallow a real error inside the model and report it as a\n        # signature mismatch.\n        if \"interpolate_pos_encoding\" in inspect.signature(DINO.forward).parameters:\n            DINO_KW[\"interpolate_pos_encoding\"] = True\n    except Exception as exc:\n        print(\"could not load DINOv2:\", type(exc).__name__, exc)\n        DINO = None\n        CAN_ENCODE = False\n\nif DINO is not None:\n    try:\n        n_par = sum(p.numel() for p in DINO.parameters())\n        hidden = int(getattr(DINO.config, \"hidden_size\", EMBED_DIM))\n        patch = int(getattr(DINO.config, \"patch_size\", PATCH_PX))\n        native = int(getattr(DINO.config, \"image_size\", 0))\n        DINO_INFO = \"{} params, hidden {}, patch {}, configured at {} px\".format(\n            \"{:,}\".format(n_par), hidden, patch, native)\n        print(\"DINOv2 loaded from\", DINO_DIR)\n        print(\"  \", DINO_INFO)\n        print(\"   forward kwargs:\", DINO_KW if DINO_KW else \"none needed\")\n        if patch != PATCH_PX or hidden != EMBED_DIM:\n            print(\"   WARNING: checkpoint geometry differs from the constants; the \"\n                  \"millimetre arithmetic above assumed patch {} hidden {}\".format(\n                      PATCH_PX, EMBED_DIM))\n        if native:\n            print(\"   positional grid: native {0}x{0} -> {1}x{1} here, {2}x{2} at \"\n                  \"the frontier's 126 px\".format(\n                      native // patch, IMG_SIZE // patch, 126 // patch))\n    except Exception as exc:\n        # Diagnostics only. Whatever fails here, the backbone is loaded and usable.\n        print(\"checkpoint diagnostics unavailable:\", type(exc).__name__, exc)\n\nIMAGENET_MEAN = np.asarray([0.485, 0.456, 0.406], np.float32).reshape(1, 3, 1, 1)\nIMAGENET_STD = np.asarray([0.229, 0.224, 0.225], np.float32).reshape(1, 3, 1, 1)\n\n\ndef embed_images(imgs):\n    \"\"\"Embed a batch of 2.5D images with the frozen backbone.\n\n    Only cuda:0 is used even when two GPUs are visible. The GPU phase is a small\n    share of this pipeline's wall clock and wrapping the model in DataParallel\n    buys minutes at the cost of a code path that cannot be tested here.\n\n    Args:\n        imgs: uint8 array of shape (B, 3, size, size).\n\n    Returns:\n        float32 array of shape (B, SLICE_DIM), the CLS token concatenated with\n        the mean and the maximum over patch tokens.\n    \"\"\"\n    if DINO is None or imgs.shape[0] == 0:\n        return np.zeros((imgs.shape[0], SLICE_DIM), np.float32)\n    out = np.zeros((imgs.shape[0], SLICE_DIM), np.float32)\n    with torch.no_grad():\n        for start in range(0, imgs.shape[0], GPU_BATCH):\n            chunk = imgs[start:start + GPU_BATCH]\n            x = torch.from_numpy(chunk).to(DEVICE).float().div_(255.0)\n            x = (x - torch.from_numpy(IMAGENET_MEAN).to(DEVICE)) \\\n                / torch.from_numpy(IMAGENET_STD).to(DEVICE)\n            with torch.autocast(device_type=\"cuda\", dtype=torch.float16,\n                                enabled=(DEVICE == \"cuda\")):\n                h = DINO(pixel_values=x, **DINO_KW).last_hidden_state\n            h = h.float()\n            vec = torch.cat([h[:, 0], h[:, 1:].mean(1), h[:, 1:].amax(1)], dim=1)\n            out[start:start + chunk.shape[0]] = vec.cpu().numpy()\n    return out"},{"cell_type":"markdown","id":"4794d1c725c1","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">2</span><span style=\"vertical-align:middle;\">The resolution argument, in millimetres</span></h2></div>\n\nDINOv2 ViT-S/14 cuts its input into 14-pixel squares and emits one token per square. The\nphysical size of a token is fixed by the field of view and the input size alone:\n\n```\nmm per pixel        = 150 mm / input_px\nmm per patch token  = 150 mm * 14 / input_px\n```\n\nThe sampling bound is the part that decides this. A feature of width *d* millimetres\nneeds a pixel pitch of at most *d/2* to survive resampling. A 1 mm tear therefore needs\n0.5 mm per pixel, which needs an input of at least **300 pixels**. At 126 px the pipeline\nis a factor of 2.4 short of that bound, so no amount of head capacity downstream can\nrecover the signal: it was destroyed by the resize.\n\n336 is the smallest multiple of 14 that clears the bound with margin. It gives 0.446 mm\nper pixel, a 24 by 24 patch grid, and 6.25 mm per token, so the meniscus now spans about\ntwo tokens instead of sitting inside two thirds of one.\n\nThe cost is not the obstacle people assume. The cell below computes the forward FLOPs of\nViT-S/14 exactly, from the token count, and the projection cell in section 7 measures the\nreal throughput on this machine before any bulk work starts. The binding constraint in\nthis pipeline is DICOM decode, not the GPU."},{"cell_type":"code","execution_count":null,"id":"906644a5d241","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def vit_gflops(n_tokens, dim=EMBED_DIM, blocks=N_BLOCKS, mlp_ratio=4):\n    \"\"\"Forward FLOPs of one ViT image, counting multiplies and adds.\n\n    Per block: 6*n*d^2 for the qkv projection, 4*n^2*d for the two attention\n    matmuls, 2*n*d^2 for the output projection and 2*mlp_ratio*2*n*d^2 for the\n    MLP. Patch embedding and the head are ignored; both are under 1%.\n\n    Args:\n        n_tokens: Sequence length including the CLS token.\n        dim: Hidden width.\n        blocks: Number of transformer blocks.\n        mlp_ratio: MLP expansion factor.\n\n    Returns:\n        Forward GFLOPs for one image.\n    \"\"\"\n    per_block = (6 + 2 + 4 * mlp_ratio) * n_tokens * dim * dim \\\n        + 4 * n_tokens * n_tokens * dim\n    return blocks * per_block / 1e9\n\n\nLESIONS = [(\"meniscal tear\", 1.0, 3.0), (\"meniscus width\", 9.0, 12.0),\n           (\"cartilage thickness\", 2.0, 4.0)]\n\nrows = []\nfor size in [126, 168, 224, 280, 336, 392, 448, 518]:\n    grid = size // PATCH_PX\n    n_tok = grid * grid + 1\n    mm_px = FOV_MM / size\n    rows.append({\n        \"input_px\": size,\n        \"grid\": \"{}x{}\".format(grid, grid),\n        \"patch_tokens\": grid * grid,\n        \"mm_per_px\": mm_px,\n        \"mm_per_token\": FOV_MM * PATCH_PX / size,\n        \"gflops\": vit_gflops(n_tok),\n        \"resolves_1mm\": \"yes\" if mm_px <= 0.5 else \"no\",\n        \"resolves_3mm\": \"yes\" if mm_px <= 1.5 else \"no\",\n    })\ngeom = pd.DataFrame(rows)\n\nprint(\"DINOv2 ViT-S/14 over a {:.0f} mm field of view\".format(FOV_MM))\nprint()\nprint(geom.to_string(index=False, float_format=lambda v: \"{:.3f}\".format(v)))\nprint()\nprint(\"Nyquist: a feature d mm wide needs a pitch of at most d/2 mm per pixel.\")\nfor name, lo, hi in LESIONS:\n    need = lo / 2.0\n    min_px = FOV_MM / need\n    print(\"  {:20s} {:>4.1f}-{:<4.1f} mm  needs <= {:.2f} mm/px  ->  input >= {:.0f} px\"\n          .format(name, lo, hi, need, math.ceil(min_px)))\nprint()\nfrontier = geom.loc[geom[\"input_px\"] == 126].iloc[0]\nours = geom.loc[geom[\"input_px\"] == IMG_SIZE].iloc[0]\nprint(\"frontier at {:.0f} px : {:.3f} mm/px, {:.2f} mm per token, {} tokens, {:.1f} GFLOPs\"\n      .format(frontier[\"input_px\"], frontier[\"mm_per_px\"], frontier[\"mm_per_token\"],\n              int(frontier[\"patch_tokens\"]), frontier[\"gflops\"]))\nprint(\"here at     {:.0f} px : {:.3f} mm/px, {:.2f} mm per token, {} tokens, {:.1f} GFLOPs\"\n      .format(ours[\"input_px\"], ours[\"mm_per_px\"], ours[\"mm_per_token\"],\n              int(ours[\"patch_tokens\"]), ours[\"gflops\"]))\nprint(\"ratio            : {:.1f}x the tokens, {:.1f}x the FLOPs, {:.2f}x the millimetres \"\n      \"per token\".format(ours[\"patch_tokens\"] / frontier[\"patch_tokens\"],\n                         ours[\"gflops\"] / frontier[\"gflops\"],\n                         ours[\"mm_per_token\"] / frontier[\"mm_per_token\"]))\nprint()\nprint(\"a 1 mm tear is {:.2f} px wide at 126 px and {:.2f} px wide at {} px\".format(\n    1.0 / frontier[\"mm_per_px\"], 1.0 / ours[\"mm_per_px\"], IMG_SIZE))"},{"cell_type":"code","execution_count":null,"id":"ce150ecb8005","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# Two panels, stacked, sharing one x axis so a reader can drop a vertical line at\n# any input size and read both consequences off it. The upper panel is the physical\n# argument and the lower one is what it costs.\nsizes = np.arange(112, 532, 14)\nmm_token = FOV_MM * PATCH_PX / sizes\nmm_px = FOV_MM / sizes\ntokens = (sizes // PATCH_PX) ** 2\nflops = np.array([vit_gflops(t + 1) for t in tokens])\n\nfig = plt.figure(figsize=(13.6, 11.4))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.25, 1.0], hspace=0.26)\n\nax0 = fig.add_subplot(gs[0, 0])\nax0.axhspan(9.0, 12.0, color=PRIMARY, alpha=0.13, zorder=0)\nax0.axhspan(1.0, 3.0, color=ACCENT, alpha=0.16, zorder=0)\nax0.plot(sizes, mm_token, color=INK, zorder=3)\nax0.text(520, 10.5, \"meniscus width, 9-12 mm\", ha=\"right\", va=\"center\",\n         fontsize=11.5, color=PRIMARY_DEEP)\nax0.text(520, 2.0, \"meniscal tear, 1-3 mm\", ha=\"right\", va=\"center\",\n         fontsize=11.5, color=ACCENT_TEXT)\nfor size, colour, tag in [(126, ACCENT, \"frontier\"), (IMG_SIZE, PRIMARY_DEEP, \"here\")]:\n    val = FOV_MM * PATCH_PX / size\n    ax0.plot([size], [val], marker=\"o\", color=colour, zorder=5)\n    ax0.annotate(\"{} px\\n{:.1f} mm/token\".format(size, val), (size, val),\n                 textcoords=\"offset points\", xytext=(10, 14), fontsize=12,\n                 color=colour, fontweight=\"bold\")\nax0.set_ylabel(\"millimetres covered by one 14 px patch token\")\nax0.set_ylim(0, 19)\nstyle(ax0, \"y\")\n\nax1 = fig.add_subplot(gs[1, 0], sharex=ax0)\nax1.plot(sizes, flops, color=PRIMARY, zorder=3)\nfor size, colour in [(126, ACCENT), (IMG_SIZE, PRIMARY_DEEP)]:\n    val = vit_gflops((size // PATCH_PX) ** 2 + 1)\n    ax1.plot([size], [val], marker=\"o\", color=colour, zorder=5)\n    ax1.annotate(\"{:.1f} GFLOPs\".format(val), (size, val),\n                 textcoords=\"offset points\", xytext=(10, -4), fontsize=12,\n                 color=colour, fontweight=\"bold\")\nax1.set_ylabel(\"forward GFLOPs per image, ViT-S/14\")\nax1.set_xlabel(\"input size, pixels\")\nax1.set_xlim(112, 532)\nstyle(ax1, \"y\")\n\nfinish(fig,\n       \"At the frontier's 126 px one patch token is wider than the whole meniscus\",\n       detail=\"Upper: millimetres per patch token against input size, over the corpus's \"\n              \"150 mm field of view. Lower: what the extra tokens cost.\",\n       caption=\"Field of view verified at 150.0 mm across four acquisition matrices \"\n               \"(960/896/800/640 px at 0.1562/0.1674/0.1875/0.2344 mm). FLOPs are \"\n               \"counted from the token sequence length, multiplies and adds, patch \"\n               \"embedding and head excluded.\")"},{"cell_type":"markdown","id":"edf794746153","metadata":{},"source":"### The arithmetic is a prediction. This is the measurement.\n\nSampling theory says the resize destroys a 1 mm feature at 126 px and preserves it at\n336. That is a prediction about a specific network, and it is cheap to test against the\nfrozen backbone itself: paint a bright disc of known physical width onto a textured field\nat the corpus's native 0.1875 mm per pixel, resize the whole field to each candidate\ninput size exactly as the pipeline does, and measure how far the pooled embedding moves.\n\nIf the resolution argument is right the response should separate by lesion size. Large\nfeatures should be equally visible at every input size, and small ones should appear only\nonce the pitch clears their scale. The cell below runs it and prints its own numbers;\nnone are quoted here, because the exact values depend on the resize backend and on the\nrandom fields, and a table copied from a previous run would drift away from the one the\nreader is looking at.\n\nTwo things about the result are stable, and both were reproduced across independent runs\nwith different textures, different seed counts and different resize backends. **A 1 to\n3 mm feature moves the pooled embedding roughly two to five times further at 336 px than\nat 126 px, and a 10 mm feature gains nothing at all** and in some runs loses slightly.\nThat is the signature the sampling argument predicts, and nothing else predicts it: a\nuniform gain across all lesion sizes would have meant the higher input was simply feeding\nthe network more pixels rather than resolving anything.\n\nThe second stable result decides where to stop. **448 px does not beat 336 px on any\nlesion width tested**, in either run. The compute for 448 is available, and the\nmeasurement says not to spend it. The ladder stops at 336."},{"cell_type":"code","execution_count":null,"id":"e0e8c019506c","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"LESION_MM = [1.0, 2.0, 3.0, 10.0]\nPROBE_SIZES = [126, 224, 336, 448]\nPROBE_SEEDS = 5\nNATIVE_PX = 800                 # one of the four acquisition matrices in this corpus\nNATIVE_MM_PER_PX = FOV_MM / NATIVE_PX\n\n\ndef _knee_field(seed):\n    \"\"\"Build a smooth textured field at native resolution, standing in for anatomy.\n\n    Args:\n        seed: Random seed.\n\n    Returns:\n        float32 array of shape (NATIVE_PX, NATIVE_PX) valued in 0-255.\n    \"\"\"\n    rng = np.random.default_rng(seed)\n    lo = rng.random((14, 14)).astype(np.float32)\n    big = resize_2d(lo, NATIVE_PX)\n    big = big * 0.55 + rng.random((NATIVE_PX, NATIVE_PX)).astype(np.float32) * 0.12 + 0.25\n    return np.clip(big * 255.0, 0.0, 255.0).astype(np.float32)\n\n\ndef _paint_lesion(field, mm, seed):\n    \"\"\"Paint a bright disc of a known physical width onto a copy of the field.\n\n    Args:\n        field: Native-resolution field.\n        mm: Lesion diameter in millimetres.\n        seed: Random seed for the lesion position.\n\n    Returns:\n        Tuple of (field carrying the lesion, native pixels it covers).\n    \"\"\"\n    rng = np.random.default_rng(seed + 991)\n    r = max(1, int(round(mm / NATIVE_MM_PER_PX / 2.0)))\n    cy, cx = rng.integers(NATIVE_PX // 4, 3 * NATIVE_PX // 4, size=2)\n    yy, xx = np.ogrid[:NATIVE_PX, :NATIVE_PX]\n    m = (yy - cy) ** 2 + (xx - cx) ** 2 <= r * r\n    out = field.copy()\n    out[m] = np.clip(out[m] + 110.0, 0.0, 255.0)\n    return out, int(m.sum())\n\n\ndef _embed_field(field, size):\n    \"\"\"Resize a native field to one input size and return its pooled embedding.\n\n    Args:\n        field: Native-resolution float field valued in 0-255.\n        size: Input side length in pixels.\n\n    Returns:\n        float32 vector of length SLICE_DIM.\n    \"\"\"\n    small = np.clip(resize_2d(field, size), 0.0, 255.0).astype(np.uint8)\n    return embed_images(np.repeat(small[None, None], 3, axis=1))[0]\n\n\nlesion_tbl = None\nif DINO is not None:\n    t0 = time.time()\n    rows = []\n    for mm in LESION_MM:\n        rec = {\"lesion_mm\": mm}\n        npx = 0\n        for size in PROBE_SIZES:\n            shifts = []\n            for seed in range(PROBE_SEEDS):\n                base = _knee_field(seed)\n                spot, npx = _paint_lesion(base, mm, seed)\n                v0 = _embed_field(base, size)\n                v1 = _embed_field(spot, size)\n                shifts.append(float(np.linalg.norm(v1 - v0))\n                              / max(float(np.linalg.norm(v0)), 1e-9))\n            rec[\"px{}\".format(size)] = float(np.mean(shifts))\n        rec[\"native_px\"] = npx\n        rows.append(rec)\n    lesion_tbl = pd.DataFrame(rows)\n    ref = \"px{}\".format(min(PROBE_SIZES))\n    cur = \"px{}\".format(IMG_SIZE if IMG_SIZE in PROBE_SIZES else 336)\n    lesion_tbl[\"gain\"] = lesion_tbl[cur] / lesion_tbl[ref].clip(lower=1e-9)\n    print(\"lesion-response probe: {} widths x {} sizes x {} seeds in {:.1f}s\".format(\n        len(LESION_MM), len(PROBE_SIZES), PROBE_SEEDS, time.time() - t0))\n    print()\n    print(\"relative shift of the pooled embedding when a lesion of that physical width\")\n    print(\"is painted on the field, per input size ('gain' is {} px over 126 px)\".format(\n        cur[2:]))\n    print(lesion_tbl.to_string(index=False, float_format=lambda v: \"{:.4f}\".format(v)))\n    print()\n    big = lesion_tbl.loc[lesion_tbl[\"lesion_mm\"].idxmax()]\n    small = lesion_tbl.loc[lesion_tbl[\"lesion_mm\"].idxmin()]\n    print(\"a {:.0f} mm feature gains {:.2f}x from the higher input size; \"\n          \"a {:.0f} mm feature gains {:.2f}x\".format(\n              big[\"lesion_mm\"], big[\"gain\"], small[\"lesion_mm\"], small[\"gain\"]))\n    print(\"Resolution buys nothing for large structures and a great deal for small ones,\")\n    print(\"which is the signature the sampling argument predicts.\")\nelse:\n    print(\"lesion-response probe skipped: the backbone is not loaded here\")"},{"cell_type":"code","execution_count":null,"id":"65cb8280bd4d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"if lesion_tbl is not None and len(lesion_tbl):\n    # Two panels, stacked. Upper: the response curve per lesion width, which is the\n    # measurement. Lower: the gain the shipped input size buys over the frontier's,\n    # which is the decision that measurement supports.\n    fig = plt.figure(figsize=(13.2, 11.0))\n    gs = fig.add_gridspec(2, 1, height_ratios=[1.2, 0.85], hspace=0.34)\n\n    ax0 = fig.add_subplot(gs[0, 0])\n    ordered = lesion_tbl.sort_values(\"lesion_mm\").reset_index(drop=True)\n    for i, r in ordered.iterrows():\n        ys = [r[\"px{}\".format(s)] for s in PROBE_SIZES]\n        colour = ACCENT if r[\"lesion_mm\"] <= 1.0 else (\n            PRIMARY_DEEP if r[\"lesion_mm\"] <= 3.0 else NEUTRAL)\n        ax0.plot(PROBE_SIZES, ys, marker=\"o\", color=colour, zorder=3)\n        ax0.annotate(\"  {:.0f} mm\".format(r[\"lesion_mm\"]),\n                     (PROBE_SIZES[-1], ys[-1]), fontsize=12, va=\"center\",\n                     fontweight=\"bold\",\n                     color=ACCENT_TEXT if r[\"lesion_mm\"] <= 1.0 else (\n                         PRIMARY_DEEP if r[\"lesion_mm\"] <= 3.0 else MUTED))\n    ax0.axvline(IMG_SIZE, color=AXIS, linewidth=1.4, zorder=1)\n    ax0.set_xticks(PROBE_SIZES)\n    ax0.set_xlim(PROBE_SIZES[0] - 20, PROBE_SIZES[-1] + 80)\n    ax0.set_ylim(0, None)\n    ax0.set_ylabel(\"relative shift of the pooled embedding\")\n    ax0.set_xlabel(\"input size, pixels\")\n    style(ax0, \"y\")\n\n    ax1 = fig.add_subplot(gs[1, 0])\n    ypos = np.arange(len(ordered))\n    gains = ordered[\"gain\"].values.astype(float)\n    ax1.barh(ypos, gains, height=0.62,\n             color=[ACCENT if g >= 1.5 else NEUTRAL for g in gains], zorder=3)\n    label_bars(ax1, gains, ypos, fmt=\"{:.2f}x\", pad=max(0.03, gains.max() * 0.02))\n    ax1.axvline(1.0, color=AXIS, linewidth=1.4, zorder=2)\n    ax1.set_yticks(ypos)\n    ax1.set_yticklabels([\"{:.0f} mm\".format(v) for v in ordered[\"lesion_mm\"]],\n                        fontsize=12.5, color=INK)\n    ax1.set_xlim(0, max(1.25, float(gains.max()) * 1.22))\n    ax1.set_xlabel(\"response at {} px divided by response at 126 px\".format(IMG_SIZE))\n    style(ax1, \"x\")\n    ax1.spines[\"left\"].set_visible(False)\n\n    finish(fig,\n           \"Resolution buys nothing for a meniscus and everything for a tear\",\n           detail=\"Relative shift of the frozen backbone's pooled embedding when a \"\n                  \"lesion of known physical width is painted on the field and the \"\n                  \"whole field is then resized to each input size.\",\n           caption=\"{} seeds per point, lesion placed at a random position on a \"\n                   \"textured field at the corpus's native 0.1875 mm per pixel. The \"\n                   \"vertical rule marks the input size this notebook ships.\"\n                   .format(PROBE_SEEDS))"},{"cell_type":"markdown","id":"919a01b2088e","metadata":{},"source":"### Why the pooled embedding keeps a maximum, and what that is not\n\nEach encoded slice contributes three vectors, not one: the CLS token, the mean over patch\ntokens, and the maximum over patch tokens, 1,152 dimensions in total.\n\nThe tempting argument for the maximum is the convolutional one. A lesion lives in 1 of\n576 tokens, so mean pooling dilutes it 576-fold and max pooling does not. **That argument\nis wrong for a transformer, and measuring it says so.** Perturbing exactly one patch and\ncomparing the two blocks gives a max-to-mean sensitivity ratio of 1.01x at 126 px and\n1.05x at 336 px. Self-attention is global: one altered patch shifts every token in the\nsequence, so the patch mean is not the dilutive bottleneck a convolutional intuition\npredicts.\n\nThe maximum is kept anyway, for a smaller and more honest reason. It costs one `amax`\nover an activation already in memory, it is not a linear function of the mean so it is\nnot redundant with it, and the heads downstream weight the three blocks separately. It is\na cheap third view, not the mechanism. The mechanism is the resize, and that is what the\ncell above measures."},{"cell_type":"markdown","id":"32a282efb535","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">3</span><span style=\"vertical-align:middle;\">Labels: one switchable source</span></h2></div>\n\nOnly 58 of the 4,407 training studies carry per-condition labels. The other 4,349 ship a\nfree-text radiology report, and the labels have to be derived from it. That extraction is\na separate piece of work with its own audit; this notebook consumes its output.\n\n`LABEL_SOURCE` is the single constant that selects it. Left at `\"auto\"` it prefers, in\norder: LLM-extracted soft labels from an attached dataset, a rule-label file from an\nattached dataset, then the multilingual rule extractor run inline. The inline extractor\nis the same code as the companion baseline notebook, embedded from the same builder\nmodule so the two cannot drift apart, and it always works on Kaggle because `train.csv`\ncarries the report text.\n\nThe measured gap between the two sources, on the 58 gold studies, is the reason the\nswitch exists. Rule extractor: macro agreement 0.7960, precision 0.6802, recall 0.7648.\nClaude Opus 5 at low effort: 0.8233, 0.6910, 0.8602, a paired improvement of +0.0273\nagreement with a 95% interval of [+0.0029, +0.0532] and McNemar p = 0.0204. The gap is\nalmost entirely non-English recall: the rule extractor recalls 0.8611 on English reports\nand 0.6742 on the rest, while the LLM holds 0.8796 and 0.8333. Swapping the source is a\none-line change here and one line in `kernel-metadata.json`.\n\nThe head trains on **soft** targets. A rule label is 0 or 1; an LLM label is a\nprobability. Both are fed as floats, so the switch changes the values and not the\nplumbing."},{"cell_type":"code","execution_count":null,"id":"eee15ce89f82","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# The rule extractor below is not written here. It is embedded verbatim from the\n# builder module of the companion metadata-baseline notebook, so the two notebooks\n# consume one extractor and cannot drift apart. Its audit against the gold studies\n# is printed two cells down.\n# --------------------------------------------------------------------------- #\n# Multilingual rule-based label extractor.\n# Weak on purpose. This is the artifact an LLM extraction has to beat.\n# --------------------------------------------------------------------------- #\n\n_CHAR_MAP = str.maketrans({\"ı\": \"i\", \"İ\": \"i\", \"ß\": \"ss\",\n                           \"ø\": \"o\", \"Ø\": \"o\",\n                           \"đ\": \"d\", \"Đ\": \"d\"})\n\nTEAR = [\n    \"tear\", \"torn\", \"tearing\", \"rupture\", \"ruptured\", \"disruption\", \"discontinuity\",\n    \"rotura\", \"ruptura\", \"desgarro\", \"roto\", \"rota\",\n    \"dechirure\", \"dechire\", \"lesion meniscale\",\n    \"scheur\", \"ruptuur\", \"gescheurd\",\n    \"riss\", \"rissbildung\", \"ruptur\", \"zerreissung\", \"einriss\",\n    \"yirtik\", \"yirtig\", \"kopma\", \"butunluk kaybi\",\n    \"ρηξη\", \"ρηξις\", \"ρηγμα\",\n    \"руптура\", \"разкъсв\",\n    \"разрив\", \"puknuce\", \"prekid\",\n]\nSPRAIN = [\n    \"sprain\", \"esguince\", \"entorse\", \"verstauchung\", \"verstuiking\",\n    \"distorsiyon\", \"burkulma\", \"διαστρεμμα\",\n    \"навяхван\",\n    \"injury\", \"lesion\", \"letsel\", \"verletzung\", \"zedelenme\",\n]\nOA_FIND = [\n    \"osteoarthritis\", \"osteoarthrosis\", \"arthrosis\", \"arthritic\",\n    \"artrosis\", \"artrose\", \"arthrose\", \"gonartrose\", \"gonartroz\", \"gonarthrose\",\n    \"gonartro\", \"artroz\",\n    \"chondrosis\", \"chondral\", \"chondropathy\", \"chondropathie\", \"chondropatie\",\n    \"chondromalacia\", \"kondromalazi\", \"condropatia\", \"condral\", \"chondropatia\",\n    \"kraakbeenlijden\", \"kraakbeenverlies\", \"kraakbeenschade\",\n    \"knorpelschaden\", \"knorpeldefekt\", \"knorpelverlust\", \"knorpelbelag\",\n    \"cartilage loss\", \"cartilage thinning\", \"cartilage fissuring\",\n    \"cartilage defect\", \"cartilage fissure\", \"chondral defect\", \"chondral loss\",\n    \"osteophyt\", \"osteofit\", \"osteofyt\", \"osteophyte\", \"spurring\", \"osteofyte\",\n    \"joint space narrowing\", \"gelenkspaltverschmalerung\",\n    \"kikirdak kaybi\", \"kikirdak incelme\", \"kondral\",\n    \"ulcera condral\", \"hrskavice\", \"denudacija\",\n    \"χονδροπαθ\", \"οστεοαρθρ\",\n    \"οστεοφυτ\", \"αρθριτ\",\n    \"артроз\", \"хондропат\",\n    \"остеофит\", \"хрущялн\",\n]\nMEDIAL_Q = [\"medial\", \"interno\", \"interna\", \"interne\", \"mediaal\", \"mediale\",\n            \"innen\", \"medyal\", \"medijaln\", \"εσω\",\n            \"медиал\", \"вътреш\"]\nLATERAL_Q = [\"lateral\", \"externo\", \"externa\", \"externe\", \"buiten\", \"aussen\",\n             \"lateraln\", \"εξω\", \"латерал\",\n             \"външ\"]\nPF_Q = [\"patellofemoral\", \"patelofemoral\", \"femoropatellar\", \"femoropatelar\",\n        \"femoropatellair\", \"retropatellar\", \"retrorotulian\", \"patellar facet\",\n        \"trochlea\", \"troclea\", \"trochlee\", \"rotula\", \"rotulian\", \"patella\",\n        \"patellaire\", \"patellar\", \"diz kapagi\", \"patelofemoraln\",\n        \"επιγονατιδ\", \"τροχιλ\",\n        \"пател\", \"ретропател\"]\nTRICOMP = [\"tricompartmental\", \"tricompartimental\", \"three compartments\",\n           \"three compartmens\", \"all compartments\", \"gonartrose\", \"gonartroz\",\n           \"gonarthrose\", \"gonartrosis\", \"gonartro\", \"pangonartro\"]\n\nSPECIFIC = {\n    \"ACL\": [\n        \"acl\", \"anterior cruciate\", \"lca\", \"ligamento cruzado anterior\",\n        \"ligament croise anterieur\", \"croise anterieur\",\n        \"voorste kruisband\", \"vkb\", \"vorderes kreuzband\", \"vorderen kreuzband\",\n        \"vordere kreuzband\", \"on capraz bag\", \"anterior capraz bag\",\n        \"προσθιο χιαστ\",\n        \"προσθιου χιαστ\",\n        \"προσθιος χιαστ\",\n        \"предна кръстна\",\n        \"предната кръстна\",\n        \"prednji ukrizeni\",\n    ],\n    \"MCL\": [\n        \"mcl\", \"medial collateral\", \"lcm\", \"ligamento colateral medial\",\n        \"ligamento colateral interno\", \"ligament collateral medial\",\n        \"collateral medial\", \"mediale collaterale\", \"mediaal collateraal\",\n        \"innenband\", \"mediales kollateralband\", \"medialen kollateralband\",\n        \"mediale kollateralband\", \"medyal kollateral\", \"ic yan bag\",\n        \"εσω πλαγιο\",\n        \"медиален колатерален\",\n        \"вътрешна колатерална\",\n    ],\n    \"Medial Meniscus\": [\n        \"medial meniscus\", \"meniscus medialis\", \"menisco medial\", \"menisco interno\",\n        \"mediale meniscus\", \"binnenmeniscus\", \"innenmeniskus\", \"meniskus medialis\",\n        \"menisque interne\", \"menisque medial\", \"medial menisk\", \"medyal menisk\",\n        \"εσω μηνισκ\",\n        \"медиалния менискус\",\n        \"медиален мениск\",\n        \"вътрешния мениск\",\n        \"medial and lateral menisc\", \"medial ve lateral menisk\",\n        \"menisco medial y lateral\", \"menisco interno y externo\",\n        \"menisco interno y lateral\", \"mediale en laterale meniscus\",\n        \"innen und aussenmeniskus\", \"medijalnog meniskusa\",\n    ],\n    \"Lateral Meniscus\": [\n        \"lateral meniscus\", \"meniscus lateralis\", \"menisco lateral\", \"menisco externo\",\n        \"laterale meniscus\", \"buitenmeniscus\", \"aussenmeniskus\", \"meniskus lateralis\",\n        \"menisque externe\", \"menisque lateral\", \"lateral menisk\",\n        \"εξω μηνισκ\",\n        \"латералния менискус\",\n        \"латерален мениск\",\n        \"външния мениск\",\n        \"medial and lateral menisc\", \"medial ve lateral menisk\",\n        \"menisco medial y lateral\", \"menisco interno y externo\",\n        \"menisco interno y lateral\", \"mediale en laterale meniscus\",\n        \"innen und aussenmeniskus\", \"lateralnog meniskusa\",\n    ],\n}\n\nGENERIC = {\n    \"ACL\": [\"cruciate\", \"cruzado\", \"croise\", \"kruisband\", \"kreuzband\",\n            \"χιαστ\", \"кръстн\",\n            \"capraz bag\", \"ukrizen\"],\n    \"MCL\": [\"collateral\", \"colateral\", \"kollateral\", \"collaterale\",\n            \"πλαγι\", \"колатерал\",\n            \"yan bag\"],\n    \"Medial Meniscus\": [\"menisc\", \"menisk\", \"μηνισκ\",\n                        \"мениск\"],\n    \"Lateral Meniscus\": [\"menisc\", \"menisk\", \"μηνισκ\",\n                         \"мениск\"],\n}\nSIDE_Q = {\n    \"ACL\": [\"anterior\", \"anterieur\", \"voorste\", \"vorder\",\n            \"προσθι\", \"предн\", \"prednj\"],\n    \"MCL\": MEDIAL_Q,\n    \"Medial Meniscus\": MEDIAL_Q,\n    \"Lateral Meniscus\": LATERAL_Q,\n}\n\nDIRECT = {\n    \"Effusion\": [\n        \"effusion\", \"joint fluid\", \"hemarthrosis\", \"haemarthrosis\", \"hydrops\",\n        \"derrame\", \"epanchement\", \"gewrichtsvocht\", \"vocht in het gewricht\",\n        \"erguss\", \"gelenkerguss\", \"gelenkserguss\",\n        \"eklem ici sivi\", \"sivi artisi\", \"efuzyon\", \"eklem mesafesinde sivi\",\n        \"eklem ici serbest sivi\", \"eklem sivisi\",\n        \"ενδαρθρικ\",\n        \"αρθρικο υγρο\",\n        \"συλλογη υγρου\",\n        \"ставен излив\",\n        \"излив\", \"хидропс\",\n        \"zglobni izljev\", \"izljev\",\n    ],\n    \"Synovitis\": [\n        \"synovitis\", \"synovial thickening\", \"synovial hypertrophy\",\n        \"thickened synovial\", \"hypertrophy of the synovium\",\n        \"synovial proliferation\", \"proliferation of the synovium\",\n        \"sinovitis\", \"synovite\", \"synovitiden\", \"synovialitis\",\n        \"verdikking van het synovium\", \"verdikkingen van het synovium\",\n        \"synoviale verdikking\", \"synovialisverdickung\", \"synovialverdickung\",\n        \"sinovit\", \"hoffitis\", \"sinovyal kalinlasma\", \"sinovyal proliferasyon\",\n        \"συνοβιτ\", \"υμενιτ\",\n        \"синовит\", \"sinovij\", \"sinovije\",\n    ],\n    \"Baker's\": [\n        \"baker\", \"popliteal cyst\", \"poplitealcyst\", \"popliteal cysts\",\n        \"quiste popliteo\", \"quistes popliteos\", \"quiste de baker\",\n        \"kyste de baker\", \"kyste poplite\", \"popliteale cyste\", \"popliteale cyst\",\n        \"bakercyste\", \"bakerzyste\", \"poplitealzyste\", \"popliteazyste\",\n        \"baker kisti\", \"popliteal kist\",\n        \"κυστη baker\", \"κυστη του baker\",\n        \"киста на бейкър\",\n        \"бейкърова киста\",\n        \"poplitealna cista\",\n    ],\n    \"Contusion\": [\n        \"contusion\", \"contusiones\", \"contusie\", \"kontusyon\", \"kontuzyon\",\n        \"bone bruise\", \"bone bruising\", \"knochenprellung\", \"prellung\",\n        \"botcontusie\", \"μωλωπ\", \"θλαση\",\n        \"контузи\", \"kontuzij\", \"nagnjecen\",\n    ],\n    \"Fracture\": [\n        \"fracture\", \"fractur\", \"fractura\", \"fraktur\", \"fractuur\", \"breuk\",\n        \"kirik\", \"kirig\", \"καταγμα\",\n        \"καταγματ\",\n        \"фрактура\", \"счупван\",\n        \"avulsion\", \"avulsie\", \"avulsiyon\", \"prijelom\",\n    ],\n}\n\n# Negation cues, word-boundary anchored. A bare \"no \" substring matches inside the\n# Spanish word \"cuerno\", which silently killed most Spanish positives before this\n# was anchored.\nNEG_SRC = [\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bnon\\b\", r\"\\bwithout\\b\", r\"\\babsen\\w*\",\n    r\"\\bnegative for\\b\", r\"\\bunremarkable\\b\", r\"\\bintact\\w*\", r\"\\bnormal\\w*\",\n    r\"\\bpreserved\\b\", r\"\\bfree of\\b\", r\"\\bexcluded\\b\",\n    r\"\\bno hay\\b\", r\"\\bsin\\b\", r\"\\bausen\\w*\", r\"\\bno se\\b\", r\"\\bconservad\\w*\",\n    r\"\\bintegr\\w*\", r\"\\bindemne\\b\",\n    r\"\\bpas de\\b\", r\"\\baucun\\w*\", r\"\\bsans\\b\",\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bonopvallend\\w*\", r\"\\bnormaal\\b\",\n    r\"\\bvrij van\\b\", r\"\\bintacte\\b\",\n    r\"\\bkein\\w*\", r\"\\bohne\\b\", r\"\\bunauffallig\\w*\", r\"\\bregelrecht\\w*\",\n    r\"\\bnicht\\b\", r\"\\bintakt\\w*\",\n    r\"\\byok\\w*\", r\"\\bizlenmemis\\w*\", r\"\\bizlenmedi\\w*\", r\"\\bsaptanmamis\\w*\",\n    r\"\\bgozlenmemis\\w*\", r\"\\bgorulmemis\\w*\", r\"\\bkorunmus\\w*\",\n    r\"\\bmevcut degil\\b\", r\"\\bnormaldir\\b\", r\"\\bdogaldir\\b\",\n    r\"\\bδεν\\b\", r\"\\bχωρις\\b\",\n    r\"\\bφυσιολογικ\\w*\",\n    r\"\\bακεραι\\w*\",\n    r\"\\bбез\\b\", r\"\\bняма\\b\",\n    r\"\\bне се\\b\", r\"\\bнормал\\w*\",\n    r\"\\bзапазен\\w*\",\n    r\"\\bсъхранен\\w*\",\n    r\"\\bsenza\\b\", r\"\\bsem\\b\", r\"\\bnao\\b\", r\"\\bbez\\b\", r\"\\buredn\\w*\",\n]\nNEG_RE = re.compile(\"|\".join(NEG_SRC))\n\nNEG_BACK = 70\nNEG_FWD = 45\nQUAL_WINDOW = 90\nPAIR_WINDOW = 160\nCLAUSE_SPLIT = re.compile(r\"[.;:\\n\\r•·]+|\\s-\\s|\\s>\\s|\\s\\*\\s\")\nPF_MASK = re.compile(r\"(medial|lateral)\\s+(patellar|patella|facet|trochlea|trochlear|retinac)\\w*\")\n\n\ndef fold_text(text):\n    \"\"\"Lowercase, normalise script-specific letters, and strip diacritics.\n\n    Turkish dotless i, the German sharp s and the Croatian barred d have no\n    combining-mark decomposition, so they are mapped explicitly before NFKD.\n\n    Args:\n        text: Any report text.\n\n    Returns:\n        A lowercase, accent-free string safe for substring matching.\n    \"\"\"\n    t = str(text).lower().translate(_CHAR_MAP)\n    t = unicodedata.normalize(\"NFKD\", t)\n    return \"\".join(c for c in t if not unicodedata.combining(c))\n\n\ndef _find_any(clause, terms):\n    \"\"\"Return the earliest index at which any term occurs, or -1.\n\n    Args:\n        clause: Folded clause text.\n        terms: Iterable of folded surface forms.\n\n    Returns:\n        Character index of the earliest hit, or -1 when none matches.\n    \"\"\"\n    best = -1\n    for t in terms:\n        i = clause.find(t)\n        if i >= 0 and (best < 0 or i < best):\n            best = i\n    return best\n\n\ndef _negated(clause, pos, span):\n    \"\"\"Test whether a negation cue sits near a matched finding term.\n\n    Args:\n        clause: Folded clause text.\n        pos: Start index of the finding term.\n        span: Length of the finding term.\n\n    Returns:\n        True when a cue appears in the preceding or following window.\n    \"\"\"\n    back = clause[max(0, pos - NEG_BACK):pos]\n    fwd = clause[pos + span:pos + span + NEG_FWD]\n    return bool(NEG_RE.search(back) or NEG_RE.search(fwd))\n\n\ndef _fire_direct(clause, terms):\n    \"\"\"Test a direct finding term, retrying later occurrences past a negation.\n\n    Args:\n        clause: Folded clause text.\n        terms: Surface forms that are themselves the finding.\n\n    Returns:\n        True when at least one occurrence is unnegated.\n    \"\"\"\n    for t in terms:\n        start = 0\n        while True:\n            i = clause.find(t, start)\n            if i < 0:\n                break\n            if not _negated(clause, i, len(t)):\n                return True\n            start = i + 1\n    return False\n\n\ndef _anatomy_index(clause, label):\n    \"\"\"Locate the anatomy mention for a paired label.\n\n    Falls back to a generic organ term paired with a side qualifier, which is what\n    catches Greek and Bulgarian phrasing that names the compartment rather than the\n    structure.\n\n    Args:\n        clause: Folded clause text.\n        label: One of the four paired label names.\n\n    Returns:\n        Character index of the anatomy mention, or -1.\n    \"\"\"\n    i = _find_any(clause, SPECIFIC[label])\n    if i >= 0:\n        return i\n    g = _find_any(clause, GENERIC[label])\n    if g >= 0 and _find_any(clause, SIDE_Q[label]) >= 0:\n        return g\n    return -1\n\n\ndef extract_labels(report):\n    \"\"\"Extract twelve binary findings from one free-text radiology report.\n\n    Args:\n        report: Report text in any of the languages present in the corpus.\n\n    Returns:\n        Dict mapping each of the twelve label names to 0 or 1.\n    \"\"\"\n    out = {lab: 0 for lab in LABELS}\n    if not isinstance(report, str) or not report.strip():\n        return out\n    text = fold_text(report)\n    for raw in CLAUSE_SPLIT.split(text):\n        clause = raw.strip()\n        if len(clause) < 3:\n            continue\n        for lab in (\"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"):\n            if not out[lab] and _fire_direct(clause, DIRECT[lab]):\n                out[lab] = 1\n        for lab in SPECIFIC:\n            if out[lab]:\n                continue\n            ai = _anatomy_index(clause, lab)\n            if ai < 0:\n                continue\n            finds = TEAR + SPRAIN if lab in (\"ACL\", \"MCL\") else TEAR\n            for t in finds:\n                fi = clause.find(t)\n                if fi < 0 or abs(fi - ai) > PAIR_WINDOW:\n                    continue\n                if not _negated(clause, fi, len(t)):\n                    out[lab] = 1\n                    break\n        oi = _find_any(clause, OA_FIND)\n        if oi >= 0 and not _negated(clause, oi, 8):\n            if _find_any(clause, TRICOMP) >= 0:\n                out[\"Medial OA\"] = 1\n                out[\"Lateral OA\"] = 1\n                out[\"PF OA\"] = 1\n            win = clause[max(0, oi - QUAL_WINDOW):oi + QUAL_WINDOW]\n            if _find_any(win, PF_Q) >= 0:\n                out[\"PF OA\"] = 1\n            masked = PF_MASK.sub(\" \", win)\n            if _find_any(masked, MEDIAL_Q) >= 0:\n                out[\"Medial OA\"] = 1\n            if _find_any(masked, LATERAL_Q) >= 0:\n                out[\"Lateral OA\"] = 1\n    return out\n\n\nprint(\"extractor ready:\", len(SPECIFIC), \"paired labels,\",\n      len(DIRECT), \"direct labels,\", len(NEG_SRC), \"negation cues\")"},{"cell_type":"code","execution_count":null,"id":"fd03abebb6fa","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def _label_search_dirs():\n    \"\"\"Directories that may hold a pre-extracted label table.\n\n    Kaggle nests every input under a category directory rather than mounting it flat.\n    The competition lands at /kaggle/input/competitions/<slug>, the model at\n    /kaggle/input/models/<owner>/<name>/<framework>/<variation>/<version>, and an\n    attached dataset at /kaggle/input/datasets/<owner>/<slug>. Version 2 of this\n    notebook looked only at the direct children of /kaggle/input, found no label file,\n    and fell back to the regex extractor while reporting a clean run, which measured\n    the exact artifact the LLM extraction was built to beat.\n\n    So this walks a bounded depth instead of listing one level, and prunes the\n    competition tree, which holds roughly 730,000 DICOM files and would otherwise make\n    the walk cost more than the encoding.\n\n    Returns:\n        List of Path objects, existing directories only, shallowest first.\n    \"\"\"\n    out = []\n    kin = Path(\"/kaggle/input\")\n    if kin.is_dir():\n        out.append(kin)\n        stack = [(kin, 0)]\n        while stack:\n            d, depth = stack.pop()\n            if depth >= 5:\n                continue\n            try:\n                children = sorted(p for p in d.iterdir() if p.is_dir())\n            except OSError:\n                continue\n            for c in children:\n                # Never descend into the DICOM corpus or a git/cache tree.\n                if c.name in (\"competitions\", \"train_series\", \"test_series\",\n                              \".git\", \"__pycache__\"):\n                    continue\n                out.append(c)\n                stack.append((c, depth + 1))\n    out.append(Path(\"/mnt/data/Github/Kaggle/Competitions\"\n                    \"/rsna_knee_abnormality_detection/data/interim\"))\n    out.append(Path(\"data/interim\"))\n    seen, uniq = set(), []\n    for d in out:\n        if str(d) not in seen and d.is_dir():\n            seen.add(str(d))\n            uniq.append(d)\n    return uniq\n\n\ndef _find_label_file(patterns):\n    \"\"\"Find the newest file matching any of the given glob patterns.\n\n    Args:\n        patterns: Glob patterns, tried in order across every search directory.\n\n    Returns:\n        A Path, or None when nothing matches.\n    \"\"\"\n    hits = []\n    for d in _label_search_dirs():\n        for pat in patterns:\n            try:\n                hits.extend(sorted(d.glob(pat)))\n            except OSError:\n                continue\n    if not hits:\n        return None\n    return sorted(hits, key=lambda p: (p.name, str(p)))[-1]\n\n\ndef _read_label_table(path):\n    \"\"\"Read a label table and keep only the id column and the twelve findings.\n\n    Args:\n        path: Parquet or CSV path.\n\n    Returns:\n        DataFrame indexed by StudyInstanceUID with the twelve label columns as\n        floats, or None when the file cannot be read or lacks the columns.\n    \"\"\"\n    try:\n        df = pd.read_parquet(path) if path.suffix == \".parquet\" else pd.read_csv(path)\n    except Exception as exc:\n        print(\"could not read\", path, \":\", exc)\n        return None\n    if \"StudyInstanceUID\" not in df.columns or not set(LABELS).issubset(df.columns):\n        print(\"label file\", path.name, \"is missing required columns; ignored\")\n        return None\n    out = df[[\"StudyInstanceUID\"] + LABELS].copy()\n    out[\"StudyInstanceUID\"] = out[\"StudyInstanceUID\"].astype(str)\n    out = out.drop_duplicates(\"StudyInstanceUID\").set_index(\"StudyInstanceUID\")\n    return out[LABELS].astype(float).clip(0.0, 1.0)\n\n\ntrain_ids = train[\"StudyInstanceUID\"].astype(str).tolist()\nreport_col = \"Report\" if \"Report\" in train.columns else None\n\n# ---- 1. Always compute the rule labels; they are the floor and the fallback ----\nrule_df = None\nif LABEL_SOURCE in (\"auto\", \"rule_file\"):\n    p = _find_label_file(RULE_LABEL_PATTERNS)\n    if p is not None:\n        rule_df = _read_label_table(p)\n        if rule_df is not None:\n            print(\"rule labels from file :\", p.name, \"|\", len(rule_df), \"studies\")\nif rule_df is None:\n    if report_col is None:\n        raise KeyError(\"train.csv has no Report column and no label file was found\")\n    t0 = time.time()\n    rule_df = pd.DataFrame([extract_labels(r) for r in train[report_col]],\n                           index=pd.Index(train_ids, name=\"StudyInstanceUID\"))\n    rule_df = rule_df[LABELS].astype(float)\n    print(\"rule labels extracted inline: {} reports in {:.1f}s\".format(\n        len(rule_df), time.time() - t0))\n\nY_soft = rule_df.reindex(train_ids).astype(float)\nY_soft = Y_soft.fillna(0.0)\nLABEL_PROVENANCE = \"rule\"\nn_llm = 0\n\n# ---- 2. Overwrite with LLM soft labels wherever they exist --------------------\nif LABEL_SOURCE in (\"auto\", \"llm\"):\n    p = _find_label_file(LLM_LABEL_PATTERNS)\n    if p is not None:\n        llm = _read_label_table(p)\n        if llm is not None:\n            common = [s for s in train_ids if s in llm.index]\n            if common:\n                Y_soft.loc[common, LABELS] = llm.loc[common, LABELS].values\n                n_llm = len(common)\n                LABEL_PROVENANCE = \"llm+rule\" if n_llm < len(train_ids) else \"llm\"\n                print(\"LLM labels from file  :\", p.name, \"|\", n_llm,\n                      \"of\", len(train_ids), \"studies overwritten\")\n    elif LABEL_SOURCE == \"llm\":\n        print(\"LABEL_SOURCE='llm' but no LLM label file is mounted; using rule labels\")\n\n# ---- 2b. Fail closed when a declared label dataset did not arrive -------------\n# Falling back to the regex extractor is the right behaviour when no label dataset was\n# ever declared. It is the WRONG behaviour when one was declared and did not mount:\n# the run then measures the artifact the extraction exists to beat, and reports a clean\n# result while doing it. Version 2 of this notebook did exactly that. A guard that\n# treats \"the good input is missing\" as permission to continue is fail-open.\nif LABEL_DATASET_SOURCES and LABEL_PROVENANCE not in (\"llm\", \"llm+rule\"):\n    _listing = []\n    _kin = Path(\"/kaggle/input\")\n    if _kin.is_dir():\n        for _d in _label_search_dirs()[:40]:\n            try:\n                _listing.append(f\"{_d} -> \"\n                                + \", \".join(sorted(p.name for p in _d.iterdir())[:6]))\n            except OSError:\n                continue\n    raise RuntimeError(\n        \"Declared label dataset(s) \" + \", \".join(LABEL_DATASET_SOURCES)\n        + \" but no file matched \" + \" / \".join(LLM_LABEL_PATTERNS)\n        + f\", so labels resolved to '{LABEL_PROVENANCE}'. Refusing to train on the \"\n          \"fallback labels, because the whole point of this run is the extracted ones. \"\n          \"Directories searched:\\n  \" + \"\\n  \".join(_listing)\n    )\n\n# ---- 3. Gold labels are ground truth and override anything extracted ----------\nif N_GOLD:\n    gold_vals = (train.loc[gold_mask, [\"StudyInstanceUID\"] + LABELS]\n                 .set_index(train.loc[gold_mask, \"StudyInstanceUID\"].astype(str)))\n    Y_soft.loc[gold_ids, LABELS] = gold_vals[LABELS].astype(float).values\n\nIS_GOLD = pd.Series([s in set(gold_ids) for s in train_ids], index=train_ids)\nSAMPLE_WEIGHT = np.where(IS_GOLD.values, GOLD_WEIGHT, 1.0).astype(np.float32)\n\nprev = pd.DataFrame({\n    \"mean_soft_target\": Y_soft.mean().round(4),\n    \"positives_at_0.5\": (Y_soft >= 0.5).sum().astype(int),\n    \"prevalence_%\": ((Y_soft >= 0.5).mean() * 100).round(1),\n})\nprint()\nprint(\"label source      :\", LABEL_SOURCE, \"-> resolved to\", LABEL_PROVENANCE)\nprint(\"gold overrides    :\", N_GOLD, \"studies, weighted x{:.0f} in the loss\".format(\n    GOLD_WEIGHT))\nprint()\nprint(prev.to_string())"},{"cell_type":"code","execution_count":null,"id":"fde180681990","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# The gold studies are the only ground truth available and there are 58 of them.\n# What that supports is checking whether the extractor agrees with a radiologist,\n# per label. It does not support ranking two models; the standard error on a\n# per-label AUC at this n is near 0.07.\nif N_GOLD and report_col is not None:\n    gold_true = train.loc[gold_mask, LABELS].astype(int)\n    pred_gold = (rule_df.reindex(gold_ids)[LABELS].values >= 0.5).astype(int)\n    rows = []\n    for j, lab in enumerate(LABELS):\n        yt = gold_true[lab].values\n        yp = pred_gold[:, j]\n        tp = int(((yt == 1) & (yp == 1)).sum())\n        fp = int(((yt == 0) & (yp == 1)).sum())\n        fn = int(((yt == 1) & (yp == 0)).sum())\n        rows.append({\n            \"label\": lab,\n            \"gold_pos\": int(yt.sum()),\n            \"extractor_pos\": int(yp.sum()),\n            \"agreement\": float((yt == yp).mean()),\n            \"precision\": tp / (tp + fp) if (tp + fp) else np.nan,\n            \"recall\": tp / (tp + fn) if (tp + fn) else np.nan,\n        })\n    audit = pd.DataFrame(rows)\n    print(\"Rule extractor vs gold labels, n = {} studies\".format(N_GOLD))\n    print(audit.round(3).to_string(index=False))\n    print()\n    print(\"mean agreement {:.4f} | mean precision {:.4f} | mean recall {:.4f}\".format(\n        audit[\"agreement\"].mean(), audit[\"precision\"].mean(), audit[\"recall\"].mean()))\n    print()\n    print(\"Reference, measured off-notebook on the same 58 studies:\")\n    print(\"  rule extractor      agreement 0.7960  precision 0.6802  recall 0.7648\")\n    print(\"  Claude Opus 5 (low) agreement 0.8233  precision 0.6910  recall 0.8602\")\n    print(\"  paired delta +0.0273 agreement, 95% CI [+0.0029, +0.0532], McNemar p=0.0204\")\nelse:\n    audit = pd.DataFrame(columns=[\"label\", \"agreement\", \"precision\", \"recall\"])\n    print(\"no gold labels present in this input; the extractor audit is skipped\")"},{"cell_type":"markdown","id":"01d0bd738929","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">4</span><span style=\"vertical-align:middle;\">Series selection: five slots, not six</span></h2></div>\n\n`train_series.csv` gives three flags per series: `Anatomical_Plane`, `Fluid_Sensitive`\nand `Fat_Suppression`. The first thing to check is whether those last two are two\nflags. They are not. `Fluid_Sensitive` and `Fat_Suppression` are **identical on all\n24,371 rows**, so the corpus carries exactly two contrast classes, 14,010 fluid-sensitive\nseries and 10,361 not. Any design that spends a dimension on \"fluid-sensitive and\nfat-suppressed\" is spending it on a copy.\n\nThe frontier fills six slots, three planes by two contrasts, by scoring the series in a\nplane and taking the best unused one for each contrast. Read that rule carefully: the\nsecond slot in a plane fills whenever the plane has **two or more series**, whatever\ntheir contrast. Measured on the manifest, that happens for 97.3% of studies in Sagittal,\n74.5% in Coronal and 24.8% in Axial, which is where the 4.97 of 6 mean fill comes from.\nIt also means that in 237 of the 1,094 studies where the Axial second slot fills, it\nfills with a series of the *same* contrast as the first. The slot's name is not what is\nin it.\n\nDefining the slots by contrast instead, and refusing to fill a slot with the wrong\ncontrast, gives these fill rates over the 4,407 training studies:\n\n| Slot | Studies with a genuine series | Fill |\n|---|---|---|\n| Axial fluid-sensitive | 4,407 | 100.0% |\n| Sagittal non-fluid | 4,266 | 96.8% |\n| Coronal fluid-sensitive | 4,248 | 96.4% |\n| Sagittal fluid-sensitive | 4,150 | 94.2% |\n| Coronal non-fluid | 3,406 | 77.3% |\n| **Axial non-fluid** | **857** | **19.4%** |\n\nThe last row is the decision. An Axial non-fluid slot is empty for four studies in five,\nso a sixth of the encoding budget buys a feature that exists for 19.4% of the corpus and\nis a padded zero for the rest. That budget is worth more as slices in the slots that\nalways fill. Five slots it is, with a mean genuine fill of 4.65 of 5, or 92.9% per slot,\nagainst the frontier's 82.8%.\n\nThe slice budget is then weighted by what each view reads. Sagittal carries the ACL, both\nmeniscal horns, the patellofemoral cartilage and the popliteal fossa, so it takes twice\nwhat Axial takes. Coronal carries the meniscal bodies, the MCL and the medial and lateral\ncompartments. Axial carries the patellofemoral joint, synovium and the suprapatellar\nrecess. The numbers are in the config cell and the expected slices per study is printed\nfrom the manifest, not assumed."},{"cell_type":"code","execution_count":null,"id":"1804cd446b55","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"PLANES = [\"Sagittal\", \"Coronal\", \"Axial\"]\n\n\ndef normalise_manifest(df):\n    \"\"\"Coerce a series manifest into the columns the slot builder needs.\n\n    Args:\n        df: A raw train_series or test_series frame.\n\n    Returns:\n        A copy with string study/series ids, a title-cased plane and an integer\n        fluid flag. Empty in, empty out.\n    \"\"\"\n    out = df.copy()\n    for c in SERIES_COLS:\n        if c not in out.columns:\n            out[c] = np.nan\n    out[\"StudyInstanceUID\"] = out[\"StudyInstanceUID\"].astype(str)\n    out[\"SeriesInstanceUID\"] = out[\"SeriesInstanceUID\"].astype(str)\n    out[\"plane\"] = out[\"Anatomical_Plane\"].astype(str).str.strip().str.title()\n    out[\"fluid\"] = pd.to_numeric(out[\"Fluid_Sensitive\"], errors=\"coerce\") \\\n        .fillna(0).astype(int)\n    return out\n\n\ntr_man = normalise_manifest(train_series)\nte_man = normalise_manifest(test_series)\n\n# The flag check comes first, because the whole six-slot design downstream rests on\n# it. Fluid_Sensitive and Fat_Suppression are one column wearing two names.\nif len(tr_man) and \"Fat_Suppression\" in train_series.columns:\n    fs = pd.to_numeric(train_series[\"Fluid_Sensitive\"], errors=\"coerce\")\n    fat = pd.to_numeric(train_series[\"Fat_Suppression\"], errors=\"coerce\")\n    n_disagree = int((fs != fat).sum())\n    print(\"Fluid_Sensitive vs Fat_Suppression over {:,} series: {} rows disagree\"\n          .format(len(train_series), n_disagree))\n    print(\"  -> {} contrast classes, not four\".format(2 if n_disagree == 0 else 4))\n    print(\"  fluid-sensitive series {:,} | non-fluid {:,}\".format(\n        int((fs == 1).sum()), int((fs == 0).sum())))\n    print()\n\n\ndef slot_fill_table(man):\n    \"\"\"Measure how often each plane-by-contrast slot has a genuine series.\n\n    Args:\n        man: A normalised manifest.\n\n    Returns:\n        DataFrame with one row per candidate slot, sorted by fill rate.\n    \"\"\"\n    n = man[\"StudyInstanceUID\"].nunique()\n    rows = []\n    for plane in PLANES:\n        for f in (1, 0):\n            sub = man[(man[\"plane\"] == plane) & (man[\"fluid\"] == f)]\n            have = sub[\"StudyInstanceUID\"].nunique()\n            rows.append({\n                \"slot\": \"{}_{}\".format(plane, \"Fluid\" if f else \"Structural\"),\n                \"series\": len(sub),\n                \"studies\": have,\n                \"fill\": have / max(1, n),\n                \"kept\": \"{}_{}\".format(plane, \"Fluid\" if f else \"Structural\")\n                        in SLOT_NAMES,\n            })\n    return pd.DataFrame(rows).sort_values(\"fill\", ascending=False)\n\n\nfill = slot_fill_table(tr_man)\nprint(\"Six candidate slots, measured over {:,} training studies\".format(\n    tr_man[\"StudyInstanceUID\"].nunique()))\nprint(fill.to_string(index=False, float_format=lambda v: \"{:.4f}\".format(v)))\nprint()\n\n# The frontier's rule fills a plane's second slot whenever the plane holds two or\n# more series, whatever contrast they are. Reproduce that so the comparison is a\n# measurement rather than a claim.\ncnt = (tr_man.groupby([\"StudyInstanceUID\", \"plane\"]).size().unstack(fill_value=0)\n       .reindex(columns=PLANES, fill_value=0))\nn_studies_tr = max(1, len(cnt))\nrival_fill = 3.0 + sum((cnt[p] >= 2).mean() for p in PLANES)\nprint(\"frontier's six-slot rule, second slot per plane fills when the plane has >=2 series:\")\nfor p in PLANES:\n    print(\"  {:9s} {:5.1f}% of studies\".format(p, 100.0 * (cnt[p] >= 2).mean()))\nax_two = set(cnt.index[cnt[\"Axial\"] >= 2])\nax_true = set(tr_man.loc[(tr_man[\"plane\"] == \"Axial\") & (tr_man[\"fluid\"] == 0),\n                         \"StudyInstanceUID\"])\nprint(\"  mean slots filled {:.3f} / 6 = {:.1%} per slot\".format(\n    rival_fill, rival_fill / 6.0))\nprint(\"  of the {} studies whose Axial second slot fills, {} have no genuine non-fluid\"\n      \" axial series, so the slot holds a duplicate contrast\".format(\n          len(ax_two), len(ax_two - ax_true)))\nprint()\n\nkept = fill[fill[\"kept\"]]\nmean_fill = kept[\"fill\"].sum()\nprint(\"this notebook's {} slots, contrast enforced:\".format(N_SLOTS))\nfor name, (_, _, k) in zip(SLOT_NAMES, SLOT_SPEC):\n    row = fill.loc[fill[\"slot\"] == name].iloc[0]\n    print(\"  {:22s} fill {:6.1%}  slice budget {:2d}\".format(name, row[\"fill\"], k))\nprint(\"  mean slots filled {:.3f} / {} = {:.1%} per slot\".format(\n    mean_fill, N_SLOTS, mean_fill / N_SLOTS))\nexp_slices = sum(fill.loc[fill[\"slot\"] == n, \"fill\"].iloc[0] * k\n                 for n, (_, _, k) in zip(SLOT_NAMES, SLOT_SPEC))\nprint(\"  expected slices encoded per study: {:.1f}\".format(exp_slices))\nprint()\n\n\ndef build_slots(man):\n    \"\"\"Assign one series to each contrast-defined slot, per study.\n\n    A slot is filled only by a series whose plane and fluid flag both match. A\n    slot with no matching series stays empty and is carried as a masked, zero\n    embedding rather than being back-filled with the wrong contrast.\n\n    Args:\n        man: A normalised manifest.\n\n    Returns:\n        Dict mapping study id to a dict of slot name -> series id.\n    \"\"\"\n    out = {}\n    if not len(man):\n        return out\n    for sid, rows in man.groupby(\"StudyInstanceUID\", sort=False):\n        chosen = {}\n        used = set()\n        for name, (plane, flag, _) in zip(SLOT_NAMES, SLOT_SPEC):\n            part = rows[(rows[\"plane\"] == plane) & (rows[\"fluid\"] == flag)]\n            for series_id in part[\"SeriesInstanceUID\"].tolist():\n                if series_id not in used:\n                    chosen[name] = series_id\n                    used.add(series_id)\n                    break\n        out[str(sid)] = chosen\n    return out\n\n\nSLOTS_TRAIN = build_slots(tr_man)\nSLOTS_TEST = build_slots(te_man)\nfilled_tr = np.array([len(SLOTS_TRAIN.get(s, {})) for s in train_ids])\nprint(\"slot assignment    : train {} studies, mean {:.3f} slots filled\".format(\n    len(SLOTS_TRAIN), filled_tr.mean() if len(filled_tr) else 0.0))\nprint(\"                     test  {} studies\".format(len(SLOTS_TEST)))"},{"cell_type":"code","execution_count":null,"id":"3295c24c4d48","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# Two panels, stacked, one label order shared, so a slot can be tracked straight\n# down the figure. Upper: how often each candidate slot holds a genuine series.\n# Lower: what each kept slot is given to spend.\norder = fill.sort_values(\"fill\", ascending=True).reset_index(drop=True)\nypos = np.arange(len(order))\ncolours = [PRIMARY if k else NEUTRAL for k in order[\"kept\"]]\n\nfig = plt.figure(figsize=(13.2, 10.8))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.0, 0.85], hspace=0.34)\n\nax0 = fig.add_subplot(gs[0, 0])\nax0.barh(ypos, order[\"fill\"].values, height=0.68, color=colours, zorder=3)\nlabel_bars(ax0, order[\"fill\"].values, ypos, fmt=\"{:.1%}\", pad=0.012)\nax0.set_yticks(ypos)\nax0.set_yticklabels(order[\"slot\"].tolist(), fontsize=12.5, color=INK)\nax0.set_xlim(0, 1.16)\nax0.set_xticks([0.0, 0.25, 0.5, 0.75, 1.0])\nax0.set_xticklabels([\"0%\", \"25%\", \"50%\", \"75%\", \"100%\"])\nax0.set_xlabel(\"share of training studies holding a series of that plane and contrast\")\nstyle(ax0, \"x\")\nax0.spines[\"left\"].set_visible(False)\ndropped = order.loc[~order[\"kept\"], \"slot\"].tolist()\nif dropped:\n    ax0.text(1.14, ypos[0], \"dropped\", ha=\"right\", va=\"center\", fontsize=12,\n             color=ACCENT_TEXT, fontweight=\"bold\")\n\nax1 = fig.add_subplot(gs[1, 0])\nbudgets = []\nfor name in order[\"slot\"]:\n    b = dict(zip(SLOT_NAMES, [k for _, _, k in SLOT_SPEC])).get(name, 0)\n    budgets.append(b)\nax1.barh(ypos, budgets, height=0.68,\n         color=[PRIMARY if b else NEUTRAL for b in budgets], zorder=3)\nlabel_bars(ax1, budgets, ypos, fmt=\"{:.0f}\", pad=0.14)\nax1.set_yticks(ypos)\nax1.set_yticklabels(order[\"slot\"].tolist(), fontsize=12.5, color=INK)\nax1.set_xlim(0, max(budgets + [1]) * 1.22)\nax1.set_xlabel(\"slices encoded per study from that slot\")\nstyle(ax1, \"x\")\nax1.spines[\"left\"].set_visible(False)\n\nfinish(fig,\n       \"An Axial non-fluid slot is empty for four studies in five, so its budget \"\n       \"goes to slices instead\",\n       detail=\"Upper: measured fill over {:,} training studies. Lower: the slice \"\n              \"budget each kept slot receives.\".format(\n                  tr_man[\"StudyInstanceUID\"].nunique()),\n       caption=\"A slot is counted filled only when a series matches both the plane \"\n               \"and the contrast. The frontier's rule fills a plane's second slot \"\n               \"from any spare series, which is why its reported fill is higher and \"\n               \"its Axial second slot holds a duplicate contrast in {} of {} cases.\"\n               .format(len(ax_two - ax_true), len(ax_two)))"},{"cell_type":"markdown","id":"fb8a1b07299f","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">5</span><span style=\"vertical-align:middle;\">Slice order is a geometry problem, not a sort</span></h2></div>\n\nThe frontier orders slices by filename. Filenames here are SOP Instance UIDs, which are\nrandom. It then samples 11 of them, reads their headers, sorts *those* by position and\npicks three fractional positions from the sample. So the three slices it encodes are\nthree positions within a random 11-slice subset, and the physical extent they span is\nwhatever that subset happened to cover.\n\nPosition is in the files. `ImageOrientationPatient` gives the two in-plane direction\ncosines; their cross product is the slice normal; the dot product of\n`ImagePositionPatient` with that normal is the slice's signed position along the stack.\nThat is exact, it is invariant to how the scanner numbered the instances, and it is the\nonly ordering that makes \"the 30% slice\" mean the same thing in two different studies.\n\nHeaders are read with `stop_before_pixels=True` and a four-tag allowlist, so no pixel\ndecoder is involved and no transfer syntax can fail the read. The corpus mixes\nuncompressed Explicit VR, Implicit VR, JPEG Lossless and JPEG 2000, and that mix is a\nreal hazard for the pixel path, not for this one. Every series has its headers read up to\na cap; the median series holds 30 slices and the cap is well above it, so ordering is\nexact for almost every series and falls back to an evenly spaced filename subset only in\nthe long tail.\n\nSlices are then drawn at evenly spaced fractional positions across a per-plane window of\nthe ordered stack, and the physical millimetres covered is reported. **Each encoded image\nis a three-channel stack of the decoded slice and its two decoded neighbours**, so\nneighbourhood context costs no extra DICOM reads: every decoded slice appears in up to\nthree images."},{"cell_type":"code","execution_count":null,"id":"d38239ab6a8d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"GEOM_TAGS = [\"ImagePositionPatient\", \"ImageOrientationPatient\",\n             \"SliceLocation\", \"InstanceNumber\", \"Laterality\", \"ImageLaterality\"]\n\n# Counters, so a degraded run is visible in the log rather than silent. They are\n# incremented from worker threads, and `d[k] += 1` is three bytecodes, not one, so\n# the lock is not decoration: an undercounted failure reads as a clean run.\nimport threading\n\nSTATS = {\"series_missing\": 0, \"header_fail\": 0, \"pixel_fail\": 0,\n         \"series_truncated\": 0, \"slots_empty\": 0, \"slices_decoded\": 0,\n         \"exact_order\": 0, \"fallback_order\": 0,\n         \"lat_left\": 0, \"lat_right_mirrored\": 0, \"lat_right_reversed\": 0,\n         \"lat_absent\": 0, \"lat_conflict\": 0}\n_STATS_LOCK = threading.Lock()\n\n\ndef bump(key, n=1):\n    \"\"\"Increment one diagnostic counter under a lock.\n\n    Args:\n        key: Counter name.\n        n: Amount to add.\n\n    Returns:\n        None.\n    \"\"\"\n    with _STATS_LOCK:\n        STATS[key] = STATS.get(key, 0) + n\n\n\ndef list_dicom_paths(folder):\n    \"\"\"List the .dcm files in one series directory.\n\n    Args:\n        folder: Series directory.\n\n    Returns:\n        A list of Paths sorted by filename, empty when the directory is missing\n        or unreadable. Filename order is SOP UID order here, which is arbitrary;\n        it is used only to take a deterministic subset, never as slice order.\n    \"\"\"\n    try:\n        return sorted((Path(e.path) for e in os.scandir(folder)\n                       if e.is_file() and e.name.lower().endswith(\".dcm\")),\n                      key=lambda p: p.name)\n    except Exception:\n        return []\n\n\ndef projected_position(ds, fallback):\n    \"\"\"Signed position of a slice along its own stack normal, in millimetres.\n\n    The two in-plane direction cosines in ImageOrientationPatient span the slice\n    plane; their cross product is the slice normal. Projecting\n    ImagePositionPatient onto that normal gives a scalar that increases\n    monotonically through the stack and is invariant to how the scanner numbered\n    the instances. SliceLocation and InstanceNumber are fallbacks only.\n\n    Args:\n        ds: A pydicom dataset read with stop_before_pixels.\n        fallback: Value to return when no geometry tag is usable.\n\n    Returns:\n        Float position.\n    \"\"\"\n    try:\n        iop = np.asarray(ds.ImageOrientationPatient, dtype=np.float64)\n        ipp = np.asarray(ds.ImagePositionPatient, dtype=np.float64)\n        if iop.size == 6 and ipp.size == 3:\n            normal = np.cross(iop[:3], iop[3:])\n            norm = float(np.linalg.norm(normal))\n            if norm > 1e-8:\n                return float(np.dot(ipp, normal / norm))\n    except Exception:\n        pass\n    for attr in (\"SliceLocation\", \"InstanceNumber\"):\n        try:\n            v = float(getattr(ds, attr))\n            if np.isfinite(v):\n                return v\n        except Exception:\n            pass\n    return float(fallback)\n\n\ndef read_laterality(ds):\n    \"\"\"Read which knee this slice belongs to, as \"L\", \"R\" or \"\" for unknown.\n\n    Two tags carry it and neither is mandatory. `Laterality` is the general one;\n    `ImageLaterality` is what some vendors write instead. The corpus contains at\n    least the spellings L, R and LEFT, so the value is folded to its first\n    character rather than compared whole, and anything that is not L or R is\n    treated as unknown rather than guessed at.\n\n    Args:\n        ds: A pydicom dataset read with the GEOM_TAGS allowlist.\n\n    Returns:\n        \"L\", \"R\", or \"\" when neither tag is present or the value is unrecognised.\n    \"\"\"\n    for tag in (\"Laterality\", \"ImageLaterality\"):\n        v = getattr(ds, tag, None)\n        if v is None:\n            continue\n        s = str(v).strip().upper()\n        if s[:1] in (\"L\", \"R\"):\n            return s[:1]\n    return \"\"\n\n\ndef resolve_laterality(votes):\n    \"\"\"Collapse per-slice laterality readings into one value for the series.\n\n    Args:\n        votes: Iterable of per-slice values from `read_laterality`.\n\n    Returns:\n        \"L\", \"R\", or \"\" when the series is unlabelled or self-contradictory. A\n        series that claims both sides is not a series whose geometry is\n        understood, so it is left alone rather than resolved by majority.\n    \"\"\"\n    seen = {v for v in votes if v}\n    if len(seen) == 1:\n        return seen.pop()\n    if len(seen) > 1:\n        bump(\"lat_conflict\")\n    return \"\"\n\n\ndef order_series(folder):\n    \"\"\"Order one series' slices by projected physical position.\n\n    Headers are read with stop_before_pixels and a four-tag allowlist, so no\n    pixel decoder is involved and none of the four transfer syntaxes in this\n    corpus can fail the read. Series longer than the cap are subsampled by\n    filename first, which is a uniform draw over positions because filenames are\n    SOP UIDs.\n\n    Args:\n        folder: Series directory.\n\n    Returns:\n        Tuple of (ordered paths, ordered positions, number of files on disk,\n        whether every file on disk was ordered, series laterality as \"L\", \"R\"\n        or \"\").\n    \"\"\"\n    paths = list_dicom_paths(folder)\n    n_files = len(paths)\n    if n_files == 0:\n        bump(\"series_missing\")\n        return [], np.zeros(0), 0, False, \"\"\n\n    exact = n_files <= MAX_HEADER_READS_PER_SERIES\n    if not exact:\n        bump(\"series_truncated\")\n        idx = np.linspace(0, n_files - 1, MAX_HEADER_READS_PER_SERIES)\n        idx = np.unique(np.round(idx).astype(int))\n        paths = [paths[i] for i in idx]\n\n    positioned, lat_votes = [], []\n    for i, p in enumerate(paths):\n        try:\n            ds = pydicom.dcmread(str(p), stop_before_pixels=True,\n                                 specific_tags=GEOM_TAGS, force=True)\n            positioned.append((projected_position(ds, i), p))\n            lat_votes.append(read_laterality(ds))\n        except Exception:\n            bump(\"header_fail\")\n            positioned.append((float(i), p))\n    positioned.sort(key=lambda t: t[0])\n    bump(\"exact_order\" if exact else \"fallback_order\")\n    return ([t[1] for t in positioned],\n            np.asarray([t[0] for t in positioned], dtype=np.float64),\n            n_files, exact, resolve_laterality(lat_votes))\n\n\ndef pick_indices(n, k, lo, hi):\n    \"\"\"Choose k evenly spaced indices across a fractional window of n slices.\n\n    Args:\n        n: Number of ordered slices available.\n        k: How many to take.\n        lo: Lower fractional bound of the window, in [0, 1].\n        hi: Upper fractional bound.\n\n    Returns:\n        A sorted array of unique integer indices, shorter than k when n is small.\n    \"\"\"\n    if n <= 0 or k <= 0:\n        return np.zeros(0, dtype=int)\n    if n <= k:\n        return np.arange(n)\n    fracs = np.linspace(lo, hi, k)\n    idx = np.round(fracs * (n - 1)).astype(int)\n    return np.unique(np.clip(idx, 0, n - 1))"},{"cell_type":"markdown","id":"d9daf8507fd5","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">6</span><span style=\"vertical-align:middle;\">Preprocessing: the window belongs to the series</span></h2></div>\n\nPixels are 12-bit and `RescaleSlope`/`RescaleIntercept` are applied first. Then the\nquestion is what to normalise against, and the EDA already answered it: intensity\ndistributions differ sharply from series to series, so a global window is wrong.\n\nA per-*slice* window, which is what the frontier uses, is wrong in the other direction.\nIt rescales every slice to its own 1st and 99th percentile, which means a slice through a\nlarge effusion and a slice through dry bone come out with the same dynamic range. The\nbrightness of fluid on a fluid-sensitive sequence is the finding. Normalising it away\nper slice destroys exactly the signal that `Effusion`, `Synovitis` and `Baker's` are\nmade of, and it also gives the three channels of one 2.5D image three different windows,\nwhich paints colour fringes on every anatomical edge.\n\nSo the percentile window is computed once **per series**, from the pooled pixels of the\nslices that series contributes, and applied unchanged to all of them. The body bounding\nbox is likewise computed once per series from the union of its slices, so the crop does\nnot jitter between neighbouring slices.\n\nOne thing the crop does *not* do here is buy resolution. On every series available\noffline the knee fills the frame: the tissue mask reaches the image edge at every\nthreshold from 0.03 to 0.20, and the bounding box keeps 98 to 100% of the frame. These\nacquisitions are already coned tightly to the joint, so the effective pitch stays at the\nfull 150 mm over the input size and there is no free magnification to collect. The crop\nis kept as a guard for the series that are framed loosely, not as a lever. Whatever it\nachieves on the full corpus is measured and printed after encoding rather than assumed\nhere."},{"cell_type":"markdown","id":"1d02bda693f3","metadata":{},"source":"### 6b. Which knee is this, and why four labels depend on the answer\n\nFour of the twelve targets name a side: `Medial Meniscus`, `Lateral Meniscus`,\n`Medial OA` and `Lateral OA`. Medial and lateral are defined against the midline of the\nbody, not against the image. Which side of the frame the medial compartment lands on\ntherefore depends on which knee was scanned, and that is a fact about the patient rather\nthan about the anatomy the label describes.\n\nThe EDA read the `Laterality` tag across the studies it sampled and found the corpus\nmixes sides: 96 left, 84 right, 5 spelled `LEFT`, and 162 with the tag absent. So close\nto a quarter of studies are mirror images of the rest, and a third of the metric is\nbeing asked to learn a distinction from an axis the model has no way to observe. This is\nnot a subtle modelling choice. It is a third of the score sitting on a coin flip.\n\nThe correction is not one operation, because the mirror falls on a different axis in\neach plane:\n\n| Plane | Medial-lateral direction | Correction |\n|---|---|---|\n| Coronal, Axial | in the image plane | flip the last pixel axis |\n| Sagittal | along the slice axis | reverse the stack order, leave pixels alone |\n\nSagittal is the one worth stating carefully. Mirroring a sagittal slice does nothing\nuseful, because each slice is already a section through the joint at one medial-lateral\ndepth. What differs between a left and a right knee is the direction the stack travels\nthrough the joint, so the stack is reversed and the pixels are untouched. The reversal\nhappens before the 2.5D channels are assembled, so each slice still meets its true\nneighbours, and the fractional position handed to the head is inverted with it rather\nthan left running backwards against its own pixels.\n\nStudies with no tag are left exactly as they are. Forty-seven percent of the sample\ncarried no side, and a wrong flip is strictly worse than no flip: it moves a study to\nthe wrong side of an axis instead of leaving it on an unknown one. Those studies keep an\nunresolved side axis and the cost falls on the same four labels as diluted supervision.\nThe run prints how many slots were normalised, how many were left alone, and how many\nseries contradicted themselves across their own slices, so the size of that residual is\na number in the log rather than an assumption here.\n\nOne consequence for the cache. Embeddings are keyed by a tag that now carries a\npreprocessing version, because a cached embedding is the one artifact that cannot tell\nyou which pixels produced it. Reusing the previous cache here would reload right knees\nencoded in their original orientation and report them as a new result."},{"cell_type":"code","execution_count":null,"id":"5b15458209c4","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def read_slice(path):\n    \"\"\"Decode one DICOM slice into a rescaled float array.\n\n    RescaleSlope and RescaleIntercept are applied before anything else, and\n    MONOCHROME1 is inverted so bright always means high signal. No window is\n    applied here: the window belongs to the series, not the slice.\n\n    Args:\n        path: Path to a .dcm file.\n\n    Returns:\n        A 2D float32 array, or None when the file cannot be decoded.\n    \"\"\"\n    try:\n        ds = pydicom.dcmread(str(path), force=True)\n        arr = ds.pixel_array.astype(np.float32)\n    except Exception:\n        bump(\"pixel_fail\")\n        return None\n    try:\n        slope = float(getattr(ds, \"RescaleSlope\", 1.0) or 1.0)\n        intercept = float(getattr(ds, \"RescaleIntercept\", 0.0) or 0.0)\n        arr = arr * slope + intercept\n        if str(getattr(ds, \"PhotometricInterpretation\", \"\")).upper() == \"MONOCHROME1\":\n            arr = float(np.nanmax(arr)) - arr\n    except Exception:\n        pass\n    if not np.isfinite(arr).any():\n        bump(\"pixel_fail\")\n        return None\n    bump(\"slices_decoded\")\n    return arr\n\n\ndef series_window(arrays, lo_pct=1.0, hi_pct=99.5):\n    \"\"\"Outlier-resistant intensity window for a series, from its pooled pixels.\n\n    This is the correction that matters most in preprocessing. A per-slice\n    window rescales a slice through a large effusion and a slice through dry\n    bone onto the same range, which deletes the very contrast that Effusion,\n    Synovitis and Baker's are made of, and gives the three channels of one 2.5D\n    image three different windows. One window per series keeps relative\n    brightness intact within the series while still absorbing the sharp\n    between-series intensity differences the EDA found.\n\n    Args:\n        arrays: List of 2D float arrays from one series.\n        lo_pct: Lower percentile.\n        hi_pct: Upper percentile.\n\n    Returns:\n        Tuple (low, high) with high strictly greater than low.\n    \"\"\"\n    pool = []\n    for a in arrays:\n        flat = a[::3, ::3].ravel()\n        flat = flat[np.isfinite(flat)]\n        if flat.size:\n            pool.append(flat)\n    if not pool:\n        return 0.0, 1.0\n    pool = np.concatenate(pool)\n    nz = pool[np.abs(pool) > 1e-8]\n    src = nz if nz.size >= 256 else pool\n    lo, hi = np.percentile(src, [lo_pct, hi_pct])\n    if not np.isfinite(lo) or not np.isfinite(hi) or hi <= lo:\n        lo, hi = float(np.min(pool)), float(np.max(pool))\n    if hi <= lo:\n        hi = lo + 1.0\n    return float(lo), float(hi)\n\n\ndef series_bbox(arrays, low, high, thresh=0.06, margin=0.06):\n    \"\"\"Body bounding box for a whole series, from the union of its slices.\n\n    Computed once per series so the crop does not jitter between neighbouring\n    slices, which would put the same anatomy in different patch positions in the\n    three channels of one image.\n\n    Args:\n        arrays: List of 2D float arrays from one series.\n        low: Series window low.\n        high: Series window high.\n        thresh: Normalised intensity above which a pixel counts as tissue.\n        margin: Fractional margin added around the union box.\n\n    Returns:\n        Tuple (y0, y1, x0, x1). The full frame when no box can be found.\n    \"\"\"\n    h, w = arrays[0].shape\n    y0, y1, x0, x1 = h, 0, w, 0\n    span = max(high - low, 1e-6)\n    for a in arrays:\n        if a.shape != (h, w):\n            return 0, h, 0, w\n        m = ((a - low) / span) > thresh\n        if m.sum() < 0.01 * h * w:\n            continue\n        ys, xs = np.where(m)\n        y0 = min(y0, int(ys.min())); y1 = max(y1, int(ys.max()) + 1)\n        x0 = min(x0, int(xs.min())); x1 = max(x1, int(xs.max()) + 1)\n    if y1 <= y0 or x1 <= x0:\n        return 0, h, 0, w\n    my = max(2, int(margin * (y1 - y0)))\n    mx = max(2, int(margin * (x1 - x0)))\n    return (max(0, y0 - my), min(h, y1 + my), max(0, x0 - mx), min(w, x1 + mx))\n\n\nCROP_RATIOS = []\n\n\ndef normalise_laterality(planes, plane, lat):\n    \"\"\"Map a right knee onto the left-knee convention.\n\n    Four of the twelve targets are medial/lateral pairs, and medial and lateral\n    are defined against the body midline rather than against the image. On a\n    right knee they therefore fall on the opposite side of the frame from a left\n    knee, so without this the four side-specific labels are being asked to learn\n    from an axis the model cannot observe.\n\n    The correction differs by plane because the mirror acts on a different axis:\n\n    * Coronal and axial carry the medial-lateral direction in the image plane, so\n      one knee is the horizontal mirror of the other and the last axis flips.\n    * Sagittal carries it along the *slice* axis. Each slice is unchanged by the\n      mirror; what differs is the direction in which the stack traverses the\n      joint, so the stack order reverses and the pixels are left alone.\n\n    A study with no `Laterality` tag is left untouched. Roughly half the corpus\n    is unlabelled, and a wrong flip is worse than no flip: it moves a study to\n    the wrong side rather than leaving it on an unknown one.\n\n    Args:\n        planes: List of normalised 2D slices in through-plane order.\n        plane: Anatomical plane of this slot.\n        lat: Series laterality, \"L\", \"R\" or \"\".\n\n    Returns:\n        Tuple of (possibly transformed planes, whether the stack was reversed).\n    \"\"\"\n    if lat != \"R\":\n        bump(\"lat_left\" if lat == \"L\" else \"lat_absent\")\n        return planes, False\n    if plane in (\"Coronal\", \"Axial\"):\n        bump(\"lat_right_mirrored\")\n        return [p[:, ::-1] for p in planes], False\n    bump(\"lat_right_reversed\")\n    return planes[::-1], True\n\n\ndef build_slot_stack(folder, plane, k, size):\n    \"\"\"Decode, window, crop and resize one slot into 2.5D image channels.\n\n    Every encoded image is the decoded slice plus its two decoded neighbours as\n    the three channels, so neighbourhood context costs no extra DICOM reads:\n    each decoded slice is reused in up to three images.\n\n    Args:\n        folder: Series directory.\n        plane: Anatomical plane, selecting the sampling window.\n        k: How many slices to encode from this slot.\n        size: Output side length in pixels.\n\n    Returns:\n        Tuple of (images uint8 [m, 3, size, size], positions float32 [m],\n        millimetres spanned, number of files on disk). m may be 0.\n    \"\"\"\n    paths, positions, n_files, _, lat = order_series(folder)\n    if not paths:\n        return np.zeros((0, 3, size, size), np.uint8), np.zeros(0, np.float32), 0.0, 0\n    lo, hi = PLANE_WINDOW.get(plane, (0.15, 0.85))\n    idx = pick_indices(len(paths), k, lo, hi)\n    if idx.size == 0:\n        return np.zeros((0, 3, size, size), np.uint8), np.zeros(0, np.float32), 0.0, n_files\n\n    arrays, kept = [], []\n    for i in idx:\n        a = read_slice(paths[int(i)])\n        if a is not None:\n            arrays.append(a)\n            kept.append(int(i))\n    if not arrays:\n        return np.zeros((0, 3, size, size), np.uint8), np.zeros(0, np.float32), 0.0, n_files\n\n    low, high = series_window(arrays)\n    y0, y1, x0, x1 = series_bbox(arrays, low, high)\n    h, w = arrays[0].shape\n    if h * w > 0:\n        CROP_RATIOS.append(math.sqrt(max(1, (y1 - y0) * (x1 - x0)) / float(h * w)))\n\n    planes_norm = []\n    span = max(high - low, 1e-6)\n    for a in arrays:\n        if a.shape != (h, w):\n            a = resize_2d(a, max(h, w))[:h, :w]\n        c = a[y0:y1, x0:x1]\n        c = np.clip((c - low) / span, 0.0, 1.0)\n        side = max(c.shape)\n        pad = np.zeros((side, side), np.float32)\n        oy, ox = (side - c.shape[0]) // 2, (side - c.shape[1]) // 2\n        pad[oy:oy + c.shape[0], ox:ox + c.shape[1]] = c\n        planes_norm.append(np.clip(resize_2d(pad, size) * 255.0, 0, 255).astype(np.uint8))\n\n    # Section 6b. Applied before the 2.5D channels are assembled, so a reversed\n    # sagittal stack still pairs each slice with its true neighbours rather than\n    # with the neighbours it had in the un-normalised direction.\n    planes_norm, stack_reversed = normalise_laterality(planes_norm, plane, lat)\n\n    m = len(planes_norm)\n    imgs = np.zeros((m, 3, size, size), np.uint8)\n    for j in range(m):\n        imgs[j, 0] = planes_norm[max(0, j - 1)]\n        imgs[j, 1] = planes_norm[j]\n        imgs[j, 2] = planes_norm[min(m - 1, j + 1)]\n\n    pos = positions[np.asarray(kept, dtype=int)] if positions.size else np.arange(m)\n    span_mm = float(np.abs(positions.max() - positions.min())) if positions.size > 1 else 0.0\n    rng = max(1e-6, float(pos.max() - pos.min())) if m > 1 else 1.0\n    frac = ((pos - pos.min()) / rng).astype(np.float32) if m > 1 else np.zeros(m, np.float32)\n    if stack_reversed:\n        # The head reads `frac` as distance along the slot's own axis. Reversing\n        # the images without reversing this would hand every reversed study a\n        # positional encoding that runs backwards against its own pixels.\n        frac = (1.0 - frac[::-1]).astype(np.float32)\n    return imgs, frac, span_mm, n_files\n\n\nprint(\"preprocessing ready\")\nprint(\"laterality                   : right knees mapped onto the left convention, \"\n      \"coronal/axial mirrored, sagittal stack reversed, untagged left alone\")\nprint(\"per-series window percentiles: 1.0 / 99.5, pooled over the slices encoded\")\nprint(\"sampling window by plane      :\",\n      \", \".join(\"{} {:.2f}-{:.2f}\".format(k, v[0], v[1])\n                for k, v in PLANE_WINDOW.items()))"},{"cell_type":"markdown","id":"cb63784b1b03","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">7</span><span style=\"vertical-align:middle;\">Encoding with a frozen DINOv2</span></h2></div>\n\nThe backbone loaded back in section 1 is `facebook/dinov2-small` from the Kaggle model\nmount `metaresearch/dinov2/PyTorch/small/1`, read with `local_files_only=True`. Internet\nis off, which a code-competition submission requires anyway.\n\nThe backbone is frozen and never fine-tuned. That is a deliberate choice given 58 gold\nstudies and a report-derived training signal of unknown fidelity: fine-tuning a 22M\nparameter ViT against noisy pseudo-labels mostly teaches it the noise. What a frozen\nbackbone buys instead is that embeddings are computed once, cached, and every head\nexperiment afterwards costs seconds rather than hours.\n\nDINOv2 interpolates its positional grid on the fly, so a non-native input size needs no\nsurgery, and the direction of that interpolation is worth naming. The released checkpoint\nis configured at 518 px, a 37 by 37 grid. A 336 px input interpolates it down to 24 by 24.\nA 126 px input interpolates it down to 9 by 9, discarding 94% of the positional\nresolution the model was trained with. The higher input size is not an exotic setting\nhere, it is the closer one."},{"cell_type":"markdown","id":"459e6fc27a63","metadata":{},"source":"### Runtime projection before the work\n\nThe visible `test.csv` is a placeholder with three rows. At rerun it holds roughly 1,300\nstudies this notebook has never seen. Before any bulk encoding, a short probe times the\nthree phases separately (header listing and ordering, DICOM decode and preprocess, GPU\nforward) and projects each to the rerun test set and to the full training set. The\nprojection is printed first, then acted on.\n\nIf the projection exceeds the budget the pipeline steps down in a fixed order. If the GPU\nis the phase over budget, the input size drops one rung down the ladder, because that is\nthe only knob that moves GPU cost. Otherwise the per-slot slice budget is scaled back,\nand only after that is the training set subsampled. The test set is never subsampled and\nits encoding runs **first**, so a truncated run still has something to predict with."},{"cell_type":"code","execution_count":null,"id":"37e1c06ab2d4","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"from concurrent.futures import ThreadPoolExecutor\n\nSLOT_ORDER = list(zip(SLOT_NAMES, SLOT_SPEC))\n\n\ndef study_stacks(study_id, dcm_root, slots, size):\n    \"\"\"Build the 2.5D image stack for every slot of one study.\n\n    Slots are decoded in parallel; each slot is an independent series directory,\n    so five threads is five concurrent DICOM decodes. Every failure is caught and\n    counted rather than raised, and an empty slot returns an empty stack that the\n    caller carries as a masked zero.\n\n    Args:\n        study_id: Study instance UID.\n        dcm_root: Directory holding the per-study series directories.\n        slots: Mapping study id -> slot name -> series id.\n        size: Input side length in pixels.\n\n    Returns:\n        Dict mapping slot index to (images, fractional positions, span mm,\n        files on disk).\n    \"\"\"\n    chosen = slots.get(str(study_id), {})\n\n    def one(args):\n        \"\"\"Decode a single slot, returning an empty stack on any failure.\"\"\"\n        i, (name, (plane, _flag, k)) = args\n        series_id = chosen.get(name)\n        if series_id is None or dcm_root is None:\n            bump(\"slots_empty\")\n            return i, (np.zeros((0, 3, size, size), np.uint8),\n                       np.zeros(0, np.float32), 0.0, 0)\n        folder = Path(dcm_root) / str(study_id) / str(series_id)\n        try:\n            return i, build_slot_stack(folder, plane, k, size)\n        except Exception:\n            bump(\"slots_empty\")\n            return i, (np.zeros((0, 3, size, size), np.uint8),\n                       np.zeros(0, np.float32), 0.0, 0)\n\n    with ThreadPoolExecutor(max_workers=min(IO_WORKERS, N_SLOTS)) as pool:\n        return dict(pool.map(one, list(enumerate(SLOT_ORDER))))\n\n\ndef probe_cost(ids, dcm_root, slots, size, n):\n    \"\"\"Time the decode and the GPU forward on a handful of studies.\n\n    Args:\n        ids: Candidate study ids.\n        dcm_root: DICOM root for that split.\n        slots: Slot assignment for that split.\n        size: Input side length.\n        n: How many studies to probe.\n\n    Returns:\n        Dict of measured per-study seconds, per-study image counts and the GPU\n        rate in images per second. Zeros when nothing could be read.\n    \"\"\"\n    use = [s for s in ids[:max(1, n)]]\n    t_cpu = t_gpu = 0.0\n    n_img = 0\n    for sid in use:\n        t0 = time.time()\n        stacks = study_stacks(sid, dcm_root, slots, size)\n        t_cpu += time.time() - t0\n        imgs = [v[0] for v in stacks.values() if v[0].shape[0]]\n        if not imgs:\n            continue\n        batch = np.concatenate(imgs, axis=0)\n        n_img += batch.shape[0]\n        t1 = time.time()\n        embed_images(batch)\n        t_gpu += time.time() - t1\n    k = max(1, len(use))\n    return {\n        \"studies\": len(use),\n        \"cpu_per_study\": t_cpu / k,\n        \"gpu_per_study\": t_gpu / k,\n        \"images_per_study\": n_img / k,\n        \"gpu_rate\": (n_img / t_gpu) if t_gpu > 1e-6 else 0.0,\n    }\n\n\ntest_ids = test[\"StudyInstanceUID\"].astype(str).tolist()\nprobe_root, probe_ids, probe_slots = TEST_DCM, test_ids, SLOTS_TEST\nif probe_root is None or not probe_ids:\n    probe_root, probe_ids, probe_slots = TRAIN_DCM, train_ids, SLOTS_TRAIN\n\nPROBE = {\"studies\": 0, \"cpu_per_study\": 0.0, \"gpu_per_study\": 0.0,\n         \"images_per_study\": 0.0, \"gpu_rate\": 0.0}\nif CAN_ENCODE and probe_root is not None:\n    PROBE = probe_cost(probe_ids, probe_root, probe_slots, IMG_SIZE, PROBE_STUDIES)\n\nprint(\"probe: {} studies at {} px\".format(PROBE[\"studies\"], IMG_SIZE))\nprint(\"  decode + preprocess {:.2f} s/study\".format(PROBE[\"cpu_per_study\"]))\nprint(\"  GPU forward         {:.2f} s/study  ({:.0f} images/s)\".format(\n    PROBE[\"gpu_per_study\"], PROBE[\"gpu_rate\"]))\nprint(\"  images encoded      {:.1f} per study\".format(PROBE[\"images_per_study\"]))\nprint()\n\nper_study = PROBE[\"cpu_per_study\"] + PROBE[\"gpu_per_study\"]\nN_TRAIN_USE = len(train_ids)\nSLICE_SCALE = 1.0\nnotes = []\n\ndef _projected_per_study():\n    \"\"\"Seconds per study under the current input size and slice scale.\n\n    Returns:\n        Float seconds. Zero when the probe measured nothing.\n    \"\"\"\n    return (PROBE[\"cpu_per_study\"] + PROBE[\"gpu_per_study\"]) * SLICE_SCALE\n\n\nif per_study > 0:\n    for _attempt in range(6):\n        per_study = _projected_per_study()\n        t_test = per_study * RERUN_TEST_STUDIES\n        t_train = per_study * N_TRAIN_USE\n        gpu_share = ((PROBE[\"gpu_per_study\"] / (PROBE[\"cpu_per_study\"]\n                                                + PROBE[\"gpu_per_study\"]))\n                     if per_study else 0.0)\n        if t_test <= TEST_ENCODE_BUDGET_S and t_train <= TRAIN_ENCODE_BUDGET_S:\n            break\n        if gpu_share > 0.5 and IMG_SIZE in SIZE_LADDER \\\n                and SIZE_LADDER.index(IMG_SIZE) < len(SIZE_LADDER) - 1:\n            old = IMG_SIZE\n            IMG_SIZE = SIZE_LADDER[SIZE_LADDER.index(IMG_SIZE) + 1]\n            # Token count scales with the square of the input size and attention adds\n            # a further quadratic term, so the square is the conservative estimate.\n            PROBE[\"gpu_per_study\"] *= (IMG_SIZE / float(old)) ** 2\n            notes.append(\"input size {} -> {} px, GPU was {:.0%} of the cost\".format(\n                old, IMG_SIZE, gpu_share))\n            continue\n        if SLICE_SCALE > 0.5:\n            SLICE_SCALE = max(0.5, SLICE_SCALE - 0.25)\n            notes.append(\"slice budget scaled to {:.0%}\".format(SLICE_SCALE))\n            continue\n        per_study = _projected_per_study()\n        N_TRAIN_USE = max(MIN_TRAIN_STUDIES,\n                          min(N_TRAIN_USE, int(TRAIN_ENCODE_BUDGET_S / per_study)))\n        notes.append(\"training studies capped at {}\".format(N_TRAIN_USE))\n        break\n\nif SLICE_SCALE < 1.0:\n    SLOT_SPEC = [(p, f, max(3, int(round(k * SLICE_SCALE)))) for p, f, k in SLOT_SPEC]\n    MAX_SLICES = max(k for _, _, k in SLOT_SPEC)\n    SLOT_ORDER = list(zip(SLOT_NAMES, SLOT_SPEC))\n\nprint(\"projected encode cost at {} px, {:.1f} images per study\".format(\n    IMG_SIZE, PROBE[\"images_per_study\"] * SLICE_SCALE))\nprint(\"  visible test  ({:5d} studies): {:8.0f} s\".format(\n    len(test_ids), per_study * len(test_ids)))\nprint(\"  rerun test    (~{:4d} studies): {:8.0f} s   budget {:.0f} s\".format(\n    RERUN_TEST_STUDIES, per_study * RERUN_TEST_STUDIES, TEST_ENCODE_BUDGET_S))\nprint(\"  train         ({:5d} studies): {:8.0f} s   budget {:.0f} s\".format(\n    N_TRAIN_USE, per_study * N_TRAIN_USE, TRAIN_ENCODE_BUDGET_S))\nprint(\"  head + submission             : {:8.0f} s   budget\".format(HEAD_BUDGET_S))\nprint(\"  total against the 9 h kernel  : {:8.0f} s of {:.0f} s\".format(\n    per_study * (RERUN_TEST_STUDIES + N_TRAIN_USE) + HEAD_BUDGET_S, WALL_BUDGET_S))\nif notes:\n    print()\n    print(\"downshifts applied, in order:\")\n    for n in notes:\n        print(\"  -\", n)\nelse:\n    print()\n    print(\"no downshift needed\")\n\n# The ladder has a floor: the test set is never subsampled and the training set\n# never goes below MIN_TRAIN_STUDIES. When the projection is still over budget\n# after the floor is reached, say so here rather than printing numbers that\n# quietly exceed the budget. The encode loop's own wall clock is what actually\n# stops it, and it stops train before test because test runs first.\nover_test = per_study * RERUN_TEST_STUDIES > TEST_ENCODE_BUDGET_S\nover_train = per_study * N_TRAIN_USE > TRAIN_ENCODE_BUDGET_S\nif over_test or over_train:\n    print()\n    print(\"STILL OVER BUDGET after every downshift ({}{}{}).\".format(\n        \"rerun test\" if over_test else \"\",\n        \" and \" if over_test and over_train else \"\",\n        \"train\" if over_train else \"\"))\n    print(\"The encode loop is bounded by its own wall clock and will truncate,\")\n    print(\"leaving the remaining studies as masked zeros. Test is encoded first,\")\n    print(\"so a truncated run still predicts every test study.\")\n\n# Gold studies go first so they survive any truncation, then a deterministic shuffle.\nrest = [s for s in train_ids if s not in set(gold_ids)]\n_rng = np.random.default_rng(SEED)\n_rng.shuffle(rest)\nTRAIN_ENCODE_IDS = (gold_ids + rest)[:N_TRAIN_USE]\nprint()\nprint(\"studies to encode: {} train (gold first), {} test\".format(\n    len(TRAIN_ENCODE_IDS), len(test_ids)))"},{"cell_type":"markdown","id":"c6f076d066ec","metadata":{},"source":"### The cache, and how to promote it\n\nEmbeddings are written to `/kaggle/working` as a float16 memory-mapped array plus a mask\nand an id list, tagged with the input size, the slot layout and the slice budget. A rerun\nwith the same tag reloads instead of recomputing, which is what makes head iteration\ncheap inside one session.\n\nTo split this into a train notebook and an inference notebook, which is how it should end\nup once the design settles:\n\n1. Run this notebook with `SAVE_TRAIN_CACHE = True`. The train embeddings and the fitted\n   head weights land in the kernel output.\n2. Publish that output as a Kaggle dataset with\n   `kaggle datasets create -p /path/to/output`.\n3. Add the dataset to the inference notebook's `dataset_sources`, set\n   `LOAD_TRAIN_CACHE_FROM_DATASET = True`, and the rerun then encodes only the ~1,300\n   test studies. That is roughly a fifth of the work, and it moves the training cost out\n   of the nine-hour submission window entirely.\n\nThe train cache is deleted before the notebook finishes unless `SAVE_TRAIN_CACHE` is on,\nbecause a 600 MB array in the kernel output slows every commit for no benefit until step\n2 is actually wanted."},{"cell_type":"code","execution_count":null,"id":"0a93c94c360b","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"PIX = {}\nPIXOK = {}\n\n# The pixel cache is what makes fine-tuning affordable, and it is also the one\n# artifact here big enough to fill the working directory and take the whole run\n# with it. Project it before allocating anything: one 2.5D image per slot per\n# study, uint8, for both splits. If it does not fit, the fine-tune is dropped and\n# the frozen path proceeds, which is the degradation that costs least.\nFT_BYTES_PER_STUDY = N_SLOTS * 3 * IMG_SIZE * IMG_SIZE\nif FINETUNE:\n    _ft_gb = (len(TRAIN_ENCODE_IDS) + len(test_ids)) * FT_BYTES_PER_STUDY / 1e9\n    print(\"pixel cache projected: {:.1f} GB for {} train + {} test studies\".format(\n        _ft_gb, len(TRAIN_ENCODE_IDS), len(test_ids)))\n    if _ft_gb > FT_DISK_BUDGET_GB:\n        FINETUNE = False\n        print(\"  over the {:.0f} GB budget; fine-tuning disabled, frozen path only\"\n              .format(FT_DISK_BUDGET_GB))\n\nSLOT_SIG = \"-\".join(str(k) for _, _, k in SLOT_SPEC)\n# Bumped whenever the pixels handed to the backbone change, which is the only\n# thing a cached embedding cannot tell you about itself. v2 is laterality\n# normalisation (section 6b). Without this bump a rerun would silently reload\n# embeddings encoded from un-normalised right knees and report them as new.\nPREP_VERSION = \"lat2\"\nCACHE_TAG = \"dinos{}_slots{}_{}_{}\".format(IMG_SIZE, SLOT_SIG, SLICE_DIM,\n                                           PREP_VERSION)\n\n\ndef cache_files(split):\n    \"\"\"Paths of the four cache artifacts for one split.\n\n    Args:\n        split: \"train\" or \"test\".\n\n    Returns:\n        Dict of name -> Path under the working directory.\n    \"\"\"\n    stem = \"emb_{}_{}\".format(split, CACHE_TAG)\n    f = {\n        \"emb\": WORK / (stem + \".npy\"),\n        \"mask\": WORK / (stem + \"_mask.npy\"),\n        \"pos\": WORK / (stem + \"_pos.npy\"),\n        \"meta\": WORK / (stem + \"_meta.npz\"),\n        \"ids\": WORK / (stem + \"_ids.csv\"),\n    }\n    if FINETUNE:\n        # Pixels, not embeddings. Fine-tuning revisits the same images every epoch\n        # and re-decoding them per epoch would make the epoch count a function of\n        # I/O rather than of learning. One 2.5D image per slot at uint8 costs\n        # N_SLOTS * 3 * IMG_SIZE^2 bytes per study, about 1.7 MB at 336 px.\n        f[\"pix\"] = WORK / (stem + \"_pix.npy\")\n        f[\"pixok\"] = WORK / (stem + \"_pixok.npy\")\n    return f\n\n\ndef load_cached(split, ids):\n    \"\"\"Reload a cached split when it exists and covers exactly these ids.\n\n    Args:\n        split: \"train\" or \"test\".\n        ids: The study ids this run wants, in order.\n\n    Returns:\n        Tuple of (emb, mask, pos, nfiles, used, span) or None.\n    \"\"\"\n    f = cache_files(split)\n    dataset_dir = None\n    if split == \"train\" and LOAD_TRAIN_CACHE_FROM_DATASET:\n        for d in _label_search_dirs():\n            if (d / f[\"emb\"].name).is_file():\n                dataset_dir = d\n                break\n    if dataset_dir is not None:\n        f = {k: dataset_dir / v.name for k, v in f.items()}\n    if not all(p.is_file() for p in f.values()):\n        return None\n    try:\n        cached_ids = pd.read_csv(f[\"ids\"])[\"StudyInstanceUID\"].astype(str).tolist()\n        if cached_ids != list(ids):\n            print(\"cache for\", split, \"covers different studies; recomputing\")\n            return None\n        meta = np.load(f[\"meta\"])\n        print(\"reusing cached embeddings for\", split, \"from\", f[\"emb\"].parent)\n        if FINETUNE:\n            PIX[split] = np.load(f[\"pix\"], mmap_mode=\"r\")\n            PIXOK[split] = np.load(f[\"pixok\"])\n        return (np.load(f[\"emb\"], mmap_mode=\"r\"), np.load(f[\"mask\"]),\n                np.load(f[\"pos\"]), meta[\"nfiles\"], meta[\"used\"], meta[\"span\"])\n    except Exception as exc:\n        print(\"cache for\", split, \"unreadable:\", exc)\n        return None\n\n\ndef encode_split(ids, dcm_root, slots, split, budget_s):\n    \"\"\"Encode one split into a cached embedding array.\n\n    The encoding is bounded by a wall clock. If the budget runs out the loop\n    stops and every remaining study is left as a masked zero, which the head\n    handles; a truncated run still produces a submission. Test is always encoded\n    before train for exactly this reason.\n\n    Args:\n        ids: Study ids to encode, in order.\n        dcm_root: Directory holding per-study series directories.\n        slots: Slot assignment for that split.\n        split: \"train\" or \"test\".\n        budget_s: Wall-clock seconds allowed.\n\n    Returns:\n        Tuple of (emb, mask, pos, nfiles, used, span).\n    \"\"\"\n    n = len(ids)\n    hit = load_cached(split, ids)\n    if hit is not None:\n        return hit\n\n    f = cache_files(split)\n    emb = np.lib.format.open_memmap(\n        f[\"emb\"], mode=\"w+\", dtype=np.float16, shape=(n, N_SLOTS, MAX_SLICES, SLICE_DIM))\n    mask = np.zeros((n, N_SLOTS, MAX_SLICES), np.bool_)\n    pos = np.zeros((n, N_SLOTS, MAX_SLICES), np.float16)\n    nfiles = np.zeros((n, N_SLOTS), np.int32)\n    used = np.zeros((n, N_SLOTS), np.int32)\n    span = np.zeros((n, N_SLOTS), np.float32)\n\n    pix = pixok = None\n    if FINETUNE:\n        pix = np.lib.format.open_memmap(\n            f[\"pix\"], mode=\"w+\", dtype=np.uint8,\n            shape=(n, N_SLOTS, 3, IMG_SIZE, IMG_SIZE))\n        pixok = np.zeros((n, N_SLOTS), np.bool_)\n        print(\"pixel cache for {}: {:.1f} GB at {}\".format(\n            split, pix.nbytes / 1e9, f[\"pix\"].name))\n        PIX[split] = pix\n        PIXOK[split] = pixok\n\n    if not CAN_ENCODE or dcm_root is None:\n        print(\"encode\", split, \": skipped, the imaging path is unavailable here\")\n        emb.flush()\n        pd.DataFrame({\"StudyInstanceUID\": ids}).to_csv(f[\"ids\"], index=False)\n        np.save(f[\"mask\"], mask); np.save(f[\"pos\"], pos)\n        np.savez(f[\"meta\"], nfiles=nfiles, used=used, span=span)\n        if pix is not None:\n            pix.flush()\n            np.save(f[\"pixok\"], pixok)\n        return emb, mask, pos, nfiles, used, span\n\n    t0 = time.time()\n    deadline = t0 + budget_s\n    done = 0\n    report_every = max(1, n // 20)\n    for i, sid in enumerate(ids):\n        if time.time() > deadline:\n            print(\"encode {}: budget of {:.0f}s spent after {} of {} studies\".format(\n                split, budget_s, i, n))\n            break\n        try:\n            stacks = study_stacks(sid, dcm_root, slots, IMG_SIZE)\n        except Exception:\n            bump(\"slots_empty\")\n            continue\n        batch, index = [], []\n        for si, (imgs, frac, sp, nf) in stacks.items():\n            nfiles[i, si] = nf\n            span[i, si] = sp\n            m = min(imgs.shape[0], MAX_SLICES)\n            if m == 0:\n                continue\n            used[i, si] = m\n            pos[i, si, :m] = frac[:m].astype(np.float16)\n            if pix is not None:\n                # The slot's centre image. imgs[c] is already the 2.5D triplet of\n                # three consecutive slices, so this is the same input the frozen\n                # path sees at that position, kept rather than thrown away after\n                # embedding. The pixels are in memory here; storing them costs no\n                # extra DICOM read.\n                pix[i, si] = imgs[m // 2]\n                pixok[i, si] = True\n            batch.append(imgs[:m])\n            index.extend([(si, j) for j in range(m)])\n        if not batch:\n            continue\n        try:\n            vecs = embed_images(np.concatenate(batch, axis=0))\n        except Exception as exc:\n            print(\"embed failed on study\", i, \":\", type(exc).__name__, exc)\n            bump(\"pixel_fail\")\n            continue\n        for (si, j), v in zip(index, vecs):\n            emb[i, si, j] = v.astype(np.float16)\n            mask[i, si, j] = True\n        done += 1\n        if (i + 1) % report_every == 0:\n            rate = (i + 1) / max(1e-6, time.time() - t0)\n            print(\"  {} {:5d}/{:5d}  {:.2f} studies/s  eta {:.0f}s\".format(\n                split, i + 1, n, rate, (n - i - 1) / max(rate, 1e-6)))\n\n    emb.flush()\n    pd.DataFrame({\"StudyInstanceUID\": ids}).to_csv(f[\"ids\"], index=False)\n    np.save(f[\"mask\"], mask)\n    np.save(f[\"pos\"], pos)\n    np.savez(f[\"meta\"], nfiles=nfiles, used=used, span=span)\n    if pix is not None:\n        pix.flush()\n        np.save(f[\"pixok\"], pixok)\n        print(\"  pixel cache: {} of {} studies carry at least one slot\".format(\n            int(pixok.any(axis=1).sum()), n))\n    print(\"encode {}: {} of {} studies in {:.0f}s -> {} ({:.0f} MB)\".format(\n        split, done, n, time.time() - t0, f[\"emb\"].name,\n        f[\"emb\"].stat().st_size / 1e6 if f[\"emb\"].is_file() else 0.0))\n    return emb, mask, pos, nfiles, used, span\n\n\n# Test first. A truncated run must still have something to predict with.\nTE = encode_split(test_ids, TEST_DCM, SLOTS_TEST, \"test\", TEST_ENCODE_BUDGET_S)\nTR = encode_split(TRAIN_ENCODE_IDS, TRAIN_DCM, SLOTS_TRAIN, \"train\",\n                  TRAIN_ENCODE_BUDGET_S)\nemb_te, mask_te, pos_te, nf_te, used_te, span_te = TE\nemb_tr, mask_tr, pos_tr, nf_tr, used_tr, span_tr = TR\ngc.collect()"},{"cell_type":"code","execution_count":null,"id":"da5db53089ef","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def coverage_report(name, mask, nfiles, used, span, man, ids):\n    \"\"\"Print what fraction of each study's volume was actually encoded.\n\n    Args:\n        name: Split name for the log.\n        mask: Boolean array [studies, slots, slices].\n        nfiles: Files on disk per selected series.\n        used: Slices encoded per selected series.\n        span: Millimetres spanned by each selected series.\n        man: Normalised manifest for that split.\n        ids: Study ids in row order.\n\n    Returns:\n        None.\n    \"\"\"\n    n_enc = mask.sum(axis=(1, 2))\n    have = n_enc > 0\n    if not have.any():\n        print(name, \": nothing was encoded\")\n        return\n    sel_files = nfiles.sum(axis=1)\n    sel_used = used.sum(axis=1)\n    ok = sel_files > 0\n    per_series = np.where(nfiles > 0, used / np.maximum(nfiles, 1), np.nan)\n\n    series_per_study = (man.groupby(\"StudyInstanceUID\").size()\n                        .reindex(ids).fillna(0).values if len(man) else np.zeros(len(ids)))\n    mean_len = np.where(used.sum(axis=1) > 0,\n                        nfiles.sum(axis=1) / np.maximum((nfiles > 0).sum(axis=1), 1), 0.0)\n    est_study_slices = series_per_study * mean_len\n\n    print(\"{}: {} of {} studies carry at least one encoded slice\".format(\n        name, int(have.sum()), len(ids)))\n    print(\"  slices encoded per study      : mean {:.1f}, median {:.0f}, max {}\".format(\n        n_enc[have].mean(), np.median(n_enc[have]), int(n_enc.max())))\n    if ok.any():\n        print(\"  share of the SELECTED series  : {:.1%}\".format(\n            sel_used[ok].sum() / max(1, sel_files[ok].sum())))\n    good = est_study_slices > 0\n    if good.any():\n        print(\"  share of the whole study      : {:.1%} (estimated from {:.1f} series \"\n              \"per study x {:.0f} slices per series)\".format(\n                  sel_used[good].sum() / max(1.0, est_study_slices[good].sum()),\n                  series_per_study[good].mean(), mean_len[good].mean()))\n    finite = span[span > 0]\n    if finite.size:\n        print(\"  physical extent per series    : {:.0f} mm scanned, sampled across \"\n              \"the plane window\".format(np.median(finite)))\n    slot_use = (mask.sum(axis=2) > 0).mean(axis=0)\n    print(\"  slots carrying data           : \" + \", \".join(\n        \"{} {:.0%}\".format(s, v) for s, v in zip(SLOT_NAMES, slot_use)))\n\n\ncoverage_report(\"train\", mask_tr, nf_tr, used_tr, span_tr, tr_man, TRAIN_ENCODE_IDS)\nprint()\ncoverage_report(\"test \", mask_te, nf_te, used_te, span_te, te_man, test_ids)\nprint()\nif CROP_RATIOS:\n    r = float(np.median(CROP_RATIOS))\n    print(\"median body crop keeps {:.0%} of the frame, so the effective pitch is \"\n          \"{:.3f} mm/px and one patch token covers {:.2f} mm\".format(\n              r, FOV_MM * r / IMG_SIZE, FOV_MM * r * PATCH_PX / IMG_SIZE))\nprint()\nprint(\"failure counters:\", \", \".join(\"{} {}\".format(k, v) for k, v in STATS.items()))\nprint(\"  series ordered exactly (every file on disk): {} | subsampled: {}\".format(\n    STATS[\"exact_order\"], STATS[\"fallback_order\"]))\n_lat_seen = (STATS[\"lat_left\"] + STATS[\"lat_right_mirrored\"]\n             + STATS[\"lat_right_reversed\"] + STATS[\"lat_absent\"])\n_lat_fixed = STATS[\"lat_right_mirrored\"] + STATS[\"lat_right_reversed\"]\nprint(\"  laterality: {} already left, {} right mirrored (coronal/axial), {} right \"\n      \"reversed (sagittal), {} unlabelled and left alone, {} self-contradictory\".format(\n          STATS[\"lat_left\"], STATS[\"lat_right_mirrored\"],\n          STATS[\"lat_right_reversed\"], STATS[\"lat_absent\"], STATS[\"lat_conflict\"]))\nif _lat_seen:\n    print(\"  {:.1%} of slots were normalised onto the left-knee convention; \"\n          \"{:.1%} carried no side and keep an unresolved axis\".format(\n              _lat_fixed / _lat_seen, STATS[\"lat_absent\"] / _lat_seen))"},{"cell_type":"markdown","id":"53ab3a03d2b9","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">8</span><span style=\"vertical-align:middle;\">Head and validation</span></h2></div>\n\nTwo heads, fitted on the cached embeddings, reported separately and blended.\n\n**A ridge probe.** Per study, per slot, the mean and the maximum over encoded slices of\nthe 1,152-dimensional slice vector, concatenated across the five slots, reduced by a\ntruncated SVD fitted on the training rows of each fold, then a multi-output `RidgeCV`.\nIt is the control: linear, closed-form, no capacity to invent structure.\n\n**A slice-attention head.** The same cached vectors, but the pooling is learned. Twelve\nqueries, one per finding, attend over the flattened slot-by-slice token set with the\npadding mask applied, so `Medial Meniscus` can weight the medial sagittal slices and\n`PF OA` can weight the axial ones. Fixed epochs, no early stopping on the fold being\nscored, three seeds averaged. A bare linear probe cannot express \"look at this slice\",\nwhich is the whole reason for encoding 43 slices instead of 3.\n\nThe two are blended at a fixed 0.5 by rank average. The full weight curve is printed for\ninformation and deliberately not tuned: choosing a blend weight on twelve labels' worth\nof out-of-fold estimates is how a number that does not survive contact with the\nleaderboard gets manufactured.\n\n**Folds.** `GroupKFold`, grouped by duplicate report text. Forty-nine report texts in\nthis corpus are shared by 183 studies, the largest block holding 37. Any extractor gives\nevery study in such a block an identical target vector, so a block split across folds\nscores the model on a target whose source it has already seen. Note what is *not* used\nhere: `PatientID`. Patients are 1:1 with studies in this corpus, 4,407 of 4,407, so\ngrouping by patient is a no-op, and the duplicate-report grouping is the one that does\nany work. Skipping the patient grouping also means the fold construction reads no DICOM\nheaders at all."},{"cell_type":"code","execution_count":null,"id":"a88c11d2fa0c","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"from sklearn.model_selection import GroupKFold\nfrom sklearn.metrics import roc_auc_score\n\n# Patients are 1:1 with studies in this corpus, 4,407 of 4,407, so a PatientID\n# grouping is a no-op and reading DICOM PatientID to build it would be pure cost.\n# The grouping that does work is duplicate report text: 49 texts are shared by 183\n# studies, the largest block holding 37, and any extractor gives every study in a\n# block the same twelve targets. A block split across folds scores the model on a\n# target whose source it has already seen.\nreport_by_study = {}\nif report_col is not None:\n    for _sid, _txt in zip(train[\"StudyInstanceUID\"].astype(str), train[report_col]):\n        report_by_study[_sid] = _txt.strip() if isinstance(_txt, str) else \"\"\n\nfold_ids = list(TRAIN_ENCODE_IDS)\ngroup_of = {}\nnext_group = 0\ntext_group = {}\nfor sid in fold_ids:\n    txt = report_by_study.get(sid, \"\")\n    if txt:\n        if txt not in text_group:\n            text_group[txt] = next_group\n            next_group += 1\n        group_of[sid] = text_group[txt]\n    else:\n        group_of[sid] = next_group\n        next_group += 1\ngroups = np.asarray([group_of[s] for s in fold_ids])\nsizes = pd.Series(groups).value_counts()\nn_blocks = int((sizes > 1).sum())\nprint(\"fold grouping: {} studies -> {} groups\".format(len(fold_ids), len(sizes)))\nprint(\"  {} duplicate-report blocks covering {} studies, largest block {}\".format(\n    n_blocks, int(sizes[sizes > 1].sum()), int(sizes.iloc[0])))\n\nn_splits = max(2, min(N_FOLDS, len(sizes)))\ngkf = GroupKFold(n_splits=n_splits)\nFOLDS = list(gkf.split(np.zeros(len(fold_ids)), groups=groups))\nfor k, (tr_i, va_i) in enumerate(FOLDS):\n    shared = set(groups[tr_i]) & set(groups[va_i])\n    assert not shared, \"fold {} leaks {} groups across the split\".format(k, len(shared))\nprint(\"  {} folds, no group appears on both sides of any split\".format(len(FOLDS)))"},{"cell_type":"code","execution_count":null,"id":"2f6869ebc459","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def pooled_features(emb, mask):\n    \"\"\"Collapse the per-slice embeddings to one vector per study.\n\n    Mean and maximum over the encoded slices of each slot, concatenated across\n    slots. The maximum is kept for the same reason it is kept over patch tokens:\n    a finding visible on two slices out of twelve is averaged away and is not\n    maximised away.\n\n    Args:\n        emb: float16 array [studies, slots, slices, dim].\n        mask: boolean array [studies, slots, slices].\n\n    Returns:\n        float32 array [studies, slots * 2 * dim].\n    \"\"\"\n    n = emb.shape[0]\n    out = np.zeros((n, N_SLOTS * 2 * SLICE_DIM), np.float32)\n    step = max(1, 256)\n    for start in range(0, n, step):\n        stop = min(n, start + step)\n        block = np.asarray(emb[start:stop], dtype=np.float32)\n        m = mask[start:stop][..., None].astype(np.float32)\n        cnt = np.maximum(m.sum(axis=2), 1.0)\n        mean = (block * m).sum(axis=2) / cnt\n        neg = np.where(m > 0, block, -1e4)\n        mx = neg.max(axis=2)\n        mx = np.where(m.sum(axis=2) > 0, mx, 0.0)\n        out[start:stop] = np.concatenate([mean, mx], axis=2).reshape(stop - start, -1)\n    return out\n\n\ndef manifest_features(ids, man, slots, mask):\n    \"\"\"Per-study features that both splits always carry.\n\n    Everything here comes from the series manifest, which the hidden test set\n    ships, plus the encode mask. Nothing is read from a report and nothing is\n    read from a DICOM header, so train and test go through one code path with no\n    chance of a block existing on one side only.\n\n    Args:\n        ids: Study ids in row order.\n        man: Normalised manifest.\n        slots: Slot assignment.\n        mask: Encode mask [studies, slots, slices].\n\n    Returns:\n        Tuple of (float32 array [studies, n_features], feature names).\n    \"\"\"\n    by_study = {s: g for s, g in man.groupby(\"StudyInstanceUID\")} if len(man) else {}\n    rows = []\n    for i, sid in enumerate(ids):\n        g = by_study.get(sid)\n        total = len(g) if g is not None else 0\n        counts = [float((g[\"plane\"] == p).sum()) if g is not None else 0.0\n                  for p in PLANES]\n        fluid = float(g[\"fluid\"].mean()) if g is not None and total else 0.0\n        present = [1.0 if n in slots.get(sid, {}) else 0.0 for n in SLOT_NAMES]\n        enc = (mask[i].sum(axis=1) / np.maximum(\n            np.asarray([k for _, _, k in SLOT_SPEC], np.float32), 1.0)).tolist()\n        rows.append([math.log1p(total)] + counts + [fluid] + present + enc)\n    names = ([\"log_n_series\"] + [\"n_\" + p for p in PLANES] + [\"fluid_frac\"]\n             + [\"has_\" + n for n in SLOT_NAMES] + [\"enc_\" + n for n in SLOT_NAMES])\n    return np.asarray(rows, np.float32), names\n\n\nenc_rate_tr = float((mask_tr.sum(axis=(1, 2)) > 0).mean()) if len(TRAIN_ENCODE_IDS) else 0.0\nenc_rate_te = float((mask_te.sum(axis=(1, 2)) > 0).mean()) if len(test_ids) else 0.0\nEMB_OK = (enc_rate_tr >= 0.5 and enc_rate_te >= 0.5\n          and abs(enc_rate_tr - enc_rate_te) <= 0.25)\nprint(\"studies with embeddings: train {:.1%} | test {:.1%}\".format(\n    enc_rate_tr, enc_rate_te))\nif not EMB_OK:\n    print(\"coverage is absent or lopsided, so every embedding column is dropped and\")\n    print(\"the head falls back to the series manifest, which both sides always have.\")\n\nMETA_TR, META_NAMES = manifest_features(TRAIN_ENCODE_IDS, tr_man, SLOTS_TRAIN, mask_tr)\nMETA_TE, _ = manifest_features(test_ids, te_man, SLOTS_TEST, mask_te)\n\nif EMB_OK:\n    t0 = time.time()\n    POOL_TR = pooled_features(emb_tr, mask_tr)\n    POOL_TE = pooled_features(emb_te, mask_te)\n    print(\"pooled features: train {} | test {} | {:.1f}s\".format(\n        POOL_TR.shape, POOL_TE.shape, time.time() - t0))\nelse:\n    POOL_TR = np.zeros((len(TRAIN_ENCODE_IDS), 0), np.float32)\n    POOL_TE = np.zeros((len(test_ids), 0), np.float32)\n\nY = Y_soft.reindex(TRAIN_ENCODE_IDS)[LABELS].astype(np.float32).values\nW = np.asarray([GOLD_WEIGHT if s in set(gold_ids) else 1.0\n                for s in TRAIN_ENCODE_IDS], np.float32)\nY_HARD = (Y >= 0.5).astype(int)\nassert Y.shape[0] == POOL_TR.shape[0] == META_TR.shape[0], \\\n    \"label rows drifted from feature rows\"\nprint(\"target matrix  :\", Y.shape, \"| positives at 0.5:\", int(Y_HARD.sum()))"},{"cell_type":"code","execution_count":null,"id":"02d0bdad932f","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"from sklearn.decomposition import TruncatedSVD\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.preprocessing import StandardScaler\n\n# The control head. Linear, closed form, no capacity to invent structure. The SVD\n# is refitted inside every fold on the training rows alone, so no validation row\n# contributes to the basis the model is expressed in.\nRIDGE_ALPHAS = np.logspace(0.0, 5.0, 11)\n\n\ndef fit_ridge_fold(x_tr, y_tr, w_tr, x_va, x_te):\n    \"\"\"Fit the ridge probe on one fold and predict the validation and test rows.\n\n    Args:\n        x_tr: Training design matrix.\n        y_tr: Soft targets [n, 12].\n        w_tr: Per-row weights.\n        x_va: Validation design matrix.\n        x_te: Test design matrix.\n\n    Returns:\n        Tuple of (validation predictions, test predictions).\n    \"\"\"\n    if x_tr.shape[1] == 0:\n        return (np.zeros((x_va.shape[0], len(LABELS))),\n                np.zeros((x_te.shape[0], len(LABELS))))\n    comps = int(min(SVD_COMPONENTS, max(2, min(x_tr.shape) - 1)))\n    if x_tr.shape[1] > comps:\n        svd = TruncatedSVD(n_components=comps, random_state=SEED)\n        a = svd.fit_transform(x_tr)\n        b, c = svd.transform(x_va), svd.transform(x_te)\n    else:\n        a, b, c = x_tr, x_va, x_te\n    sc = StandardScaler().fit(a)\n    a, b, c = sc.transform(a), sc.transform(b), sc.transform(c)\n    model = RidgeCV(alphas=RIDGE_ALPHAS)\n    model.fit(a, y_tr, sample_weight=w_tr)\n    return model.predict(b), model.predict(c)\n\n\nX_TR = np.hstack([POOL_TR, META_TR]) if POOL_TR.shape[1] else META_TR\nX_TE = np.hstack([POOL_TE, META_TE]) if POOL_TE.shape[1] else META_TE\nassert X_TR.shape[1] == X_TE.shape[1], \"train/test feature width mismatch\"\n\nt0 = time.time()\noof_ridge = np.zeros((len(TRAIN_ENCODE_IDS), len(LABELS)), np.float32)\nte_ridge = np.zeros((len(test_ids), len(LABELS)), np.float32)\nfor k, (tr_i, va_i) in enumerate(FOLDS):\n    try:\n        p_va, p_te = fit_ridge_fold(X_TR[tr_i], Y[tr_i], W[tr_i], X_TR[va_i], X_TE)\n    except Exception as exc:\n        print(\"ridge fold\", k, \"failed:\", type(exc).__name__, exc)\n        p_va = np.tile(Y[tr_i].mean(axis=0), (len(va_i), 1))\n        p_te = np.tile(Y[tr_i].mean(axis=0), (len(test_ids), 1))\n    oof_ridge[va_i] = p_va\n    te_ridge += p_te / len(FOLDS)\nprint(\"ridge probe: {} folds in {:.1f}s, {} features\".format(\n    len(FOLDS), time.time() - t0, X_TR.shape[1]))"},{"cell_type":"code","execution_count":null,"id":"8d959fcc262a","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# The class statement itself evaluates its base, so on a machine without torch the\n# whole cell would raise NameError before anything could degrade. Binding the base\n# conditionally keeps the cell importable everywhere; HAVE_ATTN below decides\n# whether it is ever instantiated.\n_HeadBase = nn.Module if HAVE_TORCH else object\n\n\nclass SliceAttentionHead(_HeadBase):\n    \"\"\"Twelve learned queries pooling over the slot-by-slice token set.\n\n    A linear probe on a pooled study vector cannot express \"look at this slice\",\n    which is the entire reason for encoding 43 slices rather than 3. One query\n    per finding lets Medial Meniscus weight the medial sagittal slices while\n    PF OA weights the axial ones, with the padding mask applied so an empty slot\n    contributes nothing rather than contributing a zero vector.\n\n    Args:\n        d_in: Per-slice embedding width.\n        n_slots: Number of slots.\n        n_meta: Width of the manifest feature block.\n        n_labels: Number of findings.\n        hidden: Model width.\n        heads: Attention heads.\n        dropout: Dropout rate.\n    \"\"\"\n\n    def __init__(self, d_in, n_slots, n_meta, n_labels, hidden=HEAD_HIDDEN,\n                 heads=4, dropout=HEAD_DROPOUT):\n        if not HAVE_TORCH:\n            raise RuntimeError(\"the attention head needs torch\")\n        super().__init__()\n        self.norm_in = nn.LayerNorm(d_in)\n        self.proj = nn.Linear(d_in, hidden)\n        self.slot_emb = nn.Parameter(torch.zeros(n_slots, hidden))\n        self.pos_proj = nn.Linear(1, hidden)\n        self.queries = nn.Parameter(torch.randn(n_labels, hidden) * 0.02)\n        self.attn = nn.MultiheadAttention(hidden, heads, dropout=dropout,\n                                          batch_first=True)\n        self.norm_q = nn.LayerNorm(hidden)\n        self.drop = nn.Dropout(dropout)\n        self.out_w = nn.Parameter(torch.zeros(n_labels, hidden))\n        self.out_b = nn.Parameter(torch.zeros(n_labels))\n        self.meta = nn.Linear(n_meta, n_labels) if n_meta else None\n        nn.init.normal_(self.out_w, std=0.02)\n\n    def forward(self, x, mask, pos, meta):\n        \"\"\"Score one batch of studies.\n\n        Args:\n            x: float tensor [B, slots, slices, d_in].\n            mask: bool tensor [B, slots, slices], True where a slice exists.\n            pos: float tensor [B, slots, slices], fractional stack position.\n            meta: float tensor [B, n_meta].\n\n        Returns:\n            Logit tensor [B, n_labels].\n        \"\"\"\n        b, s, k, _ = x.shape\n        h = self.proj(self.norm_in(x))\n        h = h + self.slot_emb[None, :, None, :]\n        h = h + self.pos_proj(pos.unsqueeze(-1))\n        h = torch.nn.functional.gelu(h).reshape(b, s * k, -1)\n        m = mask.reshape(b, s * k)\n        # A study with no encoded slice at all would make every key padded and the\n        # attention softmax would return NaN. Let such a study attend to one zero\n        # token instead; its prediction then comes from the meta block alone.\n        empty = ~m.any(dim=1)\n        if empty.any():\n            m = m.clone()\n            m[empty, 0] = True\n        q = self.queries.unsqueeze(0).expand(b, -1, -1)\n        pooled, _ = self.attn(q, h, h, key_padding_mask=~m, need_weights=False)\n        pooled = self.drop(self.norm_q(pooled + q))\n        logits = (pooled * self.out_w[None]).sum(-1) + self.out_b\n        if self.meta is not None:\n            logits = logits + self.meta(meta)\n        return logits\n\n\ndef run_attention_head():\n    \"\"\"Fit the attention head across folds and seeds.\n\n    Fixed epochs, no early stopping on the fold being scored, because stopping on\n    the fold you then report is how an out-of-fold estimate gets inflated.\n\n    Returns:\n        Tuple of (out-of-fold predictions, test predictions).\n    \"\"\"\n    oof = np.zeros((len(TRAIN_ENCODE_IDS), len(LABELS)), np.float32)\n    tep = np.zeros((len(test_ids), len(LABELS)), np.float32)\n    dev = torch.device(DEVICE)\n    y_t = torch.from_numpy(Y)\n    w_t = torch.from_numpy(W)\n    meta_tr_t = torch.from_numpy(META_TR)\n    meta_te_t = torch.from_numpy(META_TE).to(dev)\n    n_seeds = len(HEAD_SEEDS)\n\n    def batch_tensors(source_emb, source_mask, source_pos, idx):\n        \"\"\"Gather one batch out of the memory-mapped cache.\"\"\"\n        idx = np.sort(np.asarray(idx))\n        x = torch.from_numpy(np.asarray(source_emb[idx], dtype=np.float32)).to(dev)\n        m = torch.from_numpy(source_mask[idx]).to(dev)\n        p = torch.from_numpy(source_pos[idx].astype(np.float32)).to(dev)\n        return idx, x, m, p\n\n    for k, (tr_i, va_i) in enumerate(FOLDS):\n        for seed in HEAD_SEEDS:\n            torch.manual_seed(seed + k)\n            model = SliceAttentionHead(SLICE_DIM, N_SLOTS, META_TR.shape[1],\n                                       len(LABELS)).to(dev)\n            opt = torch.optim.AdamW(model.parameters(), lr=HEAD_LR,\n                                    weight_decay=HEAD_WD)\n            steps = max(1, math.ceil(len(tr_i) / 64)) * HEAD_EPOCHS\n            sched = torch.optim.lr_scheduler.OneCycleLR(\n                opt, max_lr=HEAD_LR, total_steps=steps, pct_start=0.25)\n            rng = np.random.default_rng(seed + k)\n            model.train()\n            for _ep in range(HEAD_EPOCHS):\n                order = rng.permutation(tr_i)\n                for start in range(0, len(order), 64):\n                    sel = order[start:start + 64]\n                    idx, x, m, p = batch_tensors(emb_tr, mask_tr, pos_tr, sel)\n                    logit = model(x, m, p, meta_tr_t[idx].to(dev))\n                    loss = torch.nn.functional.binary_cross_entropy_with_logits(\n                        logit, y_t[idx].to(dev),\n                        weight=w_t[idx].to(dev).unsqueeze(1))\n                    opt.zero_grad(set_to_none=True)\n                    loss.backward()\n                    torch.nn.utils.clip_grad_norm_(model.parameters(), 5.0)\n                    opt.step()\n                    sched.step()\n            model.eval()\n            with torch.no_grad():\n                for start in range(0, len(va_i), 128):\n                    idx, x, m, p = batch_tensors(emb_tr, mask_tr, pos_tr,\n                                                 va_i[start:start + 128])\n                    oof[idx] += torch.sigmoid(\n                        model(x, m, p, meta_tr_t[idx].to(dev))).cpu().numpy() / n_seeds\n                for start in range(0, len(test_ids), 128):\n                    sel = np.arange(start, min(len(test_ids), start + 128))\n                    idx, x, m, p = batch_tensors(emb_te, mask_te, pos_te, sel)\n                    tep[idx] += torch.sigmoid(\n                        model(x, m, p, meta_te_t[idx])).cpu().numpy() \\\n                        / (n_seeds * len(FOLDS))\n        print(\"  attention head: fold {} of {} done\".format(k + 1, len(FOLDS)))\n    return oof, tep\n\n\noof_attn = np.zeros_like(oof_ridge)\nte_attn = np.zeros_like(te_ridge)\nHAVE_ATTN = bool(HAVE_TORCH and EMB_OK and len(TRAIN_ENCODE_IDS) > 50)\nif HAVE_ATTN:\n    t0 = time.time()\n    try:\n        oof_attn, te_attn = run_attention_head()\n        print(\"attention head: {} folds x {} seeds in {:.1f}s\".format(\n            len(FOLDS), len(HEAD_SEEDS), time.time() - t0))\n    except Exception as exc:\n        print(\"attention head failed, falling back to the ridge probe alone:\",\n              type(exc).__name__, exc)\n        HAVE_ATTN = False\nelse:\n    print(\"attention head skipped: torch {} | embeddings usable {}\".format(\n        HAVE_TORCH, EMB_OK))"},{"cell_type":"markdown","id":"8eabb28e7a76","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"vertical-align:middle;\">7b. Changing what the encoder looks with</span></h2></div>\n\nSection 7 froze the backbone and gave a reason: 58 gold studies, a report-derived\nsignal of unknown fidelity, and a 22M parameter transformer that would mostly learn the\nnoise. Half of that reason does not survive reading the code above it. The head does not\ntrain on 58 rows. It trains on every encoded study, several thousand of them, carrying\nreport-derived targets, with the 58 gold rows upweighted eightfold in the loss. Fitting\na transformer to thousands of noisy labels is a different proposition from fitting it to\n58 clean ones, and the argument for freezing quietly used the smaller number.\n\nThe other half of the reason stands. A frozen encoder makes embeddings a one-time cost,\nand every head afterwards runs in seconds. That is why the frozen path stays.\n\nWhat a frozen encoder cannot do is change the vocabulary it looks with. Resolution,\nslice count, slot design and pooling all change how much the model looks and how it\nsummarises what it saw; none of them adds a feature the encoder does not already have.\nDINOv2 learned its features from natural images, where nothing resembles what a torn\nmeniscus does to a proton-density sequence. So the ceiling those axes run into is the\nrepresentation, and the only way to test that claim is to move it.\n\n**The trade this section makes is coverage for adaptation.** The frozen path encodes 46\nslices per study and cannot change its features. This path takes one 2.5D image per\nslot, five per study, and opens the last four transformer blocks. Fewer looks, deeper\nones. The two are then rank-blended, so a study is scored both by a shallow model that\nsaw all of it and by a deep model that saw the middle of each slot.\n\nThree restraints, each with a reason:\n\n**Only the trailing blocks move.** Early blocks are generic edge and texture filters.\nThere may not be enough supervision here to improve them and there is certainly enough\nto damage them.\n\n**The encoder learns two orders of magnitude slower than the head.** The head is random\nand has everything to learn; the encoder starts from a good solution and needs only to\nbe nudged off it. One learning rate would either leave the head untrained or destroy the\nencoder inside the first few hundred steps.\n\n**Pixels are cached, not re-read.** Fine-tuning revisits the same images every epoch, and\nre-decoding them per epoch would make the epoch count a function of disk rather than of\nlearning. The cache is written during the encoding pass that already decodes those\npixels, so it costs no additional DICOM read: one 2.5D image per slot as `uint8`, about\n1.7 MB per study at 336 px.\n\nOne fold is held out, not five. Five fine-tunes do not fit the wall clock. That is a\nbudget decision rather than a statistical one and it is named here rather than buried:\nthe fold score below ranks this head against the frozen ones on identical rows, and it\nis not an estimate of the leaderboard. The blend weight is fixed beforehand for the same\nreason the frozen blend weight is. Selecting it on the single fold that measures it\nwould manufacture a number that does not survive the rerun.\n\nIf any of this fails, the run keeps the frozen submission it had already earned."},{"cell_type":"code","execution_count":null,"id":"bd21b89fc888","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"FT_OK = False\nFT_NOTE = \"not attempted\"\noof_ft = np.zeros((len(TRAIN_ENCODE_IDS), len(LABELS)), np.float32)\nte_ft = np.zeros((len(test_ids), len(LABELS)), np.float32)\nft_val_rows = np.zeros(0, dtype=int)\n\n\ndef _macro_auc_rows(y_true, scores):\n    \"\"\"Macro AUC over labels that carry both classes on these rows.\n\n    Args:\n        y_true: Integer array [n, 12].\n        scores: Score array [n, 12].\n\n    Returns:\n        Tuple of (per-label list with nan where undefined, macro mean).\n    \"\"\"\n    per = []\n    for j in range(y_true.shape[1]):\n        col = y_true[:, j]\n        if col.min() == col.max():\n            per.append(float(\"nan\"))\n            continue\n        per.append(float(roc_auc_score(col, scores[:, j])))\n    finite = [v for v in per if v == v]\n    return per, (float(np.mean(finite)) if finite else float(\"nan\"))\n\n\ndef open_last_blocks(backbone, n_blocks):\n    \"\"\"Freeze the backbone, then reopen its trailing blocks and final norm.\n\n    The early blocks of a vision transformer are generic edge and texture\n    filters and the late ones carry semantics. There is not obviously enough\n    supervision here to improve the early ones and there is certainly enough to\n    damage them, so the line is drawn near the top and the encoder's learning\n    rate is held two orders below the head's.\n\n    Args:\n        backbone: A HuggingFace Dinov2Model.\n        n_blocks: How many trailing blocks to train.\n\n    Returns:\n        Tuple of (number of trainable parameters, number of blocks in the model).\n    \"\"\"\n    for p in backbone.parameters():\n        p.requires_grad_(False)\n    layers = backbone.encoder.layer\n    for blk in layers[max(0, len(layers) - n_blocks):]:\n        for p in blk.parameters():\n            p.requires_grad_(True)\n    for p in backbone.layernorm.parameters():\n        p.requires_grad_(True)\n    n_train = sum(p.numel() for p in backbone.parameters() if p.requires_grad)\n    return n_train, len(layers)\n\n\nclass SlotFineTuner(_HeadBase):\n    \"\"\"The backbone plus twelve label queries pooling over the five slots.\n\n    One 2.5D image per slot goes through the encoder, and twelve learned queries\n    attend over the five resulting slot vectors. An absent slot is masked out\n    rather than contributing a zero vector. This is deliberately the same pooling\n    idea as the frozen attention head, so that the comparison between them is a\n    comparison of frozen against adapted features and not of two architectures.\n\n    Args:\n        backbone: The DINOv2 model, already opened for partial training.\n        d_in: Width of the per-image feature, CLS concatenated with patch mean.\n        n_slots: Number of slots.\n        n_labels: Number of findings.\n        width: Model width after projection.\n        heads: Attention heads.\n        dropout: Dropout rate.\n    \"\"\"\n\n    def __init__(self, backbone, d_in, n_slots, n_labels, width=FT_WIDTH,\n                 heads=4, dropout=FT_DROPOUT):\n        if not HAVE_TORCH:\n            raise RuntimeError(\"fine-tuning needs torch\")\n        super().__init__()\n        self.backbone = backbone\n        self.proj = nn.Linear(d_in, width)\n        self.slot_emb = nn.Parameter(torch.zeros(n_slots, width))\n        self.norm = nn.LayerNorm(width)\n        self.queries = nn.Parameter(torch.randn(n_labels, width) * 0.02)\n        self.attn = nn.MultiheadAttention(width, heads, dropout=dropout,\n                                          batch_first=True)\n        self.drop = nn.Dropout(dropout)\n        self.out = nn.Linear(width, 1)\n\n    def forward(self, x, slot_ok):\n        \"\"\"Score one batch of studies.\n\n        Args:\n            x: Float tensor [B, S, 3, H, W], already normalised.\n            slot_ok: Bool tensor [B, S], True where the slot carries an image.\n\n        Returns:\n            Logit tensor [B, n_labels].\n        \"\"\"\n        b, s = x.shape[0], x.shape[1]\n        h = self.backbone(pixel_values=x.reshape(b * s, *x.shape[2:]),\n                          **DINO_KW).last_hidden_state\n        feat = torch.cat([h[:, 0], h[:, 1:].mean(1)], dim=1).float()\n        f = self.proj(feat).reshape(b, s, -1) + self.slot_emb.unsqueeze(0)\n        f = self.norm(f)\n        q = self.queries.unsqueeze(0).expand(b, -1, -1)\n        pooled, _ = self.attn(q, f, f, key_padding_mask=~slot_ok)\n        return self.out(self.drop(pooled)).squeeze(-1)\n\n\ndef _ft_batch(split, rows, dev):\n    \"\"\"Gather one batch of cached pixels and normalise it for the encoder.\n\n    Args:\n        split: \"train\" or \"test\".\n        rows: Row indices into the cache.\n        dev: Torch device.\n\n    Returns:\n        Tuple of (pixel tensor [B, S, 3, H, W], slot mask [B, S]).\n    \"\"\"\n    rows = np.sort(np.asarray(rows))\n    raw = np.asarray(PIX[split][rows], dtype=np.float32) / 255.0\n    ok = PIXOK[split][rows].copy()\n    # A study with no slot at all would make every key masked, and attention over\n    # an entirely masked key set is a NaN rather than a zero. Give it slot 0,\n    # whose pixels are zeros, so it produces a prediction near the prior.\n    dead = ~ok.any(axis=1)\n    if dead.any():\n        ok[dead, 0] = True\n    x = torch.from_numpy(raw).to(dev)\n    mean = torch.from_numpy(IMAGENET_MEAN).to(dev).unsqueeze(0)\n    std = torch.from_numpy(IMAGENET_STD).to(dev).unsqueeze(0)\n    x = (x - mean) / std\n    return x, torch.from_numpy(ok).to(dev)\n\n\ndef _ft_predict(model, split, n_rows, dev, batch, rows=None):\n    \"\"\"Run the fine-tuned model over selected rows of a split.\n\n    Rows outside the selection are left at zero rather than predicted. For the\n    training split that is deliberate: this model saw four folds of it, so a\n    prediction on those rows is a fit, not an estimate, and an array that quietly\n    mixed the two would be misread as out-of-fold the first time anyone used it.\n\n    Args:\n        model: The trained model.\n        split: \"train\" or \"test\".\n        n_rows: Number of rows in that split.\n        dev: Torch device.\n        batch: Studies per forward pass.\n        rows: Row indices to predict, or None for all of them.\n\n    Returns:\n        Float array [n_rows, n_labels] of probabilities, zero where not predicted.\n    \"\"\"\n    model.eval()\n    out = np.zeros((n_rows, len(LABELS)), np.float32)\n    todo = np.arange(n_rows) if rows is None else np.sort(np.asarray(rows))\n    with torch.no_grad():\n        for start in range(0, len(todo), batch):\n            sel = todo[start:start + batch]\n            x, ok = _ft_batch(split, sel, dev)\n            with torch.autocast(device_type=\"cuda\", dtype=torch.float16,\n                                enabled=(DEVICE == \"cuda\")):\n                logit = model(x, ok)\n            out[sel] = torch.sigmoid(logit.float()).cpu().numpy()\n    return out\n\n\ndef run_finetune():\n    \"\"\"Fine-tune the trailing blocks on one grouped fold and predict test.\n\n    One fold, not five. Five folds of a fine-tune do not fit the wall clock, and\n    a fold count chosen to fit a budget rather than to answer a question is worth\n    naming rather than hiding. The held-out fold is scored against the same rows\n    the frozen blend is scored on, so the two numbers printed here are comparable.\n\n    Returns:\n        Tuple of (out-of-fold predictions on the held-out rows, test predictions,\n        held-out row indices, a one-line note for the log).\n    \"\"\"\n    dev = torch.device(DEVICE)\n    torch.manual_seed(FT_SEED)\n\n    tr_i, va_i = FOLDS[FT_VAL_FOLD]\n    has_pix = PIXOK[\"train\"].any(axis=1)\n    tr_i = np.asarray([i for i in tr_i if has_pix[i]], dtype=int)\n    va_i = np.asarray([i for i in va_i if has_pix[i]], dtype=int)\n    if len(tr_i) < 50 or len(va_i) < 20:\n        return None, None, None, \"too few studies carry cached pixels\"\n\n    n_train, n_layer = open_last_blocks(DINO, FT_UNFREEZE)\n    d_in = int(getattr(DINO.config, \"hidden_size\", EMBED_DIM)) * 2\n    model = SlotFineTuner(DINO, d_in, N_SLOTS, len(LABELS)).to(dev)\n\n    head_params = [p for n, p in model.named_parameters()\n                   if not n.startswith(\"backbone.\") and p.requires_grad]\n    enc_params = [p for p in model.backbone.parameters() if p.requires_grad]\n    # An empty parameter group raises, and FT_UNFREEZE = 0 is a legitimate way to\n    # ask for the head alone as a control arm.\n    groups = [{\"params\": head_params, \"lr\": FT_HEAD_LR}]\n    if enc_params:\n        groups.insert(0, {\"params\": enc_params, \"lr\": FT_ENCODER_LR})\n    opt = torch.optim.AdamW(groups, weight_decay=FT_WD)\n    # torch.amp.GradScaler is the current spelling and torch.cuda.amp.GradScaler\n    # the older one. Kaggle's image has moved between them before, and a run that\n    # dies here loses a fine-tune that was otherwise ready to go.\n    try:\n        scaler = torch.amp.GradScaler(\"cuda\", enabled=(DEVICE == \"cuda\"))\n    except (AttributeError, TypeError):\n        scaler = torch.cuda.amp.GradScaler(enabled=(DEVICE == \"cuda\"))\n    lossf = nn.BCEWithLogitsLoss(reduction=\"none\")\n\n    y_t = torch.from_numpy(Y).to(dev)\n    w_t = torch.from_numpy(W).to(dev)\n\n    print(\"fine-tune: {} of {} blocks open, {:,} trainable backbone params\".format(\n        FT_UNFREEZE, n_layer, n_train))\n    print(\"  train rows {} | held-out rows {} | fold {} of {}\".format(\n        len(tr_i), len(va_i), FT_VAL_FOLD, len(FOLDS)))\n\n    batch = FT_STUDY_BATCH\n    t0 = time.time()\n    steps_per_epoch = max(1, len(tr_i) // batch)\n    epochs = FT_EPOCHS\n    measured = None\n\n    for ep in range(epochs):\n        model.train()\n        order = np.random.RandomState(FT_SEED + ep).permutation(len(tr_i))\n        run_loss, seen = 0.0, 0\n        for s in range(steps_per_epoch):\n            rows = tr_i[order[s * batch:(s + 1) * batch]]\n            if len(rows) == 0:\n                continue\n            x, ok = _ft_batch(\"train\", rows, dev)\n            idx = torch.from_numpy(np.sort(rows)).to(dev)\n            with torch.autocast(device_type=\"cuda\", dtype=torch.float16,\n                                enabled=(DEVICE == \"cuda\")):\n                logit = model(x, ok)\n            loss_el = lossf(logit.float(), y_t[idx])\n            loss = (loss_el.mean(dim=1) * w_t[idx]).sum() / w_t[idx].sum()\n            opt.zero_grad(set_to_none=True)\n            scaler.scale(loss).backward()\n            scaler.step(opt)\n            scaler.update()\n            run_loss += float(loss.detach()) * len(rows)\n            seen += len(rows)\n\n            if measured is None and s == 9:\n                # Ten steps in, the per-step cost is known. Decide how many epochs\n                # actually fit rather than discovering at epoch six that they did\n                # not, which is how a run ends with no submission.\n                per_step = (time.time() - t0) / 10.0\n                fit = int(max(1, (FT_BUDGET_S * 0.8) / max(per_step * steps_per_epoch,\n                                                           1e-6)))\n                measured = per_step\n                if fit < epochs:\n                    print(\"  budget: {:.2f}s/step projects {:.0f}s per epoch; \"\n                          \"cutting {} epochs to {}\".format(\n                              per_step, per_step * steps_per_epoch, epochs, fit))\n                    epochs = fit\n\n        if time.time() - t0 > FT_BUDGET_S:\n            print(\"  budget of {:.0f}s spent after epoch {}\".format(FT_BUDGET_S, ep + 1))\n            break\n        print(\"  epoch {}/{}  loss {:.4f}  {:.0f}s elapsed\".format(\n            ep + 1, epochs, run_loss / max(seen, 1), time.time() - t0))\n        if ep + 1 >= epochs:\n            break\n\n    # Held-out rows only. Predicting the four folds this model trained on would\n    # cost a further forward pass over most of the corpus and produce numbers that\n    # cannot be compared with anything.\n    val = _ft_predict(model, \"train\", len(TRAIN_ENCODE_IDS), dev, batch, rows=va_i)\n    tep = _ft_predict(model, \"test\", len(test_ids), dev, batch)\n    note = \"{} epochs, {:.0f}s\".format(min(epochs, FT_EPOCHS), time.time() - t0)\n    return val, tep, va_i, note\n\n\nif FINETUNE and HAVE_TORCH and CAN_ENCODE and DINO is not None \\\n        and \"train\" in PIX and \"test\" in PIX and PIXOK[\"train\"].any():\n    try:\n        _val, _te, _rows, FT_NOTE = run_finetune()\n        if _val is not None:\n            oof_ft, te_ft, ft_val_rows = _val, _te, _rows\n            FT_OK = True\n    except Exception as exc:\n        # A failed fine-tune must not cost the submission the frozen path already\n        # earned. It is reported and skipped, never allowed to raise.\n        FT_NOTE = \"failed: {}: {}\".format(type(exc).__name__, exc)\n        print(\"fine-tune failed, keeping the frozen blend only:\", FT_NOTE)\nelse:\n    FT_NOTE = \"skipped: fine-tuning needs torch, a backbone and a pixel cache\"\n\nprint(\"fine-tune:\", FT_NOTE)\n\nif FT_OK and len(ft_val_rows):\n    # The same rows for every estimate, so the three numbers are comparable. The\n    # frozen heads' predictions on these rows are genuinely out of fold: fold\n    # FT_VAL_FOLD was a validation fold for them too.\n    _yv = Y_HARD[ft_val_rows]\n    _, _m_ridge = _macro_auc_rows(_yv, oof_ridge[ft_val_rows])\n    _, _m_ft = _macro_auc_rows(_yv, oof_ft[ft_val_rows])\n    print(\"held-out fold {}, {} studies\".format(FT_VAL_FOLD, len(ft_val_rows)))\n    print(\"  frozen ridge      {:.4f}\".format(_m_ridge))\n    if HAVE_ATTN:\n        _, _m_attn = _macro_auc_rows(_yv, oof_attn[ft_val_rows])\n        print(\"  frozen attention  {:.4f}\".format(_m_attn))\n    print(\"  fine-tuned        {:.4f}\".format(_m_ft))\n    print(\"Every number here is measured against extracted labels, so all of them \"\n          \"carry the same ceiling. This ranks the heads against each other on one \"\n          \"fold. It is not an estimate of the leaderboard.\")"},{"cell_type":"markdown","id":"cccd745fd0fe","metadata":{},"source":"### What the out-of-fold numbers do and do not measure\n\nThe headline out-of-fold macro AUC is computed **against the extracted labels**. Both\nsides of that comparison come from the same extractor, so it answers \"can the images\npredict what the report extractor said\" and nothing else. It is a smoke test with a\nceiling set by extraction quality, not an estimate of the leaderboard.\n\nThe second number is macro AUC **on the 58 gold studies**. That is the right target and\nthe wrong sample size. A per-label AUC on 58 studies with 9 to 35 positives carries a\nstandard error near 0.07; the macro average of twelve such estimates is better and still\nnowhere near the resolution needed to separate two models. It is a sanity check on\ndirection, not a ranking tool.\n\nNeither number selects a model. Model ranking in this competition has to come from a\nreport-derived validation set far larger than 58, which is the next thing this pipeline\nneeds."},{"cell_type":"code","execution_count":null,"id":"76542f78fe0a","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"from scipy.stats import rankdata\n\n\ndef rank01(a):\n    \"\"\"Rank-transform each column onto [0, 1].\n\n    Args:\n        a: 2D array of scores.\n\n    Returns:\n        Array of the same shape with per-column ranks scaled to [0, 1].\n    \"\"\"\n    out = np.zeros_like(a, dtype=np.float64)\n    for j in range(a.shape[1]):\n        r = rankdata(a[:, j], method=\"average\")\n        out[:, j] = (r - 1.0) / max(1.0, len(r) - 1.0)\n    return out\n\n\ndef macro_auc(y_true, scores):\n    \"\"\"Macro-average AUC over the twelve labels, skipping degenerate ones.\n\n    Args:\n        y_true: Integer array [n, 12].\n        scores: Score array [n, 12].\n\n    Returns:\n        Tuple of (per-label list with NaN where undefined, macro average).\n    \"\"\"\n    per = []\n    for j in range(y_true.shape[1]):\n        yt = y_true[:, j]\n        if len(np.unique(yt)) < 2:\n            per.append(np.nan)\n            continue\n        try:\n            per.append(float(roc_auc_score(yt, scores[:, j])))\n        except Exception:\n            per.append(np.nan)\n    vals = [v for v in per if np.isfinite(v)]\n    return per, (float(np.mean(vals)) if vals else np.nan)\n\n\nblend_oof = BLEND_W * rank01(oof_ridge) + (1.0 - BLEND_W) * rank01(oof_attn) \\\n    if HAVE_ATTN else rank01(oof_ridge)\nblend_test = BLEND_W * rank01(te_ridge) + (1.0 - BLEND_W) * rank01(te_attn) \\\n    if HAVE_ATTN else rank01(te_ridge)\n\n# The fine-tuned model joins the TEST blend only. Its out-of-fold coverage is one\n# fold, so folding it into blend_oof would compare a five-fold estimate against a\n# one-fold one and report the mixture as if it were the shipped model. The weight\n# is fixed rather than selected: choosing it on the one fold that measures it is\n# how a number stops surviving contact with the leaderboard.\nfrozen_test = blend_test.copy()\nif FT_OK:\n    blend_test = ((1.0 - FT_BLEND_W) * rank01(frozen_test)\n                  + FT_BLEND_W * rank01(te_ft))\n    print(\"test blend: frozen {:.2f} / fine-tuned {:.2f}, ranks not probabilities\"\n          .format(1.0 - FT_BLEND_W, FT_BLEND_W))\n    print(\"blend_oof below describes the FROZEN heads only; the fine-tuned model \"\n          \"is in the submission but not in that out-of-fold number.\")\nelse:\n    print(\"test blend: frozen heads only ({})\".format(FT_NOTE))\n\nper_ridge, m_ridge = macro_auc(Y_HARD, oof_ridge)\nper_attn, m_attn = macro_auc(Y_HARD, oof_attn) if HAVE_ATTN else ([np.nan] * 12, np.nan)\nper_blend, m_blend = macro_auc(Y_HARD, blend_oof)\n\nscores = pd.DataFrame({\n    \"label\": LABELS,\n    \"positives\": Y_HARD.sum(axis=0),\n    \"ridge\": per_ridge,\n    \"attention\": per_attn,\n    \"blend\": per_blend,\n})\nprint(\"Out-of-fold macro AUC against the EXTRACTED labels\")\nprint(\"(both sides come from the same extractor: a smoke test, not a leaderboard estimate)\")\nprint()\nprint(scores.round(4).to_string(index=False))\nprint()\nprint(\"macro  ridge {:.4f} | attention {:.4f} | blend {:.4f}\".format(\n    m_ridge, m_attn, m_blend))\nprint()\n\n# The weight curve is printed for information and deliberately not used to pick a\n# weight. Choosing a blend weight on twelve labels' worth of out-of-fold estimates\n# is how a number that does not survive contact with the leaderboard is manufactured.\nif HAVE_ATTN:\n    curve = []\n    for w in (0.0, 0.25, 0.5, 0.75, 1.0):\n        _, m = macro_auc(Y_HARD, w * rank01(oof_ridge) + (1 - w) * rank01(oof_attn))\n        curve.append((w, m))\n    print(\"blend weight curve (ridge share):\",\n          \", \".join(\"{:.2f}->{:.4f}\".format(w, m) for w, m in curve))\n    print(\"shipped weight is fixed at {:.2f}, not selected from this curve\".format(BLEND_W))\n    print()\n\n# The second number: gold studies only. Right target, wrong sample size.\ngold_rows = [i for i, s in enumerate(TRAIN_ENCODE_IDS) if s in set(gold_ids)]\nif len(gold_rows) >= 10 and N_GOLD:\n    gold_truth = (train.set_index(train[\"StudyInstanceUID\"].astype(str))\n                  .loc[[TRAIN_ENCODE_IDS[i] for i in gold_rows], LABELS]\n                  .astype(int).values)\n    _, g_ridge = macro_auc(gold_truth, oof_ridge[gold_rows])\n    _, g_blend = macro_auc(gold_truth, blend_oof[gold_rows])\n    print(\"Out-of-fold macro AUC on the {} GOLD studies: ridge {:.4f} | blend {:.4f}\"\n          .format(len(gold_rows), g_ridge, g_blend))\n    print(\"A per-label AUC on this many studies carries a standard error near 0.07.\")\n    print(\"It is a sanity check on direction. It cannot rank two models.\")\nelse:\n    print(\"too few gold studies in the encoded set for the gold check\")"},{"cell_type":"code","execution_count":null,"id":"66a68f87957a","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# Two panels, stacked, one label order shared so a finding can be tracked straight\n# down the figure. Upper: where the model beats chance. Lower: how many positives\n# each estimate rests on, because an AUC on 40 positives and an AUC on 2,000 are\n# not the same measurement.\norder = scores.sort_values(\"blend\", ascending=True).reset_index(drop=True)\nypos = np.arange(len(order))\n\nfig = plt.figure(figsize=(13.2, 11.6))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.15, 0.9], hspace=0.32)\n\nax0 = fig.add_subplot(gs[0, 0])\nvals = order[\"blend\"].values.astype(float)\nax0.barh(ypos, np.nan_to_num(vals, nan=0.5), height=0.66, color=PRIMARY, zorder=3)\nlabel_bars(ax0, vals, ypos, fmt=\"{:.3f}\", pad=0.006)\nif HAVE_ATTN:\n    ax0.plot(order[\"ridge\"].values.astype(float), ypos, linestyle=\"none\",\n             marker=\"D\", markersize=7, color=ACCENT, zorder=5, label=\"ridge probe alone\")\n    ax0.legend(loc=\"lower right\", labelcolor=INK_2)\nax0.axvline(0.5, color=AXIS, linewidth=1.4, zorder=2)\n# The chance rule is annotated below the lowest bar, not above the highest: at the\n# top it lands on top of the leading bar and hides the number that matters most.\nax0.text(0.5, -0.95, \" chance\", ha=\"left\", va=\"center\", fontsize=11.5, color=MUTED)\nax0.set_ylim(-1.35, len(order) - 0.4)\nax0.set_yticks(ypos)\nax0.set_yticklabels(order[\"label\"].tolist(), fontsize=12.5, color=INK)\nlo = float(np.nanmin(np.append(vals, 0.5)))\nax0.set_xlim(min(0.45, lo - 0.03), 1.02)\nax0.set_xlabel(\"out-of-fold AUC against the extracted labels\")\nstyle(ax0, \"x\")\nax0.spines[\"left\"].set_visible(False)\n\nax1 = fig.add_subplot(gs[1, 0])\npos = order[\"positives\"].values.astype(float)\nax1.barh(ypos, pos, height=0.66, color=NEUTRAL, zorder=3)\nlabel_bars(ax1, pos, ypos, fmt=\"{:,.0f}\", pad=max(1.0, pos.max() * 0.012))\nax1.set_yticks(ypos)\nax1.set_yticklabels(order[\"label\"].tolist(), fontsize=12.5, color=INK)\nax1.set_xlim(0, max(1.0, pos.max()) * 1.18)\nax1.set_xlabel(\"extracted positives the estimate above rests on\")\nstyle(ax1, \"x\")\nax1.spines[\"left\"].set_visible(False)\n\nfinish(fig,\n       \"Out-of-fold AUC per finding, and the evidence each number rests on\",\n       detail=\"Blend of the ridge probe and the slice-attention head at a fixed \"\n              \"0.5 rank weight. Macro average {:.4f}.\".format(m_blend),\n       caption=\"Targets are extractor output, not radiologist labels, so this \"\n               \"measures agreement with the extractor and carries its ceiling. \"\n               \"Folds are grouped by duplicate report text; no group appears on \"\n               \"both sides of a split.\")"},{"cell_type":"markdown","id":"70c43b87f9c1","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">9</span><span style=\"vertical-align:middle;\">Submission</span></h2></div>\n\nColumn order comes from `sample_submission.csv` rather than being assumed. Row identity\ncomes from `test.csv`, so the file covers exactly the studies the grader expects in the\norder it expects them. Row count, column set, finiteness and range are all asserted\nbefore the file is written.\n\nThe degradation path matters more here than in a metadata notebook, because there are\nmany more ways for an imaging pipeline to come up empty. If the DINOv2 weights are not\nmounted, if Torch is missing, if the DICOM directories are absent, or if the encode phase\nruns out of budget, the notebook still writes a valid submission from the series manifest\nalone. A weak submission is a score; an exception is not."},{"cell_type":"code","execution_count":null,"id":"a33c619350d2","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"sub_cols = [\"StudyInstanceUID\"] + LABELS\nif list(sample_sub.columns[:1]) == [\"StudyInstanceUID\"] and \\\n        set(LABELS).issubset(sample_sub.columns):\n    sub_cols = list(sample_sub.columns)\n\npred = np.clip(np.asarray(blend_test, dtype=float), 0.0, 1.0)\nif not np.isfinite(pred).all():\n    print(\"non-finite predictions replaced by the per-label training prior\")\n    prior = Y.mean(axis=0)\n    bad = ~np.isfinite(pred)\n    pred[bad] = np.take(prior, np.where(bad)[1])\n\nsub = pd.DataFrame(pred, columns=LABELS)\nsub.insert(0, \"StudyInstanceUID\", test_ids)\nsub = sub[sub_cols]\n\nassert len(sub) == len(test), \"row count does not match test.csv\"\nassert list(sub.columns) == sub_cols, \"column order does not match the sample submission\"\nassert sub[\"StudyInstanceUID\"].tolist() == test[\"StudyInstanceUID\"].astype(str).tolist(), \\\n    \"row identity or order drifted from test.csv\"\nassert np.isfinite(sub[LABELS].values).all(), \"non-finite value in the submission\"\nassert (sub[LABELS].values >= 0).all() and (sub[LABELS].values <= 1).all(), \\\n    \"probability out of range\"\n\nout_path = WORK / \"submission.csv\"\nsub.to_csv(out_path, index=False)\nprint(\"wrote\", out_path, \"|\", len(sub), \"rows x\", len(sub.columns), \"columns\")\nprint()\nprint(sub.head().to_string(index=False))\nprint()\nprint(\"per-label mean predicted probability\")\nprint(sub[LABELS].mean().round(4).to_string())\n\n# The train cache is 600 MB and slows every commit until it is actually wanted as a\n# dataset. Delete it unless promotion is on.\nif not SAVE_TRAIN_CACHE:\n    removed = 0\n    for p in cache_files(\"train\").values():\n        try:\n            if p.is_file() and p.parent == WORK:\n                p.unlink()\n                removed += 1\n        except Exception:\n            pass\n    if removed:\n        print()\n        print(\"removed {} train cache files from the output; set SAVE_TRAIN_CACHE=True\"\n              \" to keep them for promotion to a dataset\".format(removed))"},{"cell_type":"markdown","id":"a8fd90e5959b","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">10</span><span style=\"vertical-align:middle;\">What this is not</span></h2></div>\n\nThe backbone is frozen and was pretrained on natural images. DINOv2 features transfer\nwell to medical imaging but they are not a knee model, and nothing in this notebook\nadapts them to MR. Fine-tuning is the obvious next lever and it is deliberately held back\nuntil the label source is settled, because fine-tuning against noisy pseudo-labels\nteaches the noise.\n\nThe training targets are extracted from reports. Whatever the extractor gets wrong is\nbaked into everything downstream, and the audit in section 3 puts the honest ceiling on\nwhat a model trained on them can reach.\n\nThe resolution argument is arithmetic about sampling, and arithmetic is not a leaderboard\nscore. It says the frontier's input cannot represent the finding; it does not by itself\nsay that 336 px converts into AUC. That claim needs a submission, and this notebook has\nnot been scored at the time of writing.\n\nThe next three things worth doing, in the order they pay: replace the extracted labels\nwith the audited LLM extraction across all 4,349 unlabelled studies; build a\nreport-derived validation set large enough to rank two models; then split train and\ninference so the nine-hour budget buys inference only. The recovered time should go into\na larger backbone or more slices, not into a larger input, because section 2 measured\nthat 448 px buys nothing over 336."},{"cell_type":"markdown","id":"1f39ffcd8900","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"vertical-align:middle;\">How this notebook was made</span></h2></div>\n\nThis notebook was generated by an agentic system rather than typed by hand. The system is\nClaude Code running a custom multi-agent LLM harness: one agent drafted the pipeline while\nseparate adversarial verifier agents tried to break it and sent each defect back to the\nbuilder script instead of patching the notebook. The verifier agents are the part that\nmatters, because a notebook that has never been run is a draft.\n\nThe source of truth is a Python builder script, `notebooks/build_dino_notebook.py` in the\nworking repository, which emits this `.ipynb` through `nbformat`. Review happens on plain\nPython, and the notebook is regenerated rather than edited. The rule label extractor is\nnot copied into this builder; it is imported from the baseline notebook's builder module,\nso the two notebooks cannot drift apart. Publication uses `kaggle kernels push` and\nsubmission uses the same CLI, as with the companion EDA and baseline notebooks.\n\nThe competition facts this notebook is built on were verified from the Kaggle API and\nfrom the data itself on 2026-08-05, the day the competition opened: the 150 mm field of\nview across four acquisition matrices, the identity of `Fluid_Sensitive` and\n`Fat_Suppression` across all 24,371 series, the six-slot fill rates, the 49 report texts\nshared by 183 studies, and the 1:1 patient-to-study mapping. The slot fill rates, the\nduplicate-report blocks and the resolution arithmetic are all recomputed at run time in\nthe cells above rather than quoted from that check.\n\nThis notebook has no leaderboard score at the time of writing and claims none."},{"cell_type":"code","execution_count":null,"id":"b4cf129fd2e8","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"import platform\n\nprint(\"python      \", platform.python_version())\nfor mod in [\"numpy\", \"pandas\", \"sklearn\", \"scipy\", \"matplotlib\", \"pydicom\",\n            \"torch\", \"transformers\", \"cv2\"]:\n    try:\n        print(\"{:12s}\".format(mod), __import__(mod).__version__)\n    except Exception as exc:\n        print(\"{:12s} unavailable ({})\".format(mod, type(exc).__name__))\nprint()\nprint(\"input root        \", ROOT)\nprint(\"seed              \", SEED)\nprint(\"input size        \", IMG_SIZE, \"px ->\",\n      \"{:.3f} mm/px, {:.2f} mm per token\".format(\n          FOV_MM / IMG_SIZE, FOV_MM * PATCH_PX / IMG_SIZE))\nprint(\"slots / budgets   \", \", \".join(\n    \"{}={}\".format(n, k) for n, (_, _, k) in zip(SLOT_NAMES, SLOT_SPEC)))\nprint(\"label source      \", LABEL_SOURCE, \"->\", LABEL_PROVENANCE)\nprint(\"backbone          \", DINO_INFO)\nprint(\"cache tag         \", CACHE_TAG)\nprint(\"studies encoded   \", \"train {} / test {}\".format(\n    int((mask_tr.sum(axis=(1, 2)) > 0).sum()), int((mask_te.sum(axis=(1, 2)) > 0).sum())))\nprint(\"slices decoded    \", STATS[\"slices_decoded\"])\nprint(\"failures          \", \", \".join(\n    \"{} {}\".format(k, v) for k, v in STATS.items() if k.endswith(\"fail\")\n    or k.endswith(\"missing\") or k == \"slots_empty\"))"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"mimetype":"text/x-python","name":"python"}},"nbformat":4,"nbformat_minor":5}