{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# FixRes head — extract once, answer properly, ship a head\n\n**Up to ~5 hours, resumable.** Writes `head_384.pt`, not `submission.csv`.\n\n## What the A/B left unfinished\n\nFive seeds gave **+0.0029, se 0.0012** — t = 2.43 on 4 df, so the 95% interval\nis **[−0.0004, +0.0063]** and still contains zero. Suggestive, not established.\nThe direction was predicted in advance from the mechanism, which makes the\none-sided p of 0.036 defensible, but five seeds is not many.\n\nIt also exposed the real constraint: **88 minutes for 900 studies**, 5.9 s each,\nalmost entirely DICOM decode. Re-decoding for every new question is what makes\nprogress expensive.\n\n## So this decodes once\n\n| phase | what | cost |\n|---|---|---|\n| 1 | HI and LO features → **memmap on disk**, resumable | hours |\n| 2 | **15 seeds** on the cached features | minutes |\n| 3 | fine-tune the deployable head at 384, save it | minutes |\n\nFifteen seeds instead of five moves `t.975` from **2.776** to **2.145**, so the\nsame effect size clears the bar with room to spare. If it still does not clear,\nthat is a real answer rather than low power.\n\n**Resumable matters.** If the session dies at hour four, re-running continues\nfrom the last saved study instead of starting over. The features and their index\nare flushed to `/kaggle/working` as they are produced.\n\n> **Turn on Settings → Persistence → \"Variables and Files\"** before running.\n> Without it `/kaggle/working` is wiped between sessions and the resume has\n> nothing to resume from, which turns a crash at hour four into four lost hours.\n> Alternatively, attach the previous version's output as an input dataset.\n\n## What it still cannot claim\n\nLabels are the report teachers, so this measures *fits the teacher better*, not\n*better on gold*. And the CoAtNet arm carries about half the blend, so a gain\nhere reaches the leaderboard roughly **halved** — this route points at\n0.936–0.937, not 0.94.\n","metadata":{}},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\nCOATNET_RENDER_IMG = 384      # decode at native; the LO arm re-blurs in RAM\nCOATNET_SHARPNESS = 1.0       # set per-arm during extraction\n\nSUBSET_STUDIES = 2500         # cap. Decode measured at 5.9 s/study.\nFEATURE_BUDGET_S = 4.5 * 3600 # stop extracting here and use what was saved\nSEEDS = list(range(15))       # 15 -> t.975 = 2.145 instead of 2.776\nVAL_FRACTION = 0.30\nHEAD_EPOCHS = 60\nHEAD_LR = 3e-4\n\nFEAT_DIR = \"/kaggle/working\"  # features are flushed here so a re-run resumes\nRESUME = True                 # needs Settings -> Persistence -> Variables and\n                              # Files, otherwise /kaggle/working is wiped between\n                              # sessions and there is nothing to resume from\n\nprint(\"FixRes head | cap %d studies | budget %.1f h | %d seeds\"\n      % (SUBSET_STUDIES, FEATURE_BUDGET_S / 3600, len(SEEDS)))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# PREFLIGHT -- seconds, before any heavy import. Names exactly what is missing.\nimport os as _os, glob as _glob\nfrom pathlib import Path as _Path\n\n_LABELS12 = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n             \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n             \"Contusion\", \"Fracture\"]\n\n\ndef _nk(x):\n    return \"\".join(ch for ch in str(x).lower() if ch.isalnum())\n\n\nprint(\"=\" * 78)\nprint(\"PREFLIGHT\")\nprint(\"=\" * 78)\n_missing = []\n\n# 1) competition data, including the per-series plane metadata\n_comp = None\nfor _d, _sub, _f in _os.walk(\"/kaggle/input\"):\n    _sub[:] = [x for x in _sub if x not in (\"train_series\", \"test_series\")]\n    if \"train.csv\" in _f and \"train_series.csv\" in _f:\n        _comp = _d\n        break\nif _comp and _os.path.isdir(_comp + \"/train_series\"):\n    print(\"  competition data   : OK   %s\" % _comp)\nelse:\n    print(\"  competition data   : MISSING\")\n    _missing.append(\n        \"The competition itself. Add Input -> Competitions -> \"\n        \"'RSNA Knee Abnormality Detection'. Needs train.csv, train_series.csv \"\n        \"and the train_series/ image folder.\")\n\n# 2) the CoAtNet checkpoint this notebook encodes with\n_ck = None\nfor _d, _sub, _f in _os.walk(\"/kaggle/input\"):\n    _sub[:] = [x for x in _sub if x not in (\"train_series\", \"test_series\")]\n    if \"raptor_ft_coatnet_v5_full_swa.pt\" in _f:\n        _ck = _d + \"/raptor_ft_coatnet_v5_full_swa.pt\"\n        break\nif _ck:\n    print(\"  CoAtNet checkpoint : OK   %s (%d MB)\"\n          % (_ck, _os.path.getsize(_ck) // (1024 * 1024)))\nelse:\n    print(\"  CoAtNet checkpoint : MISSING\")\n    _missing.append(\n        \"raptor_ft_coatnet_v5_full_swa.pt -- the dataset \"\n        \"'dreaddevelopment/raptor-knee-maxspan'. This notebook encodes with \"\n        \"that backbone; it cannot substitute another.\")\n\n# 3) report-teacher labels, recognised by columns rather than filename\n_lab_hits = []\nfor _d, _sub, _f in _os.walk(\"/kaggle/input\"):\n    _sub[:] = [x for x in _sub if x not in (\"train_series\", \"test_series\")]\n    for _n in _f:\n        _l = _n.lower()\n        if not _l.endswith(\".csv\"):\n            continue\n        if _l in (\"train.csv\", \"test.csv\", \"train_series.csv\", \"test_series.csv\"):\n            continue\n        if any(k in _l for k in (\"submission\", \"sample_sub\", \"oof\")):\n            continue\n        try:\n            import pandas as _pd\n            _h = _pd.read_csv(_Path(_d) / _n, nrows=3)\n        except Exception:\n            continue\n        _keys = {_nk(c) for c in _h.columns}\n        if not ({\"studyinstanceuid\", \"studyid\", \"uid\"} & _keys):\n            continue\n        if sum(1 for t in _LABELS12 if _nk(t) in _keys) >= 6:\n            _lab_hits.append(_n)\nif _lab_hits:\n    print(\"  teacher labels     : OK   %d table(s): %s\"\n          % (len(_lab_hits), \", \".join(sorted(_lab_hits)[:4])))\nelse:\n    print(\"  teacher labels     : MISSING\")\n    _missing.append(\n        \"A per-study label table. train.csv holds only the 58 gold studies, \"\n        \"which cannot support a held-out split. Add the dataset carrying \"\n        \"llm_labels_v2.csv / llm_labels_v4_blend.csv / report_labels_v2.csv \"\n        \"(any CSV with StudyInstanceUID plus the twelve finding columns, \"\n        \"covering ~4,349 studies). This is the SAME dataset the Synovitis and \"\n        \"label-audit notebooks used.\")\n\n# 4) accelerator\ntry:\n    import torch as _t\n    _ng = _t.cuda.device_count()\n    print(\"  GPU                : %s\" % (\"OK   %d device(s)\" % _ng if _ng else \"NONE\"))\n    if _ng == 0:\n        _missing.append(\"A GPU. Settings -> Accelerator -> GPU T4 x2.\")\nexcept Exception:\n    print(\"  GPU                : could not query torch\")\n\nprint(\"=\" * 78)\nif _missing:\n    print(\"ATTACH THESE, THEN RE-RUN\")\n    for _i, _m in enumerate(_missing, 1):\n        print(\"  %d. %s\" % (_i, _m))\n    print(\"=\" * 78)\n    raise SystemExit(\n        \"preflight failed: %d input(s) missing -- see the list above. Nothing \"\n        \"was computed, so nothing was wasted.\" % len(_missing))\nprint(\"preflight OK -- every required input is attached\")\nprint(\"=\" * 78)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# COATNET_TRANSFORMER_BLEND_V1\n#!/usr/bin/env python3\n\"\"\"Knee MRI: twelve findings from a single model\n\nThis notebook takes a knee MRI study and scores twelve findings at once: ACL tear, MCL tear,\nmedial and lateral meniscus tears, osteoarthritis in the medial, lateral and patellofemoral\ncompartments, joint effusion, synovitis, a Baker's cyst, bone contusion and fracture. It scores\n0.924 on the public leaderboard using one model, with no ensembling and no test-time augmentation.\n\nThis is the inference half of the work. The model was trained separately and its weights are\nattached as a dataset, so this notebook only loads them and predicts:\nhttps://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-widedense\n\nWhere the training labels came from\n\nWorth saying up front, because it shapes everything else. The competition gives you 4,407 studies\nbut structured labels for only 58 of them. Every other study arrives with a free-text radiology\nreport and nothing more, so there is very little to train against out of the box.\n\nThe labels behind these weights were made by reading those reports with a language model and\nturning each into twelve probabilities rather than twelve yes or no answers. A report that says a\ntear is suspected becomes a number near 0.8, not a 1, which is a fairer target than forcing every\nhedged sentence into a hard label. That yields 4,349 studies to train on. The 58 studies that came\nwith real labels were never trained on and are used to check the result honestly; the model reaches\n0.9167 macro-AUC on them.\n\nBuilding a fixed input from studies that are all shaped differently\n\nThe hard part of this competition is not the network, it is that no two studies look alike. A\nstudy holds several DICOM series shot in different planes, the number of series varies, and the\nnumber of slices in a series varies more. Anything that expects a fixed-size input has to be given\none.\n\nThe approach here is to fill five fixed slots per study, always in the same order, for a stack of\n64 images:\n\n  18 slices from a sagittal series, preferring a fluid-sensitive one\n  14 slices from a second sagittal series, preferring one that is not fluid-sensitive\n  12 slices from a coronal series, preferring a fluid-sensitive one\n   8 slices from a second coronal series\n  12 slices from an axial series\n\nPreferring a fluid-sensitive series for some slots and not for others is deliberate. Fluid-\nsensitive sequences show swelling, effusion and acute injury clearly, while the other sequences\nshow anatomy and cartilage better, and the twelve findings are split across both. If a study has\nno series for a slot, the slot is left as zeros and the model is told to skip it rather than being\nfed something misleading.\n\nWithin a series, slices are taken evenly across 6 to 94 percent of the stack rather than from the\nmiddle. The outer slices are where the collateral ligaments and the lateral meniscus sit, and\ncutting them was measurably costing accuracy on exactly those findings.\n\nEvery slice is cropped to a 140 mm box around the centre of the image using the pixel spacing from\nthe DICOM header, then resized to 336 pixels. Cropping by millimetres rather than by pixel count\nmatters: it means a knee occupies the same fraction of the frame whether the scan was acquired at\n0.3 or 0.5 mm per pixel, so the model is not asked to learn scale differences that carry no medical\ninformation.\n\nHow the model reads the stack\n\nThree neighbouring slices are stacked into the three channels of one image. The network then sees\na little of what lies above and below the slice in the middle, which is most of the benefit of a 3D\nmodel at the cost of a 2D one. Each of these three-slice windows is passed through a CoAtNet\nbackbone at 384 pixels.\n\nThe windows are combined with an attention layer that has separate weights for each of the twelve\nfindings. This is the part that matters most. A cruciate tear may be visible on two sagittal slices\nwhile osteoarthritis is spread across many coronal ones, and a single pooled score forces those two\nto share one notion of which slices are important. Giving each finding its own attention weights\nlets each one draw on the slices that actually show it.\n\nRunning it\n\nScoring uses 42 windows per study. Inference runs in half precision and automatically retries a\nstudy in full precision if it fails, so no study is ever dropped from the submission. The notebook\nneeds no internet: the backbone is loaded from the attached weights rather than downloaded.\n\"\"\"\nimport os, sys, glob, time, json, gc\nfrom concurrent.futures import ThreadPoolExecutor\nos.environ.setdefault(\"HF_HUB_OFFLINE\", \"1\")\nos.environ.setdefault(\"TRANSFORMERS_OFFLINE\", \"1\")\nos.environ.setdefault(\"HF_HUB_DISABLE_TELEMETRY\", \"1\")\nimport numpy as np\nimport torch, torch.nn as nn, torch.nn.functional as F\nimport timm\n# T4 (Turing) cuDNN v9 has fp16/fp32 conv engines but NOT bf16 for these shapes\n# (\"GET was unable to find an engine...\"); benchmark lets it pick a valid algo for\n# the fixed (1,24,3,res,res) input.\ntorch.backends.cudnn.benchmark = True\ntorch.backends.cuda.matmul.allow_tf32 = True\n\n# ---- fixed config (must match training exactly) -----------------------------\nIMG = int(globals().get(\"COATNET_RENDER_IMG\", 336) or 336)\n_SHARP = float(globals().get(\"COATNET_SHARPNESS\", 1.0))\n_CPX_LOG = []          # native crop size per slice, for the resolution audit\nprint(f\"[V35] rendering the CoAtNet arm at {IMG} px, sharpness {_SHARP:.2f}\", flush=True)\nCROP_MM = 140.0\n# 64 slices per study instead of 44, same proportions. Must match the corpus the weights\n# were trained on (knee_corpus_v4.py).\nSLOTS = [(\"Sagittal\", 1, 18), (\"Sagittal\", 0, 14), (\"Coronal\", 1, 12),\n         (\"Coronal\", 0, 8), (\"Axial\", -1, 12)]\nMAXS = sum(s[2] for s in SLOTS)                     # 64\nK_EVAL = 62   # every window position the volume holds, not an evenly spaced subset\nNORM = \"imagenet\"\nLAB = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\", \"Lateral OA\",\n       \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"]\n_MEAN = torch.tensor([0.485, 0.456, 0.406]).view(3, 1, 1)\n_STD = torch.tensor([0.229, 0.224, 0.225]).view(3, 1, 1)\n\n# Three arms: (weights filename, fallback arch, fallback res). ck carries arch+res too.\n# Selected 2026-08-19 by greedy forward selection AND exhaustive subset search over a 7-arm\n# panel on the 45-study gold set (phase2/blend_panel.py); both agree on this exact set.\n# Singles: coatnet384 0.9025 | swinbase384 0.8825 | effv2l480 0.8716.\n# Blend {coatnet+swin+effv2l} = 0.9068 (2-arm {coatnet+swin} = 0.9059, coatnet alone 0.9025).\n# Dropped as redundant: cnn336 (0.8833, the former champion), cnbase384 (0.8754),\n# cnlarge384 (0.8752), maxvit384 (0.8438).\n#\n# SINGLE ARM: coatnet_rmlp_2_rw_384 retrained on the EXPANDED 4,349-study corpus.\n#\n# Why one arm and not the 3-arm blend: on the live leaderboard CoAtNet alone scored 0.914 while\n# every blend scored 0.914-0.915, so ensembling is worth ~+0.001 there -- the ~+0.010 it showed\n# on the old 45-study gold set was gold-set noise. One arm is also 1/3 the kernel runtime.\n#\n# Corpus expansion: the corpus previously held 3,200 of the 4,349 labelled studies and only 45\n# of the 58 gold studies. Rebuilt to 4,407 studies (+37.8% training data, 58-study gate).\n#\n# Measured on the 58-study gate (the incumbent re-scored on the SAME gate for a fair compare):\n#   incumbent CoAtNet (3,155-study corpus) 0.8923\n#   this model       (4,349-study corpus) 0.9054   (+0.0131, better in 92.7% of 2000 bootstraps)\n# Biggest gains land on the findings that were capping us: Lateral Meniscus +0.071,\n# Fracture +0.057, Lateral OA +0.048, Medial Meniscus +0.035, ACL +0.028.\nARMS = [\n    {\n        \"file\": \"raptor_ft_coatnet_v5_full_swa.pt\",\n        \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\",\n        \"res\": 384,\n        \"w\": 1.0,\n    },\n]\n\nLEGACY_ARM = {\n    \"file\": \"raptor_ft_coatnet_v4_full.pt\",\n    \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\",\n    \"res\": 384,\n    \"k_eval\": 42,\n    \"span_lo\": 0.06,\n    \"span_hi\": 0.94,\n}\n\nPRIMARY_SPAN_LO = 0.02\nPRIMARY_SPAN_HI = 0.98\n\n\n# ============================================================================\n# Model -- verbatim from finetune_raptor.py\n# ============================================================================\ndef build_backbone(arch, pretrained=False):\n    # maxvit/maxxvit/coatnet are conv-attention hybrids: NO CLS token, NO interpolatable\n    # pos-embed -> avg pool. The \"vit\" substring in \"coatnet\"/\"maxvit\" must NOT route them\n    # down the ViT path (mirrors finetune_raptor.py exactly).\n    hybrid = arch.startswith((\"maxvit\", \"maxxvit\", \"coatnet\", \"coat_\", \"convnext\"))\n    is_vit = (not hybrid) and any(k in arch for k in (\"vit\", \"deit\", \"dinov2\", \"eva\", \"beit\"))\n    kw = dict(pretrained=pretrained, num_classes=0, in_chans=3)\n    if is_vit:\n        kw.update(global_pool=\"token\", dynamic_img_size=True)\n    else:\n        kw.update(global_pool=\"avg\")\n    return timm.create_model(arch, **kw)\n\n\nclass RaptorClassifier(nn.Module):\n    def __init__(self, backbone, F_dim=768, n=12, drop=0.2):\n        super().__init__()\n        self.backbone = backbone\n        self.norm = nn.LayerNorm(F_dim)\n        self.att = nn.Sequential(nn.Linear(F_dim, 256), nn.Tanh(), nn.Dropout(drop),\n                                 nn.Linear(256, n))\n        self.clsW = nn.Parameter(torch.zeros(n, F_dim))\n        self.clsb = nn.Parameter(torch.zeros(n))\n        nn.init.trunc_normal_(self.clsW, std=0.02)\n        self.n = n\n\n    def encode(self, x):\n        B, K = x.shape[:2]\n        f = self.backbone(x.flatten(0, 1))\n        return f.view(B, K, -1)\n\n    def head(self, feats):\n        h = self.norm(feats)\n        a = self.att(h)\n        a = torch.softmax(a, dim=1)\n        pooled = torch.einsum(\"bkn,bkf->bnf\", a, h)\n        logits = (pooled * self.clsW).sum(-1) + self.clsb\n        return logits\n\n    def forward(self, x):\n        return self.head(self.encode(x))\n\n\ndef load_model(pt_path, arch_default, res_default, device, ngpu=1):\n    ck = torch.load(pt_path, map_location=\"cpu\", weights_only=False)\n    arch = ck.get(\"arch\", arch_default)\n    ck_res = int(ck.get(\"res\", res_default))\n    bb = build_backbone(arch, pretrained=False)\n    model = RaptorClassifier(bb, F_dim=bb.num_features)\n    model.load_state_dict(ck[\"model\"], strict=True)\n    model.eval().to(device)\n    # NOTE: DataParallel removed on purpose. On the full hidden test it drove a system-RAM OOM\n    # (per-forward module replication over many studies); a single T4 handles K_EVAL=24 windows\n    # fine. Arms are also run SEQUENTIALLY (see main) so peak RAM == one model, not two.\n    del ck\n    gc.collect()\n    return model, ck_res\n\n\n# ============================================================================\n# Eval windowing -- verbatim from finetune_raptor.py StudyWindows (train=False)\n# ============================================================================\ndef _eval_centers(mask, D, k):\n    valid = np.where(mask > 0)[0]\n    if len(valid) < 3:\n        valid = np.arange(min(3, D))\n    lo, hi = int(valid.min()), int(valid.max())\n    cs = [c for c in range(lo + 1, hi) if c - 1 >= lo and c + 1 <= hi]\n    if not cs:\n        cs = [max(1, min((lo + hi) // 2, D - 2))]\n    idx = np.linspace(0, len(cs) - 1, k).round().astype(int)\n    return [cs[i] for i in idx]\n\n\ndef eval_windows(vol, mask, k, res, norm=NORM):\n    D = vol.shape[0]\n    cs = _eval_centers(mask, D, k)\n    wins = np.empty((len(cs), 3, res, res), np.float32)\n    for j, c in enumerate(cs):\n        c = max(1, min(c, D - 2))\n        tri = np.stack([vol[c - 1], vol[c], vol[c + 1]], 0).astype(np.float32) / 255.0\n        t = torch.from_numpy(tri)\n        if _SHARP < 1.0 and t.shape[-1] > 336:\n            # Re-blur toward the input the weights were trained on: down to 336\n            # by area, back up by bilinear, then mix. _SHARP = 0 reproduces the\n            # old input, 1 keeps the native detail, between the two hedges.\n            _lo = F.interpolate(t[None], size=(336, 336), mode=\"area\")\n            _lo = F.interpolate(_lo, size=(res, res), mode=\"bilinear\",\n                                align_corners=False)[0]\n            _hi = (t if t.shape[-1] == res else\n                   F.interpolate(t[None], size=(res, res), mode=\"bilinear\",\n                                 align_corners=False)[0])\n            t = _SHARP * _hi + (1.0 - _SHARP) * _lo\n        elif t.shape[-1] != res:\n            t = F.interpolate(t[None], size=(res, res), mode=\"bilinear\",\n                              align_corners=False)[0]\n        wins[j] = t.numpy()\n    x = torch.from_numpy(wins)\n    if norm == \"imagenet\":\n        x = (x - _MEAN) / _STD\n    return x\n\n\n@torch.no_grad()\ndef infer_probs(model, xwins, device):\n    x = xwins.unsqueeze(0).to(\n        device,\n        non_blocking=True,\n    )\n\n    use_cuda = (\n        str(device).startswith(\"cuda\")\n    )\n\n    def _forward():\n        return torch.sigmoid(\n            model(x).float()\n        )[0].cpu().numpy()\n\n    if use_cuda:\n        try:\n            with torch.autocast(\n                \"cuda\",\n                dtype=torch.float16,\n            ):\n                return _forward()\n\n        except RuntimeError as error:\n            try:\n                with torch.cuda.device(device):\n                    torch.cuda.empty_cache()\n            except Exception:\n                pass\n\n            print(\n                \"[DINOsaur V4.2] \"\n                f\"{device} fp16 retry in fp32: \"\n                f\"{type(error).__name__}\",\n                flush=True,\n            )\n\n            return _forward()\n\n    return _forward()\n\n\ndef rankpct(x):                                   # per-column percentile rank in [0,1]\n    order = x.argsort(0).argsort(0).astype(np.float64)\n    return order / max(1, (x.shape[0] - 1))\n\n\n# ============================================================================\n# Preprocessing -- verbatim from kprep2/dino_preprocess.py, retargeted to TEST\n# ============================================================================\ndef _make_reader():\n    import pydicom, cv2\n    from pydicom.pixel_data_handlers.util import apply_modality_lut\n\n    def order_and_meta(sdir):\n        fs = glob.glob(sdir + \"/*.dcm\"); recs = []; ps_list = []\n        for f in fs:\n            try:\n                h = pydicom.dcmread(f, stop_before_pixels=True)\n                iop = getattr(h, 'ImageOrientationPatient', None)\n                ipp = getattr(h, 'ImagePositionPatient', None)\n                if iop is not None and ipp is not None and len(iop) == 6:\n                    r = np.array(iop[:3], float); c = np.array(iop[3:], float)\n                    n = np.cross(r, c); pos = float(np.dot(np.array(ipp, float), n))\n                else:\n                    pos = float(getattr(h, 'InstanceNumber', 0) or 0)\n                ps = getattr(h, 'PixelSpacing', None); ps = float(ps[0]) if ps is not None else 0.5\n                ps_list.append(ps); recs.append((pos, f, ps))\n            except Exception:\n                recs.append((0.0, f, 0.5))\n        recs.sort(key=lambda x: x[0])\n        med_ps = float(np.median(ps_list)) if ps_list else 0.5\n        return [(f, ps) for _, f, ps in recs], med_ps\n\n    def read_px(f):\n        d = pydicom.dcmread(f)\n        a = apply_modality_lut(d.pixel_array, d).astype(np.float32)\n        if str(getattr(d, 'PhotometricInterpretation', '')) == 'MONOCHROME1':\n            a = a.max() - a\n        return a\n\n    def mm_crop_resize(a, ps):\n        h, w = a.shape; cpx = int(round(CROP_MM / max(ps, 1e-3)))\n        cpx = min(cpx, min(h, w)); y0 = (h - cpx) // 2; x0 = (w - cpx) // 2\n        a = a[y0:y0 + cpx, x0:x0 + cpx]\n        _CPX_LOG.append(cpx)      # what the native crop actually held\n        return cv2.resize(a, (IMG, IMG), interpolation=cv2.INTER_AREA)\n\n    return order_and_meta, read_px, mm_crop_resize\n\n\ndef _pick_series_for_slot(rows, plane, fluid, used):\n    cands = [r for r in rows if r['Anatomical_Plane'] == plane and r['SeriesInstanceUID'] not in used]\n    if fluid in (0, 1):\n        pref = [r for r in cands if int(r.get('Fluid_Sensitive', 0) or 0) == fluid]\n        if pref:\n            return pref[0]\n    return cands[0] if cands else None\n\n\ndef _fill_variant_volume(\n    target_volume,\n    offset,\n    picks,\n    pixel_cache,\n    files,\n    med_ps,\n    mm_crop_resize,\n):\n    arrays = []\n    spacings = []\n\n    for position in picks:\n        position = min(\n            int(position),\n            len(files) - 1,\n        )\n        file_path, spacing = files[\n            position\n        ]\n        arrays.append(\n            pixel_cache.get(\n                position\n            )\n        )\n        spacings.append(\n            spacing\n            if spacing > 0\n            else med_ps\n        )\n\n    valid = [\n        array\n        for array in arrays\n        if array is not None\n    ]\n\n    if valid:\n        all_pixels = np.concatenate(\n            [\n                array.ravel()\n                for array in valid\n            ]\n        )\n        low, high = np.percentile(\n            all_pixels,\n            [2.0, 98.0],\n        )\n    else:\n        low, high = 0.0, 1.0\n\n    for local_index, (\n        array,\n        spacing,\n    ) in enumerate(\n        zip(\n            arrays,\n            spacings,\n        )\n    ):\n        output_index = (\n            offset\n            + local_index\n        )\n\n        if (\n            output_index\n            >= MAXS\n        ):\n            break\n\n        if array is None:\n            continue\n\n        normalized = np.clip(\n            (\n                array - low\n            )\n            / (\n                high - low\n                + 1e-6\n            ),\n            0,\n            1,\n        )\n\n        normalized = mm_crop_resize(\n            normalized,\n            spacing,\n        )\n\n        target_volume[\n            output_index\n        ] = (\n            normalized\n            * 255\n        ).astype(\n            np.uint8\n        )\n\n\ndef build_study_pair(\n    sid,\n    ser_records,\n    tsdir,\n    reader,\n):\n    \"\"\"\n    Produce exact MaxSpan and legacy-span volumes while reading every required\n    DICOM only once. Both checkpoints keep their own percentile normalization.\n    \"\"\"\n    (\n        order_and_meta,\n        read_px,\n        mm_crop_resize,\n    ) = reader\n\n    rows = ser_records.get(\n        sid,\n        [],\n    )\n\n    primary_volume = np.zeros(\n        (\n            MAXS,\n            IMG,\n            IMG,\n        ),\n        np.uint8,\n    )\n    legacy_volume = np.zeros_like(\n        primary_volume\n    )\n\n    used = set()\n    offset = 0\n\n    for plane, fluid, count in SLOTS:\n        record = _pick_series_for_slot(\n            rows,\n            plane,\n            fluid,\n            used,\n        )\n\n        if record is None:\n            offset += count\n            continue\n\n        used.add(\n            record[\n                \"SeriesInstanceUID\"\n            ]\n        )\n\n        files, med_ps = order_and_meta(\n            f\"{tsdir}/{sid}/\"\n            f\"{record['SeriesInstanceUID']}\"\n        )\n\n        if not files:\n            offset += count\n            continue\n\n        number = len(files)\n\n        primary_low = int(\n            number\n            * PRIMARY_SPAN_LO\n        )\n        primary_high = int(\n            number\n            * PRIMARY_SPAN_HI\n        ) - 1\n        primary_high = max(\n            primary_high,\n            primary_low,\n        )\n\n        legacy_low = int(\n            number\n            * float(\n                LEGACY_ARM[\n                    \"span_lo\"\n                ]\n            )\n        )\n        legacy_high = int(\n            number\n            * float(\n                LEGACY_ARM[\n                    \"span_hi\"\n                ]\n            )\n        ) - 1\n        legacy_high = max(\n            legacy_high,\n            legacy_low,\n        )\n\n        if number > 1:\n            primary_picks = np.linspace(\n                primary_low,\n                primary_high,\n                count,\n            ).round().astype(int)\n\n            legacy_picks = np.linspace(\n                legacy_low,\n                legacy_high,\n                count,\n            ).round().astype(int)\n        else:\n            primary_picks = np.zeros(\n                count,\n                dtype=int,\n            )\n            legacy_picks = np.zeros(\n                count,\n                dtype=int,\n            )\n\n        required_positions = sorted(\n            set(\n                primary_picks.tolist()\n                + legacy_picks.tolist()\n            )\n        )\n\n        pixel_cache = {}\n\n        for position in required_positions:\n            position = min(\n                int(position),\n                number - 1,\n            )\n\n            file_path, _ = files[\n                position\n            ]\n\n            try:\n                pixel_cache[\n                    position\n                ] = read_px(\n                    file_path\n                )\n            except Exception:\n                pixel_cache[\n                    position\n                ] = None\n\n        _fill_variant_volume(\n            primary_volume,\n            offset,\n            primary_picks,\n            pixel_cache,\n            files,\n            med_ps,\n            mm_crop_resize,\n        )\n\n        _fill_variant_volume(\n            legacy_volume,\n            offset,\n            legacy_picks,\n            pixel_cache,\n            files,\n            med_ps,\n            mm_crop_resize,\n        )\n\n        offset += count\n\n        if offset >= MAXS:\n            break\n\n    primary_mask = (\n        primary_volume.reshape(\n            MAXS,\n            -1,\n        ).sum(1)\n        > 0\n    ).astype(\n        np.uint8\n    )\n\n    legacy_mask = (\n        legacy_volume.reshape(\n            MAXS,\n            -1,\n        ).sum(1)\n        > 0\n    ).astype(\n        np.uint8\n    )\n\n    return (\n        primary_volume,\n        primary_mask,\n        legacy_volume,\n        legacy_mask,\n    )\n\n\n# ============================================================================\n# Test-root discovery + weights + main\n# ============================================================================\ndef find_test_root():\n    cands = [\"/kaggle/input/competitions/rsna-knee-abnormality-detection\",\n             \"/kaggle/input/rsna-knee-abnormality-detection\"]\n    for b in cands:\n        if os.path.exists(b + \"/test.csv\"):\n            return b\n    for d, _, f in os.walk(\"/kaggle/input\"):\n        if \"test.csv\" in f and (os.path.isdir(d + \"/test_series\") or os.path.isdir(d + \"/test_images\")):\n            return d\n    for d, _, f in os.walk(\"/kaggle/input\"):\n        if \"test.csv\" in f:\n            return d\n    raise RuntimeError(\"no test root under /kaggle/input\")\n\n\ndef find_weight_file(\n    fname,\n    required=True,\n):\n    direct = [\n        f\"/kaggle/input/raptor-knee-arms/{fname}\",\n        f\"/kaggle/input/raptor-knee-arms/1/{fname}\",\n        f\"/kaggle/input/raptor-cnn336/{fname}\",\n    ]\n\n    for path in direct:\n        if os.path.exists(path):\n            return path\n\n    for directory in sorted(\n        glob.glob(\n            \"/kaggle/input/*/\"\n        )\n    ):\n        if (\n            \"competition\"\n            in directory.lower()\n        ):\n            continue\n\n        hits = glob.glob(\n            os.path.join(\n                directory,\n                \"**\",\n                fname,\n            ),\n            recursive=True,\n        )\n\n        if hits:\n            return hits[0]\n\n    if required:\n        raise RuntimeError(\n            f\"{fname} not found \"\n            \"under /kaggle/input\"\n        )\n\n    return None\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Competition tables, teacher labels, and the study subset.\nimport os, re, time, json, random\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nfrom sklearn.metrics import roc_auc_score\n\nROOT = find_test_root()\ntrain = pd.read_csv(ROOT + \"/train.csv\", dtype={\"StudyInstanceUID\": str})\nUIDS_ALL = train[\"StudyInstanceUID\"].astype(str).tolist()\nGOLD_ALL = train[LAB].apply(pd.to_numeric, errors=\"coerce\").to_numpy(np.float64)\nprint(\"train studies: %d\" % len(UIDS_ALL))\n\n_tsp = ROOT + \"/train_series.csv\"\nif not os.path.exists(_tsp):\n    raise FileNotFoundError(\n        \"train_series.csv not found at \" + _tsp + \" -- this notebook needs \"\n        \"the per-series plane metadata to pick slot series the same way \"\n        \"inference does. Without it the volumes would not match the corpus \"\n        \"the weights were trained on.\")\ntr_ser = pd.read_csv(_tsp)\ntr_ser[\"StudyInstanceUID\"] = tr_ser[\"StudyInstanceUID\"].astype(str)\ntr_ser[\"SeriesInstanceUID\"] = tr_ser[\"SeriesInstanceUID\"].astype(str)\nSER_TRAIN = {k: v.to_dict(\"records\") for k, v in tr_ser.groupby(\"StudyInstanceUID\")}\nTRDIR = ROOT + \"/train_series\"\nprint(\"train series: %d rows, %d studies with series\"\n      % (len(tr_ser), len(SER_TRAIN)))\n\n\ndef norm_key(s):\n    return re.sub(r\"[^a-z0-9]+\", \"\", str(s).lower())\n\n\ndef auc(y, p):\n    m = np.isfinite(y) & np.isfinite(p)\n    if m.sum() < 20 or len(np.unique(y[m] >= 0.5)) < 2:\n        return np.nan\n    return float(roc_auc_score((y[m] >= 0.5).astype(int), p[m]))\n\n\n# best teacher per finding, chosen on gold - the arrangement the label audit\n# found beats the reliability consensus under cross-validation\ntables = {}\n_seen_csv, _rejected = [], []\nfor root, dirs, files in os.walk(\"/kaggle/input\"):\n    dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n    for name in sorted(files):\n        low = name.lower()\n        if not low.endswith(\".csv\"):\n            continue\n        if low in (\"train.csv\", \"test.csv\", \"train_series.csv\", \"test_series.csv\"):\n            continue\n        if any(k in low for k in (\"submission\", \"sample_sub\", \"oof\")):\n            continue\n        path = Path(root) / name\n        _seen_csv.append(str(path))\n        # Judged by its columns, not its filename: a label table is anything\n        # carrying a study id and most of the twelve findings.\n        try:\n            head = pd.read_csv(path, nrows=3)\n        except Exception:\n            continue\n        keys = {norm_key(c): c for c in head.columns}\n        uid = next((keys[k] for k in (\"studyinstanceuid\", \"studyid\", \"uid\")\n                    if k in keys), None)\n        hit = [t for t in LAB if norm_key(t) in keys]\n        if uid is None or len(hit) < 6:\n            _rejected.append((name, \"no study id\" if uid is None\n                              else \"only %d of 12 findings\" % len(hit)))\n            continue\n        try:\n            raw = pd.read_csv(path)\n        except Exception:\n            continue\n        raw[uid] = raw[uid].astype(str)\n        raw = raw.drop_duplicates(uid, keep=\"last\").set_index(uid).reindex(UIDS_ALL)\n        tab = {}\n        for t in hit:\n            v = pd.to_numeric(raw[keys[norm_key(t)]], errors=\"coerce\").to_numpy(np.float64).copy()\n            v[(v < 0) | (v > 1)] = np.nan\n            # scale with the corpus so preflight and this loader agree on\n            # what counts as a usable table\n            if np.isfinite(v).sum() >= max(100, int(0.02 * len(UIDS_ALL))):\n                tab[t] = v\n        if tab:\n            tables[name] = tab\n            print(\"  label table: %-34s %2d findings, %d studies covered\"\n                  % (name[:34], len(tab),\n                     int(np.isfinite(next(iter(tab.values()))).sum())))\n\nif not tables:\n    print(\"\")\n    print(\"CSV files seen under /kaggle/input (%d):\" % len(_seen_csv))\n    for _p in _seen_csv[:25]:\n        print(\"   \", _p)\n    for _n, _why in _rejected[:15]:\n        print(\"    rejected %-32s %s\" % (_n[:32], _why))\n    raise RuntimeError(\n        \"No per-study label table found. train.csv carries the 58 gold studies \"\n        \"only, which cannot support a held-out split, so this notebook needs \"\n        \"the report-teacher labels: a CSV with StudyInstanceUID and the twelve \"\n        \"finding columns covering ~4,349 studies (llm_labels_v2.csv, \"\n        \"llm_labels_v4_blend.csv, report_labels_v2.csv or equivalent). \"\n        \"Add that dataset under Notebook -> Add Input, then re-run.\")\n\nY_ALL = np.full((len(UIDS_ALL), 12), np.nan)\nprint(\"\\nlabel source per finding (fidelity on the 58 gold studies):\")\nfor j, t in enumerate(LAB):\n    g = GOLD_ALL[:, j]\n    best, bname = -1.0, None\n    for name, tab in tables.items():\n        if t in tab:\n            a = auc(g, tab[t])\n            if np.isfinite(a) and a > best:\n                best, bname = a, name\n    if bname is not None:\n        Y_ALL[:, j] = tables[bname][t]\n        gm = np.isfinite(g)\n        Y_ALL[gm, j] = g[gm]\n        print(\"  %-18s %-26s %.3f\" % (t, bname[:26], best))\n\n# studies we can actually build: have series, and have at least one label\nROW = {u: i for i, u in enumerate(UIDS_ALL)}\nhave = [u for u in UIDS_ALL\n        if u in SER_TRAIN and np.isfinite(Y_ALL[ROW[u]]).any()]\nprint(\"\\nstudies with series and labels: %d\" % len(have))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Decode once, encode both arms, stream to disk. A re-run resumes rather than\n# repeating hours of DICOM work.\nimport torch\n\nHI_PATH = FEAT_DIR + \"/feat_hi.npy\"\nLO_PATH = FEAT_DIR + \"/feat_lo.npy\"\nIX_PATH = FEAT_DIR + \"/feat_index.json\"\n\n# Deterministic study order, so a resume continues the same sequence.\n_r = np.random.default_rng(12345)\npool = list(have)\n_r.shuffle(pool)\npool = pool[:int(SUBSET_STUDIES)]\n\nwpath = find_weight_file(ARMS[0][\"file\"], required=True)\ndev = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\nmodel, RES = load_model(wpath, ARMS[0][\"arch\"], ARMS[0][\"res\"], dev)\nmodel.eval()\ncore = model.module if hasattr(model, \"module\") else model\nF_DIM = int(core.clsW.shape[1])\nprint(\"backbone ready | res %d | feature dim %d\" % (RES, F_DIM))\n\nCFG_SIG = {\"n\": len(pool), \"K\": int(K_EVAL), \"F\": F_DIM, \"res\": int(RES),\n           \"img\": int(COATNET_RENDER_IMG)}\nKEPT, start = [], 0\n_resume = False\nif RESUME and all(os.path.exists(p) for p in (HI_PATH, LO_PATH, IX_PATH)):\n    try:\n        _ix = json.load(open(IX_PATH))\n        if _ix.get(\"cfg\") == CFG_SIG and _ix.get(\"pool\") == pool:\n            KEPT = list(_ix[\"kept\"])\n            start = int(_ix[\"next\"])\n            _resume = True\n            print(\"resuming: %d studies already extracted, continuing at %d\"\n                  % (len(KEPT), start))\n        else:\n            print(\"existing cache does not match this configuration; restarting\")\n    except Exception as _e:\n        print(\"could not read the cache index (%s); restarting\" % type(_e).__name__)\n\nif _resume:\n    hi = np.load(HI_PATH, mmap_mode=\"r+\")\n    lo = np.load(LO_PATH, mmap_mode=\"r+\")\nelse:\n    # Drop any stale cache explicitly rather than writing over it: the old file\n    # may be a different length, and silently reusing its bytes is exactly the\n    # kind of thing that produces a confident wrong answer.\n    for _p in (HI_PATH, LO_PATH, IX_PATH):\n        try:\n            os.remove(_p)\n        except OSError:\n            pass\n    hi = np.lib.format.open_memmap(HI_PATH, mode=\"w+\", dtype=np.float16,\n                                   shape=(len(pool), int(K_EVAL), F_DIM))\n    lo = np.lib.format.open_memmap(LO_PATH, mode=\"w+\", dtype=np.float16,\n                                   shape=(len(pool), int(K_EVAL), F_DIM))\n    print(\"allocated %.1f GB of feature cache on disk\"\n          % (2 * hi.nbytes / 1e9))\n\n\ndef _save_index(nxt):\n    json.dump({\"cfg\": CFG_SIG, \"pool\": pool, \"kept\": KEPT, \"next\": int(nxt)},\n              open(IX_PATH, \"w\"))\n\n\nreader = _make_reader()\nt0 = time.time()\nn_done = 0\nfor i in range(start, len(pool)):\n    if time.time() - t0 > FEATURE_BUDGET_S:\n        print(\"  feature budget reached at study %d; using what is saved\" % i,\n              flush=True)\n        _save_index(i)\n        break\n    sid = pool[i]\n    try:\n        vol, mask, _, _ = build_study_pair(sid, SER_TRAIN, TRDIR, reader)\n        row = {}\n        for tag, sharp in ((\"hi\", 1.0), (\"lo\", 0.0)):\n            globals()[\"_SHARP\"] = sharp          # eval_windows reads this\n            xw = eval_windows(vol, mask, k=K_EVAL, res=RES, norm=NORM)\n            with torch.no_grad(), torch.autocast(\n                    \"cuda\", dtype=torch.float16,\n                    enabled=str(dev).startswith(\"cuda\")):\n                row[tag] = core.encode(xw.unsqueeze(0).to(dev))[0]\n        slot = len(KEPT)\n        hi[slot] = row[\"hi\"].float().cpu().numpy().astype(np.float16)\n        lo[slot] = row[\"lo\"].float().cpu().numpy().astype(np.float16)\n        KEPT.append(sid)\n        n_done += 1\n    except Exception as exc:\n        if n_done < 5:\n            print(\"  skipped %s: %s\" % (sid[-12:], type(exc).__name__), flush=True)\n        continue\n    if (i + 1) % 100 == 0:\n        hi.flush()\n        lo.flush()\n        _save_index(i + 1)\n        el = time.time() - t0\n        rate = el / max(n_done, 1)\n        print(\"  %4d/%d kept %4d  %.1f min  %.1f s/study  eta %.1f min\"\n              % (i + 1, len(pool), len(KEPT), el / 60, rate,\n                 (len(pool) - i - 1) * rate / 60), flush=True)\nelse:\n    _save_index(len(pool))\nglobals()[\"_SHARP\"] = 1.0\nhi.flush()\nlo.flush()\n\nif len(KEPT) < 300:\n    raise RuntimeError(\n        \"only %d studies produced features - too few for 15 seeds. The cache is \"\n        \"saved, so re-running this notebook will continue from here rather than \"\n        \"start over.\" % len(KEPT))\n\nHI = np.asarray(hi[:len(KEPT)], dtype=np.float32)\nLO = np.asarray(lo[:len(KEPT)], dtype=np.float32)\nidx = [ROW[u] for u in KEPT]\nY = Y_ALL[idx]\nprint(\"\")\nprint(\"features ready: %s  (%.0f MB per arm, cached at %s)\"\n      % (HI.shape, HI.nbytes / 1e6, FEAT_DIR))\nprint(\"mean |HI - LO| per feature: %.4f   (0 would mean the arms are identical)\"\n      % float(np.abs(HI - LO).mean()))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Fifteen seeds on the cached features. Same split and same head init per arm.\nimport torch.nn as nn\nfrom scipy import stats\n\n\nclass Head(nn.Module):\n    \"\"\"RaptorClassifier.head, detached so it can train on cached features.\"\"\"\n\n    def __init__(self, F_dim, n=12, drop=0.2):\n        super().__init__()\n        self.norm = nn.LayerNorm(F_dim)\n        self.att = nn.Sequential(nn.Linear(F_dim, 256), nn.Tanh(),\n                                 nn.Dropout(drop), nn.Linear(256, n))\n        self.clsW = nn.Parameter(torch.zeros(n, F_dim))\n        self.clsb = nn.Parameter(torch.zeros(n))\n\n    def forward(self, feats):\n        h = self.norm(feats)\n        a = torch.softmax(self.att(h), dim=1)\n        pooled = torch.einsum(\"bkn,bkf->bnf\", a, h)\n        return (pooled * self.clsW).sum(-1) + self.clsb\n\n\ndef pretrained_head():\n    h = Head(F_DIM).to(dev)\n    h.norm.load_state_dict(core.norm.state_dict())\n    h.att.load_state_dict(core.att.state_dict())\n    with torch.no_grad():\n        h.clsW.copy_(core.clsW)\n        h.clsb.copy_(core.clsb)\n    return h\n\n\nBASE = {k: v.clone() for k, v in pretrained_head().state_dict().items()}\nT_ALL = torch.from_numpy(np.nan_to_num(Y, nan=0.5)).float().to(dev)\nM_ALL = torch.from_numpy(np.isfinite(Y).astype(np.float32)).to(dev)\nX = {\"hi\": torch.from_numpy(HI).to(dev), \"lo\": torch.from_numpy(LO).to(dev)}\nlossf = nn.BCEWithLogitsLoss(reduction=\"none\")\n\n\ndef train_head(arm, tr, seed, epochs=None):\n    torch.manual_seed(seed)\n    h = Head(F_DIM).to(dev)\n    h.load_state_dict(BASE)\n    opt = torch.optim.AdamW(h.parameters(), lr=HEAD_LR, weight_decay=1e-4)\n    tr_t = torch.as_tensor(np.asarray(tr), device=dev)\n    g = torch.Generator().manual_seed(seed)\n    for _ in range(int(epochs or HEAD_EPOCHS)):\n        h.train()\n        perm = tr_t[torch.randperm(len(tr_t), generator=g).to(dev)]\n        for k in range(0, len(perm), 64):\n            b = perm[k:k + 64]\n            opt.zero_grad(set_to_none=True)\n            l = (lossf(h(X[arm][b]), T_ALL[b]) * M_ALL[b]).sum() / M_ALL[b].sum().clamp(min=1)\n            l.backward()\n            opt.step()\n    h.eval()\n    return h\n\n\ndef predict(h, arm, rows):\n    with torch.no_grad():\n        return torch.sigmoid(h(X[arm][torch.as_tensor(np.asarray(rows), device=dev)])).cpu().numpy()\n\n\ndef macro(y, p):\n    per = [auc(y[:, j], p[:, j]) for j in range(12)]\n    return float(np.nanmean(per)), per\n\n\nN = len(KEPT)\nprint(\"=\" * 78)\nprint(\"PAIRED A/B OVER %d SEEDS   (%d studies, %.0f%% held out)\"\n      % (len(SEEDS), N, 100 * VAL_FRACTION))\nprint(\"=\" * 78)\ndiffs, per_hi, per_lo = [], [], []\nfor seed in SEEDS:\n    order = np.random.default_rng(seed).permutation(N)\n    nv = max(60, int(round(VAL_FRACTION * N)))\n    va, tr = order[:nv], order[nv:]\n    m_lo, l_lo = macro(Y[va], predict(train_head(\"lo\", tr, seed), \"lo\", va))\n    m_hi, l_hi = macro(Y[va], predict(train_head(\"hi\", tr, seed), \"hi\", va))\n    diffs.append(m_hi - m_lo)\n    per_lo.append(l_lo)\n    per_hi.append(l_hi)\n    print(\"  seed %-3d  LO %.4f   HI %.4f   %+.4f\"\n          % (seed, m_lo, m_hi, m_hi - m_lo), flush=True)\n\nd = np.asarray(diffs, float)\nn = len(d)\nse = float(d.std(ddof=1) / np.sqrt(n))\ntstat = float(d.mean() / se) if se > 0 else float(\"nan\")\np_two = float(2 * (1 - stats.t.cdf(abs(tstat), n - 1)))\np_one = float(1 - stats.t.cdf(tstat, n - 1))\nlo_ci, hi_ci = stats.t.interval(0.95, n - 1, loc=d.mean(), scale=se)\nCLEAR = bool(lo_ci > 0)\n\nprint(\"\")\nprint(\"  mean paired difference : %+.5f\" % d.mean())\nprint(\"  standard error         : %.5f  over %d seeds\" % (se, n))\nprint(\"  t (df %-2d)              : %.3f\" % (n - 1, tstat))\nprint(\"  95%% interval           : [%+.5f, %+.5f]\" % (lo_ci, hi_ci))\nprint(\"  p two-sided            : %.4f\" % p_two)\nprint(\"  p one-sided            : %.4f   (direction was predicted in advance)\"\n      % p_one)\nprint(\"  seeds where 384 won    : %d of %d\" % (int((d > 0).sum()), n))\nprint(\"  interval clears zero   : %s\" % CLEAR)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Per-finding breakdown, then the deployable head.\nph = np.nanmean(np.asarray(per_hi, float), axis=0)\npl = np.nanmean(np.asarray(per_lo, float), axis=0)\nFINE = {\"PF OA\", \"Lateral Meniscus\", \"Fracture\"}\nprint(\"=\" * 78)\nprint(\"PER-FINDING, averaged over %d seeds\" % len(SEEDS))\nprint(\"=\" * 78)\nprint(\"  %-20s%10s%10s%10s\" % (\"\", \"LO(336)\", \"HI(384)\", \"diff\"))\nfor j, t in enumerate(LAB):\n    print(\"  %-20s%10.4f%10.4f%+10.4f%s\"\n          % (t, pl[j], ph[j], ph[j] - pl[j],\n             \"  <- fine detail\" if t in FINE else \"\"))\nfine = float(np.nanmean([ph[LAB.index(t)] - pl[LAB.index(t)] for t in FINE]))\nrest = float(np.nanmean([ph[j] - pl[j] for j, t in enumerate(LAB) if t not in FINE]))\nprint(\"\")\nprint(\"  fine-detail findings : %+.4f\" % fine)\nprint(\"  everything else      : %+.4f\" % rest)\nprint(\"  difference           : %+.4f\" % (fine - rest))\n_spread = [(ph[LAB.index(t)] - pl[LAB.index(t)]) for t in FINE]\nif fine > rest and max(_spread) > 3 * sorted(_spread)[1]:\n    print(\"  CAUTION: the fine-detail gain rests on one finding, not three.\")\n    print(\"           Read that as weak support for the mechanism, not proof.\")\n\n# --- the deployable head ----------------------------------------------------\nprint(\"\")\nprint(\"=\" * 78)\nprint(\"DEPLOYABLE HEAD\")\nprint(\"=\" * 78)\nfinal = train_head(\"hi\", np.arange(N), 20250830)\ninsample, _ = macro(Y, predict(final, \"hi\", np.arange(N)))\nbase_head = pretrained_head()\nbase_hi, _ = macro(Y, predict(base_head, \"hi\", np.arange(N)))\nprint(\"  trained on all %d cached studies at 384\" % N)\nprint(\"  shipped head on HI features : %.4f  (in-sample, for reference only)\"\n      % base_hi)\nprint(\"  new head on HI features     : %.4f  (in-sample, optimistic by \"\n      \"construction)\" % insample)\nprint(\"  the honest number is the held-out one above: %+.5f\" % d.mean())\n\ntorch.save({\"head\": {k: v.cpu() for k, v in final.state_dict().items()},\n            \"F_dim\": F_DIM, \"res\": int(RES), \"img\": int(COATNET_RENDER_IMG),\n            \"studies\": int(N), \"epochs\": int(HEAD_EPOCHS), \"lr\": float(HEAD_LR),\n            \"seeds\": list(SEEDS), \"paired_diff\": d.tolist(),\n            \"paired_mean\": float(d.mean()), \"paired_se\": se,\n            \"ci95\": [float(lo_ci), float(hi_ci)], \"p_one_sided\": p_one,\n            \"clears_zero\": CLEAR, \"labels\": LAB},\n           \"/kaggle/working/head_384.pt\")\n\nprint(\"\")\nprint(\"#\" * 78)\nprint(\"# STUDIES CACHED    : %d   (features kept at %s)\" % (N, FEAT_DIR))\nprint(\"# PAIRED DIFFERENCE : %+.5f   95%% [%+.5f, %+.5f]\" % (d.mean(), lo_ci, hi_ci))\nprint(\"# INTERVAL CLEARS 0 : %s\" % CLEAR)\nprint(\"# HEAD WRITTEN      : /kaggle/working/head_384.pt\")\nprint(\"#\" * 78)\nif CLEAR:\n    print(\"\"\"\n  Established at 95%% with %d seeds. Expect roughly HALF of this on the\n  leaderboard, because the CoAtNet arm carries about half the blend - so\n  0.936-0.937, not 0.94. Download head_384.pt, publish it as a dataset, and\n  attach it to the submission notebook.\n\"\"\" % len(SEEDS))\nelse:\n    print(\"\"\"\n  The interval still contains zero after %d seeds. With this much power that is\n  a real answer rather than low power: the detail above 336 is not worth\n  deploying. Set COATNET_RENDER_IMG = 336 and close the resolution line.\n\n  head_384.pt is written either way, but do not ship it on this evidence.\n\"\"\" % len(SEEDS))\nprint(\"\"\"\nNOT A SUBMISSION NOTEBOOK\n  It writes head_384.pt and the feature cache, not submission.csv.\n\nTHE FEATURE CACHE IS THE ASSET\n  feat_hi.npy, feat_lo.npy and feat_index.json in /kaggle/working hold the\n  expensive part - 5.9 s of DICOM decode per study. Any further question about\n  the head can now be answered in seconds instead of hours, and a re-run of\n  this notebook continues the extraction rather than repeating it.\n\"\"\")\n","metadata":{},"outputs":[],"execution_count":null}]}