{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"b72cee1e","cell_type":"markdown","source":"<div class=\"knee-title\">\n<h1>&#129517; Raptor CoAtNet, three training snapshots blended</h1>\n<p>Three fine-tunes of the same public, CC0-licensed CoAtNet-384 checkpoint, each trained on a differently-preprocessed corpus, rank-blended together score 0.9288 on the 58 gold-labeled studies - up from 0.9255 for the strongest one alone. Getting the base recipe right took fixing five real bugs in an earlier from-scratch replication; getting the ensemble right took finding the exact per-checkpoint preprocessing config each snapshot was actually trained on, not assuming they all share one recipe.</p>\n</div>","metadata":{}},{"id":"596a7deb","cell_type":"code","source":"from IPython.display import HTML\nHTML(\"\"\"\n<style>\n.knee-title{background:linear-gradient(90deg,#1C6E8C 0%,#2E8B57 100%);color:#ffffff;\n  font-family:'Segoe UI',sans-serif;padding:16px 22px;border-radius:12px;margin:6px 0 10px 0;}\n.knee-title h1{margin:0;font-size:22px;font-weight:700;}\n.knee-title p{margin:6px 0 0 0;font-size:13px;opacity:0.92;}\n.section{color:#1C6E8C;font-family:'Segoe UI',sans-serif;font-size:15px;font-weight:700;\n  margin:10px 0 2px 0;padding-bottom:3px;border-bottom:2px solid #DCEEF0;}\n.note{background:#EAF4F4;border-left:5px solid #1C6E8C;padding:9px 15px;border-radius:0 8px 8px 0;\n  color:#22333B;font-family:'Segoe UI',sans-serif;font-size:13.5px;line-height:1.55;margin:4px 0 10px 0;}\n.note.good{background:#EAF7EF;border-left-color:#2E8B57;}\n.note.warn{background:#FFF6E9;border-left-color:#F2A541;}\n.note code{background:#ffffffa8;padding:1px 5px;border-radius:4px;}\n</style>\n\"\"\")","metadata":{},"outputs":[],"execution_count":null},{"id":"2972fd81","cell_type":"code","source":"PALETTE = {\"teal\": \"#1C6E8C\", \"green\": \"#2E8B57\", \"amber\": \"#F2A541\", \"rose\": \"#C1443C\", \"ink\": \"#22333B\"}\nimport matplotlib.pyplot as plt\nplt.rcParams.update({\n    \"figure.facecolor\": \"white\", \"axes.facecolor\": \"white\", \"axes.edgecolor\": \"#B9C7C9\",\n    \"axes.labelcolor\": PALETTE[\"ink\"], \"axes.titlecolor\": PALETTE[\"teal\"], \"axes.titleweight\": \"bold\",\n    \"xtick.color\": PALETTE[\"ink\"], \"ytick.color\": PALETTE[\"ink\"], \"font.family\": \"sans-serif\",\n    \"axes.grid\": True, \"grid.color\": \"#E7EEEF\", \"grid.linewidth\": 0.8,\n    \"axes.spines.top\": False, \"axes.spines.right\": False,\n})","metadata":{},"outputs":[],"execution_count":null},{"id":"d01b7348","cell_type":"markdown","source":"<div class=\"section\">&#128218; What changed, twice</div>\n<div class=\"note warn\">An earlier version of this notebook reconstructed the checkpoint's architecture from its state dict alone and reached 0.8527 on the gold set - a full 0.07 below the 0.9214 the checkpoint's own metadata claims. Pulling the actual public source behind the checkpoint (<code>dreaddevelopment/knee-mri-twelve-findings-from-a-single-model</code>, a 0.924-public-LB notebook) turned up five concrete bugs: wrong normalization (buffers guessed inside the model instead of the real external ImageNet mean/std), a wrong attention-head activation and an unnecessary presence-masking path, missing DICOM modality LUT and MONOCHROME1 handling, percentile normalization computed per slice instead of per slot, and a 15-window sample instead of the real 42 drawn from a 64-slice volume. Fixing all five and reloading with <code>strict=True</code> brought the gold-only score to 0.9213, then a window-count sweep found a 56-60 plateau above the shipped 42, landing on 0.9255.</div>\n<div class=\"note good\">This version adds a second real gain. The checkpoint author published seven fine-tunes of the identical CoAtNet architecture, one per differently-preprocessed corpus (named <code>widedense</code>, <code>maxspan</code> - the one used above, <code>widefov</code>, <code>fullspan</code>, <code>native384</code>, <code>finespacing</code>, <code>native384dense</code>). An earlier attempt at blending four of these (\"arms\") failed - solo AUCs came back near random - because every arm was fed through the <code>maxspan</code> recipe regardless of what it was actually trained on. Reading a competitor's public notebook that blends these same checkpoints turned up their exact per-checkpoint build resolution, slot layout and depth-span settings. Rebuilding two arms (<code>native384</code> and <code>native384dense</code>) with their own correct settings brought both to genuinely strong solo scores (0.9116 and 0.9170), and a rank-blend of all three reaches <b>0.9288</b>.</div>","metadata":{}},{"id":"8a630836","cell_type":"markdown","source":"<div class=\"section\">&#9881;&#65039; Setup</div>","metadata":{}},{"id":"ec46aac0","cell_type":"code","source":"import os, sys, time, random, hashlib, warnings, gc\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport pydicom\nimport cv2\nimport timm\nfrom pydicom.pixel_data_handlers.util import apply_modality_lut\n\nfrom sklearn.metrics import roc_auc_score\n\nwarnings.filterwarnings(\"ignore\")\nSEED = 42\nrandom.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)\ntorch.backends.cudnn.benchmark = True\ntorch.backends.cuda.matmul.allow_tf32 = True\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(\"device:\", DEVICE, \"| gpu count:\", torch.cuda.device_count())","metadata":{},"outputs":[],"execution_count":null},{"id":"b5560b9a","cell_type":"markdown","source":"<div class=\"section\">&#128193; Data</div>\n<div class=\"note\">No labels are needed here - every model below is a finished, already-trained checkpoint, so this notebook only loads the competition's study/series metadata to know which DICOM series belong to which study.</div>","metadata":{}},{"id":"986168a9","cell_type":"code","source":"ROOT = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\", \"Lateral OA\",\n           \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"]\nN_TARGETS = len(TARGETS)\n\ntrain = pd.read_csv(ROOT / \"train.csv\")\ntest = pd.read_csv(ROOT / \"test.csv\")\ntrain_series = pd.read_csv(ROOT / \"train_series.csv\")\ntest_series = pd.read_csv(ROOT / \"test_series.csv\")\nsample_sub = pd.read_csv(ROOT / \"sample_submission.csv\")\ngold_mask = train[TARGETS].notna().any(axis=1)\ngold_idx = np.where(gold_mask.values)[0]\ngold_ids = train.loc[gold_mask, \"StudyInstanceUID\"].tolist()\ngold_y_true = train[TARGETS].values[gold_idx]\ntrain_ids = train.StudyInstanceUID.tolist()\ntest_ids = test.StudyInstanceUID.tolist()\ntrain_ser_records = {k: v.to_dict(\"records\") for k, v in train_series.groupby(\"StudyInstanceUID\")}\ntest_ser_records = {k: v.to_dict(\"records\") for k, v in test_series.groupby(\"StudyInstanceUID\")}\nprint(\"train:\", train.shape, \"| test:\", test.shape, \"| gold-labeled:\", int(gold_mask.sum()))","metadata":{},"outputs":[],"execution_count":null},{"id":"9e9e2c92","cell_type":"markdown","source":"<div class=\"section\">&#129517; Shared model class and DICOM preprocessing</div>\n<div class=\"note good\">Every arm below is the same architecture - a CoAtNet-384 backbone with a per-finding attention-pooling head - so one class and one set of DICOM-reading functions serve all three. <code>strict=True</code> on each checkpoint load only succeeds on an exact architecture match; it isn't a formality, it's the actual proof each checkpoint fits this class.</div>","metadata":{}},{"id":"dc29fd2f","cell_type":"code","source":"def build_backbone(arch, pretrained=False):\n    # CoAtNet/MaxViT are conv-attention hybrids: no CLS token, no interpolatable pos-embed,\n    # so they need average pooling rather than the token pooling a plain ViT would use.\n    hybrid = arch.startswith((\"maxvit\", \"maxxvit\", \"coatnet\", \"coat_\", \"convnext\"))\n    is_vit = (not hybrid) and any(k in arch for k in (\"vit\", \"deit\", \"dinov2\", \"eva\", \"beit\"))\n    kw = dict(pretrained=pretrained, num_classes=0, in_chans=3)\n    if is_vit:\n        kw.update(global_pool=\"token\", dynamic_img_size=True)\n    else:\n        kw.update(global_pool=\"avg\")\n    return timm.create_model(arch, **kw)\n\n\nclass RaptorClassifier(nn.Module):\n    \"\"\"Per-finding attention pooling over the study's windows - each of the twelve findings\n    gets its own attention weights, so a cruciate tear visible on two sagittal slices and\n    osteoarthritis spread across many coronal ones don't have to share one pooled score.\"\"\"\n    def __init__(self, backbone, F_dim, n=12, drop=0.2):\n        super().__init__()\n        self.backbone = backbone\n        self.norm = nn.LayerNorm(F_dim)\n        self.att = nn.Sequential(nn.Linear(F_dim, 256), nn.Tanh(), nn.Dropout(drop), nn.Linear(256, n))\n        self.clsW = nn.Parameter(torch.zeros(n, F_dim))\n        self.clsb = nn.Parameter(torch.zeros(n))\n        nn.init.trunc_normal_(self.clsW, std=0.02)\n        self.n = n\n\n    def encode(self, x):\n        B, K = x.shape[:2]\n        f = self.backbone(x.flatten(0, 1))\n        return f.view(B, K, -1)\n\n    def head(self, feats):\n        h = self.norm(feats)\n        a = torch.softmax(self.att(h), dim=1)\n        pooled = torch.einsum(\"bkn,bkf->bnf\", a, h)\n        return (pooled * self.clsW).sum(-1) + self.clsb\n\n    def forward(self, x):\n        return self.head(self.encode(x))\n\n\ndef load_arm(ckpt_path):\n    ck = torch.load(ckpt_path, map_location=\"cpu\", weights_only=False)\n    bb = build_backbone(ck.get(\"arch\", \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\"), pretrained=False)\n    model = RaptorClassifier(bb, F_dim=bb.num_features)\n    model.load_state_dict(ck[\"model\"], strict=True)\n    model.eval().to(DEVICE)\n    print(\"loaded\", ckpt_path, \"| strict load OK | res:\", ck.get(\"res\"), \"| its own claimed gold_auc:\", ck.get(\"gold_auc\"))\n    return model, int(ck.get(\"res\", 384))","metadata":{},"outputs":[],"execution_count":null},{"id":"6a56dd68","cell_type":"markdown","source":"<div class=\"section\">&#129513; Building a fixed input, per checkpoint's own recipe</div>\n<div class=\"note\">No two studies look alike - the number of series and slices per series both vary - so every recipe here fills a fixed-size volume in a fixed slot order, preferring a different series per slot when more than one is available, leaving a missing slot at zero and telling the model to skip it. What differs <i>between checkpoints</i> is the exact shape of that fixed volume: the main <code>maxspan</code> checkpoint wants a 64-slice volume across five slots (18/14/12/8/12) built at 336px; <code>native384</code> wants a smaller 44-slice volume (12/10/8/6/8) built natively at 384px; <code>native384dense</code> keeps the 64-slice layout but also builds natively at 384px. Every slice is still cropped to a 140&nbsp;mm box from the DICOM pixel spacing, modality-LUT corrected, and MONOCHROME1-inverted where needed, and every slot's slices share one percentile normalization rather than being rescaled independently.</div>","metadata":{}},{"id":"153e5e59","cell_type":"code","source":"CROP_MM = 140.0\n_MEAN = torch.tensor([0.485, 0.456, 0.406]).view(3, 1, 1)\n_STD = torch.tensor([0.229, 0.224, 0.225]).view(3, 1, 1)\n\n# Slot layouts: (plane, fluid-sensitive flag, slice count). -1 fluid means \"no preference\".\nSLOTS_64 = [(\"Sagittal\", 1, 18), (\"Sagittal\", 0, 14), (\"Coronal\", 1, 12),\n            (\"Coronal\", 0, 8), (\"Axial\", -1, 12)]\nSLOTS_44 = [(\"Sagittal\", 1, 12), (\"Sagittal\", 0, 10), (\"Coronal\", 1, 8),\n            (\"Coronal\", 0, 6), (\"Axial\", -1, 8)]\n\n\ndef _order_and_meta(series_dir):\n    fs = sorted(os.listdir(series_dir))\n    recs, ps_list = [], []\n    for fn in fs:\n        f = os.path.join(series_dir, fn)\n        try:\n            h = pydicom.dcmread(f, stop_before_pixels=True)\n            iop = getattr(h, \"ImageOrientationPatient\", None)\n            ipp = getattr(h, \"ImagePositionPatient\", None)\n            if iop is not None and ipp is not None and len(iop) == 6:\n                r = np.array(iop[:3], float); c = np.array(iop[3:], float)\n                n = np.cross(r, c); pos = float(np.dot(np.array(ipp, float), n))\n            else:\n                pos = float(getattr(h, \"InstanceNumber\", 0) or 0)\n            ps = getattr(h, \"PixelSpacing\", None)\n            ps = float(ps[0]) if ps is not None else 0.5\n            ps_list.append(ps); recs.append((pos, f, ps))\n        except Exception:\n            recs.append((0.0, f, 0.5))\n    recs.sort(key=lambda x: x[0])\n    med_ps = float(np.median(ps_list)) if ps_list else 0.5\n    return [(f, ps) for _, f, ps in recs], med_ps\n\n\ndef _read_px(f):\n    d = pydicom.dcmread(f)\n    a = apply_modality_lut(d.pixel_array, d).astype(np.float32)\n    if str(getattr(d, \"PhotometricInterpretation\", \"\")) == \"MONOCHROME1\":\n        a = a.max() - a\n    return a\n\n\ndef _mm_crop_resize(a, ps, img_build):\n    h, w = a.shape\n    cpx = int(round(CROP_MM / max(ps, 1e-3)))\n    cpx = min(cpx, min(h, w))\n    y0 = (h - cpx) // 2; x0 = (w - cpx) // 2\n    a = a[y0:y0 + cpx, x0:x0 + cpx]\n    return cv2.resize(a, (img_build, img_build), interpolation=cv2.INTER_AREA)\n\n\ndef _pick_series_for_slot(rows, plane, fluid, used):\n    cands = [r for r in rows if r[\"Anatomical_Plane\"] == plane and r[\"SeriesInstanceUID\"] not in used]\n    if fluid in (0, 1):\n        pref = [r for r in cands if int(r.get(\"Fluid_Sensitive\", 0) or 0) == fluid]\n        if pref:\n            return pref[0]\n    return cands[0] if cands else None\n\n\ndef build_study(study_id, ser_records, split_dir, img_build, slots, span_lo, span_hi):\n    maxs = sum(s[2] for s in slots)\n    rows = ser_records.get(study_id, [])\n    vol = np.zeros((maxs, img_build, img_build), np.uint8)\n    idx = 0\n    used = set()\n    for plane, fluid, k in slots:\n        r = _pick_series_for_slot(rows, plane, fluid, used)\n        if r is None:\n            idx += k\n            continue\n        used.add(r[\"SeriesInstanceUID\"])\n        series_dir = ROOT / split_dir / study_id / r[\"SeriesInstanceUID\"]\n        try:\n            files, med_ps = _order_and_meta(series_dir)\n        except FileNotFoundError:\n            files, med_ps = [], 0.5\n        if not files:\n            idx += k\n            continue\n        n = len(files)\n        lo, hi = int(n * span_lo), int(n * span_hi) - 1\n        hi = max(hi, lo)\n        picks = np.linspace(lo, hi, k).round().astype(int) if n > 1 else [0] * k\n        arrs, pss = [], []\n        for p in picks:\n            fp, ps = files[min(p, n - 1)]\n            try:\n                arrs.append(_read_px(fp)); pss.append(ps)\n            except Exception:\n                arrs.append(None); pss.append(med_ps)\n        valid = [a for a in arrs if a is not None]\n        if valid:\n            allpx = np.concatenate([a.ravel() for a in valid])\n            loq, hiq = np.percentile(allpx, [2.0, 98.0])\n        else:\n            loq, hiq = 0.0, 1.0\n        for a, ps in zip(arrs, pss):\n            if idx >= maxs:\n                break\n            if a is None:\n                idx += 1\n                continue\n            aw = np.clip((a - loq) / (hiq - loq + 1e-6), 0, 1)\n            aw = _mm_crop_resize(aw, ps if ps > 0 else med_ps, img_build)\n            vol[idx] = (aw * 255).astype(np.uint8)\n            idx += 1\n        if idx >= maxs:\n            break\n    mask = (vol.reshape(maxs, -1).sum(1) > 0).astype(np.uint8)\n    return vol, mask","metadata":{},"outputs":[],"execution_count":null},{"id":"f06e21f5","cell_type":"markdown","source":"<div class=\"section\">&#128269; Evaluation windows</div>\n<div class=\"note\">Three neighbouring slices stack into the three channels of one window, so the network sees a little of what lies just above and below the slice in the middle. Window centres are drawn evenly from every valid interior position the volume holds - a window can span a slot boundary and the model still gets a dense, overlapping view of the whole study.</div>","metadata":{}},{"id":"f68834a9","cell_type":"code","source":"def _eval_centers(mask, D, k):\n    valid = np.where(mask > 0)[0]\n    if len(valid) < 3:\n        valid = np.arange(min(3, D))\n    lo, hi = int(valid.min()), int(valid.max())\n    cs = [c for c in range(lo + 1, hi) if c - 1 >= lo and c + 1 <= hi]\n    if not cs:\n        cs = [max(1, min((lo + hi) // 2, D - 2))]\n    idx = np.linspace(0, len(cs) - 1, k).round().astype(int)\n    return [cs[i] for i in idx]\n\n\ndef eval_windows(vol, mask, k, res):\n    D = vol.shape[0]\n    cs = _eval_centers(mask, D, k)\n    wins = np.empty((len(cs), 3, res, res), np.float32)\n    for j, c in enumerate(cs):\n        c = max(1, min(c, D - 2))\n        tri = np.stack([vol[c - 1], vol[c], vol[c + 1]], 0).astype(np.float32) / 255.0\n        t = torch.from_numpy(tri)\n        if t.shape[-1] != res:\n            t = F.interpolate(t[None], size=(res, res), mode=\"bilinear\", align_corners=False)[0]\n        wins[j] = t.numpy()\n    x = torch.from_numpy(wins)\n    return (x - _MEAN) / _STD\n\n\n@torch.no_grad()\ndef infer_probs(model, xwins, device):\n    x = xwins.unsqueeze(0).to(device)\n    if str(device).startswith(\"cuda\"):\n        try:\n            with torch.autocast(\"cuda\", dtype=torch.float16):\n                o = torch.sigmoid(model(x).float())\n            return o[0].cpu().numpy()\n        except RuntimeError:\n            torch.cuda.empty_cache()\n    o = torch.sigmoid(model(x).float())\n    return o[0].cpu().numpy()\n\n\ndef raptor_predict(study_ids, ser_records, split_dir, model, device, img_build, slots, span_lo, span_hi, k_eval, res):\n    out = np.full((len(study_ids), N_TARGETS), 0.5, np.float32)\n    for i, sid in enumerate(study_ids):\n        vol, mask = build_study(sid, ser_records, split_dir, img_build, slots, span_lo, span_hi)\n        xw = eval_windows(vol, mask, k=k_eval, res=res)\n        out[i] = infer_probs(model, xw, device)\n    return out","metadata":{},"outputs":[],"execution_count":null},{"id":"90532087","cell_type":"markdown","source":"<div class=\"section\">&#128193; Loading the three arms</div>\n<div class=\"note\">Requires three datasets attached: <code>dreaddevelopment/raptor-knee-maxspan</code>, <code>dreaddevelopment/raptor-knee-native384dense</code>, and <code>dreaddevelopment/raptor-knee-native384</code> (all CC0-1.0).</div>\n<div class=\"note\">Each arm's config below - build resolution, slot layout, depth span, window count - matches exactly what that checkpoint was trained on, not a shared guess. The main <code>maxspan</code> settings come from the checkpoint author's own published inference notebook; the <code>native384</code> and <code>native384dense</code> settings come from a competitor's public notebook that blends the same three checkpoints, read directly rather than reverse-engineered. An earlier attempt to blend four <i>other</i> arms without their correct per-checkpoint settings produced near-random solo predictions - the lesson held here too, just with the right settings in hand this time.</div>","metadata":{}},{"id":"944d74a1","cell_type":"code","source":"ARMS = [\n    {\"name\": \"maxspan-v5\", \"path\": \"/kaggle/input/datasets/dreaddevelopment/raptor-knee-maxspan/raptor_ft_coatnet_v5_full_swa.pt\",\n     \"img_build\": 336, \"slots\": SLOTS_64, \"span\": (0.06, 0.94), \"k_eval\": 60},\n    {\"name\": \"native384dense-v10\", \"path\": \"/kaggle/input/datasets/dreaddevelopment/raptor-knee-native384dense/raptor_ft_coatnet_v10_full.pt\",\n     \"img_build\": 384, \"slots\": SLOTS_64, \"span\": (0.02, 0.98), \"k_eval\": 62},\n    {\"name\": \"native384-v8\", \"path\": \"/kaggle/input/datasets/dreaddevelopment/raptor-knee-native384/raptor_ft_coatnet_v8_full_swa.pt\",\n     \"img_build\": 384, \"slots\": SLOTS_44, \"span\": (0.06, 0.94), \"k_eval\": 42},\n]\n\nfor arm in ARMS:\n    model, res = load_arm(arm[\"path\"])\n    arm[\"model\"] = model\n    arm[\"res\"] = res","metadata":{},"outputs":[],"execution_count":null},{"id":"8206e9b1","cell_type":"markdown","source":"<div class=\"section\">&#9989; Validation</div>\n<div class=\"note\">Scored only on the 58 real gold-labeled studies, which none of these checkpoints' training ever saw. Each arm's solo score is computed first, then a rank-percentile joint weight search finds the best blend - rank blending (not raw probability averaging) is what the checkpoint author's own multi-arm code uses, and it came out slightly ahead here too.</div>","metadata":{}},{"id":"10703ff2","cell_type":"code","source":"def compute_masked_auc(y_true, y_pred, cols):\n    aucs = {}\n    for i, col in enumerate(cols):\n        y, p = y_true[:, i], y_pred[:, i]\n        mask = ~np.isnan(y)\n        yb = (y[mask] >= 0.5).astype(int)\n        if mask.sum() > 1 and len(set(yb)) > 1:\n            aucs[col] = roc_auc_score(yb, p[mask])\n    return aucs, (float(np.mean(list(aucs.values()))) if aucs else float(\"nan\"))\n\n\ndef rankpct(x):\n    order = x.argsort(0).argsort(0).astype(np.float64)\n    return order / max(1, (x.shape[0] - 1))\n\n\nfor arm in ARMS:\n    t0 = time.time()\n    span_lo, span_hi = arm[\"span\"]\n    preds = raptor_predict(gold_ids, train_ser_records, \"train_series\", arm[\"model\"], DEVICE,\n                            arm[\"img_build\"], arm[\"slots\"], span_lo, span_hi, arm[\"k_eval\"], arm[\"res\"])\n    _, auc = compute_masked_auc(gold_y_true, preds, TARGETS)\n    arm[\"gold_preds\"] = preds\n    arm[\"gold_auc\"] = auc\n    print(f\"{arm['name']:20s} solo gold-only AUC: {auc:.4f}  ({time.time()-t0:.0f}s)\")\n\nbest_auc, best_w = -1.0, None\nstep = 0.05\ngrid = np.arange(0.0, 1.0 + 1e-9, step)\nranked = [rankpct(arm[\"gold_preds\"]) for arm in ARMS]\nfor w0 in grid:\n    for w1 in grid:\n        if w0 + w1 > 1.0 + 1e-9:\n            continue\n        w2 = 1.0 - w0 - w1\n        if w2 < -1e-9:\n            continue\n        w2 = max(w2, 0.0)\n        ens = w0 * ranked[0] + w1 * ranked[1] + w2 * ranked[2]\n        _, auc = compute_masked_auc(gold_y_true, ens, TARGETS)\n        if auc > best_auc:\n            best_auc, best_w = auc, (w0, w1, w2)\n\nprint(f\"\\nbest blend: {ARMS[0]['name']}={best_w[0]:.2f} {ARMS[1]['name']}={best_w[1]:.2f} {ARMS[2]['name']}={best_w[2]:.2f} -> {best_auc:.4f}\")\nprint(f\"delta over strongest solo arm: {best_auc - max(a['gold_auc'] for a in ARMS):+.4f}\")","metadata":{},"outputs":[],"execution_count":null},{"id":"58554da0","cell_type":"markdown","source":"<div class=\"section\">&#128203; Submission</div>","metadata":{}},{"id":"10c68b03","cell_type":"code","source":"for arm in ARMS:\n    t0 = time.time()\n    span_lo, span_hi = arm[\"span\"]\n    arm[\"test_preds\"] = raptor_predict(test_ids, test_ser_records, \"test_series\", arm[\"model\"], DEVICE,\n                                        arm[\"img_build\"], arm[\"slots\"], span_lo, span_hi, arm[\"k_eval\"], arm[\"res\"])\n    print(f\"{arm['name']:20s} test predict elapsed: {time.time()-t0:.0f}s\")\n\nranked_test = [rankpct(arm[\"test_preds\"]) for arm in ARMS]\npred_test = best_w[0] * ranked_test[0] + best_w[1] * ranked_test[1] + best_w[2] * ranked_test[2]\n\npred_df = pd.DataFrame(pred_test, columns=TARGETS)\npred_df.insert(0, \"StudyInstanceUID\", test_ids)\n\nsubmission = sample_sub[[\"StudyInstanceUID\"]].merge(pred_df, on=\"StudyInstanceUID\", how=\"left\")\nsubmission[TARGETS] = submission[TARGETS].fillna(0.5)\nsubmission = submission[sample_sub.columns.tolist()]\n\nassert list(submission.columns) == list(sample_sub.columns)\nassert len(submission) == len(sample_sub)\nassert set(submission.StudyInstanceUID) == set(sample_sub.StudyInstanceUID)\nassert submission[TARGETS].isna().sum().sum() == 0\nassert ((submission[TARGETS] >= 0) & (submission[TARGETS] <= 1)).all().all()\n\nsubmission.to_csv(\"submission.csv\", index=False)\nprint(\"submission.csv written:\", submission.shape)\nsubmission","metadata":{},"outputs":[],"execution_count":null},{"id":"ec56d985","cell_type":"markdown","source":"<div class=\"section\">&#128204; Notes</div>\n<div class=\"note good\">Solo gold-only AUC: <code>maxspan-v5</code> 0.9255 (the same checkpoint the previous version of this notebook shipped, scoring 0.928 on the real leaderboard at its 42-window setting), <code>native384dense-v10</code> 0.9170, <code>native384-v8</code> 0.9116. Rank-blended at the weights found above: <b>0.9288</b>. This is the first ensembling gain of any real size found in this project since the base recipe was fixed - every other attempt at adding a second model added at most a few thousandths of an AUC point, because none of those secondary models were anywhere near raptor's own accuracy. These three checkpoints all share raptor's own architecture and general training recipe, just fit on differently-preprocessed corpora - close enough in strength that combining them adds real information instead of diluting it.</div>\n<div class=\"note\">Every other real alternative model found or built was tested properly and none added more than a few thousandths of an AUC point once blended: three more checkpoints in the same architecture family with recipes we could verify (<code>widefov</code> v6 0.9108, <code>fullspan</code> v7 0.9063, <code>finespacing</code> v9 0.8985 solo - all decent, all driven to zero weight by the joint search); the competitor's own real out-of-fold predictions for a 5-fold DINOv2 model (0.8400 solo, +0.0000 blended) and three RadImageNet-ResNet50 variants (0.83-0.86 solo, +0.0006 blended, best case); and a properly task-fine-tuned OrthoFoundation-L (a DINOv3 model pretrained on ~1.25M knee images), taken from a 0.75 frozen-embedding probe to a real 0.8124 solo AUC via three epochs of attention-pooling fine-tuning - still only +0.0004 blended. Four checkpoints in the raptor family remain untested (<code>widedense</code> v4/v4-swa and two others) for lack of a verified per-checkpoint recipe - not attempted, to avoid repeating the near-random-prediction failure that a wrong config produces.</div>\n<div class=\"note warn\">All of the above is validated against 58 gold-labeled studies - informative, but a small enough sample that the real test is the leaderboard once this gets submitted. Two competing public notebooks scoring higher than this (0.936, 0.937) were read in full: their edge comes mainly from blending a second, independently strong DINOv2 model (not a trick this notebook could just copy in), and their own code shows real signs of small-sample overfitting on their internal validation set - a caution worth applying to the weights found in this notebook too. Fresh discussion-thread research points at continued self-supervised pretraining before fine-tuning as the more likely path past this ceiling, not more models to blend in.</div>","metadata":{}}]}