{"cells":[{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"\"\"\"[E53] raptor 互換 native3 コーパスの生成（shard 3/5）。\"\"\"\nimport json, os, sys, time\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\n\nSHARD, N_SHARD, SMOKE, N_SPLIT = 3, 5, False, 8\n\nPath(\"/kaggle/working/native3_worker.py\").write_text('\"\"\"[E53] native3 コーパスの worker（子プロセスから import される）。\"\"\"\\nimport os\\n\\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\"):\\n    os.environ[_v] = \"1\"\\n\\nimport numpy as np\\n\\n# --- INLINE-BEGIN: refs/raptor_training/infer_knee_dreaddevelopment.py (未改変) ---\\n#!/usr/bin/env python3\\n\"\"\"Knee MRI: twelve findings from a single model\\n\\nThis notebook takes a knee MRI study and scores twelve findings at once: ACL tear, MCL tear,\\nmedial and lateral meniscus tears, osteoarthritis in the medial, lateral and patellofemoral\\ncompartments, joint effusion, synovitis, a Baker\\'s cyst, bone contusion and fracture. It scores\\n0.924 on the public leaderboard using one model, with no ensembling and no test-time augmentation.\\n\\nThis is the inference half of the work. The model was trained separately and its weights are\\nattached as a dataset, so this notebook only loads them and predicts:\\nhttps://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-widedense\\n\\nWhere the training labels came from\\n\\nWorth saying up front, because it shapes everything else. The competition gives you 4,407 studies\\nbut structured labels for only 58 of them. Every other study arrives with a free-text radiology\\nreport and nothing more, so there is very little to train against out of the box.\\n\\nThe labels behind these weights were made by reading those reports with a language model and\\nturning each into twelve probabilities rather than twelve yes or no answers. A report that says a\\ntear is suspected becomes a number near 0.8, not a 1, which is a fairer target than forcing every\\nhedged sentence into a hard label. That yields 4,349 studies to train on. The 58 studies that came\\nwith real labels were never trained on and are used to check the result honestly; the model reaches\\n0.9167 macro-AUC on them.\\n\\nBuilding a fixed input from studies that are all shaped differently\\n\\nThe hard part of this competition is not the network, it is that no two studies look alike. A\\nstudy holds several DICOM series shot in different planes, the number of series varies, and the\\nnumber of slices in a series varies more. Anything that expects a fixed-size input has to be given\\none.\\n\\nThe approach here is to fill five fixed slots per study, always in the same order, for a stack of\\n64 images:\\n\\n  18 slices from a sagittal series, preferring a fluid-sensitive one\\n  14 slices from a second sagittal series, preferring one that is not fluid-sensitive\\n  12 slices from a coronal series, preferring a fluid-sensitive one\\n   8 slices from a second coronal series\\n  12 slices from an axial series\\n\\nPreferring a fluid-sensitive series for some slots and not for others is deliberate. Fluid-\\nsensitive sequences show swelling, effusion and acute injury clearly, while the other sequences\\nshow anatomy and cartilage better, and the twelve findings are split across both. If a study has\\nno series for a slot, the slot is left as zeros and the model is told to skip it rather than being\\nfed something misleading.\\n\\nWithin a series, slices are taken evenly across 6 to 94 percent of the stack rather than from the\\nmiddle. The outer slices are where the collateral ligaments and the lateral meniscus sit, and\\ncutting them was measurably costing accuracy on exactly those findings.\\n\\nEvery slice is cropped to a 140 mm box around the centre of the image using the pixel spacing from\\nthe DICOM header, then resized to 336 pixels. Cropping by millimetres rather than by pixel count\\nmatters: it means a knee occupies the same fraction of the frame whether the scan was acquired at\\n0.3 or 0.5 mm per pixel, so the model is not asked to learn scale differences that carry no medical\\ninformation.\\n\\nHow the model reads the stack\\n\\nThree neighbouring slices are stacked into the three channels of one image. The network then sees\\na little of what lies above and below the slice in the middle, which is most of the benefit of a 3D\\nmodel at the cost of a 2D one. Each of these three-slice windows is passed through a CoAtNet\\nbackbone at 384 pixels.\\n\\nThe windows are combined with an attention layer that has separate weights for each of the twelve\\nfindings. This is the part that matters most. A cruciate tear may be visible on two sagittal slices\\nwhile osteoarthritis is spread across many coronal ones, and a single pooled score forces those two\\nto share one notion of which slices are important. Giving each finding its own attention weights\\nlets each one draw on the slices that actually show it.\\n\\nRunning it\\n\\nScoring uses 42 windows per study. Inference runs in half precision and automatically retries a\\nstudy in full precision if it fails, so no study is ever dropped from the submission. The notebook\\nneeds no internet: the backbone is loaded from the attached weights rather than downloaded.\\n\"\"\"\\nimport os, sys, glob, time, json, gc\\nos.environ.setdefault(\"HF_HUB_OFFLINE\", \"1\")\\nos.environ.setdefault(\"TRANSFORMERS_OFFLINE\", \"1\")\\nos.environ.setdefault(\"HF_HUB_DISABLE_TELEMETRY\", \"1\")\\nimport numpy as np\\nimport torch, torch.nn as nn, torch.nn.functional as F\\nimport timm\\n# T4 (Turing) cuDNN v9 has fp16/fp32 conv engines but NOT bf16 for these shapes\\n# (\"GET was unable to find an engine...\"); benchmark lets it pick a valid algo for\\n# the fixed (1,24,3,res,res) input.\\ntorch.backends.cudnn.benchmark = True\\ntorch.backends.cuda.matmul.allow_tf32 = True\\n\\n# ---- fixed config (must match training exactly) -----------------------------\\nIMG = 336\\nCROP_MM = 140.0\\n# 64 slices per study instead of 44, same proportions. Must match the corpus the weights\\n# were trained on (knee_corpus_v4.py).\\nSLOTS = [(\"Sagittal\", 1, 18), (\"Sagittal\", 0, 14), (\"Coronal\", 1, 12),\\n         (\"Coronal\", 0, 8), (\"Axial\", -1, 12)]\\nMAXS = sum(s[2] for s in SLOTS)                     # 64\\nK_EVAL = 42   # every window position the volume holds, not an evenly spaced subset\\nNORM = \"imagenet\"\\nLAB = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\", \"Lateral OA\",\\n       \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker\\'s\", \"Contusion\", \"Fracture\"]\\n_MEAN = torch.tensor([0.485, 0.456, 0.406]).view(3, 1, 1)\\n_STD = torch.tensor([0.229, 0.224, 0.225]).view(3, 1, 1)\\n\\n# Three arms: (weights filename, fallback arch, fallback res). ck carries arch+res too.\\n# Selected 2026-08-19 by greedy forward selection AND exhaustive subset search over a 7-arm\\n# panel on the 45-study gold set (phase2/blend_panel.py); both agree on this exact set.\\n# Singles: coatnet384 0.9025 | swinbase384 0.8825 | effv2l480 0.8716.\\n# Blend {coatnet+swin+effv2l} = 0.9068 (2-arm {coatnet+swin} = 0.9059, coatnet alone 0.9025).\\n# Dropped as redundant: cnn336 (0.8833, the former champion), cnbase384 (0.8754),\\n# cnlarge384 (0.8752), maxvit384 (0.8438).\\n#\\n# SINGLE ARM: coatnet_rmlp_2_rw_384 retrained on the EXPANDED 4,349-study corpus.\\n#\\n# Why one arm and not the 3-arm blend: on the live leaderboard CoAtNet alone scored 0.914 while\\n# every blend scored 0.914-0.915, so ensembling is worth ~+0.001 there -- the ~+0.010 it showed\\n# on the old 45-study gold set was gold-set noise. One arm is also 1/3 the kernel runtime.\\n#\\n# Corpus expansion: the corpus previously held 3,200 of the 4,349 labelled studies and only 45\\n# of the 58 gold studies. Rebuilt to 4,407 studies (+37.8% training data, 58-study gate).\\n#\\n# Measured on the 58-study gate (the incumbent re-scored on the SAME gate for a fair compare):\\n#   incumbent CoAtNet (3,155-study corpus) 0.8923\\n#   this model       (4,349-study corpus) 0.9054   (+0.0131, better in 92.7% of 2000 bootstraps)\\n# Biggest gains land on the findings that were capping us: Lateral Meniscus +0.071,\\n# Fracture +0.057, Lateral OA +0.048, Medial Meniscus +0.035, ACL +0.028.\\nARMS = [\\n    {\"file\": \"raptor_ft_coatnet_v4_full.pt\", \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\", \"res\": 384, \"w\": 1.0},\\n]\\n\\n\\n# ============================================================================\\n# Model -- verbatim from finetune_raptor.py\\n# ============================================================================\\ndef build_backbone(arch, pretrained=False):\\n    # maxvit/maxxvit/coatnet are conv-attention hybrids: NO CLS token, NO interpolatable\\n    # pos-embed -> avg pool. The \"vit\" substring in \"coatnet\"/\"maxvit\" must NOT route them\\n    # down the ViT path (mirrors finetune_raptor.py exactly).\\n    hybrid = arch.startswith((\"maxvit\", \"maxxvit\", \"coatnet\", \"coat_\", \"convnext\"))\\n    is_vit = (not hybrid) and any(k in arch for k in (\"vit\", \"deit\", \"dinov2\", \"eva\", \"beit\"))\\n    kw = dict(pretrained=pretrained, num_classes=0, in_chans=3)\\n    if is_vit:\\n        kw.update(global_pool=\"token\", dynamic_img_size=True)\\n    else:\\n        kw.update(global_pool=\"avg\")\\n    return timm.create_model(arch, **kw)\\n\\n\\nclass RaptorClassifier(nn.Module):\\n    def __init__(self, backbone, F_dim=768, n=12, drop=0.2):\\n        super().__init__()\\n        self.backbone = backbone\\n        self.norm = nn.LayerNorm(F_dim)\\n        self.att = nn.Sequential(nn.Linear(F_dim, 256), nn.Tanh(), nn.Dropout(drop),\\n                                 nn.Linear(256, n))\\n        self.clsW = nn.Parameter(torch.zeros(n, F_dim))\\n        self.clsb = nn.Parameter(torch.zeros(n))\\n        nn.init.trunc_normal_(self.clsW, std=0.02)\\n        self.n = n\\n\\n    def encode(self, x):\\n        B, K = x.shape[:2]\\n        f = self.backbone(x.flatten(0, 1))\\n        return f.view(B, K, -1)\\n\\n    def head(self, feats):\\n        h = self.norm(feats)\\n        a = self.att(h)\\n        a = torch.softmax(a, dim=1)\\n        pooled = torch.einsum(\"bkn,bkf->bnf\", a, h)\\n        logits = (pooled * self.clsW).sum(-1) + self.clsb\\n        return logits\\n\\n    def forward(self, x):\\n        return self.head(self.encode(x))\\n\\n\\ndef load_model(pt_path, arch_default, res_default, device, ngpu=1):\\n    ck = torch.load(pt_path, map_location=\"cpu\", weights_only=False)\\n    arch = ck.get(\"arch\", arch_default)\\n    ck_res = int(ck.get(\"res\", res_default))\\n    bb = build_backbone(arch, pretrained=False)\\n    model = RaptorClassifier(bb, F_dim=bb.num_features)\\n    model.load_state_dict(ck[\"model\"], strict=True)\\n    model.eval().to(device)\\n    # NOTE: DataParallel removed on purpose. On the full hidden test it drove a system-RAM OOM\\n    # (per-forward module replication over many studies); a single T4 handles K_EVAL=24 windows\\n    # fine. Arms are also run SEQUENTIALLY (see main) so peak RAM == one model, not two.\\n    del ck\\n    gc.collect()\\n    return model, ck_res\\n\\n\\n# ============================================================================\\n# Eval windowing -- verbatim from finetune_raptor.py StudyWindows (train=False)\\n# ============================================================================\\ndef _eval_centers(mask, D, k):\\n    valid = np.where(mask > 0)[0]\\n    if len(valid) < 3:\\n        valid = np.arange(min(3, D))\\n    lo, hi = int(valid.min()), int(valid.max())\\n    cs = [c for c in range(lo + 1, hi) if c - 1 >= lo and c + 1 <= hi]\\n    if not cs:\\n        cs = [max(1, min((lo + hi) // 2, D - 2))]\\n    idx = np.linspace(0, len(cs) - 1, k).round().astype(int)\\n    return [cs[i] for i in idx]\\n\\n\\ndef eval_windows(vol, mask, k, res, norm=NORM):\\n    D = vol.shape[0]\\n    cs = _eval_centers(mask, D, k)\\n    wins = np.empty((len(cs), 3, res, res), np.float32)\\n    for j, c in enumerate(cs):\\n        c = max(1, min(c, D - 2))\\n        tri = np.stack([vol[c - 1], vol[c], vol[c + 1]], 0).astype(np.float32) / 255.0\\n        t = torch.from_numpy(tri)\\n        if t.shape[-1] != res:\\n            t = F.interpolate(t[None], size=(res, res), mode=\"bilinear\",\\n                              align_corners=False)[0]\\n        wins[j] = t.numpy()\\n    x = torch.from_numpy(wins)\\n    if norm == \"imagenet\":\\n        x = (x - _MEAN) / _STD\\n    return x\\n\\n\\n@torch.no_grad()\\ndef infer_probs(model, xwins, device):\\n    x = xwins.unsqueeze(0).to(device)\\n    use_cuda = device != \"cpu\" and str(device).startswith(\"cuda\")\\n    if use_cuda:\\n        # fp16 conv on T4 is fully cuDNN-supported (bf16 is NOT -> \"no engine\").\\n        try:\\n            with torch.autocast(\"cuda\", dtype=torch.float16):\\n                o = torch.sigmoid(model(x).float())\\n            return o[0].cpu().numpy()\\n        except RuntimeError:\\n            # fp32 always has a Turing conv engine; slower but never drops a study.\\n            torch.cuda.empty_cache()\\n            o = torch.sigmoid(model(x).float())\\n            return o[0].cpu().numpy()\\n    o = torch.sigmoid(model(x).float())\\n    return o[0].cpu().numpy()\\n\\n\\ndef rankpct(x):                                   # per-column percentile rank in [0,1]\\n    order = x.argsort(0).argsort(0).astype(np.float64)\\n    return order / max(1, (x.shape[0] - 1))\\n\\n\\n# ============================================================================\\n# Preprocessing -- verbatim from kprep2/dino_preprocess.py, retargeted to TEST\\n# ============================================================================\\ndef _make_reader():\\n    import pydicom, cv2\\n    from pydicom.pixel_data_handlers.util import apply_modality_lut\\n\\n    def order_and_meta(sdir):\\n        fs = glob.glob(sdir + \"/*.dcm\"); recs = []; ps_list = []\\n        for f in fs:\\n            try:\\n                h = pydicom.dcmread(f, stop_before_pixels=True)\\n                iop = getattr(h, \\'ImageOrientationPatient\\', None)\\n                ipp = getattr(h, \\'ImagePositionPatient\\', None)\\n                if iop is not None and ipp is not None and len(iop) == 6:\\n                    r = np.array(iop[:3], float); c = np.array(iop[3:], float)\\n                    n = np.cross(r, c); pos = float(np.dot(np.array(ipp, float), n))\\n                else:\\n                    pos = float(getattr(h, \\'InstanceNumber\\', 0) or 0)\\n                ps = getattr(h, \\'PixelSpacing\\', None); ps = float(ps[0]) if ps is not None else 0.5\\n                ps_list.append(ps); recs.append((pos, f, ps))\\n            except Exception:\\n                recs.append((0.0, f, 0.5))\\n        recs.sort(key=lambda x: x[0])\\n        med_ps = float(np.median(ps_list)) if ps_list else 0.5\\n        return [(f, ps) for _, f, ps in recs], med_ps\\n\\n    def read_px(f):\\n        d = pydicom.dcmread(f)\\n        a = apply_modality_lut(d.pixel_array, d).astype(np.float32)\\n        if str(getattr(d, \\'PhotometricInterpretation\\', \\'\\')) == \\'MONOCHROME1\\':\\n            a = a.max() - a\\n        return a\\n\\n    def mm_crop_resize(a, ps):\\n        h, w = a.shape; cpx = int(round(CROP_MM / max(ps, 1e-3)))\\n        cpx = min(cpx, min(h, w)); y0 = (h - cpx) // 2; x0 = (w - cpx) // 2\\n        a = a[y0:y0 + cpx, x0:x0 + cpx]\\n        return cv2.resize(a, (IMG, IMG), interpolation=cv2.INTER_AREA)\\n\\n    return order_and_meta, read_px, mm_crop_resize\\n\\n\\ndef _pick_series_for_slot(rows, plane, fluid, used):\\n    cands = [r for r in rows if r[\\'Anatomical_Plane\\'] == plane and r[\\'SeriesInstanceUID\\'] not in used]\\n    if fluid in (0, 1):\\n        pref = [r for r in cands if int(r.get(\\'Fluid_Sensitive\\', 0) or 0) == fluid]\\n        if pref:\\n            return pref[0]\\n    return cands[0] if cands else None\\n\\n\\ndef build_study(sid, ser_records, tsdir, reader):\\n    order_and_meta, read_px, mm_crop_resize = reader\\n    rows = ser_records.get(sid, [])\\n    vol = np.zeros((MAXS, IMG, IMG), np.uint8); idx = 0; used = set()\\n    for plane, fluid, k in SLOTS:\\n        r = _pick_series_for_slot(rows, plane, fluid, used)\\n        if r is None:\\n            idx += k; continue\\n        used.add(r[\\'SeriesInstanceUID\\'])\\n        files, med_ps = order_and_meta(f\"{tsdir}/{sid}/{r[\\'SeriesInstanceUID\\']}\")\\n        if not files:\\n            idx += k; continue\\n        # wide span: the collateral ligaments and lateral meniscus live in the\\n        # peripheral slices the old 0.15-0.85 crop threw away. Must match the corpus\\n        # the weights were trained on (knee_corpus_v2.py, SPAN_LO/SPAN_HI).\\n        n = len(files); lo, hi = int(n * 0.06), int(n * 0.94) - 1; hi = max(hi, lo)\\n        picks = np.linspace(lo, hi, k).round().astype(int) if n > 1 else [0] * k\\n        arrs = []; pss = []\\n        for p in picks:\\n            fp, ps = files[min(p, n - 1)]\\n            try:\\n                arrs.append(read_px(fp)); pss.append(ps)\\n            except Exception:\\n                arrs.append(None); pss.append(med_ps)\\n        valid = [a for a in arrs if a is not None]\\n        if valid:\\n            allpx = np.concatenate([a.ravel() for a in valid])\\n            loq, hiq = np.percentile(allpx, [2.0, 98.0])\\n        else:\\n            loq, hiq = 0.0, 1.0\\n        for a, ps in zip(arrs, pss):\\n            if idx >= MAXS: break\\n            if a is None: idx += 1; continue\\n            aw = np.clip((a - loq) / (hiq - loq + 1e-6), 0, 1)\\n            aw = mm_crop_resize(aw, ps if ps > 0 else med_ps)\\n            vol[idx] = (aw * 255).astype(np.uint8); idx += 1\\n        if idx >= MAXS: break\\n    mask = (vol.reshape(MAXS, -1).sum(1) > 0).astype(np.uint8)\\n    return vol, mask\\n\\n\\n# ============================================================================\\n# Test-root discovery + weights + main\\n# ============================================================================\\ndef find_test_root():\\n    cands = [\"/kaggle/input/competitions/rsna-knee-abnormality-detection\",\\n             \"/kaggle/input/rsna-knee-abnormality-detection\"]\\n    for b in cands:\\n        if os.path.exists(b + \"/test.csv\"):\\n            return b\\n    for d, _, f in os.walk(\"/kaggle/input\"):\\n        if \"test.csv\" in f and (os.path.isdir(d + \"/test_series\") or os.path.isdir(d + \"/test_images\")):\\n            return d\\n    for d, _, f in os.walk(\"/kaggle/input\"):\\n        if \"test.csv\" in f:\\n            return d\\n    raise RuntimeError(\"no test root under /kaggle/input\")\\n\\n\\ndef find_weight_file(fname):\\n    # direct dataset mounts first; NEVER recursive-glob the competitions DICOM tree.\\n    direct = [f\"/kaggle/input/raptor-knee-arms/{fname}\",\\n              f\"/kaggle/input/raptor-knee-arms/1/{fname}\",\\n              f\"/kaggle/input/raptor-cnn336/{fname}\"]\\n    for p in direct:\\n        if os.path.exists(p):\\n            return p\\n    for d in sorted(glob.glob(\"/kaggle/input/*/\")):\\n        if \"competition\" in d.lower():\\n            continue\\n        hits = glob.glob(os.path.join(d, \"**\", fname), recursive=True)\\n        if hits:\\n            return hits[0]\\n    raise RuntimeError(f\"{fname} not found under /kaggle/input\")\\n\\n\\ndef main():\\n    import pandas as pd\\n    t0 = time.time()\\n    dev = \"cuda\" if torch.cuda.is_available() else \"cpu\"\\n    ngpu = torch.cuda.device_count()\\n    print(f\"device {dev} | gpus {ngpu} | torch {torch.__version__}\", flush=True)\\n\\n    ROOT = find_test_root()\\n    tsdir = ROOT + \"/test_series\"\\n    if not os.path.isdir(tsdir):\\n        tsdir = ROOT + \"/test_images\"\\n    print(\"test root:\", ROOT, \"| series dir:\", tsdir, flush=True)\\n\\n    test = pd.read_csv(ROOT + \"/test.csv\"); test[\"StudyInstanceUID\"] = test[\"StudyInstanceUID\"].astype(str)\\n    test_ids = test[\"StudyInstanceUID\"].tolist()\\n    tser = pd.read_csv(ROOT + \"/test_series.csv\")\\n    tser[\"StudyInstanceUID\"] = tser[\"StudyInstanceUID\"].astype(str)\\n    tser[\"SeriesInstanceUID\"] = tser[\"SeriesInstanceUID\"].astype(str)\\n    SER = {k: v.to_dict(\"records\") for k, v in tser.groupby(\"StudyInstanceUID\")}\\n    print(f\"test studies {len(test_ids)} | test series {len(tser)}\", flush=True)\\n\\n    sub_cols = [\"StudyInstanceUID\"] + LAB\\n    ssub = os.path.join(ROOT, \"sample_submission.csv\")\\n    if os.path.exists(ssub):\\n        sub_cols = list(pd.read_csv(ssub, nrows=1).columns)\\n\\n    reader = _make_reader()\\n    N = len(test_ids); A = len(ARMS)\\n    arm_probs = [np.full((N, len(LAB)), 0.5, np.float32) for _ in range(A)]\\n\\n    # SEQUENTIAL ARMS (the OOM fix): only ONE model is resident at a time, so peak system RAM ==\\n    # one model == the single-arm champion\\'s footprint (which graded fine at 0.879). Holding both\\n    # arms simultaneously OOM\\'d system RAM on the full hidden test. Each study is re-preprocessed\\n    # per arm (build_study is cheap vs inference) and every per-study buffer is freed. Same models,\\n    # same windowing, same rank-mean blend -> identical 0.8893 result, just serialized.\\n    for a, arm in enumerate(ARMS):\\n        wp = find_weight_file(arm[\"file\"])\\n        model, res = load_model(wp, arm[\"arch\"], arm[\"res\"], dev)\\n        print(f\"[arm {a}] loaded {arm[\\'file\\']} | res {res} | {time.time()-t0:.0f}s\", flush=True)\\n        for i, sid in enumerate(test_ids):\\n            try:\\n                vol, mask = build_study(sid, SER, tsdir, reader)\\n                xw = eval_windows(vol, mask, k=K_EVAL, res=res, norm=NORM)\\n                arm_probs[a][i] = infer_probs(model, xw, dev)\\n                del vol, mask, xw\\n            except Exception as e:\\n                print(f\"  [arm {a}] study {i} {sid[:16]} FALLBACK ({type(e).__name__}: {e})\", flush=True)\\n            if (i + 1) % 100 == 0 or i + 1 == N:\\n                print(f\"  [arm {a}] {i+1}/{N} | {time.time()-t0:.0f}s\", flush=True)\\n        del model\\n        gc.collect()\\n        if str(dev).startswith(\"cuda\"):\\n            torch.cuda.empty_cache()\\n        print(f\"[arm {a}] done + freed | {time.time()-t0:.0f}s\", flush=True)\\n\\n    # WEIGHTED rank-mean blend across the test set, per finding (the offline recipe).\\n    # Weights come from ARMS[*][\"w\"] and are normalised here, so dropping/adding an arm can\\n    # never silently change the scale. Falls back to equal weights if none are declared.\\n    _w = np.array([float(a.get(\"w\", 1.0)) for a in ARMS], dtype=np.float64)\\n    _w = _w / _w.sum()\\n    print(f\"[blend] weighted rank-mean w={dict(zip([a[\\'file\\'] for a in ARMS], _w.round(4)))}\", flush=True)\\n    ranks = np.tensordot(_w, np.stack([rankpct(np.clip(p, 0, 1)) for p in arm_probs]),\\n                         axes=(0, 0))                                          # (N,12) in [0,1]\\n    if not np.isfinite(ranks).all():\\n        ranks[~np.isfinite(ranks)] = 0.5\\n\\n    sub = pd.DataFrame(ranks.astype(np.float32), columns=LAB)\\n    sub.insert(0, \"StudyInstanceUID\", test_ids)\\n    sub = sub[sub_cols]\\n    assert list(sub.columns) == sub_cols, \"column order drift\"\\n    assert sub[\"StudyInstanceUID\"].tolist() == test_ids, \"row identity drift\"\\n    assert np.isfinite(sub[LAB].values).all()\\n    out = \"/kaggle/working/submission.csv\"\\n    sub.to_csv(out, index=False)\\n    print(\"wrote\", out, \"|\", len(sub), \"rows x\", len(sub.columns), \"cols\", flush=True)\\n    print(sub.head().to_string(index=False), flush=True)\\n    print(f\"DONE {time.time()-t0:.0f}s\", flush=True)\\n\\n\\n\\n# --- INLINE-END ---\\n\\n# 🔴 SLOTS/SPAN の代入は**必ず INLINE 区間の後**に置く。\\n#\\n# 前に置くと、埋め込んだ公開ソースの `SLOTS = [...合計64...]` に**後勝ちで上書きされる**。\\n# 2026-08-31 に実際にこれを踏み、44スライスのつもりで 64スライスのコーパスを\\n# 生成していた（[E32] の `_IMAGENET_MEAN` 事故と同型で、これが3例目）。\\n# 型も名前も同じなので例外にならず、`meta.json` を見るまで気づけなかった。\\n# 直後に assert して、静かに通り抜けないようにする。\\nSLOTS = [(\\'Sagittal\\', 1, 12), (\\'Sagittal\\', 0, 10), (\\'Coronal\\', 1, 8), (\\'Coronal\\', 0, 6), (\\'Axial\\', -1, 8)]\\nSPAN = (0.06, 0.94)\\nassert sum(s[2] for s in SLOTS) == 44, f\"SLOTS が上書きされている: {SLOTS}\"\\nassert SPAN == (0.06, 0.94), f\"SPAN が上書きされている: {SPAN}\"\\n\\n# --- INLINE-BEGIN: src/eval/raptor_arms.py::build_study_native3 ([E52] 検証済み) ---\\ndef build_study_native3(ns, sid, rows, tsdir, reader, slots, span):\\n    \"\"\"`build_study` と**同じ格子**で体積を作りつつ、各位置に元スライスの真の隣接3枚を持たせる。\\n\\n    ## なぜ必要か\\n\\n    公開実装は各シリーズを 6〜94% の帯から k 枚へ間引き、**その間引き後の連続3枚**を\\n    RGB に積む。gold58 の生 DICOM から実測すると、この3チャンネルの**物理距離**は\\n\\n    | スロット | チャンネル間隔（5/50/95%tile） |\\n    |---|---|\\n    | Sagittal FS | 6.7 / 8.0 / 10.6 mm |\\n    | Coronal 非FS | **15.4 / 17.7 / 24.2 mm** |\\n    | Axial | 6.5 / 16.6 / 20.1 mm |\\n\\n    と **6〜24mm** に達する（元スライス厚は 2.2〜5.5mm）。\\n    **「隣接3枚の2.5D」という説明は実態と合っていない**——冠状非FSでは3枚で 35mm、\\n    膝の前後径の大半をまたぐ。`doc/medical/` の\\n    「連続2スライス以上で関節面に接触＝断裂（two-slice touch rule）」\\n    「5mmスライスで連続3 segment 以上＝円板状半月板」\\n    「骨折線が1スライスのみなら Contusion に倒す」\\n    はいずれもこの入力では原理的に評価できない。\\n\\n    ## 何を変えるか\\n\\n    **サンプリング格子（＝被覆）は一切変えない。** 変えるのは各格子点の3チャンネルの出所だけで、\\n    間引き後の隣ではなく**元シリーズの真の隣**（`files[p-1], files[p], files[p+1]`）を使う。\\n    これでチャンネル間隔が元スライス厚（2〜5mm）になり、局所的な3D文脈になる。\\n\\n    Returns:\\n        `(MAXS, 3, IMG, IMG)` uint8 と `(MAXS,)` の mask。\\n    \"\"\"\\n    order_and_meta, read_px, mm_crop_resize = reader\\n    img = ns[\"IMG\"]\\n    maxs = sum(s[2] for s in slots)\\n    vol = np.zeros((maxs, 3, img, img), np.uint8)\\n    idx = 0\\n    used: set = set()\\n    lo_f, hi_f = span\\n    for plane, fluid, k in slots:\\n        r = ns[\"_pick_series_for_slot\"](rows, plane, fluid, used)\\n        if r is None:\\n            idx += k\\n            continue\\n        used.add(r[\"SeriesInstanceUID\"])\\n        files, med_ps = order_and_meta(f\"{tsdir}/{sid}/{r[\\'SeriesInstanceUID\\']}\")\\n        if not files:\\n            idx += k\\n            continue\\n        n = len(files)\\n        lo, hi = int(n * lo_f), int(n * hi_f) - 1\\n        hi = max(hi, lo)\\n        picks = np.linspace(lo, hi, k).round().astype(int) if n > 1 else [0] * k\\n        # 正規化の分位点は**元実装と同じ母集団**（格子上の k 枚）から取る。\\n        # 隣接ぶんも混ぜると窓の輝度分布が変わり、変更点が2つになる。\\n        center_arrs = []\\n        for p in picks:\\n            try:\\n                center_arrs.append(read_px(files[min(p, n - 1)][0]))\\n            except Exception:                                     # noqa: BLE001\\n                center_arrs.append(None)\\n        valid = [a for a in center_arrs if a is not None]\\n        if valid:\\n            loq, hiq = np.percentile(np.concatenate([a.ravel() for a in valid]), [2.0, 98.0])\\n        else:\\n            loq, hiq = 0.0, 1.0\\n        for p, ca in zip(picks, center_arrs):\\n            if idx >= maxs:\\n                break\\n            if ca is None:\\n                idx += 1\\n                continue\\n            for ch, q in enumerate((p - 1, p, p + 1)):\\n                q = int(min(max(q, 0), n - 1))\\n                fp, ps = files[q]\\n                try:\\n                    a = read_px(fp) if q != p else ca\\n                except Exception:                                 # noqa: BLE001\\n                    a = ca\\n                aw = np.clip((a - loq) / (hiq - loq + 1e-6), 0, 1)\\n                aw = mm_crop_resize(aw, ps if ps > 0 else med_ps)\\n                vol[idx, ch] = (aw * 255).astype(np.uint8)\\n            idx += 1\\n        if idx >= maxs:\\n            break\\n    mask = (vol.reshape(maxs, -1).sum(1) > 0).astype(np.uint8)\\n    return vol, mask\\n\\n# --- INLINE-END ---\\n\\nMAXS_44 = sum(s[2] for s in SLOTS)\\n_NS = {\"IMG\": IMG, \"_pick_series_for_slot\": _pick_series_for_slot}\\n_READER = None\\n\\n\\ndef one(payload):\\n    \"\"\"1 study 分の `(MAXS,3,IMG,IMG)` と mask を返す。失敗しても全体を落とさない。\"\"\"\\n    global _READER\\n    sid, rows, tsdir = payload\\n    if _READER is None:\\n        _READER = _make_reader()\\n    try:\\n        return build_study_native3(_NS, sid, rows, tsdir, _READER, SLOTS, SPAN)\\n    except Exception as e:                                        # noqa: BLE001\\n        print(f\"  FAIL {sid[:16]} {type(e).__name__}: {e}\", flush=True)\\n        return (np.zeros((MAXS_44, 3, IMG, IMG), np.uint8),\\n                np.zeros(MAXS_44, np.uint8))\\n')\nsys.path.insert(0, \"/kaggle/working\")\nimport native3_worker as W\n\nROOT = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\nif not ROOT.exists():\n    for c in Path(\"/kaggle/input\").glob(\"**/train.csv\"):\n        ROOT = c.parent\n        break\nprint(\"ROOT:\", ROOT, \"| MAXS:\", W.MAXS_44, \"| IMG:\", W.IMG, \"| CROP_MM:\", W.CROP_MM, flush=True)\n\ntrain = pd.read_csv(ROOT / \"train.csv\")\nser = pd.read_csv(ROOT / \"train_series.csv\")\nser[\"StudyInstanceUID\"] = ser[\"StudyInstanceUID\"].astype(str)\nser[\"SeriesInstanceUID\"] = ser[\"SeriesInstanceUID\"].astype(str)\nser_map = {k: v.to_dict(\"records\") for k, v in ser.groupby(\"StudyInstanceUID\")}\n\nuids = sorted(train[\"StudyInstanceUID\"].astype(str).unique().tolist())\nb = [round(i * len(uids) / N_SHARD) for i in range(N_SHARD + 1)]\nmine = uids[b[SHARD]:b[SHARD + 1]]\nif SMOKE:\n    mine = mine[:8]\nprint(f\"shard {SHARD}/{N_SHARD}: {len(mine)} study / 全{len(uids)}\", flush=True)\n\ntsdir = str(ROOT / \"train_series\")\nfrom joblib import Parallel, delayed\n\nt0 = time.time()\nn_job = max(1, (os.cpu_count() or 2) - 1)\n# 🔴 part 単位で書き出す。882 study x 44 x 3 x 336^2 = 13GB は RAM に載らない。\nsplits = [round(i * len(mine) / N_SPLIT) for i in range(N_SPLIT + 1)]\nmasks_all, ids_all = [], []\nfor k in range(N_SPLIT):\n    part = mine[splits[k]:splits[k + 1]]\n    if not part:\n        continue\n    res = Parallel(n_jobs=n_job, backend=\"loky\")(\n        delayed(W.one)((s, ser_map.get(s, []), tsdir)) for s in part)\n    np.save(f\"/kaggle/working/native3_vols_part{k}.npy\",\n            np.stack([r[0] for r in res]))\n    masks_all.append(np.stack([r[1] for r in res]))\n    ids_all.extend(part)\n    del res\n    print(f\"  part{k}: {len(part)} study | {time.time()-t0:.0f}s\", flush=True)\n\nnp.save(\"/kaggle/working/native3_masks.npy\", np.concatenate(masks_all))\nnp.save(\"/kaggle/working/native3_ids.npy\", np.array(ids_all))\njson.dump({\"shard\": SHARD, \"n_shard\": N_SHARD, \"n_study\": len(ids_all),\n           \"slots\": W.SLOTS, \"span\": W.SPAN, \"maxs\": W.MAXS_44, \"img\": W.IMG,\n           \"crop_mm\": W.CROP_MM, \"n_split\": N_SPLIT},\n          open(\"/kaggle/working/native3_meta.json\", \"w\"), indent=1)\nprint(f\"DONE {time.time()-t0:.0f}s | {len(ids_all)} study\", flush=True)\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}