{"cells":[{"cell_type":"markdown","id":"88e28ec2","metadata":{},"source":"# RSNA Knee DINO–RadImageNet–Raptor Ensemble\n\nAn inference-only, one-dataset reproduction of [Evgen Dvorkin's public RSNA Baseline](https://www.kaggle.com/code/evgendvorkin/rsna-baseline), pinned at version 15. It uses **only the public model assets used by that notebook**.\n\nThe source recipe has a verified **0.941 public LB** result. This consolidated version must be scored separately; **0.942 is a target, not a claimed result**. No training or competition submission runs here.\n\n## How many models?\n\n**41 unique checkpoint members**, across five groups:\n\n| Group | Unique members | How predictions are combined |\n|---|---:|---|\n| DINOv2-small | 20 | Per-member, per-target percentile ranks, equal mean |\n| DINOv3-small | 5 | Five fold ranks, equal mean; 45% of the transformer blend |\n| RadImageNet attention heads | 10 | Five reference + five E13 heads, sharing one ResNet-50 encoder |\n| Public Raptor CoAtNet | 3 | v5, v10, v8; v5 is reused for a reverse-channel view |\n| Public residual-gated CoAtNet | 3 | Epochs 4, 6, 8 from one training run; equal mean of checkpoint ranks |\n\nThe DINOv2 initialization asset and shared RadImageNet ResNet-50 encoder are included in the dataset. They are not extra ensemble votes. DINOv2 loads each complete competition checkpoint directly into its configured architecture, avoiding redundant initialization-weight loading. A frozen **88-input, 12-output linear calibration table** is also included. Repeated views and epoch snapshots are not independent training folds.\n\n## Workflow and blend\n\n```text\nCompetition DICOMs\n  ├─ 20 DINOv2-small → mean of ranks ─┐\n  ├─  5 DINOv3-small → mean of ranks ─┴─ 55% / 45% transformer blend\n  │                                              │\n  ├─ RadImageNet encoder → reference/E13 heads ────┤\n  │    E13 layout + second layout + flipped view  │\n  │                                  frozen calibration → T\n  │\n  ├─ Raptor v5/v10/v5-reverse/v8 → 60/10/10/20 probability mean → P\n  └─ Residual CoAtNet epochs 4/6/8 → mean of ranks → C\n\n              H = rank(0.60 × rank(P) + 0.40 × rank(C))\n  final[target] = rank((1 − w[target]) × rank(T) + w[target] × H)\n                                      │\n                               submission.csv\n```\n\n`rank` means per-target average percentile rank across **all** test studies. The public Raptor stage first applies the original ordinal rank implementation; the hybrid re-ranks it exactly as the source does. Rank positions and probability means cannot be interchanged.\n\n| Finding | Raptor hybrid H | DINO/Rad/calibration T |\n|---|---:|---:|\n| ACL, Lateral OA, Fracture | 75% | 25% |\n| Medial Meniscus | 80% | 20% |\n| Lateral Meniscus | 100% | 0% |\n| MCL, Medial OA, PF OA, Effusion, Synovitis, Baker's, Contusion | 60% | 40% |\n\nInside **T**, reference and E13 Rad ranks are mixed 50/50. That Rad branch receives a 50% vote on ten findings; Baker's and Fracture retain the transformer values at this step. The same E13 heads are then reused on the second slot layout with normal and horizontal-flip views; this branch receives 15%. Finally, the frozen calibrator adds a 40% rank vote on ACL, Medial OA, Lateral OA, PF OA, Effusion, Baker's and Contusion.\n\n## Inputs and reproducibility\n\nAttach only the competition and **one dataset**: [consolidated assets](https://www.kaggle.com/datasets/tonylica/rsna-knee-bend-dinov3-0917-repro-assets). The notebook verifies the exact manifest and every file, then runs readable Python modules from `pipeline/`. The dataset contains checkpoints, model definitions, calibration, the pinned OpenCV wheel, source provenance and workflow documentation. No private experimental model, report-label table or training cache is needed.\n\nPreprocessing remains specific to each trained family: DINO crops use 130 mm; Raptor uses its 140 mm/336–384 px settings; the residual CoAtNet uses its own 130 mm physical triplet renderer. These representations are preserved rather than forced into one image cache.\n\n## Credits\n\nBuilt from the public work of **Evgen Dvorkin, Mattia Angeli, Dread Development / Johnathan Wagner, Pilkwang, Antoine, Sofia Anjenje, Prvsiyan, Marwan and Meta**, as attributed by the source notebook and its asset providers. Exact source references and component licenses are recorded in `SOURCE_AND_LICENSES.md` and `bundle_manifest.json`. This package reproduces checkpoint inference; it does not claim that all upstream models can be retrained bit-for-bit from the material their authors published.\n"},{"cell_type":"markdown","id":"d82ade97","metadata":{},"source":"## 1. Verify the complete asset bundle\n\nEvery file is checked before loading models. This run requires an offline T4 × 2 session."},{"cell_type":"code","execution_count":null,"id":"51197ecd","metadata":{},"outputs":[],"source":"import hashlib\nimport json\nimport os\nfor key in ('OMP_NUM_THREADS', 'OPENBLAS_NUM_THREADS', 'MKL_NUM_THREADS'):\n    os.environ.setdefault(key, '4')\nfrom pathlib import Path\nimport subprocess\nimport sys\nimport time\nimport pandas as pd\n\nASSET_SLUG = 'rsna-knee-bend-dinov3-0917-repro-assets'\nASSET = next((p for p in [\n    Path('/kaggle/input/datasets/tonylica') / ASSET_SLUG,\n    Path('/kaggle/input') / ASSET_SLUG,\n] if (p / 'bundle_manifest.json').is_file()), None)\nROOT = next((p for p in [\n    Path('/kaggle/input/competitions/rsna-knee-abnormality-detection'),\n    Path('/kaggle/input/rsna-knee-abnormality-detection'),\n] if (p / 'test.csv').is_file()), None)\nassert ASSET is not None, 'Attach the consolidated reproduction dataset'\nassert ROOT is not None and (ROOT / 'test_series').is_dir(), 'Attach the competition data'\n\ndef sha256_file(path):\n    digest = hashlib.sha256()\n    with path.open('rb') as stream:\n        for block in iter(lambda: stream.read(8 << 20), b''):\n            digest.update(block)\n    return digest.hexdigest()\n\nmanifest_path = ASSET / 'bundle_manifest.json'\nobserved_manifest_sha256 = sha256_file(manifest_path)\nassert observed_manifest_sha256 == '3b6279ab50ca8c3ce100644eadda9bc14e0359c05120990e28a43edc6a5364a8', f'Dataset version or manifest mismatch: {observed_manifest_sha256}'\nmanifest = json.loads(manifest_path.read_text())\nfor record in manifest['files']:\n    path = ASSET / record['path']\n    assert path.is_relative_to(ASSET) and path.is_file(), record['path']\n    assert path.stat().st_size == record['bytes'], f\"Size mismatch: {record['path']}\"\n    assert sha256_file(path) == record['sha256'], f\"Hash mismatch: {record['path']}\"\nprint(f\"Verified {manifest['file_count']} files; all models come from public Baseline v15.\")\nos.environ.update(RSNA_ASSET_ROOT=str(ASSET), RSNA_COMPETITION_ROOT=str(ROOT),\n                  HF_HUB_OFFLINE='1', TRANSFORMERS_OFFLINE='1', HF_HUB_DISABLE_TELEMETRY='1')\nWORK = Path('/kaggle/working')\n(WORK / 'branches').mkdir(exist_ok=True)\n# Remove only a stale primary artifact from a previous interactive run.\n(WORK / 'submission.csv').unlink(missing_ok=True)\nSTAGE_TIMES = {}\n\ndef run_stage(script):\n    started = time.monotonic()\n    subprocess.run([sys.executable, '-u', str(ASSET / 'pipeline' / script)],\n                   cwd=WORK, check=True, env=dict(os.environ))\n    STAGE_TIMES[script] = round(time.monotonic() - started, 3)\n    print(f\"Completed {script} in {STAGE_TIMES[script]:.1f}s\")\n    (WORK / 'stage_times.json').write_text(json.dumps(STAGE_TIMES, indent=2))\n"},{"cell_type":"markdown","id":"11e3f043","metadata":{},"source":"## 2. DINO and RadImageNet branch\n\nRuns 20 DINOv2 models, five DINOv3 models, ten Rad heads with their prescribed views, then the frozen calibration.\n\nReadable implementation: `pipeline/transformer_rad.py` in the attached dataset."},{"cell_type":"code","execution_count":null,"id":"484fb7d5","metadata":{},"outputs":[],"source":"run_stage('transformer_rad.py')"},{"cell_type":"markdown","id":"f989734e","metadata":{},"source":"## 3. Public Raptor branch\n\nRuns three CoAtNet checkpoints in four views and combines probabilities at 60/10/10/20.\n\nReadable implementation: `pipeline/public_raptor.py` in the attached dataset."},{"cell_type":"code","execution_count":null,"id":"2a5f88f6","metadata":{},"outputs":[],"source":"run_stage('public_raptor.py')"},{"cell_type":"markdown","id":"ba095889","metadata":{},"source":"## 4. Residual CoAtNet branch\n\nRuns epochs 4, 6 and 8 with the pinned OpenCV renderer, then averages the three per-target rank vectors.\n\nReadable implementation: `pipeline/residual_coat.py` in the attached dataset."},{"cell_type":"code","execution_count":null,"id":"3449a5c7","metadata":{},"outputs":[],"source":"run_stage('residual_coat.py')"},{"cell_type":"markdown","id":"9f8ecc77","metadata":{},"source":"## 5. Final target-specific ensemble\n\nForms the 60/40 Raptor hybrid, applies the target weights above, and validates every submission row.\n\nReadable implementation: `pipeline/blend.py` in the attached dataset."},{"cell_type":"code","execution_count":null,"id":"c7243537","metadata":{},"outputs":[],"source":"run_stage('blend.py')"},{"cell_type":"markdown","id":"663eb9c5","metadata":{},"source":"## 6. Submission artifact\n\nAfter all checks pass, use this notebook version and select `submission.csv` in Kaggle’s **Submit to Competition** flow. This notebook does not submit automatically."},{"cell_type":"code","execution_count":null,"id":"9b6160c9","metadata":{},"outputs":[],"source":"submission = pd.read_csv('/kaggle/working/submission.csv', dtype={'StudyInstanceUID': str})\nprint(f'Ready: {len(submission)} studies, {len(submission.columns)-1} targets')\ndisplay(submission.head())\ndisplay(json.loads(Path('/kaggle/working/run_receipt.json').read_text()))"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Overlay Medial Meniscus T30/R60/bag10 (L1-V6 recipe) di atas fork-941 final.\nKomponen: pre-final-rerank transformer + hybrid_raptor (yang dihasilkan blend.py),\nbag dari meniscus-bag public0033. Output: overwrite submission.csv + receipt.\"\"\"\n\nimport json as _ov_json\nimport os as _ov_os\nimport numpy as _ov_np\nimport pandas as _ov_pd\n\n_W = '/kaggle/working'\n_ASSET = _ov_os.environ['RSNA_ASSET_ROOT']\n_ROOT = _ov_os.environ['RSNA_COMPETITION_ROOT']\n_TARGET = 'Medial Meniscus'\n_W_TR, _W_CR, _W_BAG = 0.3, 0.6, 0.1\nassert abs(_W_TR + _W_CR + _W_BAG - 1.0) < 1e-9\n\ndef _ov_rank_pct(v):\n    s = _ov_pd.Series(v)\n    return s.rank(method='average', pct=True).to_numpy()\n\ndef _ov_sha(path):\n    import hashlib\n    h = hashlib.sha256()\n    with open(path, 'rb') as f:\n        for b in iter(lambda: f.read(8 << 20), b''):\n            h.update(b)\n    return h.hexdigest()\n\n# --- load parent final submission (dari blend.py, sudah tervalidasi) ---\n_sub = _ov_pd.read_csv(f'{_W}/submission.csv', dtype={'StudyInstanceUID': str})\nLABELS = [c for c in _sub.columns if c != 'StudyInstanceUID']\nassert _TARGET in LABELS and len(LABELS) == 12, 'schema drift'\n_ids = _ov_pd.read_csv(f'{_ROOT}/test.csv', dtype={'StudyInstanceUID': str}).StudyInstanceUID.tolist()\nassert _sub.StudyInstanceUID.tolist() == _ids, 'uid drift'\n_backup = f'{_W}/submission_parent_exact.csv'\n_sub.to_csv(_backup, index=False)\n\n# --- komponen pre-rerank: transformer & hybrid raptor branch ---\n_tr = _ov_pd.read_csv(f'{_W}/branches/transformer_rad.csv', dtype={'StudyInstanceUID': str})\n_hy = _ov_pd.read_csv(f'{_W}/branches/hybrid_raptor.csv', dtype={'StudyInstanceUID': str})\nassert _hy.StudyInstanceUID.tolist() == _ids and _tr.StudyInstanceUID.tolist() == _ids\n\n# --- bag rank dari public0033 (glob robust path mount Kaggle 2026) ---\nimport glob as _ov_glob\nfrom pathlib import Path as _ov_Path\n_bag_cands = _ov_glob.glob('/kaggle/input/**/public0033*bag*.csv', recursive=True) or \\\n             _ov_glob.glob('/kaggle/input/**/meniscus*bag*.csv', recursive=True) or \\\n             _ov_glob.glob('/kaggle/input/**/bag_raw.csv', recursive=True)\nassert _bag_cands, 'meniscus bag csv tidak ditemukan di input'\n_bag = _ov_pd.read_csv(sorted(_bag_cands)[0], dtype={'StudyInstanceUID': str})\n_bag = _bag.set_index('StudyInstanceUID').loc[_ids].reset_index()\n_bag_rank = _ov_rank_pct(_bag[_TARGET].to_numpy())\n\n# --- rerank target: rank(0.3*tr + 0.6*raptor_hybrid + 0.1*bag) ---\n_tr_rank = _ov_rank_pct(_tr[_TARGET].to_numpy())\n_cr_rank = _ov_rank_pct(_hy[_TARGET].to_numpy())\n_blend = _W_TR * _tr_rank + _W_CR * _cr_rank + _W_BAG * _bag_rank\n_sub[_TARGET] = _ov_rank_pct(_blend)\n\n# --- validasi: 11 kolom lain TIDAK berubah (pola p33: untouched digests) ---\n_par = _ov_pd.read_csv(_backup, dtype={'StudyInstanceUID': str})\nfor c in LABELS:\n    if c == _TARGET:\n        continue\n    assert _ov_np.allclose(_sub[c].to_numpy(), _par[c].to_numpy(), atol=0), f'{c} berubah!'\n\n# --- tulis + receipt ---\n_tmp = f'{_W}/submission.csv.tmp'\n_sub.to_csv(_tmp, index=False)\n_ov_os.replace(_tmp, f'{_W}/submission.csv')\n_fin = _ov_pd.read_csv(f'{_W}/submission.csv', dtype={'StudyInstanceUID': str})\nassert _fin.StudyInstanceUID.tolist() == _ids and _fin[_TARGET].between(0, 1).all()\n_receipt = {\n    'schema_version': 'l1v6_medial_t30r60bag10_on_fork941_v1',\n    'status': 'passed',\n    'target': _TARGET, 'weights': {'transformer': _W_TR, 'raptor_hybrid': _W_CR, 'bag': _W_BAG},\n    'parent_submission_sha256': _ov_sha(_backup),\n    'output_sha256': _ov_sha(f'{_W}/submission.csv'),\n    'source_notebook': 'mtitoadhip/knee-sub-tonyrank',\n    'recipe_origin': 'p33 public0033 Medial T30/R60/bag10 overlay (terbukti COMPLETE 15/9)',\n}\n(_ov_Path(_W) / 'overlay_receipt.json').write_text(_ov_json.dumps(_receipt, indent=2) + '\\n')\nprint(f'[L1-ON-941] overlay PASSED | target={_TARGET} | out_sha={_receipt[\"output_sha256\"]}', flush=True)\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"rsna_reproduction":{"new_version_scored":false,"source_sha256":"eb850464cec6a406d36a225625cc4bf22916720ec0175c8aab1fcf6023eec739","upstream":"evgendvorkin/rsna-baseline/15","verified_source_lb":0.941}},"nbformat":4,"nbformat_minor":5}