{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"phase2_protocol":"1.0.0"},"nbformat_minor":5,"nbformat":4,"cells":[{"id":"834b081c","cell_type":"markdown","source":"# Phase II — Performance-Qualified Adult Chest-Radiograph Proof of Concept\n\n## Abstract\n\nPhase I demonstrated high internal discrimination but substantial external degradation, poor calibration,\nand near-zero specificity for AP acquisitions. Phase II therefore evaluates a narrower proposition: whether\na compact adult chest-radiograph model can be selected under a prespecified qualification protocol,\nevaluated once on a locked independent radiologist-annotated cohort, exported without numerical drift, and integrated\ninto an offline educational proof of concept. Adult RSNA images form the development cohort. CheXpert is\nused for qualification because it was previously inspected, while the official 15,000-image VinBigData\ncompetition training cohort is reserved for a single locked proxy evaluation. The external endpoint is\nradiographic air-space opacity rather than confirmed pneumonia. The system is a research and engineering demonstration and is\nnot intended for diagnosis, triage, treatment, or clinical deployment.","metadata":{}},{"id":"c21ded4f","cell_type":"markdown","source":"### Reproducible environment and project discovery\n\nThis cell locates the clean Phase II package in either a local checkout or a Kaggle dataset mount. It records\nthe interpreter and accelerator but does not inspect any radiograph or outcome. The input is the source tree;\nthe output is a resolved project root placed on `sys.path`. Resolution must be unique so that an older package\ncannot be imported accidentally. Failure to locate exactly one `phase2_adult_cxr_poc` project stops execution.\nThis administrative cell cannot influence model or artifact selection.","metadata":{}},{"id":"857f00cc","cell_type":"code","source":"\"\"\"Resolve the clean Phase II source tree and record the execution environment.\"\"\"\nfrom pathlib import Path\nimport hashlib\nimport os\nimport platform\nimport sys\nimport zipfile\n\nos.environ.setdefault(\"CUBLAS_WORKSPACE_CONFIG\", \":4096:8\")\nLOCAL = Path.cwd().resolve()\nKAGGLE_INPUT = Path(\"/kaggle/input\")\ncandidates = [LOCAL] if (LOCAL / \"adult_cxr\" / \"__init__.py\").is_file() else []\nif KAGGLE_INPUT.exists():\n    candidates += [path.parent.parent for path in KAGGLE_INPUT.rglob(\"adult_cxr/__init__.py\")]\n    if not candidates:\n        matching_archives = []\n        for archive_path in sorted(KAGGLE_INPUT.rglob(\"*.zip\")):\n            try:\n                with zipfile.ZipFile(archive_path) as archive:\n                    members = archive.namelist()\n                    if any(name.replace(\"\\\\\", \"/\").endswith(\"adult_cxr/__init__.py\") for name in members):\n                        matching_archives.append(archive_path)\n            except zipfile.BadZipFile:\n                continue\n        if len(matching_archives) == 1:\n            digest = hashlib.sha256(matching_archives[0].read_bytes()).hexdigest()[:12]\n            extraction_root = Path(\"/kaggle/working/phase2_source\") / digest\n            with zipfile.ZipFile(matching_archives[0]) as archive:\n                archive.extractall(extraction_root)\n            candidates += [\n                path.parent.parent\n                for path in extraction_root.rglob(\"adult_cxr/__init__.py\")\n            ]\ncandidates = sorted(set(path.resolve() for path in candidates))\nif len(candidates) != 1:\n    raise RuntimeError(f\"Expected one clean Phase II project; found {candidates}\")\n\nPROJECT = candidates[0]\nsys.path.insert(0, str(PROJECT))\n# Absolute script execution places ``scripts/`` rather than the project root on\n# the child interpreter's import path. Export the root once so every later\n# subprocess can import the sibling ``adult_cxr`` package on Kaggle.\nos.environ[\"PYTHONPATH\"] = str(PROJECT) + os.pathsep + os.environ.get(\"PYTHONPATH\", \"\")\nimport importlib\n\nfor module_name in tuple(sys.modules):\n    if module_name == \"adult_cxr\" or module_name.startswith(\"adult_cxr.\"):\n        del sys.modules[module_name]\nimportlib.invalidate_caches()\nimport adult_cxr\n\nEXPECTED_SOURCE_VERSION = \"0.4.1\"\npackage_path = Path(adult_cxr.__file__).resolve()\nif adult_cxr.__version__ != EXPECTED_SOURCE_VERSION:\n    raise RuntimeError(\n        f\"Expected Phase II source {EXPECTED_SOURCE_VERSION}, loaded {adult_cxr.__version__}. \"\n        \"Replace the attached project dataset and rerun this cell.\"\n    )\nif not package_path.is_relative_to(PROJECT):\n    raise RuntimeError(f\"adult_cxr was imported outside the selected project: {package_path}\")\n\nprint({\n    \"project\": str(PROJECT),\n    \"package\": str(package_path),\n    \"source_version\": adult_cxr.__version__,\n    \"python\": sys.version.split()[0],\n    \"platform\": platform.platform(),\n})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"820fd7dd","cell_type":"markdown","source":"### Runtime preflight before expensive computation\n\nThe export and locked external stages require ONNX and ONNX Runtime. They are checked before any dataset is\nindexed or model is trained so a missing package cannot first appear after hours of computation. On a Kaggle\nGPU session, the cell attempts to install the GPU-enabled ONNX Runtime when CUDA execution is not already\navailable; otherwise it records a CPU fallback. This affects execution speed only: the same locked ONNX graphs,\ninputs, calibrated logits, metrics, and gates are used. Package installation failure is surfaced immediately.","metadata":{}},{"id":"068d62d9","cell_type":"code","source":"\"\"\"Validate late-stage dependencies and expose the fastest available ONNX provider.\"\"\"\nimport importlib.util\nimport json\nimport subprocess\n\nrequired_packages = {\n    \"onnx\": \"onnx>=1.16,<2\",\n    \"onnxruntime\": \"onnxruntime>=1.18,<2\",\n    \"psutil\": \"psutil>=6,<8\",\n}\nmissing_packages = [requirement for module, requirement in required_packages.items()\n                    if importlib.util.find_spec(module) is None]\nif missing_packages:\n    print(\"Installing missing late-stage dependencies before computation:\", missing_packages, flush=True)\n    subprocess.run([sys.executable, \"-m\", \"pip\", \"install\", \"--quiet\", *missing_packages], check=True)\n\nimport onnx\nimport onnxruntime as ort\nimport psutil\nimport torch\n\nif torch.cuda.is_available() and \"CUDAExecutionProvider\" not in ort.get_available_providers():\n    print(\"CUDA is available but ONNX Runtime is CPU-only; attempting the GPU runtime now.\", flush=True)\n    subprocess.run(\n        [sys.executable, \"-m\", \"pip\", \"install\", \"--quiet\", \"--upgrade\", \"onnxruntime-gpu>=1.18,<2\"],\n        check=True,\n    )\n    # The current interpreter may retain an imported binary module. A child process gives the authoritative\n    # provider list that later standalone evaluators will receive.\nprovider_check = subprocess.run(\n    [sys.executable, \"-c\", \"import json, onnxruntime as o; print(json.dumps(o.get_available_providers()))\"],\n    check=True, capture_output=True, text=True,\n)\nonnx_providers = json.loads(provider_check.stdout.strip().splitlines()[-1])\nexternal_onnx_provider = (\n    \"CUDAExecutionProvider\" if \"CUDAExecutionProvider\" in onnx_providers\n    else \"CPUExecutionProvider\"\n)\nprint({\"onnx\": onnx.__version__, \"onnxruntime_providers\": onnx_providers,\n       \"locked_external_provider\": external_onnx_provider})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"a804b5d6","cell_type":"markdown","source":"## 1. Phase I motivation and frozen protocol\n\nPhase I is treated as immutable prior evidence. Its principal limitations—external discrimination loss,\ncalibration error, and acquisition-view sensitivity—motivate the Phase II endpoints without allowing its\nresults to be rewritten. This cell loads the versioned YAML protocol and creates a new artifact directory.\nThe protocol fixes seeds, cohorts, model registry, epoch horizon, gate definitions, and engineering limits.\nNo result-dependent default is introduced. The displayed digest makes any later protocol modification\ndetectable; a missing section or malformed mapping is a fatal error and prevents all downstream stages.","metadata":{}},{"id":"83089407","cell_type":"code","source":"\"\"\"Load and fingerprint the frozen Phase II protocol.\"\"\"\nfrom datetime import datetime, timezone\nimport json\nimport pandas as pd\nimport numpy as np\nimport torch\n\nfrom adult_cxr.config import load_protocol\nfrom adult_cxr.reproducibility import sha256_file\n\nCONFIG_PATH = PROJECT / \"configs\" / \"protocol.yaml\"\nPROTOCOL = load_protocol(CONFIG_PATH)\nOUTPUT = (\n    Path(\"/kaggle/working/phase2_adult_cxr_poc_v041/artifacts\")\n    if Path(\"/kaggle/working\").exists()\n    else PROJECT / \"artifacts\"\n)\nOUTPUT.mkdir(parents=True, exist_ok=True)\nresume_archive = os.environ.get(\"PHASE2_RESUME_ARCHIVE\")\nif resume_archive:\n    subprocess.run([\n        sys.executable, str(PROJECT / \"scripts\" / \"snapshot_artifacts.py\"), \"--restore\",\n        \"--artifacts\", str(OUTPUT), \"--archive\", resume_archive,\n    ], cwd=PROJECT, check=True)\nprint({\n    \"protocol\": PROTOCOL.section(\"protocol\")[\"version\"],\n    \"sha256\": sha256_file(CONFIG_PATH),\n    \"started_utc\": datetime.now(timezone.utc).isoformat(),\n    \"cuda_available\": torch.cuda.is_available(),\n})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"c5e9f19d","cell_type":"markdown","source":"## 2. Study design and evidence boundaries\n\nThe diagram encodes the direction of information flow. RSNA training, validation, calibration, and\ndevelopment-test partitions support fitting and internal diagnosis. CheXpert report and expert cohorts can\nqualify A1 versus A2 but cannot be described as untouched. VinBigData remains inaccessible until the exported\nartifact and protocol have been hashed. This separation prevents final-cohort information from changing the\nmodel. The figure is generated from protocol text, saved as PNG and SVG, and reloaded from the PNG so a broken\nsaved illustration fails visibly. It is explanatory only and cannot influence selection.","metadata":{}},{"id":"728a1507","cell_type":"code","source":"\"\"\"Draw and persist the Phase I-to-deployment evidence flow.\"\"\"\nimport matplotlib.pyplot as plt\nfrom matplotlib.patches import FancyBboxPatch, FancyArrowPatch\nfrom PIL import Image\n\nFIGURES = OUTPUT / \"figures\"\nFIGURES.mkdir(parents=True, exist_ok=True)\nfig, ax = plt.subplots(figsize=(12, 3.5))\nax.set_xlim(0, 12)\nax.set_ylim(0, 3.5)\nax.axis(\"off\")\nlabels = [\"Phase I\\nimmutable evidence\", \"Adult RSNA\\ndevelopment\", \"CheXpert\\nqualification\",\n          \"Artifact lock\", \"VinBigData\\nproxy evaluation\", \"Offline PoC\"]\ncolors = [\"#4C78A8\", \"#72B7B2\", \"#F2CF5B\", \"#B279A2\", \"#E45756\", \"#54A24B\"]\nfor index, (label, color) in enumerate(zip(labels, colors)):\n    x = 0.15 + index * 1.98\n    ax.add_patch(FancyBboxPatch((x, 1.1), 1.6, 1.25, boxstyle=\"round,pad=.08\", fc=color, ec=\"white\"))\n    ax.text(x + .8, 1.72, label, ha=\"center\", va=\"center\", fontsize=9, weight=\"bold\")\n    if index < len(labels) - 1:\n        ax.add_patch(FancyArrowPatch((x + 1.62, 1.72), (x + 1.94, 1.72), arrowstyle=\"->\", mutation_scale=13))\nax.text(8.35, .55, \"No access before lock\", color=\"#E45756\", ha=\"center\", weight=\"bold\")\nax.set_title(\"Figure 1. Prespecified evidence flow and external lock boundary\", loc=\"left\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_01_evidence_flow.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_01_evidence_flow.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4b79a405","cell_type":"markdown","source":"## 3. Methods: adult RSNA cohort and age audit\n\nThe RSNA cohort exposes age either as a DICOM `AS` value (years, months, weeks, or days) or as the challenge's\nplain numeric years. Both representations are converted explicitly to years. An adult image satisfies\n(18 \\leq a \\leq 120) years and has AP or PA `ViewPosition`. Labels are reduced\nto one row per image before decoded pixels are hashed. Exact duplicates with consistent outcomes are removed;\nconflicting outcomes stop the study. The cell writes the eligible manifest, mutually exclusive age audit, and\nduplicate audit. RSNA is development data only. If the resulting adult cohort is empty or any source image is\nmissing, execution stops rather than fabricating a placeholder table.","metadata":{}},{"id":"1bc484f1","cell_type":"code","source":"\"\"\"Construct, audit, deduplicate, and split the adult RSNA development cohort.\"\"\"\nfrom adult_cxr.data import (\n    assign_patient_splits,\n    audit_adult_rsna,\n    build_rsna_manifest,\n    deduplicate_manifest,\n    parse_dicom_age,\n)\n\nif parse_dicom_age(\"56\") != 56.0 or parse_dicom_age(\"216M\") != 18.0:\n    raise RuntimeError(\"The loaded Phase II source does not contain the frozen RSNA age parser\")\n\ndata_cfg = PROTOCOL.section(\"data\")\nmanifest_dir = OUTPUT / \"manifests\"\nmanifest_dir.mkdir(parents=True, exist_ok=True)\nrsna_all = build_rsna_manifest(data_cfg[\"rsna_root\"])\nrsna_adult, age_audit = audit_adult_rsna(\n    rsna_all, data_cfg[\"adult_age_minimum\"], data_cfg[\"plausible_age_maximum\"]\n)\nrsna_unique, duplicates = deduplicate_manifest(rsna_adult)\nrsna = assign_patient_splits(rsna_unique)\nrsna.to_csv(manifest_dir / \"rsna_adult_split.csv\", index=False)\nage_audit.to_csv(manifest_dir / \"rsna_age_audit.csv\", index=False)\nduplicates.to_csv(manifest_dir / \"rsna_duplicates_removed.csv\", index=False)\ndisplay(age_audit)\ndisplay(rsna.groupby([\"split\", \"projection\", \"target\"]).size().rename(\"images\").reset_index())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"bd86789f","cell_type":"markdown","source":"### Cohort composition and leakage verification\n\nThis cell visualizes adult-age exclusions and verifies the partition boundary at both patient and decoded-pixel\nlevels. The intended approximate proportions are 65% training, 15% validation, 10% calibration, and 10%\ndevelopment-test. Let $S(g)$ be the set of assigned partitions for group $g$; validity requires\n$|S(g)|=1$ for every patient and pixel hash. The displayed counts are descriptive and do not tune the model.\nBoth figures are based on the saved manifest, include their sample sizes, and are saved in PNG and SVG form.\nAny overlap assertion is fatal.","metadata":{}},{"id":"a875fb8d","cell_type":"code","source":"\"\"\"Verify leakage controls and illustrate age exclusions and partition composition.\"\"\"\nimport seaborn as sns\nfrom adult_cxr.data import assert_no_leakage\n\nassert_no_leakage(rsna)\nsns.set_theme(style=\"whitegrid\", palette=\"colorblind\")\nfig, axes = plt.subplots(1, 2, figsize=(12, 4))\nsns.barplot(data=age_audit, x=\"exclusion_reason\", y=\"images\", ax=axes[0], color=\"#4C78A8\")\naxes[0].tick_params(axis=\"x\", rotation=30)\naxes[0].set_title(f\"Figure 2a. RSNA age/view audit (n={len(rsna_all):,})\")\ncounts = rsna.groupby(\"split\").size().reindex([\"train\", \"validation\", \"calibration\", \"development_test\"])\naxes[1].bar(counts.index, counts.values, color=sns.color_palette(\"colorblind\", 4))\naxes[1].tick_params(axis=\"x\", rotation=25)\naxes[1].set_title(f\"Figure 2b. Patient-separated adult cohort (n={len(rsna):,})\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_02_cohort_audit.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_02_cohort_audit.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"5b722aa1","cell_type":"markdown","source":"## 4. Registered architectures and mathematical controls\n\nA0 isolates ordinary supervised learning. A1 adds RAD-DINO representation distillation while retaining average\npooling. A2 changes only the pooling operator to the residual Gaussian fuzzy mechanism. For spatial channel\nvalues $x_i$, $w_i \\propto \\exp[-(x_i-\\mu)^2/(2\\sigma^2)]$ and\n$z=\\bar{x}+\\tanh(\\alpha)(\\sum_i w_i x_i-\\bar{x})$. Initializing $\\alpha=0$ makes A2 exactly equivalent\nto average pooling before learning. DenseNet-121 seed 42 is the heavier reference. This cell displays the\ncomplete ten-fit registry and verifies the neutral-equivalence invariant numerically before training.","metadata":{}},{"id":"970b1df5","cell_type":"code","source":"\"\"\"Materialize the fixed model registry and test fuzzy-pooling neutrality.\"\"\"\nfrom adult_cxr.pooling import GlobalAveragePooling, ResidualGaussianFuzzyPooling\n\nregistry_rows = []\nfor arm in PROTOCOL.values[\"registry\"]:\n    for seed in arm[\"seeds\"]:\n        registry_rows.append({**{k: v for k, v in arm.items() if k != \"seeds\"}, \"seed\": seed})\nregistry = pd.DataFrame(registry_rows)\nassert len(registry) == 10 and set(registry[\"id\"]) == {\"A0\", \"A1\", \"A2\", \"REF\"}\nfeatures = torch.linspace(-2, 2, 2 * 8 * 7 * 7).reshape(2, 8, 7, 7)\naverage = GlobalAveragePooling()(features)\nfuzzy = ResidualGaussianFuzzyPooling(8)(features)\nneutral_error = float((average - fuzzy).abs().max())\nassert neutral_error <= 1e-7\ndisplay(registry)\nprint({\"neutral_pooling_max_error\": neutral_error, \"registered_fits\": len(registry)})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"e4a53ecd","cell_type":"markdown","source":"### Pooling and distillation mechanisms\n\nFigure 3 separates the two controlled architectural changes. Average pooling assigns equal spatial weight.\nResidual Gaussian fuzzy pooling estimates channel-specific membership weights and adds only a learned residual,\nwhich preserves a neutral average-pooling starting point. Distillation does not transfer the teacher's final\nclassification; it aligns the projected student representation with the frozen RAD-DINO representation while\nthe supervised RSNA outcome remains the classification target. The illustration is code-generated, contains no\npatient data, and has no role in candidate selection. Saved PNG and SVG files are reloaded before display.","metadata":{}},{"id":"31bb9888","cell_type":"code","source":"\"\"\"Illustrate the registered pooling and teacher-student mechanisms.\"\"\"\nfig, axes = plt.subplots(1, 2, figsize=(12, 4.2))\nspatial = np.array([[.2, .3, .1, .2], [.2, .9, .8, .1], [.1, .7, 1., .2], [.2, .1, .2, .1]])\naxes[0].imshow(spatial, cmap=\"viridis\", vmin=0, vmax=1)\nfor y in range(4):\n    for x in range(4):\n        axes[0].text(x, y, f\"{spatial[y,x]:.1f}\", ha=\"center\", va=\"center\", color=\"white\", fontsize=8)\naxes[0].set_title(\"Average: equal weights\\nFuzzy: Gaussian memberships + residual\")\naxes[0].set_xticks([])\naxes[0].set_yticks([])\naxes[1].axis(\"off\")\naxes[1].text(.08, .62, \"Adult RSNA\\nimage\", ha=\"center\", va=\"center\", bbox={\"boxstyle\":\"round\", \"fc\":\"#72B7B2\"})\naxes[1].annotate(\"\", (.34,.62), (.18,.62), arrowprops={\"arrowstyle\":\"->\"})\naxes[1].text(.45, .62, \"MobileViT-XS\\nstudent\", ha=\"center\", va=\"center\", bbox={\"boxstyle\":\"round\", \"fc\":\"#4C78A8\"})\naxes[1].text(.45, .18, \"Frozen\\nRAD-DINO\", ha=\"center\", va=\"center\", bbox={\"boxstyle\":\"round\", \"fc\":\"#F2CF5B\"})\naxes[1].annotate(\"cosine alignment\", (.48,.50), (.48,.30), arrowprops={\"arrowstyle\":\"<->\"}, ha=\"center\")\naxes[1].annotate(\"\", (.82,.62), (.58,.62), arrowprops={\"arrowstyle\":\"->\"})\naxes[1].text(.9, .62, \"Pneumonia\\nlogit\", ha=\"center\", va=\"center\", bbox={\"boxstyle\":\"round\", \"fc\":\"#54A24B\"})\naxes[1].set_title(\"Frozen teacher, supervised student\")\nfig.suptitle(\"Figure 3. Controlled Phase II architectural comparison\", x=.01, ha=\"left\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_03_architecture.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_03_architecture.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"9a78e85a","cell_type":"markdown","source":"## 5. RAD-DINO teacher cache\n\nFor distilled arms, the loss is\n$L=L_{BCE}+\\lambda[1-\\cos(Pz,t)]$, where $Pz$ is the student's projected pooled representation,\n$t$ is the frozen RAD-DINO CLS representation, and $\\lambda=0.25$. Teacher vectors are computed once for\nthe RSNA training partition, sorted by immutable image identifier, and stored independently of any student.\nThe cache reader verifies unique identifiers, finite two-dimensional vectors, and complete batch-order lookup.\nThis cell never uses CheXpert or VinBigData. A partial or mismatched cache stops training.","metadata":{}},{"id":"e5c7230a","cell_type":"code","source":"\"\"\"Create or verify the deterministic RSNA RAD-DINO teacher-feature cache.\"\"\"\nfrom adult_cxr.teacher import TeacherCache, build_rad_dino_cache, validate_teacher_cache\n\nteacher_path = OUTPUT / \"teacher\" / \"rad_dino_rsna_adult.npz\"\nteacher_manifest = rsna.loc[rsna[\"split\"].eq(\"train\")].reset_index(drop=True)\nteacher_model = PROTOCOL.section(\"models\")[\"teacher\"]\nif not validate_teacher_cache(teacher_path, teacher_manifest, teacher_model):\n    build_rad_dino_cache(\n        teacher_manifest, teacher_path, model_name=teacher_model,\n        batch_size=16, device=\"cuda\" if torch.cuda.is_available() else \"cpu\"\n    )\nteacher_cache = TeacherCache(teacher_path)\nmissing_teacher = sorted(set(teacher_manifest[\"image_id\"].astype(str)).difference(teacher_cache.index))\nif missing_teacher:\n    raise RuntimeError(f\"Teacher cache is incomplete: {len(missing_teacher)} missing images\")\nprint({\"cached_images\": len(teacher_cache.index), \"feature_dimension\": teacher_cache.features.shape[1],\n       \"sha256\": sha256_file(teacher_path)})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"0a49baaf","cell_type":"markdown","source":"## 6. Confirmatory training registry\n\nAll registered fits use the same adult RSNA partitions, conservative preprocessing, four-group AP/PA × outcome\nsampling, AdamW optimizer, cosine schedule, and eight-epoch horizon. The sampler derives a new reproducible draw\nfrom the arm's seed and epoch index, avoiding repetition of one with-replacement sample across all epochs. Each\nfull epoch is retained in the history and printed with its loss, validation NLL, and elapsed time.\nValidation negative log-likelihood, $NLL=-n^{-1}\\sum[y\\log p+(1-y)\\log(1-p)]$, identifies the stored checkpoint;\nthere is no early stopping. This cell is resumable only at the level of an already complete fit registry and\nvalid checkpoint. It does not catch training exceptions or manufacture metrics. A failed fit remains failed and\nmust be reported before the registry continues.","metadata":{}},{"id":"2162cf8a","cell_type":"code","source":"\"\"\"Train every prespecified arm and seed using one transparent registry loop.\"\"\"\nfrom torch.utils.data import DataLoader\nfrom adult_cxr.data import CXRDataset\nfrom adult_cxr.models import create_model\nfrom adult_cxr.reproducibility import seed_everything\nfrom adult_cxr.sampling import FourGroupBatchSampler\nfrom adult_cxr.training import FitSpec, completed_fit_matches, fit_model\nfrom adult_cxr.transforms import build_transform\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprotocol_cfg = PROTOCOL.section(\"protocol\")\ntraining_cfg = PROTOCOL.section(\"training\")\nmodel_cfg = PROTOCOL.section(\"models\")\nsize = protocol_cfg[\"image_size\"]\nworkers = protocol_cfg[\"workers\"]\nloader_settings = {\n    \"num_workers\": workers,\n    \"pin_memory\": device.type == \"cuda\",\n    \"persistent_workers\": workers > 0,\n}\nsingle_pass_loader_settings = {\n    \"num_workers\": workers,\n    \"pin_memory\": device.type == \"cuda\",\n}\ntrain_frame = rsna.loc[rsna[\"split\"].eq(\"train\")].reset_index(drop=True)\nvalidation_frame = rsna.loc[rsna[\"split\"].eq(\"validation\")].reset_index(drop=True)\nvalidation_loader = DataLoader(\n    CXRDataset(validation_frame, build_transform(size, False)),\n    batch_size=64,\n    **loader_settings,\n)\nfor row in registry.to_dict(\"records\"):\n    spec = FitSpec(\n        arm=row[\"id\"],\n        backbone=row[\"backbone\"],\n        pooling=row[\"pooling\"],\n        distillation=row[\"distillation\"],\n        seed=row[\"seed\"],\n        epochs=protocol_cfg[\"epochs\"],\n        learning_rate=training_cfg[\"learning_rate\"],\n        weight_decay=training_cfg[\"weight_decay\"],\n        distillation_weight=model_cfg[\"distillation_weight\"],\n    )\n    if completed_fit_matches(OUTPUT / \"fits\", spec):\n        print(f\"Skipping verified fit {row['id']} seed={row['seed']}\", flush=True)\n        continue\n    train_data = CXRDataset(train_frame, build_transform(size, True))\n    batches = FourGroupBatchSampler(train_frame, protocol_cfg[\"batch_size\"], row[\"seed\"])\n    train_loader = DataLoader(train_data, batch_sampler=batches, **loader_settings)\n    seed_everything(row[\"seed\"])\n    model = create_model(row[\"backbone\"], row[\"pooling\"], pretrained=True)\n    print(\n        f\"Starting {row['id']} seed={row['seed']}: \"\n        f\"{len(train_loader)} training batches per epoch on {device}\",\n        flush=True,\n    )\n    fit_model(model, spec, train_loader, validation_loader, OUTPUT / \"fits\", device,\n              teacher_cache if row[\"distillation\"] else None)\nprint(\"All ten registered fits have complete checkpoint and registry artifacts.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"8c2a1600","cell_type":"markdown","source":"## 7. Internal prediction and mandatory embedding evidence\n\nEach selected checkpoint is reloaded into a fresh model object. Predictions and embeddings are generated for\nvalidation, calibration, and development-test partitions, then joined to projection using the immutable image\nidentifier. Every arm-seed-partition unit is atomically committed with checkpoint, manifest, configuration,\nand source hashes. A rerun validates and reuses complete units while rejecting stale or partial files. The\nvalidation and development-test embeddings are mandatory for the acquisition-view probe:\nvalidation embeddings train the linear classifier and development-test embeddings evaluate it. This prevents\ntraining and testing the probe on the same observations. Every expected image must have exactly one prediction;\nmissing, duplicated, or non-finite evidence stops execution. These internal results precede any CheXpert access.","metadata":{}},{"id":"ae4ec356","cell_type":"code","source":"\"\"\"Generate or resume complete internal predictions and embeddings.\"\"\"\n\nprediction_dir = OUTPUT / \"predictions\" / \"rsna\"\nembedding_dir = OUTPUT / \"embeddings\"\nsubprocess.run([\n    sys.executable, str(PROJECT / \"scripts\" / \"generate_internal_evidence.py\"),\n    \"--config\", str(CONFIG_PATH), \"--artifacts\", str(OUTPUT),\n], cwd=PROJECT, check=True)\nrequired_internal = [\n    prediction_dir / f\"{row['id']}_seed{row['seed']}_{partition}.csv\"\n    for row in registry.to_dict(\"records\")\n    for partition in (\"validation\", \"calibration\", \"development_test\")\n]\nmissing_internal = [str(path) for path in required_internal if not path.is_file()]\nif missing_internal:\n    raise RuntimeError(\"Internal evidence remains incomplete:\\n\" + \"\\n\".join(missing_internal))\nprint(f\"Verified {len(required_internal)} internal prediction units and their embedding archives.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"976d75f1","cell_type":"markdown","source":"### Deterministic representative radiographs\n\nFigure 4 displays the first lexicographically sorted development-test image from each AP/PA × outcome group.\nSelection is deterministic and independent of model scores, preventing visually favorable examples from being\nchosen after prediction. Images are deidentified public RSNA radiographs, normalized with the same percentile\nrule used by the model, and shown only to clarify acquisition and label strata. The target identifies an RSNA\nopacity annotation rather than a universally adjudicated clinical pneumonia diagnosis. An empty group is a\nprotocol failure because all four groups are required for macro-view evaluation.","metadata":{}},{"id":"bd42e158","cell_type":"code","source":"\"\"\"Create a deterministic AP/PA by outcome RSNA montage.\"\"\"\nfrom adult_cxr.data import decode_image\n\ntest_frame = rsna.loc[rsna[\"split\"].eq(\"development_test\")].sort_values(\"image_id\")\nexamples = []\nfor projection in (\"AP\", \"PA\"):\n    for target in (0, 1):\n        group = test_frame.loc[test_frame[\"projection\"].eq(projection) & test_frame[\"target\"].eq(target)]\n        if group.empty:\n            raise RuntimeError(f\"Empty montage group: {projection}, target={target}\")\n        examples.append(group.iloc[0])\nfig, axes = plt.subplots(2, 2, figsize=(7, 7))\nfor axis, row in zip(axes.ravel(), examples, strict=True):\n    axis.imshow(decode_image(row[\"path\"]), cmap=\"gray\", vmin=0, vmax=1)\n    axis.set_title(f\"{row['projection']} — target {int(row['target'])}\")\n    axis.axis(\"off\")\nfig.suptitle(f\"Figure 4. Deterministic RSNA development-test montage (n={len(test_frame):,})\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_04_rsna_montage.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_04_rsna_montage.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"1fae6791","cell_type":"markdown","source":"## 8. Shortcut falsification and localization diagnostics\n\nA projection probe measures how readily AP versus PA acquisition can be decoded from frozen representations;\nAUROC 0.5 corresponds to chance and higher values indicate retained view information, not necessarily causal\nshortcut use. Border masking measures prediction sensitivity to peripheral text and acquisition artifacts.\nFor positive RSNA cases, pointing-game accuracy and thresholded IoU compare activation maxima with annotated\nopacity boxes. These are diagnostic localization measures rather than validated explanations. This cell must\nproduce one probe record for every A1 and A2 seed. Missing embedding files are errors, never conditional skips.","metadata":{}},{"id":"214e6ab5","cell_type":"code","source":"\"\"\"Run and verify every held-out shortcut and localization diagnostic.\"\"\"\nimport subprocess\n\ndiagnostic_command = [\n    sys.executable, str(PROJECT / \"scripts\" / \"run_internal_diagnostics.py\"),\n    \"--config\", str(CONFIG_PATH), \"--artifacts\", str(OUTPUT),\n    \"--localization-cases\", \"150\",\n]\nsubprocess.run(diagnostic_command, cwd=PROJECT, check=True)\nprobe_results = pd.read_csv(OUTPUT / \"diagnostics_projection_probe.csv\")\nmasking_results = pd.read_csv(OUTPUT / \"diagnostics_border_masking.csv\")\nlocalization_results = pd.read_csv(OUTPUT / \"diagnostics_localization.csv\")\nif len(probe_results) != 6 or len(masking_results) != 6 or len(localization_results) != 900:\n    raise RuntimeError(\"Mandatory shortcut or localization evidence is incomplete\")\ndisplay(probe_results)\ndisplay(masking_results)\ndisplay(localization_results.groupby(\"arm\")[[\"pointing_game\", \"threshold_iou\"]].agg([\"mean\", \"std\", \"count\"]))\nfig, axes = plt.subplots(1, 3, figsize=(13, 3.8))\nsns.pointplot(data=probe_results, x=\"arm\", y=\"probe_auroc\", errorbar=\"sd\", ax=axes[0], color=\"#4C78A8\")\naxes[0].axhline(.5, color=\"black\", ls=\"--\")\naxes[0].set_title(\"Projection-probe AUROC\")\nsns.pointplot(data=masking_results, x=\"arm\", y=\"mean_absolute_probability_change\", errorbar=\"sd\",\n              ax=axes[1], color=\"#E45756\")\naxes[1].set_title(\"Border-mask sensitivity\")\nsns.pointplot(data=localization_results, x=\"arm\", y=\"pointing_game\", errorbar=(\"ci\",95),\n              ax=axes[2], color=\"#54A24B\")\naxes[2].set_title(\"RSNA pointing-game accuracy\")\nfig.suptitle(\"Figure 5. Mandatory shortcut and localization diagnostics\", x=.01, ha=\"left\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_05_diagnostics.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_05_diagnostics.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"d6d48a53","cell_type":"markdown","source":"## 9. Calibration and empirical tolerance\n\nTemperature scaling fits $T>0$ on RSNA calibration logits by minimizing\n$NLL(\\sigma(z/T),y)$; no base-model parameter changes. Calibration is fitted separately for each registered\nseed because each checkpoint has a distinct logit scale. The empirical non-inferiority tolerance is\n$\\delta=\\max[\\delta_{I},1.96\\sqrt{2}\\hat\\sigma_{II}]$, where the first term is extracted from the immutable\nPhase I registry and the second is the pooled within-configuration seed variability. The calculation is a\nbest-effort empirical anchor and its small configuration count is reported. Missing replicated seed metrics\nprevent qualification.","metadata":{}},{"id":"8e481bc7","cell_type":"code","source":"\"\"\"Fit calibration parameters and derive the frozen seed-variability tolerance.\"\"\"\nfrom adult_cxr.calibration import TemperatureScaler\nfrom adult_cxr.metrics import metrics_by_view\nfrom adult_cxr.statistics import empirical_seed_tolerance\nfrom adult_cxr.reproducibility import write_json_atomic\n\ncalibration_dir = OUTPUT / \"calibration\"\ncalibration_dir.mkdir(parents=True, exist_ok=True)\nseed_rows = []\nfor row in registry.loc[registry[\"id\"].isin([\"A1\", \"A2\"])].to_dict(\"records\"):\n    calibration_predictions = pd.read_csv(prediction_dir / f\"{row['id']}_seed{row['seed']}_calibration.csv\")\n    scaler = TemperatureScaler().fit(calibration_predictions[\"logit\"].to_numpy(),\n                                     calibration_predictions[\"target\"].to_numpy())\n    write_json_atomic(scaler.to_dict(), calibration_dir / f\"{row['id']}_seed{row['seed']}.json\")\n    test_predictions = pd.read_csv(prediction_dir / f\"{row['id']}_seed{row['seed']}_development_test.csv\")\n    test_predictions[\"probability\"] = scaler.predict_proba(test_predictions[\"logit\"].to_numpy())\n    overall = metrics_by_view(test_predictions).query(\"projection == 'OVERALL'\").iloc[0]\n    seed_rows.append({\"configuration\": row[\"id\"], \"seed\": row[\"seed\"], \"auroc\": overall[\"auroc\"]})\nphase1_evidence = json.loads((PROJECT / \"configs\" / \"phase1_evidence.json\").read_text(encoding=\"utf-8\"))\nphase1_delta = float(phase1_evidence[\"prespecified_auroc_relevance_margin\"])\ntolerance = empirical_seed_tolerance(phase1_delta, pd.DataFrame(seed_rows))\nwrite_json_atomic(tolerance, OUTPUT / \"seed_tolerance.json\")\ndisplay(pd.DataFrame(seed_rows))\nprint(tolerance)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4ebb87e1","cell_type":"markdown","source":"### Internal discrimination and calibration curves\n\nFigure 6 summarizes all three calibrated seeds as an equal-weight probability ensemble on the untouched RSNA\ndevelopment-test partition. ROC curves show ranking across thresholds, precision–recall curves expose behavior\nunder the observed prevalence, and reliability curves compare predicted and observed event frequencies. These\nplots diagnose A1 and A2 before CheXpert access. They do not select an operating threshold and do not imply a\nclinical decision boundary. Curves require complete one-to-one seed alignment; any missing identifier stops the\ncell. Calibration has been fitted only on the separate RSNA calibration partition.","metadata":{}},{"id":"3319977a","cell_type":"code","source":"\"\"\"Plot calibrated internal ROC, precision-recall, and reliability curves.\"\"\"\nfrom sklearn.calibration import calibration_curve\nfrom sklearn.metrics import precision_recall_curve, roc_curve\n\nfig, axes = plt.subplots(1, 3, figsize=(13, 4))\nfor arm, color in ((\"A1\", \"#4C78A8\"), (\"A2\", \"#E45756\")):\n    tables = []\n    for seed in PROTOCOL.seeds:\n        table = pd.read_csv(prediction_dir / f\"{arm}_seed{seed}_development_test.csv\")\n        parameter = json.loads((calibration_dir / f\"{arm}_seed{seed}.json\").read_text())\n        scaler = TemperatureScaler()\n        scaler.temperature = parameter[\"temperature\"]\n        table[f\"p_{seed}\"] = scaler.predict_proba(table[\"logit\"].to_numpy())\n        tables.append(table.set_index(\"image_id\")[[\"target\", f\"p_{seed}\"]])\n    probability_columns = [f\"p_{seed}\" for seed in PROTOCOL.seeds]\n    aligned = pd.concat([table.drop(columns=\"target\") for table in tables], axis=1, join=\"inner\")\n    target = tables[0].loc[aligned.index, \"target\"].to_numpy()\n    if len(aligned) != len(test_frame) or aligned.isna().any().any():\n        raise RuntimeError(\"Seed predictions are not one-to-one aligned\")\n    probability = aligned[probability_columns].mean(axis=1).to_numpy()\n    fpr, tpr, _ = roc_curve(target, probability)\n    precision, recall, _ = precision_recall_curve(target, probability)\n    observed, predicted = calibration_curve(target, probability, n_bins=10, strategy=\"quantile\")\n    axes[0].plot(fpr, tpr, label=arm, color=color)\n    axes[1].plot(recall, precision, label=arm, color=color)\n    axes[2].plot(predicted, observed, marker=\"o\", label=arm, color=color)\nfor axis, title in zip(axes, (\"ROC\", \"Precision–recall\", \"Reliability\"), strict=True):\n    axis.set_title(title)\n    axis.legend()\naxes[0].plot([0,1],[0,1],\"k--\",alpha=.5)\naxes[2].plot([0,1],[0,1],\"k--\",alpha=.5)\nfig.suptitle(f\"Figure 6. RSNA development-test curves (n={len(test_frame):,})\", x=.01, ha=\"left\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_06_internal_curves.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_06_internal_curves.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"31009ab9","cell_type":"markdown","source":"### Recoverable checkpoint of all expensive internal work\n\nAt this boundary, training, teacher features, internal predictions, embeddings, diagnostics, calibration, and\ntheir integrity registries are complete. The following administrative cell writes a stored (non-recompressed)\nZIP with a SHA-256 manifest. Download this file or preserve it as a Kaggle output before opening the qualification\nboundary. In a later session, attach the archive and set `PHASE2_RESUME_ARCHIVE` to its path before running the\nprotocol cell. Restoration rejects a different source version and never overwrites conflicting local evidence.\nThis continuity archive does not alter a model or statistic.","metadata":{}},{"id":"efa7fba9","cell_type":"code","source":"\"\"\"Create a portable integrity-checked resume point for all completed internal work.\"\"\"\nresume_destination = (\n    Path(\"/kaggle/working/phase2_resume_v041.zip\")\n    if Path(\"/kaggle/working\").exists()\n    else PROJECT / \"phase2_resume_v041.zip\"\n)\nsubprocess.run([\n    sys.executable, str(PROJECT / \"scripts\" / \"snapshot_artifacts.py\"), \"--create\",\n    \"--artifacts\", str(OUTPUT), \"--archive\", str(resume_destination),\n], cwd=PROJECT, check=True)\nprint(f\"Resume archive ready: {resume_destination}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"83e36e11","cell_type":"markdown","source":"## 10. CheXpert qualification boundary\n\nCheXpert is used only after internal artifacts are complete. Blank report labels are mapped to negative unless\neither pneumonia or lung opacity is positive; uncertain labels are excluded. The expert validation set provides\na separate label-quality fallback. A2 is eligible only if the lower paired confidence bounds for macro-view and\nworst-view AUROC differences versus A1 are both at least $-\\delta$. Eligible configurations are ranked\nlexicographically by macro NLL, macro Brier score, then CPU latency. A2 falls back to A1 if the expert overall\nAUROC lower bound is below $-\\delta$. This cell constructs manifests; it does not train or calibrate a model.","metadata":{}},{"id":"54509072","cell_type":"code","source":"\"\"\"Construct the two CheXpert qualification cohorts with explicit label policy.\"\"\"\nfrom adult_cxr.data import build_chexpert_manifest\n\nchexpert_report = build_chexpert_manifest(data_cfg[\"chexpert_root\"], \"report\")\nchexpert_expert = build_chexpert_manifest(data_cfg[\"chexpert_root\"], \"expert\")\nchexpert_report.to_csv(manifest_dir / \"chexpert_report.csv\", index=False)\nchexpert_expert.to_csv(manifest_dir / \"chexpert_expert.csv\", index=False)\nqualification_counts = pd.concat([\n    chexpert_report.assign(tier=\"report\"), chexpert_expert.assign(tier=\"expert\")\n]).groupby([\"tier\", \"projection\", \"target\"]).size().rename(\"images\").reset_index()\ndisplay(qualification_counts)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"c85b627b","cell_type":"markdown","source":"### CheXpert prediction and frozen configuration decision\n\nAll A1 and A2 seeds are evaluated on both qualification cohorts. Seed probabilities are averaged by image,\nmaintaining paired identity alignment. Paired stratified bootstrap intervals preserve both classes while\nresampling the same observations for A1 and A2. The resulting decision function accepts only the recorded\nconfidence bounds, calibration metrics, and latency measures; it cannot inspect VinBigData. Every manifest image\nmust receive one prediction from every seed. This stage writes a human-readable decision record before any\nexport or locked external access.","metadata":{}},{"id":"7497f444","cell_type":"code","source":"\"\"\"Run complete CheXpert qualification inference and freeze the A1/A2 decision.\"\"\"\nfrom adult_cxr.selection import CandidateEvidence, select_configuration\n# The evaluator validates its source/configuration/checkpoint/manifests signature before reusing any output.\nrequired_files = [OUTPUT / \"predictions\" / \"chexpert\" / f\"{arm}_seed{seed}_{tier}.csv\"\n                  for arm in (\"A1\", \"A2\") for seed in PROTOCOL.seeds for tier in (\"report\", \"expert\")]\nsubprocess.run([\n    sys.executable, str(PROJECT / \"scripts\" / \"evaluate_chexpert.py\"),\n    \"--config\", str(CONFIG_PATH), \"--artifacts\", str(OUTPUT),\n], cwd=PROJECT, check=True)\nmissing = [str(path) for path in required_files if not path.exists()]\nif missing:\n    raise RuntimeError(\"CheXpert inference finished without mandatory files:\\n\" + \"\\n\".join(missing))\ndecision_inputs = json.loads((OUTPUT / \"chexpert_decision_inputs.json\").read_text(encoding=\"utf-8\"))\nevidence = CandidateEvidence(**decision_inputs)\nselected_configuration = select_configuration(evidence, tolerance[\"delta\"])\nwrite_json_atomic({\"selected_configuration\": selected_configuration, \"evidence\": decision_inputs,\n                   \"delta\": tolerance[\"delta\"]}, OUTPUT / \"configuration_decision.json\")\nprint({\"selected_configuration\": selected_configuration, \"delta\": tolerance[\"delta\"]})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"0cceef86","cell_type":"markdown","source":"## 11. ONNX export, execution-environment CPU benchmark, and artifact lock\n\nThe selected configuration is exported for seeds 17, 42, and 73. PyTorch and ONNX probabilities must differ by\nat most (10^{-4}). The ensemble is selected only if sequential CPU inference satisfies median ≤500 ms,\np95 ≤1,000 ms, loaded-session peak memory ≤2 GB, and total size ≤150 MB; otherwise seed 42 is selected. This is an\nengineering choice and cannot use external accuracy. Configuration, selected ONNX file(s), calibration,\npreprocessing metadata, protocol, and evaluation code are SHA-256 locked. The exact locked model is subsequently\nused by both VinBigData evaluation and the web service.","metadata":{}},{"id":"07dd070c","cell_type":"code","source":"\"\"\"Verify the completed export evidence and create the external-access lock.\"\"\"\nimport shutil\n\nfrom adult_cxr.artifacts import create_artifact_lock, verify_artifact_lock\nfrom adult_cxr.selection import select_artifact\n\nexport_dir = OUTPUT / \"export\" / selected_configuration\nlock_path = OUTPUT / \"artifact_lock.json\"\nif lock_path.is_file():\n    existing_lock = verify_artifact_lock(lock_path)\n    if existing_lock[\"configuration\"] != selected_configuration:\n        raise RuntimeError(\"Existing lock names a different selected configuration\")\n    artifact_policy = existing_lock[\"artifact_policy\"]\n    print(\"Reused the existing byte-identical external artifact lock.\", flush=True)\nelse:\n    subprocess.run([\n        sys.executable, str(PROJECT / \"scripts\" / \"export_candidate.py\"),\n        \"--config\", str(CONFIG_PATH), \"--artifacts\", str(OUTPUT),\n    ], cwd=PROJECT, check=True)\n    ensemble_path = export_dir / \"ensemble_benchmark.json\"\n    single_path = export_dir / \"seed_42_benchmark.json\"\n    if not ensemble_path.exists() or not single_path.exists():\n        raise RuntimeError(\"Candidate export did not produce both mandatory benchmarks\")\n    ensemble_benchmark = json.loads(ensemble_path.read_text(encoding=\"utf-8\"))\n    single_benchmark = json.loads(single_path.read_text(encoding=\"utf-8\"))\n    artifact_policy = select_artifact(\n        ensemble_benchmark, single_benchmark, PROTOCOL.section(\"deployment\")\n    )\n    model_files = [export_dir / f\"seed_{seed}.onnx\" for seed in PROTOCOL.seeds]\n    selected_models = model_files if artifact_policy == \"ensemble\" else [export_dir / \"seed_42.onnx\"]\n    calibration_path = export_dir / f\"calibration_{artifact_policy}.json\"\n    metadata_path = export_dir / \"metadata.json\"\n    locked_inputs = OUTPUT / \"locked_inputs\"\n    locked_inputs.mkdir(parents=True, exist_ok=True)\n    provider_command = [\n        sys.executable, str(PROJECT / \"scripts\" / \"qualify_onnx_provider.py\"),\n        \"--preferred\", external_onnx_provider, \"--image-size\", str(size),\n    ]\n    for provider_model in [*selected_models, export_dir / \"densenet121_seed_42.onnx\"]:\n        provider_command.extend([\"--model\", str(provider_model)])\n    provider_result = subprocess.run(\n        provider_command, cwd=PROJECT, check=True, capture_output=True, text=True,\n    )\n    provider_evidence = json.loads(provider_result.stdout.strip().splitlines()[-1])\n    external_runtime = {\n        **provider_evidence,\n        \"decode_workers\": max(1, min(int(PROTOCOL.section(\"protocol\")[\"workers\"]), 4)),\n        \"batch_size\": 64,\n    }\n    write_json_atomic(external_runtime, locked_inputs / \"external_runtime.json\")\n    shutil.copy2(CONFIG_PATH, locked_inputs / \"protocol.yaml\")\n    shutil.copy2(PROJECT / \"configs\" / \"phase1_evidence.json\", locked_inputs / \"phase1_evidence.json\")\n    shutil.copy2(PROJECT / \"scripts\" / \"evaluate_vinbigdata_locked.py\",\n                 locked_inputs / \"evaluate_vinbigdata_locked.py\")\n    shutil.copy2(PROJECT / \"adult_cxr\" / \"data.py\", locked_inputs / \"data.py\")\n    shutil.copy2(PROJECT / \"adult_cxr\" / \"inference.py\", locked_inputs / \"inference.py\")\n    shutil.copy2(OUTPUT / \"seed_tolerance.json\", locked_inputs / \"seed_tolerance.json\")\n    lock_files = [*selected_models, export_dir / \"densenet121_seed_42.onnx\",\n                  export_dir / \"calibration_dense_seed_42.json\", calibration_path, metadata_path,\n                  ensemble_path, single_path, export_dir / \"onnx_parity.json\",\n                  locked_inputs / \"protocol.yaml\", locked_inputs / \"phase1_evidence.json\",\n                  locked_inputs / \"seed_tolerance.json\", locked_inputs / \"external_runtime.json\",\n                  locked_inputs / \"evaluate_vinbigdata_locked.py\", locked_inputs / \"data.py\",\n                  locked_inputs / \"inference.py\"]\n    lock_path = create_artifact_lock(lock_files, lock_path,\n                                     selected_configuration, artifact_policy)\nensemble_path = export_dir / \"ensemble_benchmark.json\"\nsingle_path = export_dir / \"seed_42_benchmark.json\"\nensemble_benchmark = json.loads(ensemble_path.read_text(encoding=\"utf-8\"))\nsingle_benchmark = json.loads(single_path.read_text(encoding=\"utf-8\"))\nbenchmark = ensemble_benchmark if artifact_policy == \"ensemble\" else single_benchmark\nprint({\"artifact_policy\": artifact_policy, \"artifact_lock\": str(lock_path), **benchmark})\nratio_names = [\"median_latency_ms\", \"p95_latency_ms\", \"peak_memory_mb\", \"artifact_size_mb\",\n               \"onnx_max_probability_error\"]\nratios = [benchmark[name] / PROTOCOL.section(\"deployment\")[name] for name in ratio_names]\nfig, ax = plt.subplots(figsize=(9, 3.8))\nax.barh(ratio_names, ratios, color=[\"#54A24B\" if value <= 1 else \"#E45756\" for value in ratios])\nax.axvline(1, color=\"black\", ls=\"--\", label=\"prespecified limit\")\nax.set_xlabel(\"Observed value / engineering limit\")\nax.set_title(f\"Figure 7. {artifact_policy} engineering constraints\")\nax.legend()\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_07_engineering_constraints.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_07_engineering_constraints.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"01bd1462","cell_type":"markdown","source":"## 12. Locked VinBigData external proxy evaluation\n\nThe official VinBigData competition training cohort contains 15,000 adult PA radiographs annotated independently\nby three radiologists. It is sourced from the VinDr-CXR collection but is not the unavailable five-reader consensus\ntest set. Before any outcomes are inspected, the endpoint is frozen as consolidation, infiltration, or lung opacity.\nFor image $i$, let $r_i$ be the number of readers who marked at least one endpoint class. The binary label is\n$y_i=1$ when $r_i\\geq2$ and $y_i=0$ when $r_i=0$; cases with $r_i=1$ are excluded as label-ambiguous. Consequently,\nthis stage evaluates radiographic air-space-opacity discrimination, not confirmed pneumonia diagnosis. Access is\npermitted only after every artifact-lock hash verifies.\n`PERFORMANCE_QUALIFIED` requires an AUROC lower bound above 0.50, a Brier-skill lower bound above zero, and a\npaired AUROC lower-bound difference versus DenseNet-121 of at least $-\\delta$, with complete predictions.\nNo failure may trigger retraining, recalibration, endpoint changes, or reselection.","metadata":{}},{"id":"d95e73a0","cell_type":"code","source":"\"\"\"Verify and display the one-time locked VinBigData proxy evaluation.\"\"\"\nfrom adult_cxr.artifacts import verify_artifact_lock\n\nverify_artifact_lock(OUTPUT / \"artifact_lock.json\")\nexternal_evidence_path = OUTPUT / \"vinbigdata_locked_evidence.json\"\nif not external_evidence_path.exists():\n    external_command = [\n        sys.executable, str(PROJECT / \"scripts\" / \"evaluate_vinbigdata_locked.py\"),\n        \"--config\", str(CONFIG_PATH), \"--artifacts\", str(OUTPUT),\n    ]\n    if os.environ.get(\"VINBIGDATA_ROOT\"):\n        external_command.extend([\"--vinbigdata-root\", os.environ[\"VINBIGDATA_ROOT\"]])\n    subprocess.run(external_command, cwd=PROJECT, check=True)\nif not external_evidence_path.exists():\n    raise RuntimeError(\"Locked VinBigData evaluation did not produce its evidence record\")\nexternal_evidence = json.loads(external_evidence_path.read_text(encoding=\"utf-8\"))\nrequired = {\"qualification_status\", \"auroc\", \"auroc_ci_lower\", \"auroc_ci_upper\",\n            \"brier_skill\", \"brier_skill_ci_lower\", \"dense_difference_ci_lower\", \"n\",\n            \"ambiguous_images_excluded\", \"endpoint_scope\"}\nif required.difference(external_evidence):\n    raise RuntimeError(\"Locked VinBigData evidence is incomplete\")\nif external_evidence[\"n\"] + external_evidence[\"ambiguous_images_excluded\"] != 15000:\n    raise RuntimeError(\"Locked VinBigData evidence does not account for all 15,000 source images\")\ndisplay(pd.DataFrame([external_evidence]))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"09ed58a6","cell_type":"markdown","source":"## 13. Engineering proof of concept\n\nThe FastAPI service exposes health, version, evidence, and DICOM inference endpoints. It accepts only single-frame\nAP or PA DICOM images with a usable intensity range. Corrupt, blank, lateral, multiframe, or unknown-view inputs\nare rejected or receive an explicit abstention; unsupported inputs never receive reassuring negative language.\nThe returned scalar is labeled a research score and is accompanied by the model version, protocol version,\nqualification status, warnings, and a non-clinical disclaimer. Scientific performance status and engineering\nreadiness are assessed separately so a functioning interface cannot conceal a failed external gate.","metadata":{}},{"id":"eb42077a","cell_type":"code","source":"\"\"\"Display the independently generated engineering-gate evidence.\"\"\"\nengineering_path = OUTPUT / \"engineering_gate.json\"\nsubprocess.run([\n    sys.executable, str(PROJECT / \"scripts\" / \"run_engineering_gate.py\"),\n    \"--artifacts\", str(OUTPUT),\n], cwd=PROJECT, check=True)\nif not engineering_path.exists():\n    raise RuntimeError(\"Engineering gate did not produce its evidence record\")\nengineering = json.loads(engineering_path.read_text(encoding=\"utf-8\"))\nrequired_engineering = {\"status\", \"tests_passed\", \"onnx_parity_passed\", \"dicom_gate_passed\"}\nif required_engineering.difference(engineering):\n    raise RuntimeError(\"Engineering-gate evidence is incomplete\")\ndisplay(pd.DataFrame([engineering]))\nfig, ax = plt.subplots(figsize=(10, 2.8))\nax.axis(\"off\")\nnodes = [(\"Adult AP/PA\\nDICOM\", .08, \"#72B7B2\"), (\"Quality gate\", .30, \"#F2CF5B\"),\n         (\"Locked ONNX\", .52, \"#4C78A8\"), (\"FastAPI\", .72, \"#B279A2\"),\n         (\"Educational UI\", .91, \"#54A24B\")]\nfor index, (label, x, color) in enumerate(nodes):\n    ax.text(x, .5, label, ha=\"center\", va=\"center\", bbox={\"boxstyle\":\"round,pad=.5\", \"fc\":color})\n    if index:\n        ax.annotate(\"\", (x-.09,.5), (nodes[index-1][1]+.09,.5), arrowprops={\"arrowstyle\":\"->\"})\nax.set_title(\"Figure 8. Offline inference and interface boundary\", loc=\"left\")\nfor suffix in (\"png\", \"svg\"):\n    fig.savefig(FIGURES / f\"figure_08_deployment_flow.{suffix}\", dpi=180, bbox_inches=\"tight\")\nplt.close(fig)\ndisplay(Image.open(FIGURES / \"figure_08_deployment_flow.png\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"b10dcd47","cell_type":"markdown","source":"## 14. Limitations and conclusion\n\nThe work is restricted to adult frontal radiographs and an educational research setting. RSNA acquisition\ngroups may remain imbalanced after patient separation. CheXpert labels are partly report-derived and CheXpert\nwas inspected previously, so it supports qualification rather than untouched confirmation. VinBigData contains\nthree-reader annotations rather than the unavailable five-reader consensus labels. Its frozen air-space-opacity\nproxy is not a clinical pneumonia diagnosis, and its PA-only composition cannot establish external AP performance.\nThe empirical tolerance is based on a small number of replicated configurations. Projection probes identify\nencoded acquisition information but do not prove causal shortcut use, while activation overlap is not a\nclinically validated explanation. Passing the gates supports only the narrow performance-qualification and\nengineering-reproducibility claim stated in the abstract.\n\n## References\n\n1. Zech JR et al. *PLOS Medicine* 2018;15:e1002683. https://doi.org/10.1371/journal.pmed.1002683\n2. DeGrave AJ et al. *Nature Machine Intelligence* 2021;3:610–619. https://doi.org/10.1038/s42256-021-00338-7\n3. Guo C et al. *Proceedings of ICML* 2017:1321–1330. https://proceedings.mlr.press/v70/guo17a.html\n4. Pérez-García F et al. RAD-DINO. arXiv:2401.10815. https://doi.org/10.48550/arXiv.2401.10815\n5. Nguyen HQ et al. *Scientific Data* 2022;9:429. https://doi.org/10.1038/s41597-022-01498-w\n6. Rajpurkar P et al. CheXpert. arXiv:1901.07031. https://doi.org/10.48550/arXiv.1901.07031","metadata":{}}]}