{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# VinDr-CXR Nodule/Mass 分辨率对照实验 E1\n\n这是项目第一组正式对照实验中的 **1280 分辨率实验组**。\n\n- 基线 E0：YOLO11n，640，完整胸片，当前保守增强。\n- 实验 E1：YOLO11n，1280，完整胸片，当前保守增强。\n- 唯一主要变化：端到端图像分辨率由 640 提高到 1280。\n- 任务仍为单类目标检测：`Nodule/Mass`。\n- 数据划分必须与 E0 完全一致。\n- 1280 图像必须直接从原始 DICOM 重新生成，禁止由 640 PNG 放大。\n\n本 Notebook 默认 `RUN_MODE=\"train\"`，只负责：\n\n1. 定位或生成 1280 预处理数据；\n2. 执行 V6 全量清洗；\n3. 与 E0 的 640 数据划分逐图核对；\n4. 训练 YOLO11n 1280；\n5. 保存 `best.pt`、`last.pt`、`results.csv` 和训练合同。\n\n逐图 TP/FP/FN、FROC、阈值扫描和测试集评价应放在独立评估 Notebook 中，避免训练成功后因评估显存溢出导致整个版本显示失败。\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Kaggle 运行前检查\n\n正式运行前添加以下 Input：\n\n1. **VinBigData 原始 DICOM 与 `train.csv`**  \n   只有尚未生成 1280 清洗数据时才会读取。\n2. **E0 使用的 640 V6 清洗 Output**  \n   程序只读取其中的 `manifest.csv` 来锁定相同的 train/val/test 划分，不会用 640 图片训练 E1。\n3. 已保存的 **1280 V6 清洗 Output**（若之前已经完成）  \n   程序会验证数据签名后自动跳过重复预处理与清洗。\n4. **Settings → Accelerator → GPU**。\n\n建议先在交互会话中把 `SMOKE_TEST=True` 跑 1 epoch；确认数据、标签、显存和保存逻辑正常后，再改回 `False` 并执行正式 60 epoch。\n","metadata":{}},{"cell_type":"code","source":"# 1. 用户配置：E1 1280 分辨率对照实验\nRUN_MODE = \"train\"  # 正式训练默认只训练；可选：\"train\"、\"full\"、\"evaluate\"\nAUTO_PREPARE_IF_NEEDED = True\nSMOKE_TEST = False  # 首次建议 True；通过后改回 False 正式训练\n\nEXPERIMENT_ID = \"E1\"\nEXPERIMENT_FACTOR = \"input_resolution\"\nBASELINE_640_SOURCE_SIGNATURE = \"b68b716ab287\"\nREQUIRE_BASELINE_SPLIT_MATCH = True\n\nMODEL_WEIGHTS = \"yolo11n.pt\"\nTRAIN_IMGSZ = 1280\nDEPLOY_IMGSZ = 1280\nEPOCHS = 1 if SMOKE_TEST else 60\nPATIENCE = 1 if SMOKE_TEST else 15\nTRAIN_BATCH = -1       # 与 E0 一样使用 Ultralytics AutoBatch；nbs 默认保持 64\nEVAL_BATCH = 1         # 1280 独立评估必须使用安全小 batch\nWORKERS = 2\nSEED = 2026\n\nOPTIMIZER = \"AdamW\"\nLR0 = 1e-3\nLRF = 0.01\nWEIGHT_DECAY = 5e-4\nWARMUP_EPOCHS = 3.0\n\nTARGET_VALIDATION_SENSITIVITY = 0.90\nPREDICT_CONF_FLOOR = 0.001\nFROC_IOU_THRESHOLD = 0.50\nFROC_FP_PER_IMAGE_TARGETS = (0.125, 0.25, 0.5, 1.0, 2.0, 4.0)\nBOOTSTRAP_REPEATS = 1000\n\nif RUN_MODE not in {\"train\", \"full\", \"evaluate\"}:\n    raise ValueError('RUN_MODE 只能是 \"train\"、\"full\" 或 \"evaluate\"。')\nif TRAIN_IMGSZ != 1280 or DEPLOY_IMGSZ != 1280:\n    raise ValueError(\"本 Notebook 已锁定为 E1 1280 实验，不能改回其他分辨率。\")\nif not SMOKE_TEST and EPOCHS != 60:\n    raise ValueError(\"E1 正式筛选实验必须保持 60 epoch，与 E0 一致。\")\nif EPOCHS < 1:\n    raise ValueError(\"epoch 必须至少为 1。\")\nif not 0 < TARGET_VALIDATION_SENSITIVITY <= 1:\n    raise ValueError(\"目标敏感度必须在 (0, 1]。\")\n","metadata":{"tags":["parameters"]},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 2. 依赖：只补装缺失或不兼容的包，不重装 Kaggle 自带 PyTorch\nimport importlib.util\nimport subprocess\nimport sys\nfrom importlib.metadata import PackageNotFoundError, version as package_version\n\nfrom packaging.specifiers import SpecifierSet\nfrom packaging.version import Version\n\nPACKAGE_REQUIREMENTS = {\n    \"pydicom\": (\"pydicom\", \">=3.0,<4\"),\n    \"ultralytics\": (\"ultralytics\", \">=8.3,<9\"),\n    \"onnx\": (\"onnx\", \">=1.16,<2\"),\n    \"onnxslim\": (\"onnxslim\", \">=0.1.34\"),\n    \"sklearn\": (\"scikit-learn\", \">=1.4,<2\"),\n}\n\ninstall_specs = []\nfor module_name, (distribution, specifier) in PACKAGE_REQUIREMENTS.items():\n    try:\n        installed = package_version(distribution)\n        compatible = Version(installed) in SpecifierSet(specifier)\n    except PackageNotFoundError:\n        compatible = False\n    if importlib.util.find_spec(module_name) is None or not compatible:\n        install_specs.append(distribution + specifier)\n\nif install_specs:\n    print(\"Installing:\", install_specs)\n    try:\n        subprocess.check_call(\n            [sys.executable, \"-m\", \"pip\", \"install\", \"-q\", *install_specs]\n        )\n    except subprocess.CalledProcessError as error:\n        raise RuntimeError(\n            \"依赖安装失败。请在 Kaggle Settings 中开启 Internet 后重新 Run All。\"\n        ) from error\nelse:\n    print(\"All required packages are already compatible.\")\n\nif importlib.util.find_spec(\"torch\") is None:\n    raise RuntimeError(\"当前 Kaggle 镜像没有 PyTorch，请换用 GPU Notebook 镜像。\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. 通用导入、路径、随机种子与运行环境\nimport gc\nimport hashlib\nimport json\nimport math\nimport os\nimport platform\nimport random\nimport shutil\nimport traceback\nimport warnings\nfrom datetime import datetime, timezone\nfrom pathlib import Path\nfrom time import perf_counter\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport yaml\nfrom PIL import Image, ImageDraw\nfrom tqdm.auto import tqdm\n\ntry:\n    from IPython.display import display\nexcept ImportError:\n    display = print\n\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\nINPUT_ROOT = Path(\"/kaggle/input\")\nWORKING_ROOT = Path(\"/kaggle/working\")\nMODEL_OUTPUT_ROOT = WORKING_ROOT / \"vindr_nodule_mass_E1_yolo11n_1280_v1\"\nMODEL_REPORT_ROOT = MODEL_OUTPUT_ROOT / \"reports\"\nMODEL_RUNS_ROOT = MODEL_OUTPUT_ROOT / \"runs\"\nMODEL_ARTIFACT_ROOT = MODEL_OUTPUT_ROOT / \"artifacts\"\nfor directory in (\n    MODEL_OUTPUT_ROOT,\n    MODEL_REPORT_ROOT,\n    MODEL_RUNS_ROOT,\n    MODEL_ARTIFACT_ROOT,\n):\n    directory.mkdir(parents=True, exist_ok=True)\n\nos.environ.setdefault(\"WANDB_DISABLED\", \"true\")\nos.environ.setdefault(\"PYTORCH_CUDA_ALLOC_CONF\", \"expandable_segments:True\")\n\nprint(\"Python:\", platform.python_version())\nprint(\"Run mode:\", RUN_MODE)\nprint(\"Input root:\", INPUT_ROOT)\nprint(\"Experiment:\", EXPERIMENT_ID, EXPERIMENT_FACTOR)\nprint(\"Smoke test:\", SMOKE_TEST)\nprint(\"Model output:\", MODEL_OUTPUT_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. 查找并验证已保存的 V6 清洗 Output\nEXPECTED_CLEANING_VERSION = \"6.0-release\"\nEXPECTED_SOURCE_DATASET_SIGNATURE = \"21e0ec1e3506\"\nEXPECTED_TARGET_CLASS = \"Nodule/Mass\"\nVALID_SPLITS = (\"train\", \"val\", \"test\")\n\ndef read_json(path: Path):\n    return json.loads(path.read_text(encoding=\"utf-8\"))\n\ndef parse_gate(path: Path):\n    values = {}\n    for line in path.read_text(encoding=\"utf-8\").splitlines():\n        if \"=\" in line:\n            key, value = line.split(\"=\", 1)\n            values[key.strip()] = value.strip()\n    return values\n\ndef yaml_class_name(config):\n    names = config.get(\"names\", {})\n    if isinstance(names, list):\n        return names[0] if names else None\n    return names.get(0, names.get(\"0\"))\n\ndef inspect_clean_gate(gate_path: Path):\n    gate_path = Path(gate_path)\n    run_root = gate_path.parent.parent\n    report_root = run_root / \"reports\"\n    dataset_root = run_root / \"dataset\"\n    identity_path = dataset_root / \"dataset_identity.json\"\n    acceptance_path = report_root / \"cleaning_acceptance_final.json\"\n    data_yaml_path = dataset_root / \"data.yaml\"\n    manifest_path = dataset_root / \"manifest.csv\"\n\n    reasons = []\n    for required in (\n        gate_path,\n        identity_path,\n        acceptance_path,\n        data_yaml_path,\n        manifest_path,\n    ):\n        if not required.is_file():\n            reasons.append(f\"missing:{required.name}\")\n    for split in VALID_SPLITS:\n        if not (dataset_root / \"images\" / split).is_dir():\n            reasons.append(f\"missing:images/{split}\")\n        if not (dataset_root / \"labels\" / split).is_dir():\n            reasons.append(f\"missing:labels/{split}\")\n    if reasons:\n        return {\"valid\": False, \"gate_path\": gate_path, \"reasons\": reasons}\n\n    try:\n        gate = parse_gate(gate_path)\n        identity = read_json(identity_path)\n        acceptance = read_json(acceptance_path)\n        data_config = yaml.safe_load(data_yaml_path.read_text(encoding=\"utf-8\"))\n    except Exception as error:\n        return {\n            \"valid\": False,\n            \"gate_path\": gate_path,\n            \"reasons\": [f\"parse_error:{type(error).__name__}:{error}\"],\n        }\n\n    fingerprint_values = {\n        str(gate.get(\"CLEAN_DATASET_FINGERPRINT\", \"\")).strip(),\n        str(identity.get(\"clean_dataset_fingerprint\", \"\")).strip(),\n        str(acceptance.get(\"clean_dataset_fingerprint\", \"\")).strip(),\n    }\n    fingerprint_values.discard(\"\")\n\n    checks = {\n        \"status_ready\": gate.get(\"STATUS\") == \"TRAINING_READY\",\n        \"gate_ready\": gate.get(\"TRAINING_READY\") == \"True\",\n        \"acceptance_ready\": acceptance.get(\"training_ready\") is True,\n        \"acceptance_status\": acceptance.get(\"overall_status\") == \"TRAINING_READY\",\n        \"cleaning_version\": (\n            gate.get(\"CLEANING_VERSION\") == EXPECTED_CLEANING_VERSION\n            and identity.get(\"cleaning_version\") == EXPECTED_CLEANING_VERSION\n        ),\n        \"source_signature\": (\n            gate.get(\"SOURCE_DATASET_SIGNATURE\")\n            == EXPECTED_SOURCE_DATASET_SIGNATURE\n            and identity.get(\"source_dataset_signature\")\n            == EXPECTED_SOURCE_DATASET_SIGNATURE\n        ),\n        \"target_class\": (\n            identity.get(\"target_class\") == EXPECTED_TARGET_CLASS\n            and yaml_class_name(data_config) == EXPECTED_TARGET_CLASS\n            and int(data_config.get(\"nc\", 1)) == 1\n        ),\n        \"fingerprint_present_and_consistent\": (\n            len(fingerprint_values) == 1\n            and len(next(iter(fingerprint_values), \"\")) == 64\n        ),\n        \"copy_export\": acceptance.get(\"checks\", {}).get(\n            \"portable_data_yaml_resolves\", False\n        ),\n        \"all_acceptance_checks\": all(\n            bool(value) for value in acceptance.get(\"checks\", {}).values()\n        ),\n    }\n    failed = [name for name, passed in checks.items() if not passed]\n    return {\n        \"valid\": not failed,\n        \"gate_path\": gate_path,\n        \"run_root\": run_root,\n        \"report_root\": report_root,\n        \"dataset_root\": dataset_root,\n        \"identity\": identity,\n        \"acceptance\": acceptance,\n        \"gate\": gate,\n        \"data_config\": data_config,\n        \"fingerprint\": next(iter(fingerprint_values), \"\"),\n        \"checks\": checks,\n        \"reasons\": failed,\n    }\n\ndef discover_clean_outputs():\n    gate_paths = set()\n    for search_root in (\n        INPUT_ROOT,\n        WORKING_ROOT / \"vindr_nodule_mass_cleaning_final_v6_1280\",\n    ):\n        if search_root.exists():\n            gate_paths.update(search_root.rglob(\"TRAINING_GATE.txt\"))\n    inspected = [inspect_clean_gate(path) for path in sorted(gate_paths)]\n    valid = [item for item in inspected if item.get(\"valid\")]\n    return valid, inspected\n\nclean_candidates, inspected_clean_candidates = discover_clean_outputs()\nunique_clean_fingerprints = {\n    item[\"fingerprint\"] for item in clean_candidates\n}\nif len(unique_clean_fingerprints) > 1:\n    details = \"\\n\".join(\n        f\"- {item['dataset_root']} -> {item['fingerprint']}\"\n        for item in clean_candidates\n    )\n    raise RuntimeError(\n        \"发现多个不同数据指纹的 TRAINING_READY 数据集，禁止自动选择：\\n\"\n        + details\n    )\n\ndef clean_candidate_rank(item):\n    path_text = str(item[\"dataset_root\"])\n    # 已保存的 Kaggle Input 是只读版本，优先于同指纹 working 副本。\n    return (0 if path_text.startswith(\"/kaggle/input/\") else 1, path_text)\n\nSELECTED_CLEAN_CANDIDATE = (\n    sorted(clean_candidates, key=clean_candidate_rank)[0]\n    if clean_candidates else None\n)\nCLEAN_OUTPUT_AVAILABLE = SELECTED_CLEAN_CANDIDATE is not None\n\nif CLEAN_OUTPUT_AVAILABLE:\n    print(\"Verified V6 TRAINING_READY dataset:\")\n    print(\" -\", SELECTED_CLEAN_CANDIDATE[\"dataset_root\"])\n    print(\"Fingerprint:\", SELECTED_CLEAN_CANDIDATE[\"fingerprint\"])\n    print(\"Preprocessing fallback: skipped\")\n    print(\"Cleaning fallback: skipped\")\nelse:\n    print(\"No verified V6 TRAINING_READY Output is mounted.\")\n    failed_ready_like = [\n        item for item in inspected_clean_candidates\n        if item.get(\"gate\", {}).get(\"TRAINING_READY\") == \"True\"\n        and not item.get(\"valid\")\n    ]\n    for item in failed_ready_like[:5]:\n        print(\"Rejected candidate:\", item[\"gate_path\"], item[\"reasons\"])\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4B. 在任何耗时预处理前确认 E0 的 640 划分参考已经挂载\nBASELINE_SPLIT_REFERENCE_MANIFESTS = []\n\nif INPUT_ROOT.exists():\n    for identity_path in INPUT_ROOT.rglob(\"dataset_identity.json\"):\n        try:\n            identity = read_json(identity_path)\n        except Exception:\n            continue\n        if (\n            identity.get(\"source_dataset_signature\")\n            == BASELINE_640_SOURCE_SIGNATURE\n            and identity.get(\"target_class\") == EXPECTED_TARGET_CLASS\n        ):\n            manifest_path = identity_path.parent / \"manifest.csv\"\n            if manifest_path.is_file():\n                BASELINE_SPLIT_REFERENCE_MANIFESTS.append(manifest_path)\n\nBASELINE_SPLIT_REFERENCE_MANIFESTS = sorted(\n    set(BASELINE_SPLIT_REFERENCE_MANIFESTS), key=lambda path: str(path)\n)\n\nif REQUIRE_BASELINE_SPLIT_MATCH and not BASELINE_SPLIT_REFERENCE_MANIFESTS:\n    raise FileNotFoundError(\n        \"没有找到 E0 的 640 V6 manifest.csv。请把基础模型使用的 datacleaning3 \"\n        \"TRAINING_READY Output 添加为当前 Notebook 的 Input。\"\n    )\n\nprint(\"640 split reference candidates:\")\nfor path in BASELINE_SPLIT_REFERENCE_MANIFESTS:\n    print(\" -\", path)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 5. 仅在缺少 1280 V6 Output 时定位 1280 预处理源；再缺失才运行原始 DICOM 预处理\ndef inspect_preprocessed_root(dataset_root: Path):\n    dataset_root = Path(dataset_root)\n    output_root = None\n    for parent in [dataset_root, *dataset_root.parents]:\n        if parent.name == \"vindr_nodule_mass_training_1280\":\n            output_root = parent\n            break\n    required = [\n        dataset_root / \"manifest.csv\",\n        dataset_root / \"dataset_config.json\",\n        dataset_root / \"data.yaml\",\n        *[dataset_root / \"images\" / split for split in VALID_SPLITS],\n        *[dataset_root / \"labels\" / split for split in VALID_SPLITS],\n    ]\n    if output_root is not None:\n        required.extend([\n            output_root / \"reports\" / \"fused_boxes_original_coordinates.csv\",\n            output_root / \"reports\" / \"split_plan.csv\",\n        ])\n    missing = [str(path) for path in required if not path.exists()]\n    if output_root is None:\n        missing.append(\"ancestor:vindr_nodule_mass_training_1280\")\n    if missing:\n        return None\n    try:\n        config = read_json(dataset_root / \"dataset_config.json\")\n        computed_signature = hashlib.sha256(\n            json.dumps(\n                config, sort_keys=True, ensure_ascii=False\n            ).encode(\"utf-8\")\n        ).hexdigest()[:12]\n        manifest_head = pd.read_csv(\n            dataset_root / \"manifest.csv\", usecols=[\"image_id\", \"split\"]\n        )\n    except Exception:\n        return None\n    if (\n        dataset_root.name != EXPECTED_SOURCE_DATASET_SIGNATURE\n        or computed_signature != EXPECTED_SOURCE_DATASET_SIGNATURE\n        or config.get(\"target_class\") != EXPECTED_TARGET_CLASS\n        or int(config.get(\"preprocess_max_side\", -1)) != 1280\n        or len(manifest_head) != 15000\n        or set(manifest_head[\"split\"].astype(str)) != set(VALID_SPLITS)\n    ):\n        return None\n    return dataset_root\n\npreprocessed_candidates = []\nif not CLEAN_OUTPUT_AVAILABLE:\n    exact_default = (\n        INPUT_ROOT / \"notebooks\" / \"hilarylee33\" / \"1280version1\"\n        / \"vindr_nodule_mass_training_1280\" / \"datasets\"\n        / EXPECTED_SOURCE_DATASET_SIGNATURE\n    )\n    exact_checked = inspect_preprocessed_root(exact_default)\n    if exact_checked is not None:\n        preprocessed_candidates.append(exact_checked)\n    for search_root in (\n        INPUT_ROOT,\n        WORKING_ROOT / \"vindr_nodule_mass_training_1280\" / \"datasets\",\n    ):\n        if not search_root.exists():\n            continue\n        for candidate in search_root.rglob(EXPECTED_SOURCE_DATASET_SIGNATURE):\n            if not candidate.is_dir():\n                continue\n            checked = inspect_preprocessed_root(candidate)\n            if checked is not None:\n                preprocessed_candidates.append(checked)\n\npreprocessed_candidates = sorted(\n    set(preprocessed_candidates),\n    key=lambda path: (\n        0 if \"1280\" in str(path).lower() else 1,\n        0 if str(path).startswith(\"/kaggle/input/\") else 1,\n        str(path),\n    ),\n)\nPREPROCESSED_SOURCE_ROOT = (\n    preprocessed_candidates[0] if preprocessed_candidates else None\n)\n\nNEED_CLEANING = not CLEAN_OUTPUT_AVAILABLE\nNEED_PREPROCESSING = NEED_CLEANING and PREPROCESSED_SOURCE_ROOT is None\n\nif NEED_PREPROCESSING and not AUTO_PREPARE_IF_NEEDED:\n    raise FileNotFoundError(\n        \"没有 1280 V6 Output 或 1280 预处理 Output；\"\n        \"AUTO_PREPARE_IF_NEEDED=False，因此没有自动读取原始 DICOM。\"\n    )\n\nprint(\"Need preprocessing:\", NEED_PREPROCESSING)\nprint(\"Need V6 cleaning:\", NEED_CLEANING)\nif PREPROCESSED_SOURCE_ROOT is not None:\n    print(\"Verified 1280 preprocessing source:\", PREPROCESSED_SOURCE_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## A. 内置 1280 DICOM 预处理\n\n以下代码直接从原始 DICOM 生成最长边不超过 1280 的灰度 PNG，并保留宽高比。\n\n它不是把 640 PNG 放大到 1280。只有缺少合格的 1280 预处理/清洗 Output 时才执行；否则自动跳过。\n","metadata":{}},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 2. 导入、版本记录与随机种子\n    import gc\n    import hashlib\n    import json\n    import math\n    import os\n    import platform\n    import random\n    import shutil\n    import traceback\n    from concurrent.futures import ThreadPoolExecutor, as_completed\n    from importlib.metadata import version as package_version\n    from pathlib import Path\n    from time import perf_counter\n\n    import matplotlib.pyplot as plt\n    import numpy as np\n    import pandas as pd\n    import pydicom\n    import sklearn\n    import yaml\n    from PIL import Image, ImageDraw\n    from pydicom.pixels import apply_modality_lut, apply_voi_lut\n    from sklearn.metrics import (\n        average_precision_score,\n        confusion_matrix,\n        precision_recall_curve,\n        roc_auc_score,\n        roc_curve,\n    )\n    from sklearn.model_selection import train_test_split\n    from tqdm.auto import tqdm\n\n    SEED = 2026\n    random.seed(SEED)\n    np.random.seed(SEED)\n    versions = {\n        \"python\": platform.python_version(),\n        \"numpy\": np.__version__,\n        \"pandas\": pd.__version__,\n        \"pydicom\": pydicom.__version__,\n        \"torch\": package_version(\"torch\"),\n        \"ultralytics\": package_version(\"ultralytics\"),\n        \"sklearn\": sklearn.__version__,\n        \"cuda_available\": None,  # 训练阶段导入 torch 后更新\n        \"gpu\": None,\n    }\n    print(json.dumps(versions, indent=2, ensure_ascii=False))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 3. 1280 预处理配置：仅在缺少合格 1280 Output 时执行\n    PIPELINE_STAGE = \"preprocess\"  # 可选：\"preprocess\"、\"train\"、\"full\"\n    TARGET_CLASS = \"Nodule/Mass\"\n    EXPECTED_TRAIN_DICOM_COUNT = 15000  # Kaggle VinBigData 官方训练集\n\n    # 数据与标注\n    LOWER_PERCENTILE = 0.5\n    UPPER_PERCENTILE = 99.5\n    PREPROCESS_MAX_SIDE = 1280       # 直接从原始 DICOM 生成 1280 实验图像\n    PNG_COMPRESSION = 6\n    BOX_MERGE_IOU = 0.35             # 聚合同一病灶的多医生重叠框\n    MIN_RADIOLOGISTS_PER_LESION = 1  # 保留敏感性；对照实验可改为 2\n    MIN_BOX_SIDE_PX = 2.0\n    TRAIN_RATIO = 0.80\n    VAL_RATIO = 0.10\n    TEST_RATIO = 0.10\n    MAX_TRAIN_NEGATIVE_TO_POSITIVE = None  # None = 保留训练集全部阴性图像\n    PREPROCESS_WORKERS = 1           # 正式预处理固定单张串行，避免多张大 DICOM 同时占用 RAM\n    PREPROCESS_GC_EVERY = 25\n    PROGRESS_SAVE_EVERY = 100\n    RESUME_FROM_PREVIOUS_OUTPUT = True\n    FAIL_ON_PREPROCESS_ERROR = True\n\n    # 训练（preprocess 阶段不会导入 torch/ultralytics，也不会运行这些参数）\n    MODEL_WEIGHTS = \"yolo11n.pt\"\n    TRAIN_IMGSZ = 1280\n    DEPLOY_IMGSZ = 1280\n    EPOCHS = 1 if SMOKE_TEST else 60\n    PATIENCE = 1 if SMOKE_TEST else 15\n    BATCH_CANDIDATES = (1,)          # 旧预处理模块兼容变量；正式训练使用顶部 TRAIN_BATCH\n    VAL_BATCH = 1\n    WORKERS = min(2, os.cpu_count() or 2)\n    OPTIMIZER = \"AdamW\"\n    LR0 = 1e-3\n    WEIGHT_DECAY = 5e-4\n    EXPERIMENT_NAME = \"E1_yolo11n_nm_1280_seed2026\"\n    RESUME_TRAINING = True\n\n    RUN_TRAINING = PIPELINE_STAGE in {\"train\", \"full\"}\n    RUN_DETECTION_EVALUATION = PIPELINE_STAGE == \"full\"\n    RUN_CLINICAL_EVALUATION = PIPELINE_STAGE == \"full\"\n    EXPORT_ONNX = PIPELINE_STAGE == \"full\"\n\n    # 筛查阈值与置信区间\n    TARGET_VALIDATION_SENSITIVITY = 0.90\n    PREDICT_CONF_FLOOR = 0.001\n    BOOTSTRAP_REPEATS = 1000\n\n    INPUT_ROOT = Path(\"/kaggle/input\")\n    OUTPUT_ROOT = Path(\"/kaggle/working/vindr_nodule_mass_training_1280\")\n    REPORT_ROOT = OUTPUT_ROOT / \"reports\"\n    RUNS_ROOT = OUTPUT_ROOT / \"runs\"\n    ARTIFACT_ROOT = OUTPUT_ROOT / \"artifacts\"\n    for directory in (OUTPUT_ROOT, REPORT_ROOT, RUNS_ROOT, ARTIFACT_ROOT):\n        directory.mkdir(parents=True, exist_ok=True)\n\n    if PIPELINE_STAGE not in {\"preprocess\", \"train\", \"full\"}:\n        raise ValueError('PIPELINE_STAGE 只能是 \"preprocess\"、\"train\" 或 \"full\"。')\n    if not math.isclose(TRAIN_RATIO + VAL_RATIO + TEST_RATIO, 1.0):\n        raise ValueError(\"TRAIN_RATIO + VAL_RATIO + TEST_RATIO 必须等于 1。\")\n    if not 0 <= LOWER_PERCENTILE < UPPER_PERCENTILE <= 100:\n        raise ValueError(\"百分位配置无效。\")\n\n    print(f\"Pipeline stage: {PIPELINE_STAGE}\")\n    print(\"Preprocessing device: CPU only\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 4. 自动定位 VinBigData DICOM 与 train.csv\n    REQUIRED_COLUMNS = {\n        \"image_id\", \"rad_id\", \"class_name\",\n        \"x_min\", \"y_min\", \"x_max\", \"y_max\",\n    }\n\n    def score_path(path: Path) -> int:\n        text = str(path).lower()\n        score = 0\n        score += 20 if \"vinbigdata\" in text else 0\n        score += 10 if \"chest-xray-abnormalities\" in text else 0\n        score += 3 if \"competition\" in text else 0\n        return score\n\n    csv_candidates = []\n    for path in INPUT_ROOT.rglob(\"train.csv\"):\n        try:\n            columns = set(pd.read_csv(path, nrows=0).columns)\n            if REQUIRED_COLUMNS.issubset(columns):\n                csv_candidates.append(path)\n        except Exception:\n            pass\n\n    train_dir_candidates = []\n    for candidate in INPUT_ROOT.rglob(\"train\"):\n        if candidate.is_dir() and next(candidate.glob(\"*.dicom\"), None) is not None:\n            train_dir_candidates.append(candidate)\n\n    if not csv_candidates or not train_dir_candidates:\n        raise FileNotFoundError(\n            \"没有找到 VinBigData train.csv 或 train/*.dicom。\"\n            \"请在 Kaggle 的 Add Input 中添加竞赛数据。\"\n        )\n\n    ANNOTATION_CSV = sorted(csv_candidates, key=score_path, reverse=True)[0]\n    TRAIN_DICOM_DIR = sorted(train_dir_candidates, key=score_path, reverse=True)[0]\n    DICOM_MAP = {path.stem: path for path in TRAIN_DICOM_DIR.glob(\"*.dicom\")}\n    if len(DICOM_MAP) != EXPECTED_TRAIN_DICOM_COUNT:\n        raise RuntimeError(\n            f\"检测到 {len(DICOM_MAP)} 张训练 DICOM，预期为 {EXPECTED_TRAIN_DICOM_COUNT} 张。\"\n            \"请确认 Add Input 添加的是完整 VinBigData Chest X-ray Abnormalities Detection 数据。\"\n        )\n\n    print(\"Annotation CSV:\", ANNOTATION_CSV)\n    print(\"DICOM directory:\", TRAIN_DICOM_DIR)\n    print(\"DICOM images:\", len(DICOM_MAP))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 5. 读取并清洗 Nodule/Mass 原始框\n    annotations = pd.read_csv(\n        ANNOTATION_CSV,\n        usecols=lambda column: column in REQUIRED_COLUMNS,\n    )\n    annotations[\"image_id\"] = annotations[\"image_id\"].astype(str)\n    annotations[\"rad_id\"] = annotations[\"rad_id\"].fillna(\"Unknown\").astype(str)\n    annotations[\"class_name\"] = annotations[\"class_name\"].fillna(\"\").astype(str).str.strip()\n    annotation_image_ids = set(annotations[\"image_id\"])\n    missing_annotation_ids = sorted(set(DICOM_MAP) - annotation_image_ids)\n    extra_annotation_ids = sorted(annotation_image_ids - set(DICOM_MAP))\n    if missing_annotation_ids or extra_annotation_ids:\n        raise RuntimeError(\n            \"DICOM 与 train.csv 的 image_id 不完全一致：\"\n            f\"缺少标注 {len(missing_annotation_ids)}，多余标注 {len(extra_annotation_ids)}。\"\n            \"不能在标签不完整时继续正式训练。\"\n        )\n\n    target_raw = annotations.loc[\n        annotations[\"class_name\"].eq(TARGET_CLASS)\n    ].copy()\n    coordinate_columns = [\"x_min\", \"y_min\", \"x_max\", \"y_max\"]\n    target_raw[coordinate_columns] = target_raw[coordinate_columns].apply(\n        pd.to_numeric, errors=\"coerce\"\n    )\n\n    valid_numeric = np.isfinite(target_raw[coordinate_columns]).all(axis=1)\n    valid_area = (\n        (target_raw[\"x_max\"] > target_raw[\"x_min\"])\n        & (target_raw[\"y_max\"] > target_raw[\"y_min\"])\n    )\n    invalid_target_rows = target_raw.loc[~(valid_numeric & valid_area)].copy()\n    target_raw = target_raw.loc[valid_numeric & valid_area].copy()\n\n    missing_dicom_ids = sorted(set(target_raw[\"image_id\"]) - set(DICOM_MAP))\n    if missing_dicom_ids:\n        print(f\"Warning: {len(missing_dicom_ids)} 个有标注 image_id 没有对应 DICOM，将被忽略。\")\n        target_raw = target_raw.loc[target_raw[\"image_id\"].isin(DICOM_MAP)].copy()\n\n    invalid_target_rows.to_csv(REPORT_ROOT / \"invalid_raw_target_rows.csv\", index=False)\n    print(\"All annotation rows:\", len(annotations))\n    print(\"Valid raw Nodule/Mass boxes:\", len(target_raw))\n    print(\"Images with raw Nodule/Mass boxes:\", target_raw[\"image_id\"].nunique())\n    print(\"Invalid target rows removed:\", len(invalid_target_rows))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 6. 将同一病灶的多医生重叠框聚合为一个训练框\n    def box_iou(a, b) -> float:\n        x1 = max(float(a[0]), float(b[0]))\n        y1 = max(float(a[1]), float(b[1]))\n        x2 = min(float(a[2]), float(b[2]))\n        y2 = min(float(a[3]), float(b[3]))\n        intersection = max(0.0, x2 - x1) * max(0.0, y2 - y1)\n        area_a = max(0.0, float(a[2]) - float(a[0])) * max(0.0, float(a[3]) - float(a[1]))\n        area_b = max(0.0, float(b[2]) - float(b[0])) * max(0.0, float(b[3]) - float(b[1]))\n        union = area_a + area_b - intersection\n        return intersection / union if union > 0 else 0.0\n\n    def fuse_image_boxes(group: pd.DataFrame) -> pd.DataFrame:\n        group = group.reset_index(drop=True)\n        boxes = group[coordinate_columns].to_numpy(dtype=np.float64)\n        parent = list(range(len(group)))\n\n        def find(i):\n            while parent[i] != i:\n                parent[i] = parent[parent[i]]\n                i = parent[i]\n            return i\n\n        def union(i, j):\n            root_i, root_j = find(i), find(j)\n            if root_i != root_j:\n                parent[root_j] = root_i\n\n        for i in range(len(group)):\n            for j in range(i + 1, len(group)):\n                if box_iou(boxes[i], boxes[j]) >= BOX_MERGE_IOU:\n                    union(i, j)\n\n        clusters = {}\n        for index in range(len(group)):\n            clusters.setdefault(find(index), []).append(index)\n\n        records = []\n        for lesion_index, indices in enumerate(clusters.values()):\n            cluster = group.iloc[indices]\n            radiologist_count = cluster[\"rad_id\"].nunique()\n            if radiologist_count < MIN_RADIOLOGISTS_PER_LESION:\n                continue\n            median_box = cluster[coordinate_columns].median()\n            records.append({\n                \"image_id\": str(group.loc[0, \"image_id\"]),\n                \"lesion_index\": lesion_index,\n                \"x_min\": float(median_box[\"x_min\"]),\n                \"y_min\": float(median_box[\"y_min\"]),\n                \"x_max\": float(median_box[\"x_max\"]),\n                \"y_max\": float(median_box[\"y_max\"]),\n                \"radiologist_count\": int(radiologist_count),\n                \"source_box_count\": int(len(cluster)),\n                \"radiologist_ids\": \"|\".join(sorted(cluster[\"rad_id\"].unique())),\n            })\n        return pd.DataFrame(records)\n\n    fused_frames = [\n        fuse_image_boxes(group)\n        for _, group in tqdm(target_raw.groupby(\"image_id\", sort=False), desc=\"Fusing boxes\")\n    ]\n    fused_frames = [frame for frame in fused_frames if not frame.empty]\n    fused_boxes = pd.concat(fused_frames, ignore_index=True) if fused_frames else pd.DataFrame(\n        columns=[\"image_id\", \"lesion_index\", *coordinate_columns]\n    )\n    fused_boxes.to_csv(REPORT_ROOT / \"fused_boxes_original_coordinates.csv\", index=False)\n\n    raw_positive_ids = set(target_raw[\"image_id\"])\n    invalid_target_ids = set(invalid_target_rows[\"image_id\"].astype(str))\n    positive_ids = set(fused_boxes[\"image_id\"].astype(str))\n    # 有目标类记录但无法形成有效框的图像不能被误标为阴性。\n    ambiguous_ids = (raw_positive_ids - positive_ids) | (invalid_target_ids - positive_ids)\n    usable_ids = sorted(set(DICOM_MAP) - ambiguous_ids)\n    negative_ids = set(usable_ids) - positive_ids\n    FUSED_BOX_MAP = {\n        image_id: group.reset_index(drop=True)\n        for image_id, group in fused_boxes.groupby(\"image_id\", sort=False)\n    }\n\n    if not positive_ids or not negative_ids:\n        raise RuntimeError(\"聚合后必须同时存在阳性和阴性图像。\")\n\n    print(\"Fused lesions:\", len(fused_boxes))\n    print(\"Positive images:\", len(positive_ids))\n    print(\"Negative images:\", len(negative_ids))\n    print(\"Ambiguous images excluded:\", len(ambiguous_ids))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 7. 按 PatientID（可用时）划分，防止同一患者跨集合\n    def read_patient_id(image_id: str) -> str:\n        try:\n            ds = pydicom.dcmread(\n                str(DICOM_MAP[image_id]),\n                stop_before_pixels=True,\n                specific_tags=[\"PatientID\"],\n            )\n            value = str(getattr(ds, \"PatientID\", \"\")).strip()\n            return value\n        except Exception:\n            return \"\"\n\n    patient_rows = []\n    with ThreadPoolExecutor(max_workers=PREPROCESS_WORKERS) as executor:\n        futures = {executor.submit(read_patient_id, image_id): image_id for image_id in usable_ids}\n        for future in tqdm(as_completed(futures), total=len(futures), desc=\"Reading PatientID\"):\n            image_id = futures[future]\n            patient_rows.append({\"image_id\": image_id, \"patient_id\": future.result()})\n\n    patient_df = pd.DataFrame(patient_rows)\n    nonempty_patient = patient_df[\"patient_id\"].ne(\"\")\n    patient_id_is_usable = (\n        nonempty_patient.mean() >= 0.95\n        and patient_df.loc[nonempty_patient, \"patient_id\"].nunique() >= 0.50 * len(patient_df)\n    )\n    # 即使大多数 PatientID 可用，空 ID 也必须各自回退到 image_id，不能共享同一个空组。\n    if patient_id_is_usable:\n        patient_df[\"group_id\"] = np.where(\n            nonempty_patient,\n            \"patient:\" + patient_df[\"patient_id\"].astype(str),\n            \"image:\" + patient_df[\"image_id\"].astype(str),\n        )\n    else:\n        patient_df[\"group_id\"] = \"image:\" + patient_df[\"image_id\"].astype(str)\n    patient_df[\"is_positive\"] = patient_df[\"image_id\"].isin(positive_ids).astype(int)\n\n    group_table = patient_df.groupby(\"group_id\", as_index=False).agg(\n        is_positive=(\"is_positive\", \"max\"),\n        image_count=(\"image_id\", \"size\"),\n    )\n\n    train_groups, temporary_groups = train_test_split(\n        group_table,\n        test_size=VAL_RATIO + TEST_RATIO,\n        random_state=SEED,\n        stratify=group_table[\"is_positive\"],\n    )\n    relative_test_ratio = TEST_RATIO / (VAL_RATIO + TEST_RATIO)\n    val_groups, test_groups = train_test_split(\n        temporary_groups,\n        test_size=relative_test_ratio,\n        random_state=SEED,\n        stratify=temporary_groups[\"is_positive\"],\n    )\n\n    split_group_sets = {\n        \"train\": set(train_groups[\"group_id\"]),\n        \"val\": set(val_groups[\"group_id\"]),\n        \"test\": set(test_groups[\"group_id\"]),\n    }\n    split_map = {}\n    for row in patient_df.itertuples(index=False):\n        for split_name, group_ids in split_group_sets.items():\n            if row.group_id in group_ids:\n                split_map[row.image_id] = split_name\n                break\n\n    split_ids = {\n        split_name: sorted(image_id for image_id, value in split_map.items() if value == split_name)\n        for split_name in (\"train\", \"val\", \"test\")\n    }\n\n    # 默认保留训练集全部图像。只有显式设置比例时才下采样训练阴性；val/test 始终保留原分布。\n    train_positive = [image_id for image_id in split_ids[\"train\"] if image_id in positive_ids]\n    train_negative = [image_id for image_id in split_ids[\"train\"] if image_id not in positive_ids]\n    if MAX_TRAIN_NEGATIVE_TO_POSITIVE is not None:\n        max_train_negative = int(round(len(train_positive) * MAX_TRAIN_NEGATIVE_TO_POSITIVE))\n        if len(train_negative) > max_train_negative:\n            rng = random.Random(SEED)\n            train_negative = sorted(rng.sample(train_negative, max_train_negative))\n    split_ids[\"train\"] = sorted(train_positive + train_negative)\n\n    overlap = (\n        set(split_ids[\"train\"]) & set(split_ids[\"val\"])\n        | set(split_ids[\"train\"]) & set(split_ids[\"test\"])\n        | set(split_ids[\"val\"]) & set(split_ids[\"test\"])\n    )\n    if overlap:\n        raise AssertionError(f\"发现跨集合图像泄漏：{len(overlap)}\")\n\n    image_to_group = patient_df.set_index(\"image_id\")[\"group_id\"].to_dict()\n    split_plan = pd.DataFrame([\n        {\n            \"image_id\": image_id,\n            \"split\": split_name,\n            \"is_positive\": int(image_id in positive_ids),\n            \"group_id\": image_to_group[image_id],\n        }\n        for split_name, image_ids in split_ids.items()\n        for image_id in image_ids\n    ])\n    split_plan.to_csv(REPORT_ROOT / \"split_plan.csv\", index=False)\n    print(\"Split unit:\", \"PatientID\" if patient_id_is_usable else \"image_id\")\n    print(split_plan.groupby(\"split\")[\"is_positive\"].agg([\"count\", \"sum\"]))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 8. 为本次数据配置建立稳定目录，相同配置可安全断点复用\n    DATASET_CONFIG = {\n        \"target_class\": TARGET_CLASS,\n        \"lower_percentile\": LOWER_PERCENTILE,\n        \"upper_percentile\": UPPER_PERCENTILE,\n        \"preprocess_max_side\": PREPROCESS_MAX_SIDE,\n        \"box_merge_iou\": BOX_MERGE_IOU,\n        \"min_radiologists_per_lesion\": MIN_RADIOLOGISTS_PER_LESION,\n        \"min_box_side_px\": MIN_BOX_SIDE_PX,\n        \"train_ratio\": TRAIN_RATIO,\n        \"val_ratio\": VAL_RATIO,\n        \"test_ratio\": TEST_RATIO,\n        \"max_train_negative_to_positive\": MAX_TRAIN_NEGATIVE_TO_POSITIVE,\n        \"seed\": SEED,\n    }\n    config_text = json.dumps(DATASET_CONFIG, sort_keys=True, ensure_ascii=False)\n    dataset_signature = hashlib.sha256(config_text.encode(\"utf-8\")).hexdigest()[:12]\n    if dataset_signature != EXPECTED_SOURCE_DATASET_SIGNATURE:\n        raise RuntimeError(\n            \"1280 预处理配置签名发生变化：\"\n            f\"{dataset_signature} != {EXPECTED_SOURCE_DATASET_SIGNATURE}。\"\n            \"请先更新实验协议，不允许静默生成不同数据版本。\"\n        )\n    DATASET_ROOT = OUTPUT_ROOT / \"datasets\" / dataset_signature\n    for split_name in (\"train\", \"val\", \"test\"):\n        (DATASET_ROOT / \"images\" / split_name).mkdir(parents=True, exist_ok=True)\n        (DATASET_ROOT / \"labels\" / split_name).mkdir(parents=True, exist_ok=True)\n        (DATASET_ROOT / \"records\" / split_name).mkdir(parents=True, exist_ok=True)\n    (DATASET_ROOT / \"dataset_config.json\").write_text(\n        json.dumps(DATASET_CONFIG, indent=2, ensure_ascii=False), encoding=\"utf-8\"\n    )\n    # 如果用户把上一版 Kaggle Output 添加为 Input，自动找到相同配置的断点目录。\n    RESUME_DATASET_ROOTS = []\n    if RESUME_FROM_PREVIOUS_OUTPUT:\n        for candidate in INPUT_ROOT.rglob(dataset_signature):\n            if candidate.is_dir() and candidate.parent.name == \"datasets\":\n                RESUME_DATASET_ROOTS.append(candidate)\n    RESUME_DATASET_ROOTS = sorted(set(RESUME_DATASET_ROOTS), key=lambda path: str(path))\n\n    print(\"Dataset root:\", DATASET_ROOT)\n    if RESUME_DATASET_ROOTS:\n        print(\"Previous compatible output detected:\")\n        for path in RESUME_DATASET_ROOTS:\n            print(\" -\", path)\n    else:\n        print(\"No previous compatible output mounted; local completion records will still be reused.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 9. DICOM 标准化、缩放与 YOLO 标签函数\n    def create_valid_pixel_mask(raw_array: np.ndarray, ds):\n        if not hasattr(ds, \"PixelPaddingValue\"):\n            return None\n        padding_value = float(ds.PixelPaddingValue)\n        if hasattr(ds, \"PixelPaddingRangeLimit\"):\n            padding_limit = float(ds.PixelPaddingRangeLimit)\n            lower, upper = sorted((padding_value, padding_limit))\n            return ~((raw_array >= lower) & (raw_array <= upper))\n        return raw_array != padding_value\n\n    def valid_values(array: np.ndarray, mask):\n        values = array.reshape(-1) if mask is None else array[mask]\n        finite = np.isfinite(values)\n        return values if finite.all() else values[finite]\n\n    def standardize_dicom(dicom_path: Path):\n        ds = pydicom.dcmread(str(dicom_path))\n        raw = ds.pixel_array\n        if raw.ndim != 2:\n            raise ValueError(f\"Only 2D DICOM is supported: {dicom_path.name}\")\n        original_height, original_width = raw.shape\n        mask = create_valid_pixel_mask(raw, ds)\n        if mask is not None and not np.any(mask):\n            raise ValueError(\"No valid pixels outside Pixel Padding.\")\n\n        processed = apply_modality_lut(raw, ds).astype(np.float32, copy=False)\n        has_voi = (\n            hasattr(ds, \"VOILUTSequence\")\n            or (hasattr(ds, \"WindowCenter\") and hasattr(ds, \"WindowWidth\"))\n        )\n        voi_applied = False\n        if has_voi:\n            try:\n                processed = apply_voi_lut(processed, ds, index=0).astype(np.float32, copy=False)\n                voi_applied = True\n            except Exception:\n                voi_applied = False\n\n        photometric = str(getattr(ds, \"PhotometricInterpretation\", \"\")).upper()\n        if photometric not in {\"MONOCHROME1\", \"MONOCHROME2\"}:\n            raise ValueError(f\"Unsupported PhotometricInterpretation: {photometric}\")\n\n        values = valid_values(processed, mask)\n        if values.size == 0:\n            raise ValueError(\"No finite valid pixels.\")\n        if photometric == \"MONOCHROME1\":\n            np.subtract(float(values.max()) + float(values.min()), processed, out=processed)\n            values = valid_values(processed, mask)\n\n        if voi_applied:\n            lower, upper = float(values.min()), float(values.max())\n            intensity_method = \"VOI\"\n        else:\n            lower, upper = np.percentile(values, [LOWER_PERCENTILE, UPPER_PERCENTILE])\n            lower, upper = float(lower), float(upper)\n            intensity_method = \"percentile\"\n        if not np.isfinite([lower, upper]).all() or upper <= lower:\n            raise ValueError(\"Invalid intensity range.\")\n\n        # 所有大数组运算尽量原地完成，避免单张高分辨率 DICOM 同时产生多个副本。\n        processed = np.nan_to_num(processed, copy=False, nan=lower, posinf=upper, neginf=lower)\n        np.clip(processed, lower, upper, out=processed)\n        processed -= lower\n        processed *= 255.0 / (upper - lower)\n        np.rint(processed, out=processed)\n        normalized = processed.astype(np.uint8)\n        if mask is not None:\n            normalized[~mask] = 0\n\n        metadata = {\n            \"original_width\": int(original_width),\n            \"original_height\": int(original_height),\n            \"photometric_interpretation\": photometric,\n            \"voi_applied\": bool(voi_applied),\n            \"intensity_method\": intensity_method,\n            \"lower_value\": lower,\n            \"upper_value\": upper,\n            \"pixel_padding_present\": mask is not None,\n        }\n        del ds, raw, processed, values, mask\n        return normalized, metadata\n\n    def expected_output_size(original_width: int, original_height: int):\n        longest = max(original_width, original_height)\n        if PREPROCESS_MAX_SIDE is None or longest <= PREPROCESS_MAX_SIDE:\n            return original_width, original_height\n        scale = PREPROCESS_MAX_SIDE / longest\n        return max(1, round(original_width * scale)), max(1, round(original_height * scale))\n\n    def resize_keep_aspect(image_array: np.ndarray):\n        source = Image.fromarray(image_array, mode=\"L\")\n        original_width, original_height = source.size\n        output_width, output_height = expected_output_size(original_width, original_height)\n        if (output_width, output_height) == source.size:\n            return source\n        resized = source.resize((output_width, output_height), Image.Resampling.LANCZOS)\n        source.close()\n        return resized\n\n    def build_yolo_lines(image_id: str, original_width: int, original_height: int):\n        rows = FUSED_BOX_MAP.get(image_id)\n        if rows is None or rows.empty:\n            return [], 0\n        lines = []\n        dropped = 0\n        for row in rows.itertuples(index=False):\n            x1 = float(np.clip(row.x_min, 0, original_width))\n            y1 = float(np.clip(row.y_min, 0, original_height))\n            x2 = float(np.clip(row.x_max, 0, original_width))\n            y2 = float(np.clip(row.y_max, 0, original_height))\n            width, height = x2 - x1, y2 - y1\n            if width < MIN_BOX_SIDE_PX or height < MIN_BOX_SIDE_PX:\n                dropped += 1\n                continue\n            x_center = ((x1 + x2) / 2.0) / original_width\n            y_center = ((y1 + y2) / 2.0) / original_height\n            width_norm = width / original_width\n            height_norm = height / original_height\n            values = np.array([x_center, y_center, width_norm, height_norm])\n            if not np.isfinite(values).all() or not ((values > 0).all() and (values <= 1).all()):\n                dropped += 1\n                continue\n            lines.append(f\"0 {x_center:.8f} {y_center:.8f} {width_norm:.8f} {height_norm:.8f}\")\n        return lines, dropped\n\n    def load_completion_record(\n        image_id: str,\n        split_name: str,\n        image_path: Path,\n        label_path: Path,\n        record_path: Path,\n    ):\n        # 只有完成记录、PNG 和标签三者都通过快速检查时才跳过。\n        if not (record_path.exists() and image_path.exists() and label_path.exists()):\n            return None\n        try:\n            record = json.loads(record_path.read_text(encoding=\"utf-8\"))\n            if str(record.get(\"image_id\")) != image_id or record.get(\"split\") != split_name:\n                return None\n            with Image.open(image_path) as existing:\n                expected_size = (int(record[\"output_width\"]), int(record[\"output_height\"]))\n                if existing.size != expected_size or existing.mode != \"L\":\n                    return None\n                existing.verify()\n            actual_label_count = sum(\n                bool(line.strip())\n                for line in label_path.read_text(encoding=\"utf-8\").splitlines()\n            )\n            if actual_label_count != int(record[\"label_count\"]):\n                return None\n            record[\"run_status\"] = \"skipped_completed\"\n            record[\"skipped_completed\"] = True\n            record.setdefault(\"imported_previous_output\", False)\n            return record\n        except Exception:\n            return None\n\n    def save_completion_record(record: dict, record_path: Path):\n        # 临时文件写完后再原子替换，避免中断时留下半个 JSON。\n        temporary_path = record_path.with_suffix(\".json.tmp\")\n        temporary_path.write_text(\n            json.dumps(record, ensure_ascii=False, indent=2),\n            encoding=\"utf-8\",\n        )\n        temporary_path.replace(record_path)\n\n    def copy_file_atomic(source: Path, destination: Path):\n        temporary = destination.with_suffix(destination.suffix + \".importing\")\n        shutil.copy2(source, temporary)\n        temporary.replace(destination)\n\n    def import_previous_completion(image_id: str, split_name: str, image_path: Path, label_path: Path, record_path: Path):\n        for root in RESUME_DATASET_ROOTS:\n            source_image = root / \"images\" / split_name / f\"{image_id}.png\"\n            source_label = root / \"labels\" / split_name / f\"{image_id}.txt\"\n            source_record = root / \"records\" / split_name / f\"{image_id}.json\"\n            previous = load_completion_record(\n                image_id, split_name, source_image, source_label, source_record\n            )\n            if previous is None:\n                continue\n            copy_file_atomic(source_image, image_path)\n            copy_file_atomic(source_label, label_path)\n            previous.update({\n                \"image_path\": str(image_path),\n                \"label_path\": str(label_path),\n                \"completion_record_path\": str(record_path),\n                \"run_status\": \"imported_previous_output\",\n                \"skipped_completed\": True,\n                \"imported_previous_output\": True,\n            })\n            save_completion_record(previous, record_path)\n            return previous\n        return None\n\n    def process_one(image_id: str, split_name: str):\n        image_path = DATASET_ROOT / \"images\" / split_name / f\"{image_id}.png\"\n        label_path = DATASET_ROOT / \"labels\" / split_name / f\"{image_id}.txt\"\n        record_path = DATASET_ROOT / \"records\" / split_name / f\"{image_id}.json\"\n\n        completed = load_completion_record(\n            image_id, split_name, image_path, label_path, record_path\n        )\n        if completed is not None:\n            return completed\n\n        imported = import_previous_completion(\n            image_id, split_name, image_path, label_path, record_path\n        )\n        if imported is not None:\n            return imported\n\n        dicom_path = DICOM_MAP[image_id]\n        header = pydicom.dcmread(str(dicom_path), stop_before_pixels=True)\n        original_width = int(header.Columns)\n        original_height = int(header.Rows)\n        expected_size = expected_output_size(original_width, original_height)\n        reused = False\n        metadata = {\n            \"original_width\": original_width,\n            \"original_height\": original_height,\n            \"photometric_interpretation\": str(getattr(header, \"PhotometricInterpretation\", \"\")),\n            \"voi_applied\": None,\n            \"intensity_method\": \"reused\",\n            \"lower_value\": None,\n            \"upper_value\": None,\n            \"pixel_padding_present\": hasattr(header, \"PixelPaddingValue\"),\n        }\n\n        if image_path.exists():\n            try:\n                with Image.open(image_path) as existing:\n                    reused = existing.size == expected_size and existing.mode == \"L\"\n                    if reused:\n                        existing.verify()\n            except Exception:\n                reused = False\n\n        if not reused:\n            image_array, metadata = standardize_dicom(dicom_path)\n            processed_image = resize_keep_aspect(image_array)\n            temporary_image_path = image_path.with_suffix(\".png.tmp\")\n            try:\n                processed_image.save(\n                    temporary_image_path,\n                    format=\"PNG\",\n                    compress_level=PNG_COMPRESSION,\n                    optimize=False,\n                )\n                expected_size = processed_image.size\n                temporary_image_path.replace(image_path)\n            finally:\n                processed_image.close()\n                del processed_image, image_array\n                if temporary_image_path.exists():\n                    temporary_image_path.unlink()\n\n        yolo_lines, dropped_boxes = build_yolo_lines(image_id, original_width, original_height)\n        if image_id in positive_ids and not yolo_lines:\n            raise ValueError(\"Positive image has no valid fused box after clipping.\")\n        temporary_label_path = label_path.with_suffix(\".txt.tmp\")\n        temporary_label_path.write_text(\"\\n\".join(yolo_lines) + (\"\\n\" if yolo_lines else \"\"), encoding=\"utf-8\")\n        temporary_label_path.replace(label_path)\n\n        record = {\n            \"image_id\": image_id,\n            \"split\": split_name,\n            \"is_positive\": int(bool(yolo_lines)),\n            \"label_count\": len(yolo_lines),\n            \"dropped_box_count\": dropped_boxes,\n            \"image_path\": str(image_path),\n            \"label_path\": str(label_path),\n            \"completion_record_path\": str(record_path),\n            \"output_width\": int(expected_size[0]),\n            \"output_height\": int(expected_size[1]),\n            \"reused_image\": reused,\n            \"run_status\": \"reused_image_relabelled\" if reused else \"processed\",\n            \"skipped_completed\": False,\n            \"imported_previous_output\": False,\n            **metadata,\n        }\n        save_completion_record(record, record_path)\n        return record\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 10. 批量预处理完整正式数据集：严格单张串行、周期释放、周期保存进度\n    tasks = [\n        (image_id, split_name)\n        for split_name, image_ids in split_ids.items()\n        for image_id in image_ids\n    ]\n    manifest_records = []\n    failed_records = []\n    start_time = perf_counter()\n    partial_manifest_path = DATASET_ROOT / \"manifest.partial.csv\"\n    partial_failure_path = REPORT_ROOT / f\"preprocess_failures_{dataset_signature}.partial.csv\"\n\n    print(\"Preprocessing uses CPU only; torch/ultralytics have not been imported:\", \"torch\" not in sys.modules)\n\n    for task_index, (image_id, split_name) in enumerate(\n        tqdm(tasks, total=len(tasks), desc=\"Preprocessing (CPU, one image at a time)\"),\n        start=1,\n    ):\n        try:\n            manifest_records.append(process_one(image_id, split_name))\n        except Exception as error:\n            failed_records.append({\n                \"image_id\": image_id,\n                \"split\": split_name,\n                \"error_type\": type(error).__name__,\n                \"error_message\": str(error),\n                \"traceback\": traceback.format_exc(limit=3),\n            })\n        finally:\n            if task_index % PREPROCESS_GC_EVERY == 0:\n                gc.collect()\n            if task_index % PROGRESS_SAVE_EVERY == 0:\n                pd.DataFrame(manifest_records).to_csv(partial_manifest_path, index=False)\n                pd.DataFrame(failed_records).to_csv(partial_failure_path, index=False)\n\n    manifest = pd.DataFrame(manifest_records).sort_values([\"split\", \"image_id\"]).reset_index(drop=True)\n    failures = pd.DataFrame(failed_records)\n    manifest.to_csv(DATASET_ROOT / \"manifest.csv\", index=False)\n    failures.to_csv(REPORT_ROOT / f\"preprocess_failures_{dataset_signature}.csv\", index=False)\n    if partial_manifest_path.exists():\n        partial_manifest_path.unlink()\n    if partial_failure_path.exists():\n        partial_failure_path.unlink()\n\n    elapsed = perf_counter() - start_time\n    skipped_count = int(manifest[\"skipped_completed\"].sum()) if len(manifest) else 0\n    imported_count = int(manifest[\"imported_previous_output\"].sum()) if len(manifest) else 0\n    newly_completed_count = len(manifest) - skipped_count\n    print(\n        f\"Total complete: {len(manifest)} | Newly processed: {newly_completed_count} | \"\n        f\"Skipped/imported complete: {skipped_count} | Imported from previous Output: {imported_count} | \"\n        f\"Failed: {len(failures)} | Seconds: {elapsed:.1f}\"\n    )\n    print(\"Reused existing PNG and refreshed label:\", int(manifest[\"reused_image\"].sum()) if len(manifest) else 0)\n    if len(failures) and FAIL_ON_PREPROCESS_ERROR:\n        raise RuntimeError(\n            f\"有 {len(failures)} 张图像处理失败。请先查看失败报告，不应带错误继续正式训练。\"\n        )\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 11. 标签质量检查、统计与 data.yaml\n    def validate_label_file(path: Path):\n        errors = []\n        for line_number, line in enumerate(path.read_text(encoding=\"utf-8\").splitlines(), start=1):\n            parts = line.split()\n            if len(parts) != 5:\n                errors.append(f\"line {line_number}: expected 5 fields\")\n                continue\n            try:\n                class_id = int(parts[0])\n                values = np.array([float(value) for value in parts[1:]], dtype=float)\n            except Exception:\n                errors.append(f\"line {line_number}: non-numeric field\")\n                continue\n            if class_id != 0:\n                errors.append(f\"line {line_number}: class must be 0\")\n            if not np.isfinite(values).all() or not ((values > 0).all() and (values <= 1).all()):\n                errors.append(f\"line {line_number}: normalized values outside (0, 1]\")\n        return errors\n\n    label_errors = []\n    for row in tqdm(manifest.itertuples(index=False), total=len(manifest), desc=\"Validating labels\"):\n        image_path = Path(row.image_path)\n        label_path = Path(row.label_path)\n        if not image_path.exists() or not label_path.exists():\n            label_errors.append({\"image_id\": row.image_id, \"error\": \"missing image or label\"})\n            continue\n        for error in validate_label_file(label_path):\n            label_errors.append({\"image_id\": row.image_id, \"error\": error})\n\n    if label_errors:\n        pd.DataFrame(label_errors).to_csv(REPORT_ROOT / \"label_validation_errors.csv\", index=False)\n        raise RuntimeError(f\"YOLO 标签检查失败：{len(label_errors)} 个错误。\")\n\n    actual_split_sets = {\n        split_name: set(manifest.loc[manifest[\"split\"].eq(split_name), \"image_id\"])\n        for split_name in (\"train\", \"val\", \"test\")\n    }\n    if any(actual_split_sets[a] & actual_split_sets[b] for a, b in ((\"train\", \"val\"), (\"train\", \"test\"), (\"val\", \"test\"))):\n        raise AssertionError(\"预处理后的数据存在 split 泄漏。\")\n\n    split_statistics = manifest.groupby(\"split\").agg(\n        images=(\"image_id\", \"size\"),\n        positive_images=(\"is_positive\", \"sum\"),\n        boxes=(\"label_count\", \"sum\"),\n        reused_images=(\"reused_image\", \"sum\"),\n        skipped_completed=(\"skipped_completed\", \"sum\"),\n        imported_previous_output=(\"imported_previous_output\", \"sum\"),\n    )\n    split_statistics[\"negative_images\"] = split_statistics[\"images\"] - split_statistics[\"positive_images\"]\n    split_statistics = split_statistics.reset_index()\n    split_statistics.to_csv(REPORT_ROOT / f\"split_statistics_{dataset_signature}.csv\", index=False)\n    display(split_statistics)\n\n    DATA_YAML = DATASET_ROOT / \"data.yaml\"\n    data_config = {\n        \"path\": str(DATASET_ROOT),\n        \"train\": \"images/train\",\n        \"val\": \"images/val\",\n        \"test\": \"images/test\",\n        \"names\": {0: TARGET_CLASS},\n    }\n    DATA_YAML.write_text(yaml.safe_dump(data_config, sort_keys=False), encoding=\"utf-8\")\n    print(DATA_YAML.read_text(encoding=\"utf-8\"))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_PREPROCESSING:\n    # 12. 可视化抽查：红框来自最终 YOLO 标签，不再使用原始 CSV 直接绘制\n    positive_preview = manifest.loc[(manifest[\"split\"] == \"train\") & (manifest[\"is_positive\"] == 1)].sample(\n        n=min(3, int(((manifest[\"split\"] == \"train\") & (manifest[\"is_positive\"] == 1)).sum())),\n        random_state=SEED,\n    )\n    negative_preview = manifest.loc[(manifest[\"split\"] == \"train\") & (manifest[\"is_positive\"] == 0)].sample(\n        n=min(3, int(((manifest[\"split\"] == \"train\") & (manifest[\"is_positive\"] == 0)).sum())),\n        random_state=SEED,\n    )\n    preview = pd.concat([positive_preview, negative_preview], ignore_index=True)\n    figure, axes = plt.subplots(2, 3, figsize=(18, 12))\n    axes = axes.flatten()\n\n    for axis, row in zip(axes, preview.itertuples(index=False)):\n        with Image.open(row.image_path) as source:\n            image = source.convert(\"RGB\")\n        draw = ImageDraw.Draw(image)\n        for line in Path(row.label_path).read_text(encoding=\"utf-8\").splitlines():\n            _, xc, yc, width, height = map(float, line.split())\n            x1 = (xc - width / 2) * image.width\n            y1 = (yc - height / 2) * image.height\n            x2 = (xc + width / 2) * image.width\n            y2 = (yc + height / 2) * image.height\n            draw.rectangle([x1, y1, x2, y2], outline=\"red\", width=max(2, image.width // 500))\n        axis.imshow(image)\n        axis.set_title(f\"{row.split} | positive={row.is_positive} | boxes={row.label_count}\\n{row.image_id[:18]}…\")\n        axis.axis(\"off\")\n    for axis in axes[len(preview):]:\n        axis.axis(\"off\")\n    figure.suptitle(\"Final training labels — Nodule/Mass\", fontsize=18)\n    plt.tight_layout()\n    preview_path = REPORT_ROOT / f\"label_preview_{dataset_signature}.png\"\n    figure.savefig(preview_path, dpi=150, bbox_inches=\"tight\")\n    plt.show()\n    plt.close(figure)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 1280 preprocessing 运行后把真实输出交给 V6；没有运行时保持已发现的 Input\nif NEED_PREPROCESSING:\n    PREPROCESSED_SOURCE_ROOT = Path(DATASET_ROOT)\n    checked_preprocessed = inspect_preprocessed_root(PREPROCESSED_SOURCE_ROOT)\n    if checked_preprocessed is None:\n        raise RuntimeError(\"内置 1280 preprocessing 完成后未形成合格的 21e0ec1e3506 输出。\")\n    PREPROCESSED_SOURCE_ROOT = checked_preprocessed\n    print(\"Integrated preprocessing completed:\", PREPROCESSED_SOURCE_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## B. 内置 V6 自动清洗（1280 数据）\n\n该清洗逻辑与 E0 使用的 V6 发布版保持一致，但输入源签名锁定为 1280 数据版本。\n\n它继续执行全量解码、像素哈希去重、标签几何复投影、导出一致性和最终训练闸门，不会读取 640 图片训练 E1。\n","metadata":{}},{"cell_type":"code","source":"# 为内置 V6 明确指定已验证的 1280 预处理源\nif NEED_CLEANING:\n    if PREPROCESSED_SOURCE_ROOT is None:\n        raise RuntimeError(\"需要运行 V6，但没有合格的 1280 预处理源。\")\n    os.environ[\"VINDR_PREPROCESSED_ROOT\"] = str(PREPROCESSED_SOURCE_ROOT)\n    os.environ[\"VINDR_OUTPUT_ROOT\"] = str(\n        WORKING_ROOT / \"vindr_nodule_mass_cleaning_final_v6_1280\"\n    )\n    os.environ[\"VINDR_INPUT_ROOT\"] = str(INPUT_ROOT)\n    # 内置 working 路径不带 Notebook 名称；安全性由严格配置签名负责。\n    os.environ[\"VINDR_REQUIRE_NOTEBOOK_NAME_HINT\"] = \"0\"\n    print(\"V6 1280 source:\", PREPROCESSED_SOURCE_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 1. 导入：仅 CPU 数据处理库\n    import hashlib\n    import json\n    import math\n    import os\n    import platform\n    import random\n    import shutil\n    from concurrent.futures import ThreadPoolExecutor\n    from datetime import datetime, timezone\n    from pathlib import Path\n    from time import perf_counter\n\n    import matplotlib.pyplot as plt\n    import numpy as np\n    import pandas as pd\n    import yaml\n    from PIL import Image, ImageDraw\n\n    try:\n        from tqdm.auto import tqdm\n    except ImportError:\n        def tqdm(iterable, **_):\n            return iterable\n\n    try:\n        from IPython.display import display\n    except ImportError:\n        display = print\n\n    SEED = 2026\n    random.seed(SEED)\n    np.random.seed(SEED)\n\n    print(\"Python:\", platform.python_version())\n    print(\"CPU cores:\", os.cpu_count())\n    print(\"GPU used by this notebook: False\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 2. 发布配置：锁定到 E1 的 1280 预处理输出\n    INPUT_ROOT = Path(os.environ.get(\"VINDR_INPUT_ROOT\", \"/kaggle/input\"))\n    OUTPUT_BASE_ROOT = Path(\n        os.environ.get(\n            \"VINDR_OUTPUT_ROOT\",\n            \"/kaggle/working/vindr_nodule_mass_cleaning_final_v6_1280\",\n        )\n    )\n\n    PREPROCESSED_DATASET_ROOT = Path(\n        os.environ.get(\n            \"VINDR_PREPROCESSED_ROOT\",\n            (\n                \"/kaggle/input/notebooks/hilarylee33/1280version1/\"\n                \"vindr_nodule_mass_training_1280/datasets/21e0ec1e3506\"\n            ),\n        )\n    )\n    EXPECTED_SOURCE_NOTEBOOK_ROOT = Path(\n        os.environ.get(\n            \"VINDR_1280VERSION1_ROOT\",\n            \"/kaggle/input/notebooks/hilarylee33/1280version1\",\n        )\n    )\n\n    EXPECTED_SOURCE_OUTPUT_DIRNAME = \"vindr_nodule_mass_training_1280\"\n    EXPECTED_SOURCE_NOTEBOOK_HINT = \"1280version1\"\n    EXPECTED_SOURCE_DATASET_SIGNATURE = \"21e0ec1e3506\"\n    REQUIRE_NOTEBOOK_NAME_HINT = os.environ.get('VINDR_REQUIRE_NOTEBOOK_NAME_HINT', '0') == '1'\n\n    # 这四项把 V6 锁定到本次已经审计过的完整源版本，防止误挂子集或旧输出。\n    EXPECTED_SOURCE_IMAGE_COUNT = 15000\n    EXPECTED_SOURCE_SPLIT_COUNTS = {\n        \"train\": 12000,\n        \"val\": 1500,\n        \"test\": 1500,\n    }\n    EXPECTED_SOURCE_POSITIVE_IMAGE_COUNT = 826\n    EXPECTED_SOURCE_FUSED_BOX_COUNT = 1803\n\n    TARGET_CLASS = \"Nodule/Mass\"\n    YOLO_CLASS_ID = 0\n    EXPECTED_TRAIN_RADIOLOGISTS_PER_IMAGE = 3\n\n    # 几何清洗。框完全消失或过小会被隔离，阳性图绝不会被伪装成阴性。\n    MIN_VISIBLE_FRACTION = 0.50\n    MIN_BOX_SIDE_ORIGINAL_PX = 2.0\n    MIN_BOX_SIDE_OUTPUT_PX = 1.0\n    ALGEBRAIC_GEOMETRY_TOLERANCE_PX = 1e-6\n    LABEL_SERIALIZATION_DECIMALS = 8\n    SERIALIZED_GEOMETRY_TOLERANCE_OUTPUT_PX = 1e-4\n    MAX_AMBIGUOUS_POSITIVE_FRACTION = 0.05\n\n    # 可安全自动隔离的源图片问题及其最大允许比例。\n    AUTO_EXCLUDE_UNREADABLE_OR_INVALID_IMAGES = True\n    MAX_AUTOMATIC_IMAGE_EXCLUSION_FRACTION = 0.005\n    MAX_AUTOMATIC_POSITIVE_EXCLUSION_FRACTION = 0.01\n\n    # copy：正式保存/共享；symlink：仅当前会话；none：只写报告。\n    IMAGE_EXPORT_MODE = \"copy\"\n\n    # 正式训练必须 full。sample 只能开发调试。\n    AUDIT_MODE = \"full\"\n    SAMPLE_IMAGES_PER_SPLIT = 100\n    AUDIT_WORKERS = min(8, max(1, os.cpu_count() or 1))\n    REQUIRE_GRAYSCALE = True\n    COMPUTE_SHA256 = True\n    VERIFY_EXISTING_COPY_HASH = True\n    LOW_CONTRAST_STD_REVIEW_THRESHOLD = 3.0\n    EXTREME_SATURATION_REVIEW_FRACTION = 0.98\n    REPORT_PERCEPTUAL_HASH_COLLISIONS = True\n\n    DISK_SPACE_SAFETY_FACTOR = 1.10\n    DISK_SPACE_RESERVE_GB = 0.50\n\n    FAIL_ON_NOT_TRAINING_READY = True\n    PREVIEW_POSITIVE_IMAGES = 6\n\n    VALID_SPLITS = (\"train\", \"val\", \"test\")\n    AUTO_DEDUPLICATE_EXACT_CONTENT = True\n    DUPLICATE_KEEP_PRIORITY = (\"test\", \"val\", \"train\")\n    BLOCK_ON_DUPLICATE_LABEL_CONFLICT = True\n\n    COORDINATE_COLUMNS = [\"x_min\", \"y_min\", \"x_max\", \"y_max\"]\n    REQUIRED_MANIFEST_COLUMNS = {\n        \"image_id\", \"split\", \"is_positive\", \"label_count\",\n        \"original_width\", \"original_height\", \"output_width\", \"output_height\",\n    }\n    REQUIRED_FUSED_COLUMNS = {\"image_id\", *COORDINATE_COLUMNS}\n\n    if IMAGE_EXPORT_MODE not in {\"copy\", \"symlink\", \"none\"}:\n        raise ValueError(\"IMAGE_EXPORT_MODE 必须是 copy、symlink 或 none。\")\n    if AUDIT_MODE not in {\"full\", \"sample\"}:\n        raise ValueError(\"AUDIT_MODE 必须是 full 或 sample。\")\n    if not 0 < MIN_VISIBLE_FRACTION <= 1:\n        raise ValueError(\"MIN_VISIBLE_FRACTION 必须在 (0, 1]。\")\n    if MIN_BOX_SIDE_ORIGINAL_PX <= 0 or MIN_BOX_SIDE_OUTPUT_PX <= 0:\n        raise ValueError(\"最小框边长必须大于 0。\")\n    if LABEL_SERIALIZATION_DECIMALS < 6:\n        raise ValueError(\"YOLO 标签至少保留 6 位小数。\")\n    if SERIALIZED_GEOMETRY_TOLERANCE_OUTPUT_PX <= 0:\n        raise ValueError(\"序列化几何误差阈值必须大于 0。\")\n    if not 0 <= MAX_AMBIGUOUS_POSITIVE_FRACTION <= 1:\n        raise ValueError(\"MAX_AMBIGUOUS_POSITIVE_FRACTION 必须在 [0, 1]。\")\n    if not 0 <= MAX_AUTOMATIC_IMAGE_EXCLUSION_FRACTION <= 1:\n        raise ValueError(\"自动图片隔离比例阈值必须在 [0, 1]。\")\n    if not 0 <= MAX_AUTOMATIC_POSITIVE_EXCLUSION_FRACTION <= 1:\n        raise ValueError(\"自动阳性图片隔离比例阈值必须在 [0, 1]。\")\n    if AUDIT_WORKERS <= 0:\n        raise ValueError(\"AUDIT_WORKERS 必须大于 0。\")\n    if not AUTO_DEDUPLICATE_EXACT_CONTENT:\n        raise ValueError(\"V6 发布版要求启用自动精确去重。\")\n    if set(DUPLICATE_KEEP_PRIORITY) != set(VALID_SPLITS):\n        raise ValueError(\"DUPLICATE_KEEP_PRIORITY 必须且只能包含三份 split。\")\n    if not BLOCK_ON_DUPLICATE_LABEL_CONFLICT:\n        raise ValueError(\"V6 发布版要求标签冲突严格阻断训练。\")\n    if not COMPUTE_SHA256:\n        raise ValueError(\"V6 依赖 SHA-256，COMPUTE_SHA256 必须为 True。\")\n\n    CLEANING_POLICY = {\n        \"version\": \"6.0-release\",\n        \"target_class\": TARGET_CLASS,\n        \"yolo_class_id\": YOLO_CLASS_ID,\n        \"image_export_mode\": IMAGE_EXPORT_MODE,\n        \"min_visible_fraction\": MIN_VISIBLE_FRACTION,\n        \"min_box_side_original_px\": MIN_BOX_SIDE_ORIGINAL_PX,\n        \"min_box_side_output_px\": MIN_BOX_SIDE_OUTPUT_PX,\n        \"max_ambiguous_positive_fraction\": (\n            MAX_AMBIGUOUS_POSITIVE_FRACTION\n        ),\n        \"label_serialization_decimals\": LABEL_SERIALIZATION_DECIMALS,\n        \"serialized_geometry_tolerance_output_px\": (\n            SERIALIZED_GEOMETRY_TOLERANCE_OUTPUT_PX\n        ),\n        \"auto_exclude_invalid_images\": AUTO_EXCLUDE_UNREADABLE_OR_INVALID_IMAGES,\n        \"max_auto_image_exclusion_fraction\": (\n            MAX_AUTOMATIC_IMAGE_EXCLUSION_FRACTION\n        ),\n        \"max_auto_positive_exclusion_fraction\": (\n            MAX_AUTOMATIC_POSITIVE_EXCLUSION_FRACTION\n        ),\n        \"duplicate_hash\": \"decoded_grayscale_pixels_sha256\",\n        \"duplicate_keep_priority\": list(DUPLICATE_KEEP_PRIORITY),\n        \"block_on_duplicate_label_conflict\": (\n            BLOCK_ON_DUPLICATE_LABEL_CONFLICT\n        ),\n        \"require_grayscale\": REQUIRE_GRAYSCALE,\n        \"audit_mode\": AUDIT_MODE,\n        \"seed\": SEED,\n    }\n    CLEANING_POLICY_SIGNATURE = hashlib.sha256(\n        json.dumps(\n            CLEANING_POLICY, sort_keys=True, ensure_ascii=False\n        ).encode(\"utf-8\")\n    ).hexdigest()[:12]\n\n    OUTPUT_BASE_ROOT.mkdir(parents=True, exist_ok=True)\n\n    print(\"Input root:\", INPUT_ROOT)\n    print(\"Output base root:\", OUTPUT_BASE_ROOT)\n    print(\"Image export mode:\", IMAGE_EXPORT_MODE)\n    print(\"Audit mode:\", AUDIT_MODE)\n    print(\"Audit workers:\", AUDIT_WORKERS)\n    print(\"Cleaning policy signature:\", CLEANING_POLICY_SIGNATURE)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    def source_output_root_for(dataset_root: Path):\n        for parent in [dataset_root, *dataset_root.parents]:\n            if parent.name == EXPECTED_SOURCE_OUTPUT_DIRNAME:\n                return parent\n        return None\n\n\n    def candidate_reason(dataset_root: Path):\n        required_paths = [\n            dataset_root / \"manifest.csv\",\n            dataset_root / \"dataset_config.json\",\n            dataset_root / \"data.yaml\",\n            *[dataset_root / \"images\" / split for split in VALID_SPLITS],\n            *[dataset_root / \"labels\" / split for split in VALID_SPLITS],\n        ]\n        missing = [str(path) for path in required_paths if not path.exists()]\n        if (\n            REQUIRE_NOTEBOOK_NAME_HINT\n            and EXPECTED_SOURCE_NOTEBOOK_HINT.lower() not in str(dataset_root).lower()\n        ):\n            missing.append(\n                f\"path containing notebook hint {EXPECTED_SOURCE_NOTEBOOK_HINT!r}\"\n            )\n        output_root = source_output_root_for(dataset_root)\n        if output_root is None:\n            missing.append(\n                f\"ancestor directory named {EXPECTED_SOURCE_OUTPUT_DIRNAME}\"\n            )\n        else:\n            for report_name in (\n                \"fused_boxes_original_coordinates.csv\",\n                \"split_plan.csv\",\n            ):\n                report_path = output_root / \"reports\" / report_name\n                if not report_path.exists():\n                    missing.append(str(report_path))\n        return missing\n\n\n    if PREPROCESSED_DATASET_ROOT is not None:\n        source_dataset_root = Path(PREPROCESSED_DATASET_ROOT)\n        if source_dataset_root.name == \"manifest.csv\":\n            source_dataset_root = source_dataset_root.parent\n        missing_source_parts = candidate_reason(source_dataset_root)\n        if missing_source_parts:\n            raise FileNotFoundError(\n                \"指定目录不是完整的 1280 preprocessing 预处理输出，缺少：\\n- \"\n                + \"\\n- \".join(missing_source_parts)\n            )\n    else:\n        # 仅在明确的 1280 preprocessing Notebook Output 内查找，禁止扫描整个 Kaggle Input。\n        if not EXPECTED_SOURCE_NOTEBOOK_ROOT.is_dir():\n            raise FileNotFoundError(\n                \"未找到 1280 preprocessing Output：\"\n                f\"{EXPECTED_SOURCE_NOTEBOOK_ROOT}\"\n            )\n        candidate_roots = []\n        for manifest_path in EXPECTED_SOURCE_NOTEBOOK_ROOT.rglob(\"manifest.csv\"):\n            dataset_root = manifest_path.parent\n            if not candidate_reason(dataset_root):\n                candidate_roots.append(dataset_root)\n        candidate_roots = sorted(set(candidate_roots), key=lambda path: str(path))\n        if not candidate_roots:\n            raise FileNotFoundError(\n                \"没有找到完整的 1280 preprocessing 预处理输出。请先对 1280 preprocessing Run All、\"\n                \"Save Version，再把该 Notebook Output 添加为当前 Notebook 的 Input。\"\n            )\n        if len(candidate_roots) > 1:\n            raise RuntimeError(\n                \"发现多个合格输出，为防止误读没有自动选择。\"\n                \"请填写 PREPROCESSED_DATASET_ROOT：\\n- \"\n                + \"\\n- \".join(str(path) for path in candidate_roots)\n            )\n        source_dataset_root = candidate_roots[0]\n\n    source_output_root = source_output_root_for(source_dataset_root)\n    source_manifest_path = source_dataset_root / \"manifest.csv\"\n    source_config_path = source_dataset_root / \"dataset_config.json\"\n    source_data_yaml_path = source_dataset_root / \"data.yaml\"\n    source_image_root = source_dataset_root / \"images\"\n    source_label_root = source_dataset_root / \"labels\"\n    source_report_root = source_output_root / \"reports\"\n    source_fused_boxes_path = (\n        source_report_root / \"fused_boxes_original_coordinates.csv\"\n    )\n    source_split_plan_path = source_report_root / \"split_plan.csv\"\n\n    source_config = json.loads(source_config_path.read_text(encoding=\"utf-8\"))\n    required_config_keys = {\"target_class\", \"preprocess_max_side\", \"seed\"}\n    missing_config_keys = required_config_keys - set(source_config)\n    if missing_config_keys:\n        raise RuntimeError(\n            f\"dataset_config.json 缺少关键字段：{sorted(missing_config_keys)}\"\n        )\n\n    config_text = json.dumps(source_config, sort_keys=True, ensure_ascii=False)\n    computed_signature = hashlib.sha256(config_text.encode(\"utf-8\")).hexdigest()[:12]\n    signature_matches = source_dataset_root.name == computed_signature\n    source_notebook_hint_matches = (\n        EXPECTED_SOURCE_NOTEBOOK_HINT.lower() in str(source_dataset_root).lower()\n    )\n    if computed_signature != EXPECTED_SOURCE_DATASET_SIGNATURE:\n        raise RuntimeError(\n            \"源配置签名不是本次发布版锁定值：\"\n            f\"{computed_signature} != {EXPECTED_SOURCE_DATASET_SIGNATURE}。\"\n        )\n    if not signature_matches:\n        raise RuntimeError(\n            \"预处理配置签名不匹配：目录名为 \"\n            f\"{source_dataset_root.name}，计算结果为 {computed_signature}。\"\n        )\n    if REQUIRE_NOTEBOOK_NAME_HINT and not source_notebook_hint_matches:\n        raise RuntimeError(\n            f\"路径不包含 {EXPECTED_SOURCE_NOTEBOOK_HINT!r}，禁止自动读取。\"\n        )\n    if source_config.get(\"target_class\") != TARGET_CLASS:\n        raise RuntimeError(\n            \"预处理目标类别不匹配：\"\n            f\"{source_config.get('target_class')!r} != {TARGET_CLASS!r}\"\n        )\n    source_preprocess_max_side = source_config.get(\"preprocess_max_side\")\n    if int(source_preprocess_max_side) != 1280:\n        raise RuntimeError(\n            \"E1 只允许直接从原始 DICOM 生成的 1280 数据：\"\n            f\"preprocess_max_side={source_preprocess_max_side!r}。\"\n        )\n\n    # 源配置与清洗策略共同隔离输出，避免旧版本或不同阈值残留混入。\n    RUN_ROOT = (\n        OUTPUT_BASE_ROOT / computed_signature\n        / f\"policy-{CLEANING_POLICY_SIGNATURE}\"\n    )\n    REPORT_ROOT = RUN_ROOT / \"reports\"\n    DATASET_ROOT = RUN_ROOT / \"dataset\"\n    NEW_IMAGE_ROOT = DATASET_ROOT / \"images\"\n    NEW_LABEL_ROOT = DATASET_ROOT / \"labels\"\n    for directory in (RUN_ROOT, REPORT_ROOT, DATASET_ROOT, NEW_LABEL_ROOT):\n        directory.mkdir(parents=True, exist_ok=True)\n    if IMAGE_EXPORT_MODE != \"none\":\n        NEW_IMAGE_ROOT.mkdir(parents=True, exist_ok=True)\n\n    # Fail closed：当前运行完成前，旧的成功闸门绝不能继续显示为可训练。\n    RUN_STARTED_AT = datetime.now(timezone.utc).isoformat()\n    (REPORT_ROOT / \"TRAINING_GATE.txt\").write_text(\n        \"STATUS=RUNNING\\nTRAINING_READY=False\\n\"\n        f\"RUN_STARTED_AT={RUN_STARTED_AT}\\n\",\n        encoding=\"utf-8\",\n    )\n    (REPORT_ROOT / \"cleaning_acceptance_final.json\").write_text(\n        json.dumps({\n            \"overall_status\": \"RUNNING\",\n            \"training_ready\": False,\n            \"run_started_at\": RUN_STARTED_AT,\n            \"source_dataset_signature\": computed_signature,\n            \"cleaning_policy_signature\": CLEANING_POLICY_SIGNATURE,\n        }, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    source_provenance = {\n        \"input_policy\": \"1280 preprocessing_preprocessed_output_only\",\n        \"raw_dicom_read\": False,\n        \"raw_train_csv_read\": False,\n        \"source_dataset_root\": str(source_dataset_root),\n        \"source_output_root\": str(source_output_root),\n        \"source_manifest\": str(source_manifest_path),\n        \"source_dataset_config\": str(source_config_path),\n        \"source_data_yaml\": str(source_data_yaml_path),\n        \"source_fused_boxes\": str(source_fused_boxes_path),\n        \"source_split_plan\": str(source_split_plan_path),\n        \"expected_source_notebook_hint\": EXPECTED_SOURCE_NOTEBOOK_HINT,\n        \"dataset_signature\": computed_signature,\n        \"expected_dataset_signature\": EXPECTED_SOURCE_DATASET_SIGNATURE,\n        \"cleaning_policy\": CLEANING_POLICY,\n        \"cleaning_policy_signature\": CLEANING_POLICY_SIGNATURE,\n        \"run_started_at\": RUN_STARTED_AT,\n        \"dataset_signature_verified\": signature_matches,\n        \"source_notebook_hint_verified\": source_notebook_hint_matches,\n        \"source_configuration\": source_config,\n        \"v6_run_root\": str(RUN_ROOT),\n    }\n    (REPORT_ROOT / \"source_provenance.json\").write_text(\n        json.dumps(source_provenance, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    print(\"SOURCE LOCK PASSED\")\n    print(\"Only preprocessed output will be read:\", source_dataset_root)\n    print(\"Raw DICOM read: False\")\n    print(\"Raw train.csv read: False\")\n    print(\"Verified dataset signature:\", computed_signature)\n    print(\"V6 run root:\", RUN_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 4. 验证 manifest、几何变换契约和预处理文件结构\n    manifest = pd.read_csv(source_manifest_path)\n    missing_manifest_columns = REQUIRED_MANIFEST_COLUMNS - set(manifest.columns)\n    if missing_manifest_columns:\n        raise ValueError(\n            f\"manifest 缺少 1280_preprocessing 必需字段：{sorted(missing_manifest_columns)}\"\n        )\n\n    manifest[\"image_id\"] = manifest[\"image_id\"].astype(str).str.strip()\n    manifest[\"split\"] = manifest[\"split\"].astype(str).str.strip().str.lower()\n    invalid_manifest_image_ids = manifest.loc[\n        manifest[\"image_id\"].str.lower().isin([\"\", \"nan\", \"none\", \"null\"])\n    ].copy()\n    numeric_manifest_columns = [\n        \"is_positive\", \"label_count\",\n        \"original_width\", \"original_height\", \"output_width\", \"output_height\",\n    ]\n    for column in numeric_manifest_columns:\n        manifest[column] = pd.to_numeric(manifest[column], errors=\"coerce\")\n\n    duplicate_manifest_rows = manifest.loc[\n        manifest.duplicated(\"image_id\", keep=False)\n    ].copy()\n    invalid_split_rows = manifest.loc[\n        ~manifest[\"split\"].isin(VALID_SPLITS)\n    ].copy()\n    invalid_dimension_rows = manifest.loc[\n        ~np.isfinite(\n            manifest[\n                [\"original_width\", \"original_height\", \"output_width\", \"output_height\"]\n            ]\n        ).all(axis=1)\n        | (manifest[\"original_width\"] <= 0)\n        | (manifest[\"original_height\"] <= 0)\n        | (manifest[\"output_width\"] <= 0)\n        | (manifest[\"output_height\"] <= 0)\n        | (\n            manifest[\n                [\"original_width\", \"original_height\",\n                 \"output_width\", \"output_height\"]\n            ]\n            != np.floor(manifest[\n                [\"original_width\", \"original_height\",\n                 \"output_width\", \"output_height\"]\n            ])\n        ).any(axis=1)\n    ].copy()\n    invalid_status_rows = manifest.loc[\n        ~np.isfinite(manifest[[\"is_positive\", \"label_count\"]]).all(axis=1)\n        | ~manifest[\"is_positive\"].isin([0, 1])\n        | (manifest[\"label_count\"] < 0)\n        | (manifest[\"label_count\"] != np.floor(manifest[\"label_count\"]))\n    ].copy()\n\n    invalid_manifest_image_ids.to_csv(\n        REPORT_ROOT / \"invalid_manifest_image_ids.csv\", index=False\n    )\n    duplicate_manifest_rows.to_csv(\n        REPORT_ROOT / \"duplicate_manifest_rows.csv\", index=False\n    )\n    invalid_split_rows.to_csv(REPORT_ROOT / \"invalid_split_rows.csv\", index=False)\n    invalid_dimension_rows.to_csv(\n        REPORT_ROOT / \"invalid_manifest_dimensions.csv\", index=False\n    )\n    invalid_status_rows.to_csv(\n        REPORT_ROOT / \"invalid_manifest_status.csv\", index=False\n    )\n    if len(invalid_manifest_image_ids):\n        raise RuntimeError(\"预处理 manifest 含空或非法 image_id。\")\n    if len(duplicate_manifest_rows):\n        raise RuntimeError(\"预处理 manifest 含重复 image_id，禁止继续。\")\n    if len(invalid_split_rows):\n        raise RuntimeError(\"预处理 manifest 含非法 split，禁止继续。\")\n    if len(invalid_dimension_rows):\n        raise RuntimeError(\"预处理 manifest 含非法尺寸，禁止继续。\")\n    if len(invalid_status_rows):\n        raise RuntimeError(\"预处理 manifest 含非法阳性状态或框数量，禁止继续。\")\n    if set(manifest[\"split\"]) != set(VALID_SPLITS):\n        raise RuntimeError(\"manifest 必须同时包含 train、val、test。\")\n\n    actual_source_split_counts = {\n        split: int(count)\n        for split, count in manifest.groupby(\"split\").size().items()\n    }\n    actual_source_positive_count = int(manifest[\"is_positive\"].sum())\n    source_count_contract = {\n        \"expected_image_count\": EXPECTED_SOURCE_IMAGE_COUNT,\n        \"actual_image_count\": int(len(manifest)),\n        \"expected_split_counts\": EXPECTED_SOURCE_SPLIT_COUNTS,\n        \"actual_split_counts\": actual_source_split_counts,\n        \"expected_positive_image_count\": EXPECTED_SOURCE_POSITIVE_IMAGE_COUNT,\n        \"actual_positive_image_count\": actual_source_positive_count,\n    }\n    (REPORT_ROOT / \"source_count_contract.json\").write_text(\n        json.dumps(source_count_contract, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n    source_count_contract_pass = (\n        len(manifest) == EXPECTED_SOURCE_IMAGE_COUNT\n        and actual_source_split_counts == EXPECTED_SOURCE_SPLIT_COUNTS\n        and actual_source_positive_count\n        == EXPECTED_SOURCE_POSITIVE_IMAGE_COUNT\n    )\n    if not source_count_contract_pass:\n        raise RuntimeError(\n            \"源数据数量不符合锁定版本；可能挂载了子集、旧输出或错误版本。\"\n        )\n\n\n    def expected_output_size_contract(original_width, original_height):\n        original_width = int(original_width)\n        original_height = int(original_height)\n        longest = max(original_width, original_height)\n        if (\n            source_preprocess_max_side is None\n            or longest <= float(source_preprocess_max_side)\n        ):\n            return original_width, original_height\n        scale = float(source_preprocess_max_side) / longest\n        return (\n            max(1, round(original_width * scale)),\n            max(1, round(original_height * scale)),\n        )\n\n\n    geometry_contract = manifest[\n        [\n            \"image_id\", \"split\", \"original_width\", \"original_height\",\n            \"output_width\", \"output_height\",\n        ]\n    ].copy()\n    expected_sizes = geometry_contract.apply(\n        lambda row: expected_output_size_contract(\n            row[\"original_width\"], row[\"original_height\"]\n        ),\n        axis=1,\n    )\n    geometry_contract[\"expected_output_width\"] = [\n        size[0] for size in expected_sizes\n    ]\n    geometry_contract[\"expected_output_height\"] = [\n        size[1] for size in expected_sizes\n    ]\n    geometry_contract[\"dimension_contract_matches\"] = (\n        geometry_contract[\"output_width\"].astype(int)\n        == geometry_contract[\"expected_output_width\"].astype(int)\n    ) & (\n        geometry_contract[\"output_height\"].astype(int)\n        == geometry_contract[\"expected_output_height\"].astype(int)\n    )\n    geometry_contract[\"scale_x\"] = (\n        geometry_contract[\"output_width\"] / geometry_contract[\"original_width\"]\n    )\n    geometry_contract[\"scale_y\"] = (\n        geometry_contract[\"output_height\"] / geometry_contract[\"original_height\"]\n    )\n    geometry_contract[\"aspect_ratio_relative_error\"] = (\n        (\n            (geometry_contract[\"output_width\"] / geometry_contract[\"output_height\"])\n            / (\n                geometry_contract[\"original_width\"]\n                / geometry_contract[\"original_height\"]\n            )\n        )\n        - 1\n    ).abs()\n    geometry_contract.to_csv(\n        REPORT_ROOT / \"geometry_manifest_contract.csv\", index=False\n    )\n    geometry_manifest_failures = geometry_contract.loc[\n        ~geometry_contract[\"dimension_contract_matches\"]\n    ].copy()\n    geometry_manifest_failures.to_csv(\n        REPORT_ROOT / \"geometry_manifest_contract_failures.csv\", index=False\n    )\n    geometry_manifest_contract_pass = len(geometry_manifest_failures) == 0\n    if not geometry_manifest_contract_pass:\n        raise RuntimeError(\n            \"manifest 输出尺寸不符合 1280_preprocessing 的等比例缩放契约；\"\n            \"可能存在裁剪、补边、路径混用或 manifest 错配。\"\n        )\n\n    resolved_records = []\n    for row in manifest.itertuples(index=False):\n        image_path = source_image_root / row.split / f\"{row.image_id}.png\"\n        label_path = source_label_root / row.split / f\"{row.image_id}.txt\"\n        resolved_records.append({\n            \"image_id\": row.image_id,\n            \"split\": row.split,\n            \"source_image_path\": str(image_path),\n            \"source_label_path\": str(label_path),\n            \"source_image_exists\": image_path.is_file(),\n            \"source_label_exists\": label_path.is_file(),\n        })\n    resolved_paths = pd.DataFrame(resolved_records)\n    manifest = manifest.merge(\n        resolved_paths,\n        on=[\"image_id\", \"split\"],\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n\n    # 明确生成受来源锁约束的图片清单；后续绝不遍历其他 Input 图片目录。\n    SOURCE_IMAGES = tuple(\n        Path(path) for path in manifest[\"source_image_path\"].tolist()\n    )\n    for image_path in SOURCE_IMAGES:\n        if source_image_root not in image_path.parents:\n            raise RuntimeError(\n                f\"来源锁违规：图片不在 1280_preprocessing images 目录内：{image_path}\"\n            )\n\n    missing_source_files = manifest.loc[\n        ~manifest[\"source_image_exists\"] | ~manifest[\"source_label_exists\"]\n    ].copy()\n    missing_source_files.to_csv(\n        REPORT_ROOT / \"missing_preprocessed_source_files.csv\", index=False\n    )\n    if len(missing_source_files):\n        raise FileNotFoundError(\n            f\"预处理输出缺少 {len(missing_source_files)} 组 PNG/标签文件。\"\n        )\n\n    expected_image_keys = set(zip(manifest[\"split\"], manifest[\"image_id\"]))\n    actual_image_keys = {\n        (path.parent.name, path.stem)\n        for split in VALID_SPLITS\n        for path in (source_image_root / split).glob(\"*.png\")\n    }\n    actual_label_keys = {\n        (path.parent.name, path.stem)\n        for split in VALID_SPLITS\n        for path in (source_label_root / split).glob(\"*.txt\")\n    }\n    extra_preprocessed_images = sorted(actual_image_keys - expected_image_keys)\n    extra_preprocessed_labels = sorted(actual_label_keys - expected_image_keys)\n    pd.DataFrame(\n        extra_preprocessed_images, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"extra_preprocessed_images.csv\", index=False)\n    pd.DataFrame(\n        extra_preprocessed_labels, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"extra_preprocessed_labels.csv\", index=False)\n\n    unexpected_source_files = []\n    for split in VALID_SPLITS:\n        unexpected_source_files.extend(\n            str(path)\n            for path in (source_image_root / split).iterdir()\n            if path.is_file() and path.suffix.lower() != \".png\"\n        )\n        unexpected_source_files.extend(\n            str(path)\n            for path in (source_label_root / split).iterdir()\n            if path.is_file() and path.suffix.lower() != \".txt\"\n        )\n    pd.DataFrame({\"path\": unexpected_source_files}).to_csv(\n        REPORT_ROOT / \"unexpected_source_files.csv\", index=False\n    )\n\n    dicom_inside_selected_dataset = [\n        path\n        for path in source_dataset_root.rglob(\"*\")\n        if path.is_file() and path.suffix.lower() in {\".dcm\", \".dicom\"}\n    ]\n    if dicom_inside_selected_dataset:\n        raise RuntimeError(\n            \"选中的数据集内部发现 DICOM，说明它不是纯预处理输出：\"\n            f\"{dicom_inside_selected_dataset[0]}\"\n        )\n\n    source_yaml = yaml.safe_load(\n        source_data_yaml_path.read_text(encoding=\"utf-8\")\n    )\n    source_yaml_names = source_yaml.get(\"names\", {})\n    if isinstance(source_yaml_names, list):\n        source_yaml_class_name = (\n            source_yaml_names[YOLO_CLASS_ID]\n            if len(source_yaml_names) > YOLO_CLASS_ID else None\n        )\n    else:\n        source_yaml_class_name = source_yaml_names.get(\n            YOLO_CLASS_ID, source_yaml_names.get(str(YOLO_CLASS_ID))\n        )\n    source_yaml_class_matches = source_yaml_class_name == TARGET_CLASS\n    if not source_yaml_class_matches:\n        raise RuntimeError(\n            \"源 data.yaml 类别与 Nodule/Mass 锁定目标不一致。\"\n        )\n\n    source_baseline_hash_column = next(\n        (\n            column\n            for column in (\"image_sha256\", \"sha256\", \"png_sha256\")\n            if column in manifest.columns\n        ),\n        None,\n    )\n\n    print(\"Manifest images:\", len(manifest))\n    print(\"Source count contract:\", source_count_contract_pass)\n    print(\"Source data.yaml class:\", source_yaml_class_name)\n    print(manifest.groupby(\"split\").size())\n    print(\"Geometry manifest contract:\", geometry_manifest_contract_pass)\n    print(\"Missing PNG/label pairs:\", len(missing_source_files))\n    print(\"Source baseline hash column:\", source_baseline_hash_column)\n    print(\"DICOM files read:\", 0)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    source_fused_boxes = pd.read_csv(source_fused_boxes_path)\n    missing_fused_columns = REQUIRED_FUSED_COLUMNS - set(source_fused_boxes.columns)\n    if missing_fused_columns:\n        raise ValueError(\n            f\"预处理融合框报告缺少字段：{sorted(missing_fused_columns)}\"\n        )\n\n    source_fused_boxes[\"image_id\"] = (\n        source_fused_boxes[\"image_id\"].astype(str).str.strip()\n    )\n    invalid_fused_image_ids = source_fused_boxes.loc[\n        source_fused_boxes[\"image_id\"].str.lower().isin(\n            [\"\", \"nan\", \"none\", \"null\"]\n        )\n    ].copy()\n    invalid_fused_image_ids.to_csv(\n        REPORT_ROOT / \"invalid_fused_image_ids.csv\", index=False\n    )\n    if len(invalid_fused_image_ids):\n        raise RuntimeError(\"融合框报告含空或非法 image_id。\")\n    if len(source_fused_boxes) != EXPECTED_SOURCE_FUSED_BOX_COUNT:\n        raise RuntimeError(\n            \"源融合框数量不是锁定版本的 \"\n            f\"{EXPECTED_SOURCE_FUSED_BOX_COUNT}，实际为 \"\n            f\"{len(source_fused_boxes)}。\"\n        )\n    source_fused_boxes[\"_source_row\"] = source_fused_boxes.index.astype(int)\n    source_fused_boxes[COORDINATE_COLUMNS] = (\n        source_fused_boxes[COORDINATE_COLUMNS].apply(\n            pd.to_numeric, errors=\"coerce\"\n        )\n    )\n    source_positive_ids = set(source_fused_boxes[\"image_id\"])\n    manifest_ids = set(manifest[\"image_id\"])\n    source_fused_count_by_image = (\n        source_fused_boxes.groupby(\"image_id\").size()\n        .rename(\"source_fused_box_count\")\n    )\n    source_manifest_box_contract = manifest[\n        [\"image_id\", \"is_positive\", \"label_count\"]\n    ].merge(\n        source_fused_count_by_image,\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n    source_manifest_box_contract[\"source_fused_box_count\"] = (\n        source_manifest_box_contract[\"source_fused_box_count\"]\n        .fillna(0).astype(int)\n    )\n    source_manifest_box_contract[\"count_matches\"] = (\n        source_manifest_box_contract[\"label_count\"].astype(int)\n        == source_manifest_box_contract[\"source_fused_box_count\"]\n    )\n    source_manifest_box_contract[\"status_matches\"] = (\n        source_manifest_box_contract[\"is_positive\"].astype(int)\n        == (\n            source_manifest_box_contract[\"source_fused_box_count\"] > 0\n        ).astype(int)\n    )\n    source_manifest_box_contract_failures = (\n        source_manifest_box_contract.loc[\n            ~source_manifest_box_contract[\"count_matches\"]\n            | ~source_manifest_box_contract[\"status_matches\"]\n        ].copy()\n    )\n    source_manifest_box_contract.to_csv(\n        REPORT_ROOT / \"source_manifest_box_contract.csv\", index=False\n    )\n    source_manifest_box_contract_failures.to_csv(\n        REPORT_ROOT / \"source_manifest_box_contract_failures.csv\",\n        index=False,\n    )\n    source_manifest_box_contract_pass = (\n        len(source_manifest_box_contract_failures) == 0\n    )\n\n    box_table = source_fused_boxes.merge(\n        manifest[\n            [\n                \"image_id\", \"original_width\", \"original_height\",\n                \"output_width\", \"output_height\",\n            ]\n        ],\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"many_to_one\",\n    )\n    for coordinate in COORDINATE_COLUMNS:\n        box_table[f\"raw_{coordinate}\"] = box_table[coordinate]\n\n    issue_records = []\n\n\n    def record_issues(frame, mask, issue_type, action, detail):\n        if not bool(mask.any()):\n            return\n        preferred = [\n            \"_source_row\", \"image_id\",\n            *[f\"raw_{column}\" for column in COORDINATE_COLUMNS],\n            *COORDINATE_COLUMNS,\n            \"original_width\", \"original_height\", \"output_width\", \"output_height\",\n            \"visible_fraction\", \"output_box_width\", \"output_box_height\",\n            \"radiologist_count\", \"source_box_count\", \"radiologist_ids\",\n        ]\n        columns = [column for column in preferred if column in frame.columns]\n        issue = frame.loc[mask, columns].copy()\n        issue[\"issue_type\"] = issue_type\n        issue[\"action\"] = action\n        issue[\"issue_detail\"] = detail\n        issue_records.append(issue)\n\n\n    numeric_ok = np.isfinite(box_table[COORDINATE_COLUMNS]).all(axis=1)\n    record_issues(\n        box_table, ~numeric_ok, \"non_finite_coordinate\", \"dropped\",\n        \"At least one fused coordinate is NaN, inf, or non-numeric.\",\n    )\n    box_table = box_table.loc[numeric_ok].copy()\n\n    missing_dimensions = box_table[\n        [\"original_width\", \"original_height\", \"output_width\", \"output_height\"]\n    ].isna().any(axis=1)\n    record_issues(\n        box_table, missing_dimensions, \"missing_manifest_dimensions\", \"dropped\",\n        \"No matching preprocessed manifest row.\",\n    )\n    box_table = box_table.loc[~missing_dimensions].copy()\n\n    positive_area = (\n        (box_table[\"x_max\"] > box_table[\"x_min\"])\n        & (box_table[\"y_max\"] > box_table[\"y_min\"])\n    )\n    record_issues(\n        box_table, ~positive_area, \"non_positive_area\", \"dropped\",\n        \"x_max <= x_min or y_max <= y_min.\",\n    )\n    box_table = box_table.loc[positive_area].copy()\n\n    box_table[\"raw_box_width\"] = box_table[\"x_max\"] - box_table[\"x_min\"]\n    box_table[\"raw_box_height\"] = box_table[\"y_max\"] - box_table[\"y_min\"]\n    box_table[\"raw_box_area\"] = (\n        box_table[\"raw_box_width\"] * box_table[\"raw_box_height\"]\n    )\n\n    completely_outside = (\n        (box_table[\"x_max\"] <= 0)\n        | (box_table[\"y_max\"] <= 0)\n        | (box_table[\"x_min\"] >= box_table[\"original_width\"])\n        | (box_table[\"y_min\"] >= box_table[\"original_height\"])\n    )\n    record_issues(\n        box_table, completely_outside, \"box_outside_image\", \"dropped\",\n        \"Fused box has no intersection with the original image.\",\n    )\n    box_table = box_table.loc[~completely_outside].copy()\n\n    box_table[\"clipped_x_min\"] = box_table[\"x_min\"].clip(lower=0)\n    box_table[\"clipped_y_min\"] = box_table[\"y_min\"].clip(lower=0)\n    box_table[\"clipped_x_max\"] = np.minimum(\n        box_table[\"x_max\"], box_table[\"original_width\"]\n    )\n    box_table[\"clipped_y_max\"] = np.minimum(\n        box_table[\"y_max\"], box_table[\"original_height\"]\n    )\n    box_table[\"clipped_box_width\"] = (\n        box_table[\"clipped_x_max\"] - box_table[\"clipped_x_min\"]\n    )\n    box_table[\"clipped_box_height\"] = (\n        box_table[\"clipped_y_max\"] - box_table[\"clipped_y_min\"]\n    )\n    box_table[\"clipped_box_area\"] = (\n        box_table[\"clipped_box_width\"] * box_table[\"clipped_box_height\"]\n    )\n    box_table[\"visible_fraction\"] = (\n        box_table[\"clipped_box_area\"] / box_table[\"raw_box_area\"]\n    )\n\n    low_visible_fraction = (\n        box_table[\"visible_fraction\"] < MIN_VISIBLE_FRACTION\n    )\n    record_issues(\n        box_table,\n        low_visible_fraction,\n        \"insufficient_visible_fraction\",\n        \"dropped\",\n        f\"Visible box fraction is below {MIN_VISIBLE_FRACTION:.2f}.\",\n    )\n    box_table = box_table.loc[~low_visible_fraction].copy()\n\n    partially_outside = (\n        (box_table[\"x_min\"] < 0)\n        | (box_table[\"y_min\"] < 0)\n        | (box_table[\"x_max\"] > box_table[\"original_width\"])\n        | (box_table[\"y_max\"] > box_table[\"original_height\"])\n    )\n    record_issues(\n        box_table,\n        partially_outside,\n        \"box_clipped_to_boundary\",\n        \"clipped_for_review\",\n        (\n            \"Box was clipped only after passing the visible-fraction threshold; \"\n            \"review before publication.\"\n        ),\n    )\n    for coordinate in COORDINATE_COLUMNS:\n        box_table[coordinate] = box_table[f\"clipped_{coordinate}\"]\n\n    box_table[\"output_box_width\"] = (\n        (box_table[\"x_max\"] - box_table[\"x_min\"])\n        * box_table[\"output_width\"]\n        / box_table[\"original_width\"]\n    )\n    box_table[\"output_box_height\"] = (\n        (box_table[\"y_max\"] - box_table[\"y_min\"])\n        * box_table[\"output_height\"]\n        / box_table[\"original_height\"]\n    )\n    too_small = (\n        ((box_table[\"x_max\"] - box_table[\"x_min\"]) < MIN_BOX_SIDE_ORIGINAL_PX)\n        | ((box_table[\"y_max\"] - box_table[\"y_min\"]) < MIN_BOX_SIDE_ORIGINAL_PX)\n        | (box_table[\"output_box_width\"] < MIN_BOX_SIDE_OUTPUT_PX)\n        | (box_table[\"output_box_height\"] < MIN_BOX_SIDE_OUTPUT_PX)\n    )\n    record_issues(\n        box_table, too_small, \"box_too_small_after_preprocessing\", \"dropped\",\n        (\n            f\"Box side is below {MIN_BOX_SIDE_ORIGINAL_PX} original pixels or \"\n            f\"{MIN_BOX_SIDE_OUTPUT_PX} output pixels.\"\n        ),\n    )\n    box_table = box_table.loc[~too_small].copy()\n\n    duplicate_subset = [\"image_id\", *COORDINATE_COLUMNS]\n    exact_duplicate_mask = box_table.duplicated(duplicate_subset, keep=\"first\")\n    record_issues(\n        box_table, exact_duplicate_mask, \"exact_duplicate_fused_box\", \"dropped\",\n        \"Exact duplicate fused box; first occurrence retained.\",\n    )\n    box_table = box_table.loc[~exact_duplicate_mask].copy()\n\n    clean_fused_boxes = box_table.reset_index(drop=True)\n    clean_positive_ids = set(clean_fused_boxes[\"image_id\"])\n    ambiguous_ids = sorted(source_positive_ids - clean_positive_ids)\n    positive_ids_missing_manifest = sorted(source_positive_ids - manifest_ids)\n\n    box_issues = (\n        pd.concat(issue_records, ignore_index=True)\n        if issue_records\n        else pd.DataFrame(\n            columns=[\n                \"_source_row\", \"image_id\", *COORDINATE_COLUMNS,\n                \"issue_type\", \"action\", \"issue_detail\",\n            ]\n        )\n    )\n    box_issues.to_csv(\n        REPORT_ROOT / \"preprocessed_fused_box_issues.csv\", index=False\n    )\n    clean_fused_boxes.to_csv(\n        REPORT_ROOT / \"cleaned_fused_boxes_preprocessed_only.csv\", index=False\n    )\n    box_issues.loc[\n        box_issues[\"action\"].eq(\"clipped_for_review\")\n    ].to_csv(REPORT_ROOT / \"clipped_box_review.csv\", index=False)\n    radiologist_metadata_valid = True\n    radiologist_metadata_issues = clean_fused_boxes.iloc[0:0].copy()\n    if \"radiologist_count\" in clean_fused_boxes.columns:\n        radiologist_counts = pd.to_numeric(\n            clean_fused_boxes[\"radiologist_count\"], errors=\"coerce\"\n        )\n        invalid_radiologist_mask = (\n            ~np.isfinite(radiologist_counts)\n            | (radiologist_counts < 1)\n            | (radiologist_counts\n               > EXPECTED_TRAIN_RADIOLOGISTS_PER_IMAGE)\n            | (radiologist_counts != np.floor(radiologist_counts))\n        )\n        radiologist_metadata_issues = clean_fused_boxes.loc[\n            invalid_radiologist_mask\n        ].copy()\n        radiologist_metadata_valid = len(radiologist_metadata_issues) == 0\n        low_consensus_boxes = clean_fused_boxes.loc[\n            radiologist_counts.fillna(0) < 2\n        ].copy()\n    else:\n        low_consensus_boxes = clean_fused_boxes.iloc[0:0].copy()\n    radiologist_metadata_issues.to_csv(\n        REPORT_ROOT / \"invalid_radiologist_metadata.csv\", index=False\n    )\n    low_consensus_boxes.to_csv(\n        REPORT_ROOT / \"single_radiologist_box_review.csv\", index=False\n    )\n    if not radiologist_metadata_valid:\n        raise RuntimeError(\n            \"融合框的 radiologist_count 超出 VinDr-CXR 训练集 1–3 范围。\"\n        )\n    pd.DataFrame({\"image_id\": ambiguous_ids}).to_csv(\n        REPORT_ROOT / \"ambiguous_positive_images_excluded.csv\", index=False\n    )\n    pd.DataFrame({\"image_id\": positive_ids_missing_manifest}).to_csv(\n        REPORT_ROOT / \"positive_images_missing_manifest.csv\", index=False\n    )\n\n    # 代数缩放预检查；真正的 TXT 序列化回投影在后续逐框完成。\n    geometry_box_audit = clean_fused_boxes[\n        [\n            \"image_id\", *COORDINATE_COLUMNS,\n            \"original_width\", \"original_height\", \"output_width\", \"output_height\",\n        ]\n    ].copy()\n    coordinate_axes = {\n        \"x_min\": (\"original_width\", \"output_width\"),\n        \"x_max\": (\"original_width\", \"output_width\"),\n        \"y_min\": (\"original_height\", \"output_height\"),\n        \"y_max\": (\"original_height\", \"output_height\"),\n    }\n    error_columns = []\n    for coordinate, (original_dimension, output_dimension) in coordinate_axes.items():\n        normalized = (\n            geometry_box_audit[coordinate]\n            / geometry_box_audit[original_dimension]\n        )\n        projected = normalized * geometry_box_audit[output_dimension]\n        explicitly_scaled = (\n            geometry_box_audit[coordinate]\n            * geometry_box_audit[output_dimension]\n            / geometry_box_audit[original_dimension]\n        )\n        error_column = f\"{coordinate}_roundtrip_error_px\"\n        geometry_box_audit[error_column] = (\n            projected - explicitly_scaled\n        ).abs()\n        error_columns.append(error_column)\n    geometry_box_audit[\"max_roundtrip_error_px\"] = (\n        geometry_box_audit[error_columns].max(axis=1)\n        if len(geometry_box_audit)\n        else pd.Series(dtype=float)\n    )\n    geometry_box_audit.to_csv(\n        REPORT_ROOT / \"geometry_box_roundtrip_audit.csv\", index=False\n    )\n    max_geometry_roundtrip_error = (\n        float(geometry_box_audit[\"max_roundtrip_error_px\"].max())\n        if len(geometry_box_audit)\n        else 0.0\n    )\n    geometry_roundtrip_pass = (\n        np.isfinite(max_geometry_roundtrip_error)\n        and max_geometry_roundtrip_error <= ALGEBRAIC_GEOMETRY_TOLERANCE_PX\n    )\n    if not geometry_roundtrip_pass:\n        raise RuntimeError(\n            \"原始坐标归一化与实际输出 PNG 投影不一致，禁止生成训练标签。\"\n        )\n\n    print(\"Source fused boxes:\", len(source_fused_boxes))\n    print(\"Source manifest/box contract:\",\n          source_manifest_box_contract_pass)\n    print(\"Radiologist metadata valid:\", radiologist_metadata_valid)\n    print(\"Clean fused boxes:\", len(clean_fused_boxes))\n    print(\"Box issue records:\", len(box_issues))\n    print(\"Ambiguous positive images excluded:\", len(ambiguous_ids))\n    print(\"Single-radiologist fused boxes for review:\", len(low_consensus_boxes))\n    print(\"Max geometry roundtrip error (px):\", max_geometry_roundtrip_error)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 6. 建立进入自动去重前的候选 manifest 与基础人工复核队列\n    usable_manifest = manifest.loc[\n        ~manifest[\"image_id\"].isin(ambiguous_ids)\n    ].copy()\n    usable_manifest[\"is_positive_clean\"] = (\n        usable_manifest[\"image_id\"].isin(clean_positive_ids).astype(int)\n    )\n\n    source_box_counts = (\n        source_fused_boxes.groupby(\"image_id\").size().rename(\"source_box_count\")\n    )\n    clean_box_counts = (\n        clean_fused_boxes.groupby(\"image_id\").size().rename(\"clean_box_count\")\n    )\n    dropped_issue_summary = (\n        box_issues.loc[box_issues[\"action\"].eq(\"dropped\")]\n        .groupby(\"image_id\")[\"issue_type\"]\n        .agg(lambda values: \"|\".join(sorted(set(values))))\n        .rename(\"drop_reasons\")\n    )\n    manual_review_queue = pd.DataFrame({\"image_id\": ambiguous_ids})\n    if len(manual_review_queue):\n        manual_review_queue = (\n            manual_review_queue\n            .merge(\n                manifest[[\"image_id\", \"split\", \"is_positive\"]],\n                on=\"image_id\", how=\"left\",\n            )\n            .merge(source_box_counts, on=\"image_id\", how=\"left\")\n            .merge(clean_box_counts, on=\"image_id\", how=\"left\")\n            .merge(dropped_issue_summary, on=\"image_id\", how=\"left\")\n        )\n        manual_review_queue[\"clean_box_count\"] = (\n            manual_review_queue[\"clean_box_count\"].fillna(0).astype(int)\n        )\n        manual_review_queue[\"review_action\"] = (\n            \"Keep excluded; ask a qualified reviewer to confirm/relabel. \"\n            \"Do not automatically mark negative.\"\n        )\n\n    source_positive_ids_in_manifest = source_positive_ids & manifest_ids\n    ambiguous_ids_in_manifest = set(ambiguous_ids) & manifest_ids\n    ambiguous_positive_fraction = (\n        len(ambiguous_ids_in_manifest) / len(source_positive_ids_in_manifest)\n        if source_positive_ids_in_manifest\n        else 0.0\n    )\n\n    print(\"Candidates before content deduplication:\", len(usable_manifest))\n    print(\"Manual-review images before duplicate checks:\", len(manual_review_queue))\n    print(\"Ambiguous positive fraction:\", round(ambiguous_positive_fraction, 6))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 7. 根据清洗后的融合框重新生成 YOLO 标签\n    FUSED_BOX_MAP = {\n        image_id: group.reset_index(drop=True)\n        for image_id, group in clean_fused_boxes.groupby(\"image_id\", sort=False)\n    }\n\n\n    def yolo_lines_for_image(image_id, original_width, original_height):\n        rows = FUSED_BOX_MAP.get(str(image_id))\n        if rows is None or rows.empty:\n            return []\n        lines_out = []\n        for row in rows.itertuples(index=False):\n            # 框处于原始坐标系；1280_preprocessing 仅等比例缩放、不裁剪、不补边。\n            x_center = ((row.x_min + row.x_max) / 2.0) / original_width\n            y_center = ((row.y_min + row.y_max) / 2.0) / original_height\n            width = (row.x_max - row.x_min) / original_width\n            height = (row.y_max - row.y_min) / original_height\n            values = np.array([x_center, y_center, width, height], dtype=float)\n            if not np.isfinite(values).all():\n                raise ValueError(f\"{image_id}: regenerated non-finite label\")\n            if (\n                width <= 0\n                or height <= 0\n                or x_center - width / 2 < -1e-7\n                or y_center - height / 2 < -1e-7\n                or x_center + width / 2 > 1 + 1e-7\n                or y_center + height / 2 > 1 + 1e-7\n            ):\n                raise ValueError(f\"{image_id}: regenerated box outside image\")\n            values_text = [\n                f\"{value:.{LABEL_SERIALIZATION_DECIMALS}f}\"\n                for value in (x_center, y_center, width, height)\n            ]\n            lines_out.append(\n                f\"{YOLO_CLASS_ID} \" + \" \".join(values_text)\n            )\n        return sorted(lines_out)\n\n\n    def parse_label(path):\n        boxes, errors = [], []\n        if not path.exists():\n            return boxes, [\"missing_label_file\"]\n        for line_number, line in enumerate(\n            path.read_text(encoding=\"utf-8\").splitlines(), start=1\n        ):\n            if not line.strip():\n                continue\n            parts = line.split()\n            if len(parts) != 5:\n                errors.append(f\"line_{line_number}:expected_5_fields\")\n                continue\n            try:\n                class_id = int(parts[0])\n                x_center, y_center, width, height = map(float, parts[1:])\n            except Exception:\n                errors.append(f\"line_{line_number}:non_numeric\")\n                continue\n            values = np.array([x_center, y_center, width, height], dtype=float)\n            if class_id != YOLO_CLASS_ID:\n                errors.append(f\"line_{line_number}:wrong_class_id\")\n            if not np.isfinite(values).all():\n                errors.append(f\"line_{line_number}:non_finite\")\n                continue\n            if width <= 0 or height <= 0:\n                errors.append(f\"line_{line_number}:non_positive_size\")\n            if (\n                x_center - width / 2 < -1e-7\n                or y_center - height / 2 < -1e-7\n                or x_center + width / 2 > 1 + 1e-7\n                or y_center + height / 2 > 1 + 1e-7\n            ):\n                errors.append(f\"line_{line_number}:complete_box_out_of_bounds\")\n            boxes.append((class_id, x_center, y_center, width, height))\n        return boxes, errors\n\n\n    def canonical_boxes(boxes):\n        return sorted(\n            tuple([int(box[0])] + [round(float(value), 6) for value in box[1:]])\n            for box in boxes\n        )\n\n\n    def atomic_write_text(path, text):\n        path.parent.mkdir(parents=True, exist_ok=True)\n        temporary = path.with_suffix(path.suffix + \".tmp\")\n        temporary.write_text(text, encoding=\"utf-8\")\n        temporary.replace(path)\n\n\n    def sha256_file(path, chunk_size=1024 * 1024):\n        digest = hashlib.sha256()\n        with Path(path).open(\"rb\") as handle:\n            for chunk in iter(lambda: handle.read(chunk_size), b\"\"):\n                digest.update(chunk)\n        return digest.hexdigest()\n\n\n    def export_image(source, destination, expected_source_sha256):\n        source = Path(source)\n        destination = Path(destination)\n        current_source_sha256 = sha256_file(source)\n        if current_source_sha256 != expected_source_sha256:\n            raise RuntimeError(\n                \"Source image changed after the pre-export audit.\"\n            )\n        if IMAGE_EXPORT_MODE == \"none\":\n            return None, \"not_exported\"\n        destination.parent.mkdir(parents=True, exist_ok=True)\n        if IMAGE_EXPORT_MODE == \"symlink\":\n            if destination.exists() or destination.is_symlink():\n                destination.unlink()\n            destination.symlink_to(source)\n            return destination, \"symlinked\"\n        if (\n            destination.is_file()\n            and not destination.is_symlink()\n            and destination.stat().st_size == source.stat().st_size\n        ):\n            if (\n                not VERIFY_EXISTING_COPY_HASH\n                or sha256_file(destination) == expected_source_sha256\n            ):\n                return destination, \"verified_existing_copy\"\n        temporary = destination.with_suffix(destination.suffix + \".copying\")\n        shutil.copy2(source, temporary)\n        temporary.replace(destination)\n        return destination, \"copied\"\n\n\n    # 7B. 导出前全量解码、文件/像素哈希、自动坏图隔离与精确去重\n    # 所有删除决定都先形成最终保留名单，再同步到框、manifest、split plan 和导出。\n    usable_manifest_before_source_image_cleaning = usable_manifest.copy()\n    RESAMPLE_LANCZOS = getattr(Image, \"Resampling\", Image).LANCZOS\n\n\n    def decoded_pixel_sha256(pixels):\n        contiguous = np.ascontiguousarray(pixels, dtype=np.uint8)\n        digest = hashlib.sha256()\n        digest.update(b\"L\")\n        digest.update(int(contiguous.shape[1]).to_bytes(8, \"little\"))\n        digest.update(int(contiguous.shape[0]).to_bytes(8, \"little\"))\n        digest.update(contiguous.tobytes())\n        return digest.hexdigest()\n\n\n    def perceptual_dhash64(grayscale_image):\n        tiny = grayscale_image.resize((9, 8), RESAMPLE_LANCZOS)\n        array = np.asarray(tiny, dtype=np.uint8)\n        bits = array[:, 1:] > array[:, :-1]\n        value = 0\n        for bit in bits.reshape(-1):\n            value = (value << 1) | int(bit)\n        tiny.close()\n        return f\"{value:016x}\"\n\n\n    def normalized_optional_hash(value):\n        if value is None or pd.isna(value) or str(value).strip() == \"\":\n            return None\n        return str(value).strip().lower()\n\n\n    def preexport_source_image_record(row):\n        baseline_value = (\n            getattr(row, source_baseline_hash_column)\n            if source_baseline_hash_column is not None\n            else None\n        )\n        baseline_hash = normalized_optional_hash(baseline_value)\n        record = {\n            \"image_id\": str(row.image_id),\n            \"split\": str(row.split),\n            \"source_image_path\": str(row.source_image_path),\n            \"expected_width\": int(row.output_width),\n            \"expected_height\": int(row.output_height),\n            \"readable\": False,\n            \"format\": None,\n            \"format_is_png\": False,\n            \"mode\": None,\n            \"mode_is_acceptable\": False,\n            \"width\": None,\n            \"height\": None,\n            \"size_matches_manifest\": False,\n            \"is_constant\": None,\n            \"quality_review_flag\": None,\n            \"source_file_sha256\": None,\n            \"source_pixel_sha256\": None,\n            \"perceptual_dhash64\": None,\n            \"source_baseline_available\": baseline_hash is not None,\n            \"source_baseline_sha256\": baseline_hash,\n            \"source_baseline_hash_matches\": None,\n            \"auto_excludable\": False,\n            \"auto_exclusion_reason\": \"\",\n            \"provenance_failure\": False,\n            \"error\": \"\",\n        }\n        reasons = []\n        try:\n            path = Path(row.source_image_path)\n            file_hash = sha256_file(path)\n            record[\"source_file_sha256\"] = file_hash\n            if baseline_hash is not None:\n                record[\"source_baseline_hash_matches\"] = (\n                    file_hash == baseline_hash\n                )\n                record[\"provenance_failure\"] = (\n                    record[\"source_baseline_hash_matches\"] is not True\n                )\n\n            with Image.open(path) as image:\n                image.verify()\n            with Image.open(path) as image:\n                record[\"format\"] = image.format\n                record[\"format_is_png\"] = str(image.format).upper() == \"PNG\"\n                record[\"mode\"] = image.mode\n                record[\"mode_is_acceptable\"] = (\n                    image.mode == \"L\"\n                    if REQUIRE_GRAYSCALE\n                    else image.mode in {\"L\", \"RGB\"}\n                )\n                record[\"width\"] = int(image.width)\n                record[\"height\"] = int(image.height)\n                record[\"size_matches_manifest\"] = (\n                    image.width == int(row.output_width)\n                    and image.height == int(row.output_height)\n                )\n                grayscale = image.convert(\"L\")\n                pixels = np.asarray(grayscale, dtype=np.uint8).copy()\n                extrema = grayscale.getextrema()\n                record[\"min_intensity\"] = int(extrema[0])\n                record[\"max_intensity\"] = int(extrema[1])\n                record[\"is_constant\"] = extrema[0] == extrema[1]\n                record[\"pixel_mean\"] = float(pixels.mean())\n                record[\"pixel_std\"] = float(pixels.std())\n                record[\"near_black_fraction\"] = float((pixels <= 1).mean())\n                record[\"near_white_fraction\"] = float((pixels >= 254).mean())\n                record[\"quality_review_flag\"] = bool(\n                    record[\"pixel_std\"] < LOW_CONTRAST_STD_REVIEW_THRESHOLD\n                    or record[\"near_black_fraction\"]\n                    > EXTREME_SATURATION_REVIEW_FRACTION\n                    or record[\"near_white_fraction\"]\n                    > EXTREME_SATURATION_REVIEW_FRACTION\n                )\n                record[\"source_pixel_sha256\"] = decoded_pixel_sha256(pixels)\n                record[\"perceptual_dhash64\"] = perceptual_dhash64(grayscale)\n                grayscale.close()\n                record[\"readable\"] = True\n\n            if not record[\"format_is_png\"]:\n                reasons.append(\"not_png\")\n            if not record[\"mode_is_acceptable\"]:\n                reasons.append(\"unexpected_image_mode\")\n            if not record[\"size_matches_manifest\"]:\n                reasons.append(\"size_mismatch\")\n            if record[\"is_constant\"]:\n                reasons.append(\"constant_image\")\n        except Exception as error:\n            record[\"error\"] = f\"{type(error).__name__}: {error}\"\n            reasons.append(\"unreadable_or_corrupt\")\n\n        record[\"auto_exclusion_reason\"] = \"|\".join(sorted(set(reasons)))\n        record[\"auto_excludable\"] = bool(reasons)\n        return record\n\n\n    preexport_source_rows = list(\n        usable_manifest_before_source_image_cleaning.itertuples(index=False)\n    )\n    with ThreadPoolExecutor(max_workers=AUDIT_WORKERS) as executor:\n        source_image_audit_records = list(\n            tqdm(\n                executor.map(\n                    preexport_source_image_record,\n                    preexport_source_rows,\n                ),\n                total=len(preexport_source_rows),\n                desc=\"Auditing source PNGs before export\",\n            )\n        )\n    source_image_audit = pd.DataFrame(source_image_audit_records)\n    source_image_audit.to_csv(\n        REPORT_ROOT / \"preexport_source_image_audit.csv\", index=False\n    )\n\n    source_provenance_failures = source_image_audit.loc[\n        source_image_audit[\"provenance_failure\"].eq(True)\n    ].copy()\n    source_provenance_failures.to_csv(\n        REPORT_ROOT / \"source_baseline_hash_failures.csv\", index=False\n    )\n    if len(source_provenance_failures):\n        raise RuntimeError(\n            \"源 PNG 与 1280_preprocessing manifest 中已有哈希不一致；\"\n            \"这属于来源完整性失败，不能靠删除图片继续训练。\"\n        )\n\n    automatic_source_image_exclusions = source_image_audit.loc[\n        source_image_audit[\"auto_excludable\"].eq(True)\n    ].copy()\n    automatic_source_image_exclusions = automatic_source_image_exclusions.merge(\n        manifest[[\"image_id\", \"is_positive\", \"label_count\"]],\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n    automatic_source_image_exclusions.to_csv(\n        REPORT_ROOT / \"automatic_source_image_exclusions.csv\", index=False\n    )\n    automatic_source_excluded_ids = set(\n        automatic_source_image_exclusions[\"image_id\"].astype(str)\n    )\n    automatic_source_image_exclusion_fraction = (\n        len(automatic_source_image_exclusions)\n        / len(usable_manifest_before_source_image_cleaning)\n        if len(usable_manifest_before_source_image_cleaning)\n        else 0.0\n    )\n    automatic_positive_exclusion_count = int(\n        automatic_source_image_exclusions[\"is_positive\"].fillna(0).sum()\n    )\n    automatic_positive_exclusion_fraction = (\n        automatic_positive_exclusion_count\n        / int(manifest[\"is_positive\"].sum())\n        if int(manifest[\"is_positive\"].sum())\n        else 0.0\n    )\n    automatic_source_image_exclusion_within_limit = (\n        automatic_source_image_exclusion_fraction\n        <= MAX_AUTOMATIC_IMAGE_EXCLUSION_FRACTION\n        and automatic_positive_exclusion_fraction\n        <= MAX_AUTOMATIC_POSITIVE_EXCLUSION_FRACTION\n    )\n    if (\n        len(automatic_source_image_exclusions)\n        and not AUTO_EXCLUDE_UNREADABLE_OR_INVALID_IMAGES\n    ):\n        raise RuntimeError(\n            \"发现损坏或违反 PNG 契约的源图，但自动隔离被关闭。\"\n        )\n    if not automatic_source_image_exclusion_within_limit:\n        raise RuntimeError(\n            \"需要自动隔离的坏图比例超过安全阈值；\"\n            \"这可能是预处理整体失败，禁止静默丢弃后训练。\"\n        )\n\n    usable_manifest = usable_manifest_before_source_image_cleaning.loc[\n        ~usable_manifest_before_source_image_cleaning[\"image_id\"].isin(\n            automatic_source_excluded_ids\n        )\n    ].copy()\n\n    source_audit_merge = source_image_audit.loc[\n        ~source_image_audit[\"image_id\"].isin(automatic_source_excluded_ids),\n        [\n            \"image_id\", \"source_file_sha256\", \"source_pixel_sha256\",\n            \"perceptual_dhash64\", \"pixel_mean\", \"pixel_std\",\n            \"near_black_fraction\", \"near_white_fraction\",\n            \"quality_review_flag\", \"source_baseline_available\",\n            \"source_baseline_sha256\", \"source_baseline_hash_matches\",\n        ],\n    ].rename(columns={\"source_file_sha256\": \"source_sha256_preexport\"})\n    usable_manifest = usable_manifest.merge(\n        source_audit_merge,\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n\n    preexport_source_audit_complete = (\n        len(source_image_audit)\n        == len(usable_manifest_before_source_image_cleaning)\n        and source_image_audit[\"image_id\"].nunique()\n        == len(usable_manifest_before_source_image_cleaning)\n        and set(source_image_audit[\"image_id\"])\n        == set(usable_manifest_before_source_image_cleaning[\"image_id\"])\n    )\n    kept_source_hashes_complete = bool(\n        usable_manifest[\n            [\"source_sha256_preexport\", \"source_pixel_sha256\"]\n        ].notna().all().all()\n    )\n\n    # 锁定整个源版本：控制文件 + 逐图文件 SHA-256。\n    source_control_hashes = {\n        \"manifest.csv\": sha256_file(source_manifest_path),\n        \"dataset_config.json\": sha256_file(source_config_path),\n        \"data.yaml\": sha256_file(source_data_yaml_path),\n        \"fused_boxes_original_coordinates.csv\": sha256_file(\n            source_fused_boxes_path\n        ),\n        \"split_plan.csv\": sha256_file(source_split_plan_path),\n    }\n    source_label_hash_records = []\n    for row in manifest.sort_values(\n        [\"split\", \"image_id\"]\n    ).itertuples(index=False):\n        source_label_hash_records.append({\n            \"image_id\": row.image_id,\n            \"split\": row.split,\n            \"source_label_path\": row.source_label_path,\n            \"source_label_sha256\": sha256_file(\n                Path(row.source_label_path)\n            ),\n        })\n    source_label_hash_manifest = pd.DataFrame(source_label_hash_records)\n    source_label_hash_manifest.to_csv(\n        REPORT_ROOT / \"source_label_hash_manifest.csv\", index=False\n    )\n    source_fingerprint_digest = hashlib.sha256()\n    source_fingerprint_digest.update(\n        json.dumps(\n            source_control_hashes, sort_keys=True\n        ).encode(\"utf-8\")\n    )\n    for row in source_image_audit.sort_values(\n        [\"split\", \"image_id\"]\n    ).itertuples(index=False):\n        source_fingerprint_digest.update(\n            (\n                f\"{row.image_id}\\t{row.split}\\t\"\n                f\"{row.source_file_sha256 or ''}\\t{row.error}\\n\"\n            ).encode(\"utf-8\")\n        )\n    for row in source_label_hash_manifest.itertuples(index=False):\n        source_fingerprint_digest.update(\n            (\n                f\"{row.image_id}\\t{row.split}\\t\"\n                f\"{row.source_label_sha256}\\n\"\n            ).encode(\"utf-8\")\n        )\n    SOURCE_DATASET_FINGERPRINT = source_fingerprint_digest.hexdigest()\n    (REPORT_ROOT / \"source_dataset_identity.json\").write_text(\n        json.dumps({\n            \"source_dataset_signature\": computed_signature,\n            \"source_dataset_fingerprint\": SOURCE_DATASET_FINGERPRINT,\n            \"source_control_hashes\": source_control_hashes,\n            \"source_image_count\": len(source_image_audit),\n            \"source_label_count\": len(source_label_hash_manifest),\n        }, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    # 用最终重建标签而不是旧 TXT 决定重复样本是否标签一致。\n    label_signature_records = []\n    for row in usable_manifest.itertuples(index=False):\n        try:\n            final_lines = yolo_lines_for_image(\n                row.image_id,\n                int(row.original_width),\n                int(row.original_height),\n            )\n            label_signature_records.append({\n                \"image_id\": row.image_id,\n                \"final_label_signature\": \"\\n\".join(final_lines),\n                \"final_label_count\": len(final_lines),\n                \"signature_error\": \"\",\n            })\n        except Exception as error:\n            label_signature_records.append({\n                \"image_id\": row.image_id,\n                \"final_label_signature\": None,\n                \"final_label_count\": None,\n                \"signature_error\": f\"{type(error).__name__}: {error}\",\n            })\n    label_signature_audit = pd.DataFrame(label_signature_records)\n    label_signature_errors = label_signature_audit.loc[\n        label_signature_audit[\"signature_error\"].ne(\"\")\n    ].copy()\n    label_signature_audit.to_csv(\n        REPORT_ROOT / \"preexport_final_label_signatures.csv\", index=False\n    )\n    label_signature_errors.to_csv(\n        REPORT_ROOT / \"preexport_label_signature_errors.csv\", index=False\n    )\n    if len(label_signature_errors):\n        raise RuntimeError(\n            f\"有 {len(label_signature_errors)} 张图无法生成最终标签签名。\"\n        )\n    usable_manifest = usable_manifest.merge(\n        label_signature_audit[\n            [\"image_id\", \"final_label_signature\", \"final_label_count\"]\n        ],\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n\n    # 精确重复使用解码像素哈希，不受 PNG 压缩参数或元数据差异影响。\n    pixel_group_sizes = (\n        usable_manifest.groupby(\"source_pixel_sha256\")[\"image_id\"]\n        .nunique()\n        .rename(\"content_image_count\")\n    )\n    duplicate_pixel_hashes = set(\n        pixel_group_sizes.loc[pixel_group_sizes > 1].index\n    )\n    duplicate_candidates = usable_manifest.loc[\n        usable_manifest[\"source_pixel_sha256\"].isin(duplicate_pixel_hashes)\n    ].copy()\n\n    priority_rank = {\n        split: rank for rank, split in enumerate(DUPLICATE_KEEP_PRIORITY)\n    }\n    duplicate_decision_records = []\n    for pixel_hash, group in duplicate_candidates.groupby(\n        \"source_pixel_sha256\", sort=True\n    ):\n        group = group.copy()\n        group[\"_priority_rank\"] = group[\"split\"].map(priority_rank)\n        group = group.sort_values(\n            [\"_priority_rank\", \"image_id\"], kind=\"stable\"\n        )\n        labels_match = (\n            group[\"final_label_signature\"].nunique(dropna=False) == 1\n        )\n        if labels_match:\n            kept = group.iloc[0]\n            kept_image_id = str(kept[\"image_id\"])\n            kept_split = str(kept[\"split\"])\n            for row in group.itertuples(index=False):\n                is_kept = str(row.image_id) == kept_image_id\n                duplicate_decision_records.append({\n                    \"decoded_pixel_sha256\": pixel_hash,\n                    \"source_file_sha256\": row.source_sha256_preexport,\n                    \"image_id\": row.image_id,\n                    \"split\": row.split,\n                    \"source_image_path\": row.source_image_path,\n                    \"final_label_count\": int(row.final_label_count),\n                    \"final_label_signature\": row.final_label_signature,\n                    \"labels_match_within_group\": True,\n                    \"kept_image_id\": kept_image_id,\n                    \"kept_split\": kept_split,\n                    \"action\": (\n                        \"retained_duplicate_representative\"\n                        if is_kept\n                        else \"auto_excluded_exact_pixel_duplicate\"\n                    ),\n                    \"reason\": (\n                        \"Identical decoded pixels and identical final labels; \"\n                        \"deterministic test>val>train priority, then image_id.\"\n                    ),\n                })\n        else:\n            for row in group.itertuples(index=False):\n                duplicate_decision_records.append({\n                    \"decoded_pixel_sha256\": pixel_hash,\n                    \"source_file_sha256\": row.source_sha256_preexport,\n                    \"image_id\": row.image_id,\n                    \"split\": row.split,\n                    \"source_image_path\": row.source_image_path,\n                    \"final_label_count\": int(row.final_label_count),\n                    \"final_label_signature\": row.final_label_signature,\n                    \"labels_match_within_group\": False,\n                    \"kept_image_id\": \"\",\n                    \"kept_split\": \"\",\n                    \"action\": \"quarantined_duplicate_label_conflict\",\n                    \"reason\": (\n                        \"Identical decoded pixels but conflicting final labels; \"\n                        \"all copies quarantined and training is blocked.\"\n                    ),\n                })\n\n    duplicate_decision_columns = [\n        \"decoded_pixel_sha256\", \"source_file_sha256\", \"image_id\", \"split\",\n        \"source_image_path\", \"final_label_count\", \"final_label_signature\",\n        \"labels_match_within_group\", \"kept_image_id\", \"kept_split\",\n        \"action\", \"reason\",\n    ]\n    duplicate_decisions = pd.DataFrame(\n        duplicate_decision_records,\n        columns=duplicate_decision_columns,\n    )\n    auto_duplicate_exclusions = duplicate_decisions.loc[\n        duplicate_decisions[\"action\"].eq(\n            \"auto_excluded_exact_pixel_duplicate\"\n        )\n    ].copy()\n    duplicate_label_conflicts = duplicate_decisions.loc[\n        duplicate_decisions[\"action\"].eq(\n            \"quarantined_duplicate_label_conflict\"\n        )\n    ].copy()\n    duplicate_decisions.to_csv(\n        REPORT_ROOT / \"preexport_duplicate_content_decisions.csv\", index=False\n    )\n    auto_duplicate_exclusions.to_csv(\n        REPORT_ROOT / \"auto_duplicate_exclusions.csv\", index=False\n    )\n    duplicate_label_conflicts.to_csv(\n        REPORT_ROOT / \"duplicate_label_conflicts_quarantined.csv\", index=False\n    )\n\n    auto_duplicate_excluded_ids = set(\n        auto_duplicate_exclusions[\"image_id\"].astype(str)\n    )\n    duplicate_conflict_ids = set(\n        duplicate_label_conflicts[\"image_id\"].astype(str)\n    )\n    duplicate_excluded_ids = (\n        auto_duplicate_excluded_ids | duplicate_conflict_ids\n    )\n    usable_manifest = usable_manifest.loc[\n        ~usable_manifest[\"image_id\"].isin(duplicate_excluded_ids)\n    ].copy()\n\n    # 所有下游对象只使用同一份最终 ID 集。\n    final_usable_ids = set(usable_manifest[\"image_id\"])\n    clean_fused_boxes = clean_fused_boxes.loc[\n        clean_fused_boxes[\"image_id\"].isin(final_usable_ids)\n    ].copy()\n    clean_fused_boxes = clean_fused_boxes.sort_values(\n        [\"image_id\", \"x_min\", \"y_min\", \"x_max\", \"y_max\"],\n        kind=\"stable\",\n    ).reset_index(drop=True)\n    clean_positive_ids = set(clean_fused_boxes[\"image_id\"])\n    FUSED_BOX_MAP = {\n        image_id: group.reset_index(drop=True)\n        for image_id, group in clean_fused_boxes.groupby(\n            \"image_id\", sort=False\n        )\n    }\n    low_consensus_boxes = low_consensus_boxes.loc[\n        low_consensus_boxes[\"image_id\"].isin(final_usable_ids)\n    ].copy()\n    clean_fused_boxes.to_csv(\n        REPORT_ROOT / \"cleaned_fused_boxes_preprocessed_only.csv\", index=False\n    )\n    low_consensus_boxes.to_csv(\n        REPORT_ROOT / \"single_radiologist_box_review.csv\", index=False\n    )\n\n    # 真正的标签序列化几何检查：8 位小数 TXT -> 实际输出 PNG 像素。\n    serialized_geometry_records = []\n    geometry_audit_source = clean_fused_boxes.rename(\n        columns={\"_source_row\": \"source_row_index\"}\n    )\n    for row in geometry_audit_source.itertuples(index=False):\n        original_width = float(row.original_width)\n        original_height = float(row.original_height)\n        output_width = float(row.output_width)\n        output_height = float(row.output_height)\n        normalized_values = [\n            ((row.x_min + row.x_max) / 2.0) / original_width,\n            ((row.y_min + row.y_max) / 2.0) / original_height,\n            (row.x_max - row.x_min) / original_width,\n            (row.y_max - row.y_min) / original_height,\n        ]\n        serialized_values = [\n            float(f\"{value:.{LABEL_SERIALIZATION_DECIMALS}f}\")\n            for value in normalized_values\n        ]\n        x_center, y_center, width, height = serialized_values\n        reconstructed = {\n            \"x_min\": (x_center - width / 2.0) * output_width,\n            \"y_min\": (y_center - height / 2.0) * output_height,\n            \"x_max\": (x_center + width / 2.0) * output_width,\n            \"y_max\": (y_center + height / 2.0) * output_height,\n        }\n        intended = {\n            \"x_min\": row.x_min * output_width / original_width,\n            \"y_min\": row.y_min * output_height / original_height,\n            \"x_max\": row.x_max * output_width / original_width,\n            \"y_max\": row.y_max * output_height / original_height,\n        }\n        errors = {\n            coordinate: abs(reconstructed[coordinate] - intended[coordinate])\n            for coordinate in COORDINATE_COLUMNS\n        }\n        serialized_geometry_records.append({\n            \"image_id\": row.image_id,\n            \"source_row\": int(row.source_row_index),\n            **{\n                f\"intended_{key}_output_px\": value\n                for key, value in intended.items()\n            },\n            **{\n                f\"reconstructed_{key}_output_px\": value\n                for key, value in reconstructed.items()\n            },\n            **{\n                f\"{key}_serialization_error_px\": value\n                for key, value in errors.items()\n            },\n            \"max_serialization_error_px\": max(errors.values()),\n        })\n    serialized_geometry_audit = pd.DataFrame(serialized_geometry_records)\n    serialized_geometry_audit.to_csv(\n        REPORT_ROOT / \"serialized_yolo_geometry_audit.csv\", index=False\n    )\n    max_serialized_geometry_error = (\n        float(\n            serialized_geometry_audit[\"max_serialization_error_px\"].max()\n        )\n        if len(serialized_geometry_audit)\n        else 0.0\n    )\n    serialized_geometry_pass = (\n        np.isfinite(max_serialized_geometry_error)\n        and max_serialized_geometry_error\n        <= SERIALIZED_GEOMETRY_TOLERANCE_OUTPUT_PX\n    )\n    if not serialized_geometry_pass:\n        raise RuntimeError(\n            \"最终 YOLO 文本序列化后投影到 PNG 的误差超过阈值。\"\n        )\n\n    usable_manifest[\"is_positive_clean\"] = (\n        usable_manifest[\"image_id\"].isin(clean_positive_ids).astype(int)\n    )\n    source_positive_mismatch = usable_manifest.loc[\n        usable_manifest[\"is_positive\"].astype(int)\n        != usable_manifest[\"is_positive_clean\"]\n    ].copy()\n    source_positive_mismatch.to_csv(\n        REPORT_ROOT / \"source_vs_clean_positive_mismatch.csv\", index=False\n    )\n\n    split_sets = {\n        split: set(\n            usable_manifest.loc[\n                usable_manifest[\"split\"].eq(split), \"image_id\"\n            ]\n        )\n        for split in VALID_SPLITS\n    }\n    image_split_leakage = (\n        (split_sets[\"train\"] & split_sets[\"val\"])\n        | (split_sets[\"train\"] & split_sets[\"test\"])\n        | (split_sets[\"val\"] & split_sets[\"test\"])\n    )\n\n    if len(duplicate_label_conflicts):\n        conflict_review_queue = duplicate_label_conflicts[\n            [\n                \"image_id\", \"split\", \"final_label_count\",\n                \"decoded_pixel_sha256\",\n            ]\n        ].copy()\n        conflict_review_queue = conflict_review_queue.merge(\n            manifest[[\"image_id\", \"is_positive\"]],\n            on=\"image_id\",\n            how=\"left\",\n            validate=\"one_to_one\",\n        )\n        conflict_review_queue = conflict_review_queue.merge(\n            source_box_counts,\n            on=\"image_id\",\n            how=\"left\",\n        ).merge(\n            clean_box_counts,\n            on=\"image_id\",\n            how=\"left\",\n        )\n        conflict_review_queue[\"drop_reasons\"] = (\n            \"duplicate_pixels_with_conflicting_final_labels\"\n        )\n        conflict_review_queue[\"review_action\"] = (\n            \"Every copy remains quarantined; a qualified reviewer must \"\n            \"resolve the conflict before the training gate can pass.\"\n        )\n        manual_review_queue = pd.concat(\n            [manual_review_queue, conflict_review_queue],\n            ignore_index=True,\n            sort=False,\n        ).drop_duplicates(\"image_id\", keep=\"first\")\n    manual_review_queue.to_csv(\n        REPORT_ROOT / \"manual_review_queue.csv\", index=False\n    )\n\n    duplicate_conflict_image_fraction = (\n        len(duplicate_conflict_ids)\n        / len(usable_manifest_before_source_image_cleaning)\n        if len(usable_manifest_before_source_image_cleaning)\n        else 0.0\n    )\n    no_duplicate_label_conflicts = len(duplicate_label_conflicts) == 0\n\n    class_distribution_valid = (\n        bool(clean_positive_ids & final_usable_ids)\n        and bool(final_usable_ids - clean_positive_ids)\n    )\n    split_class_distribution = (\n        usable_manifest.groupby(\"split\")[\"is_positive_clean\"]\n        .agg(image_count=\"count\", positive_images=\"sum\")\n        .reindex(VALID_SPLITS)\n        .reset_index()\n    )\n    split_class_distribution[\"negative_images\"] = (\n        split_class_distribution[\"image_count\"]\n        - split_class_distribution[\"positive_images\"]\n    )\n    split_class_distribution[\"has_both_classes\"] = (\n        (split_class_distribution[\"positive_images\"] > 0)\n        & (split_class_distribution[\"negative_images\"] > 0)\n    )\n    split_class_distribution.to_csv(\n        REPORT_ROOT / \"split_class_distribution.csv\", index=False\n    )\n    all_splits_have_both_classes = bool(\n        split_class_distribution[\"has_both_classes\"].all()\n    )\n\n    # 感知哈希只用于疑似近重复复核；绝不据此自动删医学图像。\n    perceptual_hash_collision_rows = pd.DataFrame()\n    cross_split_perceptual_hash_review = pd.DataFrame()\n    if REPORT_PERCEPTUAL_HASH_COLLISIONS:\n        perceptual_stats = (\n            usable_manifest.dropna(subset=[\"perceptual_dhash64\"])\n            .groupby(\"perceptual_dhash64\")\n            .agg(\n                image_count=(\"image_id\", \"nunique\"),\n                split_count=(\"split\", \"nunique\"),\n                exact_pixel_count=(\"source_pixel_sha256\", \"nunique\"),\n            )\n            .reset_index()\n        )\n        perceptual_collision_hashes = set(\n            perceptual_stats.loc[\n                (perceptual_stats[\"image_count\"] > 1)\n                & (perceptual_stats[\"exact_pixel_count\"] > 1),\n                \"perceptual_dhash64\",\n            ]\n        )\n        cross_split_perceptual_hashes = set(\n            perceptual_stats.loc[\n                (perceptual_stats[\"image_count\"] > 1)\n                & (perceptual_stats[\"split_count\"] > 1)\n                & (perceptual_stats[\"exact_pixel_count\"] > 1),\n                \"perceptual_dhash64\",\n            ]\n        )\n        perceptual_hash_collision_rows = usable_manifest.loc[\n            usable_manifest[\"perceptual_dhash64\"].isin(\n                perceptual_collision_hashes\n            )\n        ].sort_values([\"perceptual_dhash64\", \"split\", \"image_id\"])\n        cross_split_perceptual_hash_review = usable_manifest.loc[\n            usable_manifest[\"perceptual_dhash64\"].isin(\n                cross_split_perceptual_hashes\n            )\n        ].sort_values([\"perceptual_dhash64\", \"split\", \"image_id\"])\n    perceptual_hash_collision_rows.to_csv(\n        REPORT_ROOT / \"perceptual_hash_collision_review.csv\", index=False\n    )\n    cross_split_perceptual_hash_review.to_csv(\n        REPORT_ROOT / \"cross_split_perceptual_hash_review.csv\", index=False\n    )\n\n    # 固定最终顺序并生成与模型训练绑定的数据指纹。\n    split_sort_rank = {\n        split: rank for rank, split in enumerate(VALID_SPLITS)\n    }\n    usable_manifest[\"_split_sort_rank\"] = usable_manifest[\"split\"].map(\n        split_sort_rank\n    )\n    usable_manifest = usable_manifest.sort_values(\n        [\"_split_sort_rank\", \"image_id\"], kind=\"stable\"\n    ).drop(columns=\"_split_sort_rank\").reset_index(drop=True)\n\n    clean_fingerprint_digest = hashlib.sha256()\n    clean_fingerprint_digest.update(\n        (\n            SOURCE_DATASET_FINGERPRINT + \"\\n\"\n            + CLEANING_POLICY_SIGNATURE + \"\\n\"\n        ).encode(\"utf-8\")\n    )\n    for row in usable_manifest.itertuples(index=False):\n        clean_fingerprint_digest.update(\n            (\n                f\"{row.image_id}\\t{row.split}\\t\"\n                f\"{row.source_pixel_sha256}\\t\"\n                f\"{row.final_label_signature}\\n\"\n            ).encode(\"utf-8\")\n        )\n    CLEAN_DATASET_FINGERPRINT = clean_fingerprint_digest.hexdigest()\n\n    deduplication_summary = {\n        \"automatic_deduplication_enabled\": AUTO_DEDUPLICATE_EXACT_CONTENT,\n        \"exact_content_definition\": \"decoded_grayscale_pixels_sha256\",\n        \"keep_priority\": list(DUPLICATE_KEEP_PRIORITY),\n        \"candidate_images_after_source_image_cleaning\": len(\n            usable_manifest_before_source_image_cleaning\n        ) - len(automatic_source_image_exclusions),\n        \"duplicate_content_groups\": int(\n            duplicate_decisions[\"decoded_pixel_sha256\"].nunique()\n            if len(duplicate_decisions) else 0\n        ),\n        \"safe_redundant_copies_excluded\": len(auto_duplicate_exclusions),\n        \"label_conflict_images_quarantined\": len(duplicate_label_conflicts),\n        \"label_conflicts_block_training\": BLOCK_ON_DUPLICATE_LABEL_CONFLICT,\n        \"label_conflict_image_fraction\": duplicate_conflict_image_fraction,\n        \"final_usable_images\": len(usable_manifest),\n        \"source_files_modified\": False,\n    }\n    (REPORT_ROOT / \"deduplication_summary.json\").write_text(\n        json.dumps(deduplication_summary, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n    (REPORT_ROOT / \"source_image_exclusion_summary.json\").write_text(\n        json.dumps({\n            \"automatically_excluded_images\": len(\n                automatic_source_image_exclusions\n            ),\n            \"automatically_excluded_positive_images\": (\n                automatic_positive_exclusion_count\n            ),\n            \"image_exclusion_fraction\": (\n                automatic_source_image_exclusion_fraction\n            ),\n            \"positive_exclusion_fraction\": (\n                automatic_positive_exclusion_fraction\n            ),\n            \"within_configured_limits\": (\n                automatic_source_image_exclusion_within_limit\n            ),\n        }, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    print(\"Pre-export source images audited:\", len(source_image_audit))\n    print(\"Source image audit complete:\", preexport_source_audit_complete)\n    print(\"Automatically excluded invalid images:\", len(\n        automatic_source_image_exclusions\n    ))\n    print(\"Duplicate decoded-pixel groups:\", deduplication_summary[\n        \"duplicate_content_groups\"\n    ])\n    print(\"Safe redundant copies auto-excluded:\", len(\n        auto_duplicate_exclusions\n    ))\n    print(\"Conflicting duplicate images quarantined:\", len(\n        duplicate_label_conflicts\n    ))\n    print(\"Final usable preprocessed images:\", len(usable_manifest))\n    print(\n        usable_manifest.groupby(\"split\")[\"is_positive_clean\"].agg(\n            [\"count\", \"sum\"]\n        )\n    )\n    print(\"Source vs clean positive mismatch:\", len(source_positive_mismatch))\n    print(\"Image-ID split leakage:\", len(image_split_leakage))\n    print(\"Manual-review images:\", len(manual_review_queue))\n    print(\"Cross-split perceptual-hash review rows:\", len(\n        cross_split_perceptual_hash_review\n    ))\n    print(\"Max serialized geometry error (px):\", max_serialized_geometry_error)\n    print(\"All splits contain positive and negative images:\",\n          all_splits_have_both_classes)\n    print(\"Source dataset fingerprint:\", SOURCE_DATASET_FINGERPRINT)\n    print(\"Clean dataset fingerprint:\", CLEAN_DATASET_FINGERPRINT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 8. 清理 V6 旧残留、磁盘预检并原子导出最终数据集\n    for split in VALID_SPLITS:\n        (NEW_LABEL_ROOT / split).mkdir(parents=True, exist_ok=True)\n        if IMAGE_EXPORT_MODE != \"none\":\n            (NEW_IMAGE_ROOT / split).mkdir(parents=True, exist_ok=True)\n\n    expected_preexport_keys = set(\n        zip(usable_manifest[\"split\"], usable_manifest[\"image_id\"])\n    )\n    stale_output_records = []\n    unexpected_output_directories = []\n    for split in VALID_SPLITS:\n        generated_roots = [\n            (NEW_LABEL_ROOT / split, \".txt\", \"label\")\n        ]\n        if IMAGE_EXPORT_MODE != \"none\":\n            generated_roots.append(\n                (NEW_IMAGE_ROOT / split, \".png\", \"image\")\n            )\n        for generated_root, expected_suffix, kind in generated_roots:\n            for generated_path in generated_root.iterdir():\n                if generated_path.is_dir() and not generated_path.is_symlink():\n                    unexpected_output_directories.append({\n                        \"split\": split,\n                        \"kind\": kind,\n                        \"path\": str(generated_path),\n                        \"reason\": \"unexpected_subdirectory_not_removed\",\n                    })\n                    continue\n                key = (split, generated_path.stem)\n                is_expected = (\n                    generated_path.suffix.lower() == expected_suffix\n                    and key in expected_preexport_keys\n                    and not (\n                        IMAGE_EXPORT_MODE == \"copy\"\n                        and generated_path.is_symlink()\n                    )\n                )\n                if not is_expected:\n                    generated_path.unlink()\n                    stale_output_records.append({\n                        \"split\": split,\n                        \"image_id\": generated_path.stem,\n                        \"kind\": kind,\n                        \"path\": str(generated_path),\n                        \"action\": \"removed_stale_v6_output\",\n                    })\n    stale_output_cleanup = pd.DataFrame(\n        stale_output_records,\n        columns=[\"split\", \"image_id\", \"kind\", \"path\", \"action\"],\n    )\n    unexpected_output_directories_df = pd.DataFrame(\n        unexpected_output_directories,\n        columns=[\"split\", \"kind\", \"path\", \"reason\"],\n    )\n    stale_output_cleanup.to_csv(\n        REPORT_ROOT / \"stale_output_cleanup.csv\", index=False\n    )\n    unexpected_output_directories_df.to_csv(\n        REPORT_ROOT / \"unexpected_output_directories.csv\", index=False\n    )\n\n\n    def existing_copy_matches(row):\n        if IMAGE_EXPORT_MODE != \"copy\":\n            return False\n        destination = (\n            NEW_IMAGE_ROOT / row.split / f\"{row.image_id}.png\"\n        )\n        source = Path(row.source_image_path)\n        try:\n            return (\n                destination.is_file()\n                and not destination.is_symlink()\n                and destination.stat().st_size == source.stat().st_size\n                and (\n                    not VERIFY_EXISTING_COPY_HASH\n                    or sha256_file(destination)\n                    == row.source_sha256_preexport\n                )\n            )\n        except Exception:\n            return False\n\n\n    source_image_bytes = sum(\n        Path(path).stat().st_size\n        for path in usable_manifest[\"source_image_path\"]\n    )\n    estimated_additional_copy_bytes = 0\n    if IMAGE_EXPORT_MODE == \"copy\":\n        for row in usable_manifest.itertuples(index=False):\n            if not existing_copy_matches(row):\n                estimated_additional_copy_bytes += Path(\n                    row.source_image_path\n                ).stat().st_size\n\n    disk_usage = shutil.disk_usage(OUTPUT_BASE_ROOT)\n    required_free_bytes = int(\n        estimated_additional_copy_bytes * DISK_SPACE_SAFETY_FACTOR\n        + DISK_SPACE_RESERVE_GB * (1024 ** 3)\n    )\n    disk_preflight_pass = (\n        IMAGE_EXPORT_MODE != \"copy\"\n        or disk_usage.free >= required_free_bytes\n    )\n    disk_preflight = {\n        \"image_export_mode\": IMAGE_EXPORT_MODE,\n        \"source_image_bytes\": int(source_image_bytes),\n        \"estimated_additional_copy_bytes\": int(\n            estimated_additional_copy_bytes\n        ),\n        \"free_bytes\": int(disk_usage.free),\n        \"required_free_bytes\": int(required_free_bytes),\n        \"passed\": bool(disk_preflight_pass),\n    }\n    (REPORT_ROOT / \"disk_preflight.json\").write_text(\n        json.dumps(disk_preflight, indent=2), encoding=\"utf-8\"\n    )\n    if not disk_preflight_pass:\n        raise RuntimeError(\n            \"磁盘空间预检失败。请清理 /kaggle/working；\"\n            \"正式保存仍必须使用 copy。\"\n        )\n\n    export_record_columns = [\n        \"image_id\", \"split\", \"is_positive\", \"label_count\",\n        \"source_image_path\", \"source_label_path\",\n        \"source_baseline_sha256\", \"source_sha256_preexport\",\n        \"source_pixel_sha256\", \"perceptual_dhash64\",\n        \"final_label_signature\", \"image_path\", \"label_path\",\n        \"original_width\", \"original_height\",\n        \"output_width\", \"output_height\", \"image_export_status\",\n    ]\n    export_failure_columns = [\n        \"image_id\", \"split\", \"error_type\", \"error_message\"\n    ]\n    export_records = []\n    export_failures = []\n    start_time = perf_counter()\n    for row in tqdm(\n        usable_manifest.itertuples(index=False),\n        total=len(usable_manifest),\n        desc=\"Exporting final clean dataset\",\n    ):\n        try:\n            source_image = Path(row.source_image_path)\n            destination_image = (\n                NEW_IMAGE_ROOT / row.split / f\"{row.image_id}.png\"\n            )\n            destination_label = (\n                NEW_LABEL_ROOT / row.split / f\"{row.image_id}.txt\"\n            )\n            label_lines = yolo_lines_for_image(\n                row.image_id,\n                int(row.original_width),\n                int(row.original_height),\n            )\n            if int(row.is_positive_clean) == 1 and not label_lines:\n                raise ValueError(\n                    \"Clean positive image has no regenerated label.\"\n                )\n            if int(row.is_positive_clean) == 0 and label_lines:\n                raise ValueError(\n                    \"Clean negative image unexpectedly has boxes.\"\n                )\n            final_label_text = (\n                \"\\n\".join(label_lines) + (\"\\n\" if label_lines else \"\")\n            )\n            atomic_write_text(destination_label, final_label_text)\n            exported_image, export_status = export_image(\n                source_image,\n                destination_image,\n                row.source_sha256_preexport,\n            )\n            export_records.append({\n                \"image_id\": row.image_id,\n                \"split\": row.split,\n                \"is_positive\": int(bool(label_lines)),\n                \"label_count\": len(label_lines),\n                \"source_image_path\": str(source_image),\n                \"source_label_path\": str(row.source_label_path),\n                \"source_baseline_sha256\": (\n                    row.source_baseline_sha256\n                ),\n                \"source_sha256_preexport\": (\n                    row.source_sha256_preexport\n                ),\n                \"source_pixel_sha256\": row.source_pixel_sha256,\n                \"perceptual_dhash64\": row.perceptual_dhash64,\n                \"final_label_signature\": row.final_label_signature,\n                \"image_path\": (\n                    str(exported_image)\n                    if exported_image is not None\n                    else str(source_image)\n                ),\n                \"label_path\": str(destination_label),\n                \"original_width\": int(row.original_width),\n                \"original_height\": int(row.original_height),\n                \"output_width\": int(row.output_width),\n                \"output_height\": int(row.output_height),\n                \"image_export_status\": export_status,\n            })\n        except Exception as error:\n            export_failures.append({\n                \"image_id\": row.image_id,\n                \"split\": row.split,\n                \"error_type\": type(error).__name__,\n                \"error_message\": str(error),\n            })\n\n    manifest_clean = pd.DataFrame(\n        export_records, columns=export_record_columns\n    )\n    export_failures_df = pd.DataFrame(\n        export_failures, columns=export_failure_columns\n    )\n    manifest_clean.to_csv(DATASET_ROOT / \"manifest.csv\", index=False)\n    export_failures_df.to_csv(\n        REPORT_ROOT / \"export_failures.csv\", index=False\n    )\n\n    data_yaml_path = DATASET_ROOT / \"data.yaml\"\n    if IMAGE_EXPORT_MODE != \"none\":\n        clean_data_config = {\n            # 空 path 由 Ultralytics 解析为 data.yaml 所在目录。\n            \"path\": \"\",\n            \"train\": \"images/train\",\n            \"val\": \"images/val\",\n            \"test\": \"images/test\",\n            \"nc\": 1,\n            \"names\": {YOLO_CLASS_ID: TARGET_CLASS},\n        }\n        atomic_write_text(\n            data_yaml_path,\n            yaml.safe_dump(\n                clean_data_config,\n                sort_keys=False,\n                allow_unicode=True,\n            ),\n        )\n\n    expected_output_keys = set(\n        zip(manifest_clean[\"split\"], manifest_clean[\"image_id\"])\n    )\n    actual_output_image_keys = (\n        {\n            (path.parent.name, path.stem)\n            for split in VALID_SPLITS\n            for path in (NEW_IMAGE_ROOT / split).glob(\"*.png\")\n        }\n        if IMAGE_EXPORT_MODE != \"none\"\n        else set()\n    )\n    actual_output_label_keys = {\n        (path.parent.name, path.stem)\n        for split in VALID_SPLITS\n        for path in (NEW_LABEL_ROOT / split).glob(\"*.txt\")\n    }\n    extra_output_images = sorted(\n        actual_output_image_keys - expected_output_keys\n    )\n    extra_output_labels = sorted(\n        actual_output_label_keys - expected_output_keys\n    )\n    missing_output_images = sorted(\n        expected_output_keys - actual_output_image_keys\n    ) if IMAGE_EXPORT_MODE != \"none\" else []\n    missing_output_labels = sorted(\n        expected_output_keys - actual_output_label_keys\n    )\n    pd.DataFrame(\n        extra_output_images, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"extra_output_images.csv\", index=False)\n    pd.DataFrame(\n        extra_output_labels, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"extra_output_labels.csv\", index=False)\n    pd.DataFrame(\n        missing_output_images, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"missing_output_images.csv\", index=False)\n    pd.DataFrame(\n        missing_output_labels, columns=[\"split\", \"image_id\"]\n    ).to_csv(REPORT_ROOT / \"missing_output_labels.csv\", index=False)\n\n    postexport_unexpected_entries = []\n    for split in VALID_SPLITS:\n        roots = [(NEW_LABEL_ROOT / split, \".txt\", \"label\")]\n        if IMAGE_EXPORT_MODE != \"none\":\n            roots.append((NEW_IMAGE_ROOT / split, \".png\", \"image\"))\n        for root, suffix, kind in roots:\n            for path in root.iterdir():\n                if (\n                    path.is_dir()\n                    or (\n                        IMAGE_EXPORT_MODE == \"copy\"\n                        and path.is_symlink()\n                    )\n                    or path.suffix.lower() != suffix\n                    or (split, path.stem) not in expected_output_keys\n                ):\n                    postexport_unexpected_entries.append({\n                        \"split\": split,\n                        \"kind\": kind,\n                        \"path\": str(path),\n                    })\n    postexport_unexpected_entries_df = pd.DataFrame(\n        postexport_unexpected_entries,\n        columns=[\"split\", \"kind\", \"path\"],\n    )\n    postexport_unexpected_entries_df.to_csv(\n        REPORT_ROOT / \"unexpected_output_entries.csv\", index=False\n    )\n\n    yaml_paths_exist = False\n    yaml_class_matches = False\n    if IMAGE_EXPORT_MODE != \"none\" and data_yaml_path.exists():\n        loaded_yaml = yaml.safe_load(\n            data_yaml_path.read_text(encoding=\"utf-8\")\n        )\n        yaml_root = (\n            data_yaml_path.parent\n            if not loaded_yaml.get(\"path\")\n            else Path(loaded_yaml[\"path\"])\n        )\n        yaml_paths_exist = all(\n            (yaml_root / loaded_yaml[split]).is_dir()\n            for split in VALID_SPLITS\n        )\n        yaml_names = loaded_yaml.get(\"names\", {})\n        yaml_class_name = (\n            yaml_names[YOLO_CLASS_ID]\n            if isinstance(yaml_names, list)\n            else yaml_names.get(\n                YOLO_CLASS_ID, yaml_names.get(str(YOLO_CLASS_ID))\n            )\n        )\n        yaml_class_matches = (\n            yaml_class_name == TARGET_CLASS\n            and int(loaded_yaml.get(\"nc\", 1)) == 1\n        )\n\n    output_split_counts = {\n        split: int(count)\n        for split, count in manifest_clean.groupby(\"split\").size().items()\n    }\n    dataset_identity = {\n        \"cleaning_version\": \"6.0-release\",\n        \"source_dataset_signature\": computed_signature,\n        \"source_dataset_fingerprint\": SOURCE_DATASET_FINGERPRINT,\n        \"cleaning_policy_signature\": CLEANING_POLICY_SIGNATURE,\n        \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n        \"target_class\": TARGET_CLASS,\n        \"class_id\": YOLO_CLASS_ID,\n        \"image_count\": int(len(manifest_clean)),\n        \"split_counts\": output_split_counts,\n        \"source_files_modified\": False,\n    }\n    (DATASET_ROOT / \"dataset_identity.json\").write_text(\n        json.dumps(dataset_identity, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n    v6_dataset_config = {\n        **CLEANING_POLICY,\n        **dataset_identity,\n        \"source_dataset_root\": str(source_dataset_root),\n        \"image_export_mode\": IMAGE_EXPORT_MODE,\n        \"audit_mode\": AUDIT_MODE,\n        \"references\": [\n            \"10.1038/s41597-022-01498-w\",\n            \"10.1016/j.patter.2023.100804\",\n        ],\n    }\n    (DATASET_ROOT / \"cleaning_config_v6.json\").write_text(\n        json.dumps(v6_dataset_config, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    print(\"Disk preflight:\", disk_preflight_pass)\n    print(\"Removed stale V6 output files:\", len(stale_output_cleanup))\n    print(\"Exported images:\", len(manifest_clean))\n    print(\"Export failures:\", len(export_failures_df))\n    print(\"Missing output images/labels:\",\n          len(missing_output_images), len(missing_output_labels))\n    print(\"Extra output images/labels:\",\n          len(extra_output_images), len(extra_output_labels))\n    print(\"Unexpected output entries:\",\n          len(postexport_unexpected_entries_df))\n    print(\"Elapsed seconds:\", round(perf_counter() - start_time, 1))\n    if IMAGE_EXPORT_MODE != \"none\":\n        print(\"Portable training YAML:\", data_yaml_path)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 9. 对比 1280_preprocessing 源标签和 V6 重建标签\n    label_audit_records = []\n    label_comparison_records = []\n    for row in tqdm(\n        manifest_clean.itertuples(index=False),\n        total=len(manifest_clean),\n        desc=\"Comparing source and rebuilt labels\",\n    ):\n        source_boxes, source_errors = parse_label(Path(row.source_label_path))\n        clean_boxes, clean_errors = parse_label(Path(row.label_path))\n        label_audit_records.extend([\n            {\n                \"version\": \"1280_preprocessing_source\",\n                \"image_id\": row.image_id,\n                \"split\": row.split,\n                \"path\": row.source_label_path,\n                \"box_count\": len(source_boxes),\n                \"error_count\": len(source_errors),\n                \"errors\": \"|\".join(source_errors),\n            },\n            {\n                \"version\": \"final_v6\",\n                \"image_id\": row.image_id,\n                \"split\": row.split,\n                \"path\": row.label_path,\n                \"box_count\": len(clean_boxes),\n                \"error_count\": len(clean_errors),\n                \"errors\": \"|\".join(clean_errors),\n            },\n        ])\n        label_comparison_records.append({\n            \"image_id\": row.image_id,\n            \"split\": row.split,\n            \"manifest_label_count\": int(row.label_count),\n            \"source_box_count\": len(source_boxes),\n            \"clean_box_count\": len(clean_boxes),\n            \"source_error_count\": len(source_errors),\n            \"clean_error_count\": len(clean_errors),\n            \"same_labels_rounded_6dp\": (\n                canonical_boxes(source_boxes) == canonical_boxes(clean_boxes)\n            ),\n        })\n\n    label_audit = pd.DataFrame(label_audit_records)\n    label_comparison = pd.DataFrame(label_comparison_records)\n    label_audit.to_csv(REPORT_ROOT / \"label_audit.csv\", index=False)\n    label_comparison.to_csv(\n        REPORT_ROOT / \"source_vs_clean_label_comparison.csv\", index=False\n    )\n    source_label_errors = label_audit.loc[\n        label_audit[\"version\"].eq(\"1280_preprocessing_source\")\n        & (label_audit[\"error_count\"] > 0)\n    ].copy()\n    clean_label_errors = label_audit.loc[\n        label_audit[\"version\"].eq(\"final_v6\")\n        & (label_audit[\"error_count\"] > 0)\n    ].copy()\n    clean_label_count_mismatch = label_comparison.loc[\n        label_comparison[\"manifest_label_count\"]\n        != label_comparison[\"clean_box_count\"]\n    ].copy()\n    source_label_errors.to_csv(\n        REPORT_ROOT / \"source_label_errors.csv\", index=False\n    )\n    clean_label_errors.to_csv(\n        REPORT_ROOT / \"clean_label_errors.csv\", index=False\n    )\n    clean_label_count_mismatch.to_csv(\n        REPORT_ROOT / \"clean_label_count_mismatch.csv\", index=False\n    )\n\n    print(\"Invalid source label files:\", len(source_label_errors))\n    print(\"Invalid rebuilt label files:\", len(clean_label_errors))\n    print(\"Clean label-count mismatch:\", len(clean_label_count_mismatch))\n    print(\n        \"Labels changed by cleaning:\",\n        int((~label_comparison[\"same_labels_rounded_6dp\"]).sum()),\n    )\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 10. 对最终导出 PNG 再做全量解码、文件哈希和像素哈希审计\n    def audit_exported_image(row):\n        path = Path(row.image_path)\n        baseline_hash = normalized_optional_hash(\n            row.source_baseline_sha256\n        )\n        record = {\n            \"image_id\": str(row.image_id),\n            \"split\": str(row.split),\n            \"path\": str(path),\n            \"source_path\": str(row.source_image_path),\n            \"readable\": False,\n            \"format_is_png\": False,\n            \"size_matches_manifest\": False,\n            \"mode\": None,\n            \"is_constant\": None,\n            \"quality_review_flag\": None,\n            \"source_baseline_available\": baseline_hash is not None,\n            \"source_baseline_hash_matches\": None,\n            \"export_hash_matches_source\": None,\n            \"export_pixel_hash_matches_source\": None,\n            \"error\": \"\",\n        }\n        try:\n            with Image.open(path) as image:\n                image.verify()\n            with Image.open(path) as image:\n                record[\"format\"] = image.format\n                record[\"format_is_png\"] = (\n                    str(image.format).upper() == \"PNG\"\n                )\n                record[\"mode\"] = image.mode\n                record[\"width\"] = int(image.width)\n                record[\"height\"] = int(image.height)\n                record[\"size_matches_manifest\"] = (\n                    image.width == int(row.output_width)\n                    and image.height == int(row.output_height)\n                )\n                grayscale = image.convert(\"L\")\n                pixels = np.asarray(\n                    grayscale, dtype=np.uint8\n                ).copy()\n                extrema = grayscale.getextrema()\n                record[\"min_intensity\"] = int(extrema[0])\n                record[\"max_intensity\"] = int(extrema[1])\n                record[\"is_constant\"] = extrema[0] == extrema[1]\n                record[\"pixel_mean\"] = float(pixels.mean())\n                record[\"pixel_std\"] = float(pixels.std())\n                record[\"near_black_fraction\"] = float(\n                    (pixels <= 1).mean()\n                )\n                record[\"near_white_fraction\"] = float(\n                    (pixels >= 254).mean()\n                )\n                record[\"quality_review_flag\"] = bool(\n                    record[\"pixel_std\"]\n                    < LOW_CONTRAST_STD_REVIEW_THRESHOLD\n                    or record[\"near_black_fraction\"]\n                    > EXTREME_SATURATION_REVIEW_FRACTION\n                    or record[\"near_white_fraction\"]\n                    > EXTREME_SATURATION_REVIEW_FRACTION\n                )\n                export_pixel_hash = decoded_pixel_sha256(pixels)\n                record[\"export_pixel_sha256\"] = export_pixel_hash\n                record[\"export_pixel_hash_matches_source\"] = (\n                    export_pixel_hash == row.source_pixel_sha256\n                )\n                record[\"export_perceptual_dhash64\"] = (\n                    perceptual_dhash64(grayscale)\n                )\n                grayscale.close()\n                record[\"readable\"] = True\n\n            export_hash = sha256_file(path)\n            record[\"export_sha256\"] = export_hash\n            record[\"source_sha256\"] = row.source_sha256_preexport\n            record[\"source_pixel_sha256\"] = row.source_pixel_sha256\n            record[\"export_hash_matches_source\"] = (\n                export_hash == row.source_sha256_preexport\n            )\n            if baseline_hash is not None:\n                record[\"source_baseline_hash_matches\"] = (\n                    export_hash == baseline_hash\n                )\n        except Exception as error:\n            record[\"error\"] = f\"{type(error).__name__}: {error}\"\n        return record\n\n\n    if AUDIT_MODE == \"full\":\n        image_audit_source = manifest_clean.copy()\n    else:\n        image_audit_source = pd.concat(\n            [\n                group.sample(\n                    n=min(SAMPLE_IMAGES_PER_SPLIT, len(group)),\n                    random_state=SEED,\n                )\n                for _, group in manifest_clean.groupby(\"split\")\n            ],\n            ignore_index=True,\n        )\n\n    audit_rows = list(image_audit_source.itertuples(index=False))\n    with ThreadPoolExecutor(max_workers=AUDIT_WORKERS) as executor:\n        image_audit_records = list(\n            tqdm(\n                executor.map(audit_exported_image, audit_rows),\n                total=len(audit_rows),\n                desc=f\"Auditing exported PNGs ({AUDIT_MODE})\",\n            )\n        )\n    image_audit = pd.DataFrame(image_audit_records)\n\n    if REQUIRE_GRAYSCALE:\n        image_audit[\"mode_is_acceptable\"] = image_audit[\"mode\"].eq(\"L\")\n    else:\n        image_audit[\"mode_is_acceptable\"] = image_audit[\"mode\"].isin(\n            [\"L\", \"RGB\"]\n        )\n\n    image_audit[\"hash_passes\"] = (\n        image_audit[\"export_hash_matches_source\"].eq(True)\n        & image_audit[\"export_pixel_hash_matches_source\"].eq(True)\n        & (\n            ~image_audit[\"source_baseline_available\"]\n            | image_audit[\"source_baseline_hash_matches\"].eq(True)\n        )\n    )\n    image_audit[\"passes\"] = (\n        image_audit[\"readable\"]\n        & image_audit[\"format_is_png\"]\n        & image_audit[\"size_matches_manifest\"]\n        & image_audit[\"mode_is_acceptable\"]\n        & ~image_audit[\"is_constant\"].fillna(True)\n        & image_audit[\"hash_passes\"]\n    )\n    failed_image_audit = image_audit.loc[\n        ~image_audit[\"passes\"]\n    ].copy()\n    image_quality_review = image_audit.loc[\n        image_audit[\"quality_review_flag\"].eq(True)\n    ].copy()\n    image_audit.to_csv(\n        REPORT_ROOT / \"image_audit.csv\", index=False\n    )\n    failed_image_audit.to_csv(\n        REPORT_ROOT / \"failed_image_audit.csv\", index=False\n    )\n    image_quality_review.to_csv(\n        REPORT_ROOT / \"image_quality_review.csv\", index=False\n    )\n\n    hash_columns = [\n        \"image_id\", \"split\", \"source_sha256\", \"export_sha256\",\n        \"source_pixel_sha256\", \"export_pixel_sha256\",\n        \"source_baseline_available\", \"source_baseline_hash_matches\",\n        \"export_hash_matches_source\",\n        \"export_pixel_hash_matches_source\",\n    ]\n    image_hash_manifest = image_audit[\n        [column for column in hash_columns if column in image_audit.columns]\n    ].copy()\n    image_hash_manifest.to_csv(\n        REPORT_ROOT / \"image_hash_manifest.csv\", index=False\n    )\n\n    if {\n        \"export_sha256\", \"export_pixel_sha256\"\n    }.issubset(image_audit.columns):\n        manifest_clean = manifest_clean.merge(\n            image_audit[\n                [\n                    \"image_id\", \"export_sha256\",\n                    \"export_pixel_sha256\",\n                ]\n            ],\n            on=\"image_id\",\n            how=\"left\",\n            validate=\"one_to_one\",\n        )\n        manifest_clean.to_csv(\n            DATASET_ROOT / \"manifest.csv\", index=False\n        )\n\n    full_image_audit_completed = (\n        AUDIT_MODE == \"full\"\n        and len(image_audit) == len(manifest_clean)\n        and image_audit[\"image_id\"].nunique() == len(manifest_clean)\n        and set(image_audit[\"image_id\"])\n        == set(manifest_clean[\"image_id\"])\n    )\n    source_baseline_hash_available = bool(\n        image_audit[\"source_baseline_available\"].any()\n    )\n    baseline_hash_mismatches = image_audit.loc[\n        image_audit[\"source_baseline_available\"]\n        & ~image_audit[\"source_baseline_hash_matches\"].eq(True)\n    ].copy()\n    export_hash_mismatches = image_audit.loc[\n        ~image_audit[\"export_hash_matches_source\"].eq(True)\n    ].copy()\n    export_pixel_hash_mismatches = image_audit.loc[\n        ~image_audit[\"export_pixel_hash_matches_source\"].eq(True)\n    ].copy()\n\n    duplicate_content_rows = pd.DataFrame()\n    content_hash_split_leakage = pd.DataFrame()\n    same_split_duplicate_content = pd.DataFrame()\n    if (\n        full_image_audit_completed\n        and \"export_pixel_sha256\" in image_audit.columns\n    ):\n        pixel_group_stats = (\n            image_audit.groupby(\"export_pixel_sha256\")\n            .agg(\n                image_count=(\"image_id\", \"nunique\"),\n                split_count=(\"split\", \"nunique\"),\n            )\n            .reset_index()\n        )\n        duplicate_hashes = set(\n            pixel_group_stats.loc[\n                pixel_group_stats[\"image_count\"] > 1,\n                \"export_pixel_sha256\",\n            ]\n        )\n        cross_split_hashes = set(\n            pixel_group_stats.loc[\n                (pixel_group_stats[\"image_count\"] > 1)\n                & (pixel_group_stats[\"split_count\"] > 1),\n                \"export_pixel_sha256\",\n            ]\n        )\n        same_split_hashes = duplicate_hashes - cross_split_hashes\n        duplicate_content_rows = image_audit.loc[\n            image_audit[\"export_pixel_sha256\"].isin(\n                duplicate_hashes\n            )\n        ].sort_values(\n            [\"export_pixel_sha256\", \"split\", \"image_id\"]\n        )\n        content_hash_split_leakage = image_audit.loc[\n            image_audit[\"export_pixel_sha256\"].isin(\n                cross_split_hashes\n            )\n        ].sort_values(\n            [\"export_pixel_sha256\", \"split\", \"image_id\"]\n        )\n        same_split_duplicate_content = image_audit.loc[\n            image_audit[\"export_pixel_sha256\"].isin(\n                same_split_hashes\n            )\n        ].sort_values(\n            [\"export_pixel_sha256\", \"split\", \"image_id\"]\n        )\n\n    duplicate_content_rows.to_csv(\n        REPORT_ROOT / \"duplicate_image_content.csv\", index=False\n    )\n    content_hash_split_leakage.to_csv(\n        REPORT_ROOT / \"content_hash_split_leakage.csv\", index=False\n    )\n    same_split_duplicate_content.to_csv(\n        REPORT_ROOT / \"same_split_duplicate_content_review.csv\",\n        index=False,\n    )\n\n    print(\"Audit mode:\", AUDIT_MODE)\n    print(\"Audited PNGs:\", len(image_audit), \"/\", len(manifest_clean))\n    print(\"Failed image audit:\", len(failed_image_audit))\n    print(\"Non-blocking image-quality review rows:\",\n          len(image_quality_review))\n    print(\"Full image audit completed:\", full_image_audit_completed)\n    print(\"Source preprocessing hash baseline available:\",\n          source_baseline_hash_available)\n    print(\"Export file-hash mismatches:\",\n          len(export_hash_mismatches))\n    print(\"Export pixel-hash mismatches:\",\n          len(export_pixel_hash_mismatches))\n    print(\"Residual duplicate-pixel rows:\",\n          len(duplicate_content_rows))\n    print(\"Cross-split duplicate-pixel rows:\",\n          len(content_hash_split_leakage))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 11. 校验 split plan、分组泄漏和可用的患者级信息\n    split_plan_all = pd.read_csv(source_split_plan_path)\n    required_split_plan_columns = {\"image_id\", \"split\", \"group_id\"}\n    missing_split_plan_columns = (\n        required_split_plan_columns - set(split_plan_all.columns)\n    )\n    if missing_split_plan_columns:\n        raise ValueError(\n            f\"split_plan.csv 缺少字段：\"\n            f\"{sorted(missing_split_plan_columns)}\"\n        )\n\n    split_plan_all[\"image_id\"] = (\n        split_plan_all[\"image_id\"].astype(str).str.strip()\n    )\n    split_plan_all[\"split\"] = (\n        split_plan_all[\"split\"].astype(str).str.strip().str.lower()\n    )\n    split_plan_all[\"group_id\"] = (\n        split_plan_all[\"group_id\"].astype(str).str.strip()\n    )\n    invalid_id_tokens = {\"\", \"nan\", \"none\", \"null\"}\n    invalid_split_plan_image_ids = split_plan_all.loc[\n        split_plan_all[\"image_id\"].str.lower().isin(\n            invalid_id_tokens\n        )\n    ].copy()\n    duplicate_split_plan_rows = split_plan_all.loc[\n        split_plan_all.duplicated(\"image_id\", keep=False)\n    ].copy()\n    split_plan_unique = split_plan_all.loc[\n        ~split_plan_all.duplicated(\"image_id\", keep=False)\n    ].copy()\n    invalid_split_plan_splits = split_plan_all.loc[\n        ~split_plan_all[\"split\"].isin(VALID_SPLITS)\n    ].copy()\n    invalid_group_rows_all = split_plan_all.loc[\n        split_plan_all[\"group_id\"].str.lower().isin(\n            invalid_id_tokens\n        )\n    ].copy()\n\n    source_manifest_id_set = set(manifest[\"image_id\"])\n    source_split_plan_id_set = set(split_plan_all[\"image_id\"])\n    missing_source_split_plan_ids = sorted(\n        source_manifest_id_set - source_split_plan_id_set\n    )\n    extra_split_plan_ids = sorted(\n        source_split_plan_id_set - source_manifest_id_set\n    )\n    source_split_plan_comparison = manifest[\n        [\"image_id\", \"split\"]\n    ].merge(\n        split_plan_unique[[\"image_id\", \"split\"]].rename(\n            columns={\"split\": \"planned_split\"}\n        ),\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n    source_split_plan_mismatch = source_split_plan_comparison.loc[\n        source_split_plan_comparison[\"split\"]\n        != source_split_plan_comparison[\"planned_split\"]\n    ].copy()\n\n    split_plan = split_plan_unique.loc[\n        split_plan_unique[\"image_id\"].isin(manifest_clean[\"image_id\"])\n    ].copy()\n    missing_split_plan_ids = sorted(\n        set(manifest_clean[\"image_id\"]) - set(split_plan[\"image_id\"])\n    )\n    split_plan_comparison = manifest_clean[\n        [\"image_id\", \"split\"]\n    ].merge(\n        split_plan[[\"image_id\", \"split\"]].rename(\n            columns={\"split\": \"planned_split\"}\n        ),\n        on=\"image_id\",\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n    split_plan_mismatch = split_plan_comparison.loc[\n        split_plan_comparison[\"split\"]\n        != split_plan_comparison[\"planned_split\"]\n    ].copy()\n\n    group_split_counts = (\n        split_plan.groupby(\"group_id\")[\"split\"]\n        .agg(\n            split_count=\"nunique\",\n            splits=lambda values: \"|\".join(\n                sorted(set(values))\n            ),\n            image_count=\"size\",\n        )\n        .reset_index()\n    )\n    group_split_leakage = group_split_counts.loc[\n        group_split_counts[\"split_count\"] > 1\n    ].copy()\n    group_size_distribution = (\n        split_plan.groupby(\"group_id\").size()\n        .value_counts()\n        .sort_index()\n        .rename_axis(\"images_per_group\")\n        .reset_index(name=\"group_count\")\n    )\n    group_id_equals_image_id_fraction = float(\n        (\n            split_plan[\"group_id\"].astype(str)\n            == split_plan[\"image_id\"].astype(str)\n        ).mean()\n    ) if len(split_plan) else 0.0\n    group_id_semantics = (\n        source_config.get(\"split_group_semantics\")\n        or source_config.get(\"group_id_source\")\n        or \"source_split_plan_group_id\"\n    )\n\n    patient_id_column = next(\n        (\n            column for column in (\n                \"patient_id\", \"PatientID\", \"patient_hash\"\n            )\n            if column in split_plan.columns\n        ),\n        None,\n    )\n    patient_split_leakage = pd.DataFrame()\n    if patient_id_column is not None:\n        patient_values = (\n            split_plan[patient_id_column].astype(str).str.strip()\n        )\n        invalid_patient_rows = split_plan.loc[\n            patient_values.str.lower().isin(invalid_id_tokens)\n        ].copy()\n        patient_split_counts = (\n            split_plan.assign(_patient_id=patient_values)\n            .groupby(\"_patient_id\")[\"split\"]\n            .agg(\n                split_count=\"nunique\",\n                splits=lambda values: \"|\".join(\n                    sorted(set(values))\n                ),\n            )\n            .reset_index()\n        )\n        patient_split_leakage = patient_split_counts.loc[\n            patient_split_counts[\"split_count\"] > 1\n        ].copy()\n    else:\n        invalid_patient_rows = pd.DataFrame()\n\n    invalid_split_plan_image_ids.to_csv(\n        REPORT_ROOT / \"invalid_split_plan_image_ids.csv\",\n        index=False,\n    )\n    duplicate_split_plan_rows.to_csv(\n        REPORT_ROOT / \"duplicate_split_plan_rows.csv\", index=False\n    )\n    invalid_split_plan_splits.to_csv(\n        REPORT_ROOT / \"invalid_split_plan_splits.csv\", index=False\n    )\n    invalid_group_rows_all.to_csv(\n        REPORT_ROOT / \"invalid_split_group_ids.csv\", index=False\n    )\n    pd.DataFrame({\n        \"image_id\": missing_source_split_plan_ids\n    }).to_csv(\n        REPORT_ROOT / \"missing_source_split_plan_ids.csv\",\n        index=False,\n    )\n    pd.DataFrame({\n        \"image_id\": extra_split_plan_ids\n    }).to_csv(\n        REPORT_ROOT / \"extra_split_plan_ids.csv\", index=False\n    )\n    source_split_plan_mismatch.to_csv(\n        REPORT_ROOT / \"source_split_plan_mismatch.csv\",\n        index=False,\n    )\n    pd.DataFrame({\n        \"image_id\": missing_split_plan_ids\n    }).to_csv(\n        REPORT_ROOT / \"missing_clean_split_plan_ids.csv\",\n        index=False,\n    )\n    split_plan_mismatch.to_csv(\n        REPORT_ROOT / \"clean_split_plan_mismatch.csv\",\n        index=False,\n    )\n    group_split_leakage.to_csv(\n        REPORT_ROOT / \"group_split_leakage.csv\", index=False\n    )\n    group_size_distribution.to_csv(\n        REPORT_ROOT / \"group_size_distribution.csv\", index=False\n    )\n    patient_split_leakage.to_csv(\n        REPORT_ROOT / \"patient_split_leakage.csv\", index=False\n    )\n    invalid_patient_rows.to_csv(\n        REPORT_ROOT / \"invalid_patient_ids.csv\", index=False\n    )\n    (REPORT_ROOT / \"group_id_semantics.json\").write_text(\n        json.dumps({\n            \"group_id_semantics\": group_id_semantics,\n            \"patient_id_column_available\": patient_id_column,\n            \"group_id_equals_image_id_fraction\": (\n                group_id_equals_image_id_fraction\n            ),\n            \"group_count\": int(\n                split_plan[\"group_id\"].nunique()\n            ),\n        }, ensure_ascii=False, indent=2),\n        encoding=\"utf-8\",\n    )\n\n    clean_split_plan_path = DATASET_ROOT / \"split_plan.csv\"\n    split_plan = split_plan.sort_values(\n        [\"split\", \"image_id\"], kind=\"stable\"\n    ).reset_index(drop=True)\n    split_plan.to_csv(clean_split_plan_path, index=False)\n    split_plan.to_csv(\n        REPORT_ROOT / \"clean_split_plan.csv\", index=False\n    )\n    clean_split_plan_written = (\n        clean_split_plan_path.is_file()\n        and len(split_plan) == len(manifest_clean)\n        and split_plan[\"image_id\"].nunique() == len(manifest_clean)\n        and set(split_plan[\"image_id\"])\n        == set(manifest_clean[\"image_id\"])\n    )\n\n    print(\"Image-ID split leakage:\", len(image_split_leakage))\n    print(\"Group-level split leakage:\", len(group_split_leakage))\n    print(\"Patient ID column available:\", patient_id_column)\n    print(\"Patient-level split leakage rows:\",\n          len(patient_split_leakage))\n    print(\"Missing source/clean split-plan IDs:\",\n          len(missing_source_split_plan_ids),\n          len(missing_split_plan_ids))\n    print(\"Duplicate split-plan rows:\",\n          len(duplicate_split_plan_rows))\n    print(\"Source/clean split-plan mismatches:\",\n          len(source_split_plan_mismatch),\n          len(split_plan_mismatch))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    changed_label_images = int(\n        (~label_comparison[\"same_labels_rounded_6dp\"]).sum()\n    )\n    retained_clipped_boxes = int(\n        box_issues[\"action\"].eq(\"clipped_for_review\").sum()\n    )\n    invalid_auto_excluded_images_retained = (\n        set(automatic_source_image_exclusions[\"image_id\"])\n        & set(manifest_clean[\"image_id\"])\n    )\n\n    acceptance_checks = {\n        \"source_lock_verified\": (\n            source_output_root.name\n            == EXPECTED_SOURCE_OUTPUT_DIRNAME\n            and signature_matches\n            and computed_signature\n            == EXPECTED_SOURCE_DATASET_SIGNATURE\n            and (\n                source_notebook_hint_matches\n                or not REQUIRE_NOTEBOOK_NAME_HINT\n            )\n            and source_config.get(\"target_class\") == TARGET_CLASS\n        ),\n        \"raw_dicom_read_is_false\": True,\n        \"raw_train_csv_read_is_false\": True,\n        \"no_dicom_inside_selected_dataset\": (\n            len(dicom_inside_selected_dataset) == 0\n        ),\n        \"source_count_contract_passed\": source_count_contract_pass,\n        \"source_yaml_class_matches\": source_yaml_class_matches,\n        \"source_manifest_box_contract_passed\": (\n            source_manifest_box_contract_pass\n        ),\n        \"geometry_manifest_contract_passed\": (\n            geometry_manifest_contract_pass\n        ),\n        \"algebraic_geometry_check_passed\": geometry_roundtrip_pass,\n        \"serialized_yolo_geometry_passed\": serialized_geometry_pass,\n        \"radiologist_metadata_valid\": radiologist_metadata_valid,\n        \"no_duplicate_manifest_image_ids\": (\n            len(duplicate_manifest_rows) == 0\n        ),\n        \"valid_manifest_image_ids\": (\n            len(invalid_manifest_image_ids) == 0\n        ),\n        \"valid_manifest_status_and_dimensions\": (\n            len(invalid_status_rows) == 0\n            and len(invalid_dimension_rows) == 0\n        ),\n        \"no_missing_preprocessed_png_or_label\": (\n            len(missing_source_files) == 0\n        ),\n        \"no_extra_preprocessed_images\": (\n            len(extra_preprocessed_images) == 0\n        ),\n        \"no_extra_preprocessed_labels\": (\n            len(extra_preprocessed_labels) == 0\n        ),\n        \"no_unexpected_source_files\": (\n            len(unexpected_source_files) == 0\n        ),\n        \"preexport_source_image_audit_complete\": (\n            preexport_source_audit_complete\n        ),\n        \"kept_source_file_and_pixel_hashes_complete\": (\n            kept_source_hashes_complete\n        ),\n        \"no_source_provenance_hash_failures\": (\n            len(source_provenance_failures) == 0\n        ),\n        \"automatic_invalid_image_exclusion_within_limit\": (\n            automatic_source_image_exclusion_within_limit\n        ),\n        \"no_auto_excluded_images_retained\": (\n            len(invalid_auto_excluded_images_retained) == 0\n        ),\n        \"automatic_exact_pixel_dedup_enabled\": (\n            AUTO_DEDUPLICATE_EXACT_CONTENT\n        ),\n        \"no_unresolved_duplicate_label_conflicts\": (\n            no_duplicate_label_conflicts\n        ),\n        \"no_positive_images_missing_manifest\": (\n            len(positive_ids_missing_manifest) == 0\n        ),\n        \"source_positive_status_matches_rebuilt_boxes\": (\n            len(source_positive_mismatch) == 0\n        ),\n        \"positive_and_negative_images_present\": (\n            class_distribution_valid\n        ),\n        \"all_splits_have_positive_and_negative_images\": (\n            all_splits_have_both_classes\n        ),\n        \"ambiguous_positive_exclusion_within_limit\": (\n            ambiguous_positive_fraction\n            <= MAX_AMBIGUOUS_POSITIVE_FRACTION\n        ),\n        \"disk_preflight_passed\": disk_preflight_pass,\n        \"no_export_failures\": len(export_failures_df) == 0,\n        \"all_usable_images_exported\": (\n            len(manifest_clean) == len(usable_manifest)\n        ),\n        \"no_missing_output_images\": (\n            len(missing_output_images) == 0\n        ),\n        \"no_missing_output_labels\": (\n            len(missing_output_labels) == 0\n        ),\n        \"no_extra_output_images\": len(extra_output_images) == 0,\n        \"no_extra_output_labels\": len(extra_output_labels) == 0,\n        \"no_unexpected_output_directories\": (\n            len(unexpected_output_directories_df) == 0\n        ),\n        \"no_unexpected_output_entries\": (\n            len(postexport_unexpected_entries_df) == 0\n        ),\n        \"portable_data_yaml_resolves\": (\n            IMAGE_EXPORT_MODE == \"none\" or yaml_paths_exist\n        ),\n        \"portable_data_yaml_class_matches\": (\n            IMAGE_EXPORT_MODE == \"none\" or yaml_class_matches\n        ),\n        \"no_invalid_rebuilt_labels\": len(clean_label_errors) == 0,\n        \"rebuilt_label_counts_match_manifest\": (\n            len(clean_label_count_mismatch) == 0\n        ),\n        \"full_image_audit_completed\": full_image_audit_completed,\n        \"image_audit_passed\": len(failed_image_audit) == 0,\n        \"export_file_hashes_match_preexport_source\": (\n            len(export_hash_mismatches) == 0\n        ),\n        \"export_pixel_hashes_match_preexport_source\": (\n            len(export_pixel_hash_mismatches) == 0\n        ),\n        \"available_source_baseline_hashes_match\": (\n            len(baseline_hash_mismatches) == 0\n        ),\n        \"no_image_id_split_leakage\": len(image_split_leakage) == 0,\n        \"no_group_level_split_leakage\": (\n            len(group_split_leakage) == 0\n        ),\n        \"available_patient_ids_do_not_leak\": (\n            patient_id_column is None\n            or (\n                len(invalid_patient_rows) == 0\n                and len(patient_split_leakage) == 0\n            )\n        ),\n        \"source_split_plan_complete\": (\n            len(missing_source_split_plan_ids) == 0\n        ),\n        \"no_extra_source_split_plan_ids\": (\n            len(extra_split_plan_ids) == 0\n        ),\n        \"split_plan_image_ids_unique\": (\n            len(duplicate_split_plan_rows) == 0\n        ),\n        \"split_plan_image_ids_valid\": (\n            len(invalid_split_plan_image_ids) == 0\n        ),\n        \"split_plan_splits_valid\": (\n            len(invalid_split_plan_splits) == 0\n        ),\n        \"split_group_ids_valid\": (\n            len(invalid_group_rows_all) == 0\n        ),\n        \"source_split_plan_matches_manifest\": (\n            len(source_split_plan_mismatch) == 0\n        ),\n        \"clean_split_plan_complete\": (\n            len(missing_split_plan_ids) == 0\n        ),\n        \"clean_split_plan_matches_manifest\": (\n            len(split_plan_mismatch) == 0\n        ),\n        \"clean_split_plan_written\": clean_split_plan_written,\n        \"no_cross_split_duplicate_pixels\": (\n            len(content_hash_split_leakage) == 0\n        ),\n        \"no_residual_duplicate_pixels\": (\n            len(duplicate_content_rows) == 0\n        ),\n        \"dataset_fingerprint_created\": (\n            isinstance(CLEAN_DATASET_FINGERPRINT, str)\n            and len(CLEAN_DATASET_FINGERPRINT) == 64\n        ),\n    }\n    acceptance_checks = {\n        name: bool(passed)\n        for name, passed in acceptance_checks.items()\n    }\n\n    all_checks_pass = all(acceptance_checks.values())\n    checks_except_full_audit = {\n        name: passed\n        for name, passed in acceptance_checks.items()\n        if name != \"full_image_audit_completed\"\n    }\n    if all_checks_pass and IMAGE_EXPORT_MODE == \"copy\":\n        final_status = \"TRAINING_READY\"\n    elif all_checks_pass and IMAGE_EXPORT_MODE == \"symlink\":\n        final_status = \"SESSION_ONLY\"\n    elif all_checks_pass and IMAGE_EXPORT_MODE == \"none\":\n        final_status = \"REPORT_ONLY\"\n    elif (\n        AUDIT_MODE == \"sample\"\n        and all(checks_except_full_audit.values())\n    ):\n        final_status = \"SAMPLE_ONLY\"\n    else:\n        final_status = \"FAIL\"\n\n    training_ready = final_status == \"TRAINING_READY\"\n    manual_review_required = bool(\n        len(manual_review_queue) > 0\n        or retained_clipped_boxes > 0\n        or len(low_consensus_boxes) > 0\n        or len(image_quality_review) > 0\n        or len(cross_split_perceptual_hash_review) > 0\n    )\n\n    action_for_check = {\n        \"source_count_contract_passed\": (\n            \"确认添加的是完整 21e0ec1e3506 1280 输出；\"\n            \"源清单应为 15000 张且 split 为 12000/1500/1500。\"\n        ),\n        \"source_yaml_class_matches\": (\n            \"源 data.yaml 不是单类 Nodule/Mass，禁止混用。\"\n        ),\n        \"source_manifest_box_contract_passed\": (\n            \"源 manifest 的 label_count/is_positive 与融合框数量不一致。\"\n        ),\n        \"geometry_manifest_contract_passed\": (\n            \"确认输入确实来自 1280_preprocessing，且未做未记录的裁剪/补边。\"\n        ),\n        \"serialized_yolo_geometry_passed\": (\n            \"停止训练，核对标签小数序列化、坐标顺序和 PNG 尺寸。\"\n        ),\n        \"radiologist_metadata_valid\": (\n            \"radiologist_count 应是 1–3 的整数；核对融合框报告。\"\n        ),\n        \"preexport_source_image_audit_complete\": (\n            \"源图未完成全量解码审计，重新 Run All。\"\n        ),\n        \"no_source_provenance_hash_failures\": (\n            \"源 PNG 与已有预处理哈希不一致，重新保存可信的 1280_preprocessing。\"\n        ),\n        \"automatic_invalid_image_exclusion_within_limit\": (\n            \"坏图比例过高，说明预处理可能整体异常；\"\n            \"查看 automatic_source_image_exclusions.csv。\"\n        ),\n        \"no_unresolved_duplicate_label_conflicts\": (\n            \"查看 duplicate_label_conflicts_quarantined.csv；\"\n            \"同像素不同标签必须由合格人员解决后再训练。\"\n        ),\n        \"all_splits_have_positive_and_negative_images\": (\n            \"回到 1280_preprocessing 重新划分，保证三份 split 都有阳性和阴性。\"\n        ),\n        \"ambiguous_positive_exclusion_within_limit\": (\n            \"异常阳性排除比例过高，先复核或重新标框。\"\n        ),\n        \"disk_preflight_passed\": (\n            \"清理 /kaggle/working；正式保存仍使用 copy。\"\n        ),\n        \"no_export_failures\": (\n            \"查看 export_failures.csv，修复后重新 Run All。\"\n        ),\n        \"full_image_audit_completed\": (\n            \"把 AUDIT_MODE 改为 full 并 Run All；抽样不能正式训练。\"\n        ),\n        \"image_audit_passed\": (\n            \"查看 failed_image_audit.csv；最终复制件未通过解码或尺寸检查。\"\n        ),\n        \"export_file_hashes_match_preexport_source\": (\n            \"导出字节与预导出源图不同，删除对应 V6 复制件后重跑。\"\n        ),\n        \"export_pixel_hashes_match_preexport_source\": (\n            \"导出解码像素与源图不同，禁止训练。\"\n        ),\n        \"no_group_level_split_leakage\": (\n            \"回到 1280_preprocessing 按 source group_id 重新划分。\"\n        ),\n        \"available_patient_ids_do_not_leak\": (\n            \"发现患者 ID 跨 split；必须在预处理阶段重新分组划分。\"\n        ),\n        \"source_split_plan_complete\": (\n            \"修复 1280_preprocessing 的 split_plan.csv；不能自行补猜分组。\"\n        ),\n        \"source_split_plan_matches_manifest\": (\n            \"源 manifest 与 split_plan 的 split 不一致。\"\n        ),\n        \"no_cross_split_duplicate_pixels\": (\n            \"自动去重后仍有跨 split 同像素内容，禁止训练。\"\n        ),\n        \"no_residual_duplicate_pixels\": (\n            \"自动去重后仍有同像素副本，查看 duplicate_image_content.csv。\"\n        ),\n        \"dataset_fingerprint_created\": (\n            \"未生成稳定数据指纹，禁止训练 Notebook 接入。\"\n        ),\n    }\n    action_rows = []\n    for check_name, passed in acceptance_checks.items():\n        if not passed:\n            action_rows.append({\n                \"priority\": \"BLOCKER\",\n                \"item\": check_name,\n                \"action\": action_for_check.get(\n                    check_name,\n                    \"查看对应 reports 文件，修复来源后重新 Run All。\",\n                ),\n            })\n    if IMAGE_EXPORT_MODE != \"copy\":\n        action_rows.append({\n            \"priority\": \"BLOCKER\",\n            \"item\": \"non_portable_export\",\n            \"action\": \"正式保存/共享必须使用 IMAGE_EXPORT_MODE='copy'。\",\n        })\n    if len(automatic_source_image_exclusions):\n        action_rows.append({\n            \"priority\": \"NOTE\",\n            \"item\": \"invalid_source_images_auto_excluded\",\n            \"action\": (\n                f\"V6 已自动隔离 {len(automatic_source_image_exclusions)} \"\n                \"张损坏或违反 PNG 契约的图片；源文件未修改。\"\n            ),\n        })\n    if len(auto_duplicate_exclusions):\n        action_rows.append({\n            \"priority\": \"NOTE\",\n            \"item\": \"exact_pixel_duplicates_auto_excluded\",\n            \"action\": (\n                f\"V6 已自动排除 {len(auto_duplicate_exclusions)} \"\n                \"个像素和标签均相同的副本；源文件未修改。\"\n            ),\n        })\n    if len(duplicate_label_conflicts):\n        action_rows.append({\n            \"priority\": \"BLOCKER\",\n            \"item\": \"duplicate_label_conflicts_quarantined\",\n            \"action\": (\n                \"冲突组已全部隔离，但 V6 会保持 FAIL，\"\n                \"直到合格人员解决标签冲突。\"\n            ),\n        })\n    if len(manual_review_queue):\n        action_rows.append({\n            \"priority\": \"REVIEW\",\n            \"item\": \"manual_review_queue\",\n            \"action\": (\n                \"查看 manual_review_queue.csv；被隔离的阳性图\"\n                \"不能自动改成负样本。\"\n            ),\n        })\n    if retained_clipped_boxes:\n        action_rows.append({\n            \"priority\": \"REVIEW\",\n            \"item\": \"retained_clipped_boxes\",\n            \"action\": \"检查 clipped_box_review.csv 和最终框可视化。\",\n        })\n    if len(low_consensus_boxes):\n        action_rows.append({\n            \"priority\": \"REVIEW\",\n            \"item\": \"single_radiologist_boxes\",\n            \"action\": (\n                \"这是敏感性优先策略保留的低共识框；\"\n                \"报告时不要当作多医生共识真值。\"\n            ),\n        })\n    if len(image_quality_review):\n        action_rows.append({\n            \"priority\": \"REVIEW\",\n            \"item\": \"image_quality_outliers\",\n            \"action\": (\n                \"查看 image_quality_review.csv；\"\n                \"亮度统计异常只复核，不武断删除医学图像。\"\n            ),\n        })\n    if len(cross_split_perceptual_hash_review):\n        action_rows.append({\n            \"priority\": \"REVIEW\",\n            \"item\": \"perceptual_near_duplicate_candidates\",\n            \"action\": (\n                \"查看 cross_split_perceptual_hash_review.csv；\"\n                \"感知哈希仅作疑似近重复提示，不自动删除。\"\n            ),\n        })\n    if changed_label_images:\n        action_rows.append({\n            \"priority\": \"REQUIRED\",\n            \"item\": \"labels_changed\",\n            \"action\": (\n                \"正式实验必须使用 V6 data.yaml 从新权重重新训练；\"\n                \"旧 best.pt 不代表修正后的标签。\"\n            ),\n        })\n    if not source_baseline_hash_available:\n        action_rows.append({\n            \"priority\": \"NOTE\",\n            \"item\": \"no_preprocessing_hash_baseline\",\n            \"action\": (\n                \"V6 已为当前源输出建立完整指纹；\"\n                \"下次预处理应把 image_sha256 直接写入 manifest。\"\n            ),\n        })\n    if patient_id_column is None:\n        action_rows.append({\n            \"priority\": \"NOTE\",\n            \"item\": \"patient_id_not_explicitly_available\",\n            \"action\": (\n                \"V6 已检查 source group_id；报告中不要把它\"\n                \"无条件写成患者 ID，除非 1280_preprocessing 明确记录其来源。\"\n            ),\n        })\n    if training_ready:\n        action_rows.append({\n            \"priority\": \"NEXT\",\n            \"item\": \"start_training\",\n            \"action\": (\n                \"Save Version；训练 Notebook 读取 dataset/data.yaml，\"\n                \"并核对 CLEAN_DATASET_FINGERPRINT 后从新权重训练。\"\n            ),\n        })\n    action_plan = pd.DataFrame(\n        action_rows,\n        columns=[\"priority\", \"item\", \"action\"],\n    )\n    action_plan.to_csv(\n        REPORT_ROOT / \"action_plan.csv\", index=False\n    )\n\n    summary = pd.DataFrame([\n        {\"metric\": \"Final status\", \"value\": final_status},\n        {\"metric\": \"Training ready\", \"value\": training_ready},\n        {\"metric\": \"Input policy\",\n         \"value\": \"1280_preprocessing preprocessed output only\"},\n        {\"metric\": \"Source configuration signature\",\n         \"value\": computed_signature},\n        {\"metric\": \"Cleaning policy signature\",\n         \"value\": CLEANING_POLICY_SIGNATURE},\n        {\"metric\": \"Source dataset fingerprint\",\n         \"value\": SOURCE_DATASET_FINGERPRINT},\n        {\"metric\": \"Clean dataset fingerprint\",\n         \"value\": CLEAN_DATASET_FINGERPRINT},\n        {\"metric\": \"Source manifest images\", \"value\": len(manifest)},\n        {\"metric\": \"Automatically excluded invalid images\",\n         \"value\": len(automatic_source_image_exclusions)},\n        {\"metric\": \"Usable clean images\", \"value\": len(manifest_clean)},\n        {\"metric\": \"Duplicate decoded-pixel groups\",\n         \"value\": deduplication_summary[\"duplicate_content_groups\"]},\n        {\"metric\": \"Safe duplicate copies auto-excluded\",\n         \"value\": len(auto_duplicate_exclusions)},\n        {\"metric\": \"Duplicate label-conflict images quarantined\",\n         \"value\": len(duplicate_label_conflicts)},\n        {\"metric\": \"Source fused boxes\", \"value\": len(source_fused_boxes)},\n        {\"metric\": \"Clean fused boxes\", \"value\": len(clean_fused_boxes)},\n        {\"metric\": \"Excluded ambiguous positive images\",\n         \"value\": len(ambiguous_ids)},\n        {\"metric\": \"Maximum serialized geometry error (px)\",\n         \"value\": max_serialized_geometry_error},\n        {\"metric\": \"Single-radiologist boxes for review\",\n         \"value\": len(low_consensus_boxes)},\n        {\"metric\": \"Labels changed by cleaning\",\n         \"value\": changed_label_images},\n        {\"metric\": \"Audited final images\", \"value\": len(image_audit)},\n        {\"metric\": \"Failed final image audits\",\n         \"value\": len(failed_image_audit)},\n        {\"metric\": \"Cross-split perceptual review rows\",\n         \"value\": len(cross_split_perceptual_hash_review)},\n        {\"metric\": \"Manual review required\",\n         \"value\": manual_review_required},\n    ])\n    summary.to_csv(\n        REPORT_ROOT / \"cleaning_summary.csv\", index=False\n    )\n    display(summary)\n    display(action_plan)\n\n    RUN_COMPLETED_AT = datetime.now(timezone.utc).isoformat()\n    acceptance_report = {\n        \"overall_status\": final_status,\n        \"training_ready\": bool(training_ready),\n        \"manual_review_required\": bool(manual_review_required),\n        \"input_policy\": \"1280_preprocessing_preprocessed_output_only\",\n        \"cleaning_version\": \"6.0-release\",\n        \"gpu_used\": False,\n        \"raw_dicom_read\": False,\n        \"raw_train_csv_read\": False,\n        \"run_started_at\": RUN_STARTED_AT,\n        \"run_completed_at\": RUN_COMPLETED_AT,\n        \"source_dataset_root\": str(source_dataset_root),\n        \"source_dataset_signature\": computed_signature,\n        \"source_dataset_fingerprint\": SOURCE_DATASET_FINGERPRINT,\n        \"cleaning_policy_signature\": CLEANING_POLICY_SIGNATURE,\n        \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n        \"run_root\": str(RUN_ROOT),\n        \"dataset_root\": str(DATASET_ROOT),\n        \"data_yaml\": (\n            str(data_yaml_path) if data_yaml_path.exists() else None\n        ),\n        \"checks\": acceptance_checks,\n        \"counts\": {\n            \"source_manifest_images\": int(len(manifest)),\n            \"automatically_excluded_invalid_images\": int(\n                len(automatic_source_image_exclusions)\n            ),\n            \"usable_clean_images\": int(len(manifest_clean)),\n            \"duplicate_content_groups\": int(\n                deduplication_summary[\"duplicate_content_groups\"]\n            ),\n            \"safe_duplicate_copies_auto_excluded\": int(\n                len(auto_duplicate_exclusions)\n            ),\n            \"duplicate_label_conflict_images_quarantined\": int(\n                len(duplicate_label_conflicts)\n            ),\n            \"source_fused_boxes\": int(len(source_fused_boxes)),\n            \"clean_fused_boxes\": int(len(clean_fused_boxes)),\n            \"ambiguous_positive_images_excluded\": int(\n                len(ambiguous_ids)\n            ),\n            \"single_radiologist_boxes_for_review\": int(\n                len(low_consensus_boxes)\n            ),\n            \"changed_label_images\": int(changed_label_images),\n            \"failed_image_audits\": int(len(failed_image_audit)),\n            \"image_quality_review_rows\": int(\n                len(image_quality_review)\n            ),\n            \"cross_split_perceptual_hash_review_rows\": int(\n                len(cross_split_perceptual_hash_review)\n            ),\n        },\n    }\n    (REPORT_ROOT / \"cleaning_acceptance_final.json\").write_text(\n        json.dumps(\n            acceptance_report, ensure_ascii=False, indent=2\n        ),\n        encoding=\"utf-8\",\n    )\n    (REPORT_ROOT / \"TRAINING_GATE.txt\").write_text(\n        (\n            f\"STATUS={final_status}\\n\"\n            f\"TRAINING_READY={training_ready}\\n\"\n            f\"MANUAL_REVIEW_REQUIRED={manual_review_required}\\n\"\n            f\"CLEANING_VERSION=6.0-release\\n\"\n            f\"SOURCE_DATASET_SIGNATURE={computed_signature}\\n\"\n            f\"SOURCE_DATASET_FINGERPRINT=\"\n            f\"{SOURCE_DATASET_FINGERPRINT}\\n\"\n            f\"CLEAN_DATASET_FINGERPRINT=\"\n            f\"{CLEAN_DATASET_FINGERPRINT}\\n\"\n            f\"DATA_YAML=\"\n            f\"{data_yaml_path if data_yaml_path.exists() else ''}\\n\"\n            f\"RUN_STARTED_AT={RUN_STARTED_AT}\\n\"\n            f\"RUN_COMPLETED_AT={RUN_COMPLETED_AT}\\n\"\n        ),\n        encoding=\"utf-8\",\n    )\n\n    next_steps_lines = [\n        \"# V6 next steps\",\n        \"\",\n        f\"- Final status: `{final_status}`\",\n        f\"- Training ready: `{training_ready}`\",\n        f\"- Manual review required: `{manual_review_required}`\",\n        f\"- Clean dataset fingerprint: \"\n        f\"`{CLEAN_DATASET_FINGERPRINT}`\",\n        \"\",\n    ]\n    for row in action_plan.itertuples(index=False):\n        next_steps_lines.append(\n            f\"- [{row.priority}] `{row.item}`: {row.action}\"\n        )\n    (RUN_ROOT / \"README_NEXT_STEPS.md\").write_text(\n        \"\\n\".join(next_steps_lines) + \"\\n\",\n        encoding=\"utf-8\",\n    )\n\n    print(json.dumps(\n        acceptance_report, ensure_ascii=False, indent=2\n    ))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 13. 在实际预处理 PNG 上可视化最终重建标签\n    positive_preview = manifest_clean.loc[\n        manifest_clean[\"is_positive\"].eq(1)\n    ].sample(\n        n=min(\n            PREVIEW_POSITIVE_IMAGES,\n            int(manifest_clean[\"is_positive\"].sum()),\n        ),\n        random_state=SEED,\n    )\n\n    if len(positive_preview):\n        columns = min(3, len(positive_preview))\n        rows = math.ceil(len(positive_preview) / columns)\n        figure, axes = plt.subplots(rows, columns, figsize=(6 * columns, 7 * rows))\n        axes = np.atleast_1d(axes).reshape(-1)\n        for axis, row in zip(axes, positive_preview.itertuples(index=False)):\n            with Image.open(row.image_path) as source:\n                image = source.convert(\"RGB\")\n            draw = ImageDraw.Draw(image)\n            boxes, _ = parse_label(Path(row.label_path))\n            for _, x_center, y_center, width, height in boxes:\n                x1 = (x_center - width / 2) * image.width\n                y1 = (y_center - height / 2) * image.height\n                x2 = (x_center + width / 2) * image.width\n                y2 = (y_center + height / 2) * image.height\n                draw.rectangle(\n                    [x1, y1, x2, y2],\n                    outline=\"red\",\n                    width=max(2, image.width // 400),\n                )\n            axis.imshow(image)\n            axis.set_title(\n                f\"{row.split} | boxes={row.label_count}\\n\"\n                f\"{row.image_id[:20]}…\"\n            )\n            axis.axis(\"off\")\n        for axis in axes[len(positive_preview):]:\n            axis.axis(\"off\")\n        figure.suptitle(\"Final V6 YOLO labels on preprocessed PNGs\", fontsize=18)\n        figure.tight_layout()\n        preview_path = REPORT_ROOT / \"final_label_preview.png\"\n        figure.savefig(preview_path, dpi=160, bbox_inches=\"tight\")\n        plt.show()\n        plt.close(figure)\n        print(\"Preview:\", preview_path)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if NEED_CLEANING:\n    # 14. 最终结论：只有 TRAINING_READY 才允许进入正式训练\n    print(\"=\" * 92)\n    print(\"FINAL CLEANING STATUS:\",\n          acceptance_report[\"overall_status\"])\n    print(\"TRAINING READY:\",\n          acceptance_report[\"training_ready\"])\n    print(\"INPUT POLICY: 1280_preprocessing PREPROCESSED OUTPUT ONLY\")\n    print(\"Raw DICOM read: False\")\n    print(\"Raw train.csv read: False\")\n    print(\"Clean dataset fingerprint:\",\n          acceptance_report[\"clean_dataset_fingerprint\"])\n    print(\"Acceptance:\",\n          REPORT_ROOT / \"cleaning_acceptance_final.json\")\n    print(\"Action plan:\", REPORT_ROOT / \"action_plan.csv\")\n    print(\"Training gate:\", REPORT_ROOT / \"TRAINING_GATE.txt\")\n    if data_yaml_path.exists():\n        print(\"Training YAML:\", data_yaml_path)\n    print(\"=\" * 92)\n\n    if (\n        FAIL_ON_NOT_TRAINING_READY\n        and not acceptance_report[\"training_ready\"]\n    ):\n        failed_checks = [\n            name for name, passed in acceptance_checks.items()\n            if not passed\n        ]\n        raise RuntimeError(\n            \"最终训练闸门未通过，禁止开始正式训练。状态：\"\n            f\"{acceptance_report['overall_status']}。失败项：\"\n            + \", \".join(failed_checks)\n        )\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 6. 无论来源是挂载 Output 还是刚刚内置生成，统一锁定最终清洗数据\nif NEED_CLEANING:\n    current_gate_path = Path(REPORT_ROOT) / \"TRAINING_GATE.txt\"\n    FINAL_CLEAN_CANDIDATE = inspect_clean_gate(current_gate_path)\n    if not FINAL_CLEAN_CANDIDATE.get(\"valid\"):\n        raise RuntimeError(\n            \"内置 V6 没有产生合格训练数据：\"\n            + \", \".join(FINAL_CLEAN_CANDIDATE.get(\"reasons\", []))\n        )\nelse:\n    FINAL_CLEAN_CANDIDATE = SELECTED_CLEAN_CANDIDATE\n\nCLEAN_DATASET_ROOT = Path(FINAL_CLEAN_CANDIDATE[\"dataset_root\"])\nCLEAN_REPORT_ROOT = Path(FINAL_CLEAN_CANDIDATE[\"report_root\"])\nCLEAN_RUN_ROOT = Path(FINAL_CLEAN_CANDIDATE[\"run_root\"])\nCLEAN_DATASET_FINGERPRINT = FINAL_CLEAN_CANDIDATE[\"fingerprint\"]\nCLEAN_DATASET_IDENTITY = FINAL_CLEAN_CANDIDATE[\"identity\"]\nCLEAN_ACCEPTANCE = FINAL_CLEAN_CANDIDATE[\"acceptance\"]\nCLEAN_GATE = FINAL_CLEAN_CANDIDATE[\"gate\"]\n\nif CLEAN_GATE.get(\"TRAINING_READY\") != \"True\":\n    raise RuntimeError(\"训练闸门不是 True，禁止训练。\")\nif len(CLEAN_DATASET_FINGERPRINT) != 64:\n    raise RuntimeError(\"最终清洗数据指纹无效。\")\n\nprint(\"=\" * 88)\nprint(\"FINAL DATASET LOCK PASSED\")\nprint(\"Dataset:\", CLEAN_DATASET_ROOT)\nprint(\"Fingerprint:\", CLEAN_DATASET_FINGERPRINT)\nprint(\"Training ready:\", CLEAN_GATE[\"TRAINING_READY\"])\nprint(\"=\" * 88)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## C. E1 数据预检、划分锁与训练合同\n\n程序不信任旧绝对路径，而是根据当前挂载位置重建训练路径，并再次核对：\n\n- V6 数据指纹；\n- 1280 图像与标签完整性；\n- E0/E1 的 train/val/test 逐图一致性；\n- 单类别 `Nodule/Mass`；\n- 训练参数合同与断点兼容性。\n\n任何一项不一致都会在训练前停止。\n","metadata":{}},{"cell_type":"code","source":"# 7. 模型侧全量数据预检\ndef sha256_file(path: Path, chunk_size: int = 1024 * 1024) -> str:\n    digest = hashlib.sha256()\n    with Path(path).open(\"rb\") as stream:\n        for chunk in iter(lambda: stream.read(chunk_size), b\"\"):\n            digest.update(chunk)\n    return digest.hexdigest()\n\nclean_manifest_path = CLEAN_DATASET_ROOT / \"manifest.csv\"\nclean_identity_path = CLEAN_DATASET_ROOT / \"dataset_identity.json\"\nclean_data_yaml_path = CLEAN_DATASET_ROOT / \"data.yaml\"\nclean_manifest = pd.read_csv(clean_manifest_path)\n\nrequired_manifest_columns = {\n    \"image_id\", \"split\", \"is_positive\", \"label_count\",\n    \"output_width\", \"output_height\",\n}\nmissing_columns = required_manifest_columns - set(clean_manifest.columns)\nif missing_columns:\n    raise RuntimeError(\n        f\"V6 manifest 缺少训练必需字段：{sorted(missing_columns)}\"\n    )\n\nclean_manifest[\"image_id\"] = clean_manifest[\"image_id\"].astype(str).str.strip()\nclean_manifest[\"split\"] = clean_manifest[\"split\"].astype(str).str.lower().str.strip()\nclean_manifest[\"is_positive\"] = pd.to_numeric(\n    clean_manifest[\"is_positive\"], errors=\"raise\"\n).astype(int)\nclean_manifest[\"label_count\"] = pd.to_numeric(\n    clean_manifest[\"label_count\"], errors=\"raise\"\n).astype(int)\n\nif clean_manifest.empty:\n    raise RuntimeError(\"V6 manifest 为空。\")\nif clean_manifest[\"image_id\"].duplicated().any():\n    raise RuntimeError(\"V6 manifest 出现重复 image_id。\")\nif set(clean_manifest[\"split\"]) != set(VALID_SPLITS):\n    raise RuntimeError(\"V6 manifest 必须同时包含 train/val/test。\")\nif not clean_manifest[\"is_positive\"].isin([0, 1]).all():\n    raise RuntimeError(\"is_positive 只能是 0/1。\")\nif (\n    (clean_manifest[\"is_positive\"] == 1)\n    != (clean_manifest[\"label_count\"] > 0)\n).any():\n    raise RuntimeError(\"is_positive 与 label_count 不一致。\")\n\nclean_manifest[\"runtime_image_path\"] = clean_manifest.apply(\n    lambda row: str(\n        CLEAN_DATASET_ROOT / \"images\" / row[\"split\"]\n        / f\"{row['image_id']}.png\"\n    ),\n    axis=1,\n)\nclean_manifest[\"runtime_label_path\"] = clean_manifest.apply(\n    lambda row: str(\n        CLEAN_DATASET_ROOT / \"labels\" / row[\"split\"]\n        / f\"{row['image_id']}.txt\"\n    ),\n    axis=1,\n)\n\nmissing_images = []\nmissing_labels = []\nlabel_errors = []\nobserved_label_counts = {}\nfor row in tqdm(\n    clean_manifest.itertuples(index=False),\n    total=len(clean_manifest),\n    desc=\"Model preflight: images and labels\",\n):\n    image_path = Path(row.runtime_image_path)\n    label_path = Path(row.runtime_label_path)\n    if not image_path.is_file():\n        missing_images.append(row.image_id)\n    if not label_path.is_file():\n        missing_labels.append(row.image_id)\n        continue\n    lines = [\n        line.strip() for line in label_path.read_text(\n            encoding=\"utf-8\"\n        ).splitlines() if line.strip()\n    ]\n    observed_label_counts[row.image_id] = len(lines)\n    if len(lines) != int(row.label_count):\n        label_errors.append(\n            f\"{row.image_id}: manifest={row.label_count}, file={len(lines)}\"\n        )\n    for line_number, line in enumerate(lines, 1):\n        parts = line.split()\n        if len(parts) != 5:\n            label_errors.append(\n                f\"{row.image_id}:{line_number}: expected 5 fields\"\n            )\n            continue\n        try:\n            class_id = int(parts[0])\n            x, y, w, h = [float(value) for value in parts[1:]]\n        except Exception:\n            label_errors.append(\n                f\"{row.image_id}:{line_number}: non-numeric\"\n            )\n            continue\n        values = np.array([x, y, w, h], dtype=float)\n        if class_id != 0 or not np.isfinite(values).all():\n            label_errors.append(\n                f\"{row.image_id}:{line_number}: class/finite failure\"\n            )\n            continue\n        if not (0 < x <= 1 and 0 < y <= 1 and 0 < w <= 1 and 0 < h <= 1):\n            label_errors.append(\n                f\"{row.image_id}:{line_number}: normalized range\"\n            )\n        if (\n            x - w / 2 < -1e-7 or x + w / 2 > 1 + 1e-7\n            or y - h / 2 < -1e-7 or y + h / 2 > 1 + 1e-7\n        ):\n            label_errors.append(\n                f\"{row.image_id}:{line_number}: box crosses image\"\n            )\n\nexpected_by_split = {\n    split: set(\n        clean_manifest.loc[\n            clean_manifest[\"split\"].eq(split), \"image_id\"\n        ]\n    )\n    for split in VALID_SPLITS\n}\nactual_images_by_split = {\n    split: {\n        path.stem for path in (\n            CLEAN_DATASET_ROOT / \"images\" / split\n        ).glob(\"*.png\")\n    }\n    for split in VALID_SPLITS\n}\nactual_labels_by_split = {\n    split: {\n        path.stem for path in (\n            CLEAN_DATASET_ROOT / \"labels\" / split\n        ).glob(\"*.txt\")\n    }\n    for split in VALID_SPLITS\n}\nset_mismatches = {\n    split: {\n        \"missing_images\": sorted(\n            expected_by_split[split] - actual_images_by_split[split]\n        ),\n        \"extra_images\": sorted(\n            actual_images_by_split[split] - expected_by_split[split]\n        ),\n        \"missing_labels\": sorted(\n            expected_by_split[split] - actual_labels_by_split[split]\n        ),\n        \"extra_labels\": sorted(\n            actual_labels_by_split[split] - expected_by_split[split]\n        ),\n    }\n    for split in VALID_SPLITS\n}\nhas_set_mismatch = any(\n    values\n    for split_values in set_mismatches.values()\n    for values in split_values.values()\n)\nif missing_images or missing_labels or label_errors or has_set_mismatch:\n    failure_report = {\n        \"missing_images\": missing_images[:100],\n        \"missing_labels\": missing_labels[:100],\n        \"label_errors\": label_errors[:500],\n        \"set_mismatches\": set_mismatches,\n    }\n    (MODEL_REPORT_ROOT / \"model_data_preflight_failures.json\").write_text(\n        json.dumps(failure_report, indent=2, ensure_ascii=False),\n        encoding=\"utf-8\",\n    )\n    raise RuntimeError(\n        \"模型侧数据预检失败；查看 model_data_preflight_failures.json。\"\n    )\n\nif any(\n    path.is_symlink()\n    for split in VALID_SPLITS\n    for path in (CLEAN_DATASET_ROOT / \"images\" / split).glob(\"*.png\")\n):\n    raise RuntimeError(\"最终训练图片中出现软链接；正式 Output 必须是 copy。\")\n\noriginal_yaml = yaml.safe_load(\n    clean_data_yaml_path.read_text(encoding=\"utf-8\")\n)\nif (\n    yaml_class_name(original_yaml) != EXPECTED_TARGET_CLASS\n    or int(original_yaml.get(\"nc\", 1)) != 1\n):\n    raise RuntimeError(\"V6 data.yaml 不是单类 Nodule/Mass。\")\n\nRUNTIME_DATA_YAML = MODEL_OUTPUT_ROOT / \"runtime_data.yaml\"\nruntime_yaml = {\n    \"path\": str(CLEAN_DATASET_ROOT),\n    \"train\": \"images/train\",\n    \"val\": \"images/val\",\n    \"test\": \"images/test\",\n    \"nc\": 1,\n    \"names\": {0: EXPECTED_TARGET_CLASS},\n}\nRUNTIME_DATA_YAML.write_text(\n    yaml.safe_dump(\n        runtime_yaml, sort_keys=False, allow_unicode=True\n    ),\n    encoding=\"utf-8\",\n)\n\nsplit_statistics = clean_manifest.groupby(\"split\").agg(\n    images=(\"image_id\", \"size\"),\n    positive_images=(\"is_positive\", \"sum\"),\n    boxes=(\"label_count\", \"sum\"),\n)\nsplit_statistics[\"negative_images\"] = (\n    split_statistics[\"images\"] - split_statistics[\"positive_images\"]\n)\nsplit_statistics = split_statistics.reset_index()\nsplit_statistics.to_csv(\n    MODEL_REPORT_ROOT / \"locked_dataset_split_statistics.csv\", index=False\n)\ndisplay(split_statistics)\n\npreflight_report = {\n    \"status\": \"PASS\",\n    \"training_ready\": True,\n    \"dataset_root\": str(CLEAN_DATASET_ROOT),\n    \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n    \"manifest_sha256\": sha256_file(clean_manifest_path),\n    \"identity_sha256\": sha256_file(clean_identity_path),\n    \"image_count\": int(len(clean_manifest)),\n    \"positive_image_count\": int(clean_manifest[\"is_positive\"].sum()),\n    \"box_count\": int(clean_manifest[\"label_count\"].sum()),\n    \"split_counts\": {\n        row.split: int(row.images)\n        for row in split_statistics.itertuples(index=False)\n    },\n    \"target_class\": EXPECTED_TARGET_CLASS,\n    \"raw_dicom_read_by_training_stage\": False,\n    \"raw_train_csv_read_by_training_stage\": False,\n}\n(MODEL_REPORT_ROOT / \"model_data_preflight.json\").write_text(\n    json.dumps(preflight_report, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\nprint(json.dumps(preflight_report, indent=2, ensure_ascii=False))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 7B. 分辨率实验划分锁：1280 必须逐图复用 E0 的 train/val/test\ndef normalized_split_frame(frame: pd.DataFrame) -> pd.DataFrame:\n    required = {\"image_id\", \"split\"}\n    missing = required - set(frame.columns)\n    if missing:\n        raise RuntimeError(f\"split manifest 缺少字段：{sorted(missing)}\")\n    result = frame[[\"image_id\", \"split\"]].copy()\n    result[\"image_id\"] = result[\"image_id\"].astype(str).str.strip()\n    result[\"split\"] = result[\"split\"].astype(str).str.lower().str.strip()\n    if result[\"image_id\"].duplicated().any():\n        raise RuntimeError(\"split manifest 出现重复 image_id。\")\n    return result.sort_values(\"image_id\").reset_index(drop=True)\n\ndef split_mapping_sha256(frame: pd.DataFrame) -> str:\n    normalized = normalized_split_frame(frame)\n    payload = normalized.to_csv(index=False, lineterminator=\"\\n\")\n    return hashlib.sha256(payload.encode(\"utf-8\")).hexdigest()\n\ncurrent_split_frame = normalized_split_frame(clean_manifest)\ncurrent_split_hash = split_mapping_sha256(current_split_frame)\nsplit_comparison_reports = []\nmatching_references = []\n\nfor reference_manifest_path in BASELINE_SPLIT_REFERENCE_MANIFESTS:\n    try:\n        reference_frame = normalized_split_frame(\n            pd.read_csv(reference_manifest_path)\n        )\n        reference_hash = split_mapping_sha256(reference_frame)\n        current_ids = set(current_split_frame[\"image_id\"])\n        reference_ids = set(reference_frame[\"image_id\"])\n        missing_from_1280 = sorted(reference_ids - current_ids)\n        extra_in_1280 = sorted(current_ids - reference_ids)\n\n        merged = reference_frame.merge(\n            current_split_frame,\n            on=\"image_id\",\n            how=\"inner\",\n            suffixes=(\"_640\", \"_1280\"),\n        )\n        split_mismatches = merged.loc[\n            merged[\"split_640\"] != merged[\"split_1280\"],\n            [\"image_id\", \"split_640\", \"split_1280\"],\n        ]\n\n        report = {\n            \"reference_manifest\": str(reference_manifest_path),\n            \"reference_split_sha256\": reference_hash,\n            \"current_1280_split_sha256\": current_split_hash,\n            \"reference_images\": int(len(reference_frame)),\n            \"current_1280_images\": int(len(current_split_frame)),\n            \"missing_from_1280_count\": int(len(missing_from_1280)),\n            \"extra_in_1280_count\": int(len(extra_in_1280)),\n            \"split_mismatch_count\": int(len(split_mismatches)),\n            \"missing_from_1280_examples\": missing_from_1280[:50],\n            \"extra_in_1280_examples\": extra_in_1280[:50],\n            \"split_mismatch_examples\": split_mismatches.head(50).to_dict(\"records\"),\n        }\n        split_comparison_reports.append(report)\n\n        if (\n            not missing_from_1280\n            and not extra_in_1280\n            and split_mismatches.empty\n        ):\n            matching_references.append(report)\n    except Exception as error:\n        split_comparison_reports.append({\n            \"reference_manifest\": str(reference_manifest_path),\n            \"error\": f\"{type(error).__name__}: {error}\",\n        })\n\nif REQUIRE_BASELINE_SPLIT_MATCH and not matching_references:\n    failure_path = MODEL_REPORT_ROOT / \"resolution_split_lock_failure.json\"\n    failure_path.write_text(\n        json.dumps(\n            {\n                \"status\": \"FAIL\",\n                \"experiment\": EXPERIMENT_ID,\n                \"comparisons\": split_comparison_reports,\n            },\n            indent=2,\n            ensure_ascii=False,\n        ),\n        encoding=\"utf-8\",\n    )\n    raise RuntimeError(\n        \"1280 数据划分与 E0 640 基线不一致，禁止训练。\"\n        \"查看 resolution_split_lock_failure.json。\"\n    )\n\nselected_split_reference = (\n    sorted(\n        matching_references,\n        key=lambda item: item[\"reference_manifest\"],\n    )[0]\n    if matching_references else None\n)\n\nSPLIT_LOCK_REPORT = {\n    \"status\": \"PASS\" if selected_split_reference else \"NOT_REQUIRED\",\n    \"experiment\": EXPERIMENT_ID,\n    \"factor\": EXPERIMENT_FACTOR,\n    \"baseline_source_signature\": BASELINE_640_SOURCE_SIGNATURE,\n    \"current_1280_source_signature\": EXPECTED_SOURCE_DATASET_SIGNATURE,\n    \"baseline_manifest\": (\n        selected_split_reference[\"reference_manifest\"]\n        if selected_split_reference else None\n    ),\n    \"baseline_split_sha256\": (\n        selected_split_reference[\"reference_split_sha256\"]\n        if selected_split_reference else None\n    ),\n    \"current_1280_split_sha256\": current_split_hash,\n    \"exact_image_and_split_match\": bool(selected_split_reference),\n    \"comparisons\": split_comparison_reports,\n}\n(MODEL_REPORT_ROOT / \"resolution_split_lock.json\").write_text(\n    json.dumps(SPLIT_LOCK_REPORT, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\nprint(json.dumps(SPLIT_LOCK_REPORT, indent=2, ensure_ascii=False))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 8. 训练合同：数据、模型和超参数一一绑定\nTRAINING_HYPERPARAMETERS = {\n    \"model_weights\": MODEL_WEIGHTS,\n    \"architecture\": \"YOLO11n\",\n    \"experiment_id\": EXPERIMENT_ID,\n    \"experiment_factor\": EXPERIMENT_FACTOR,\n    \"task\": \"single_class_object_detection\",\n    \"target_class\": EXPECTED_TARGET_CLASS,\n    \"imgsz\": TRAIN_IMGSZ,\n    \"epochs\": EPOCHS,\n    \"patience\": PATIENCE,\n    \"batch\": TRAIN_BATCH,\n    \"optimizer\": OPTIMIZER,\n    \"lr0\": LR0,\n    \"lrf\": LRF,\n    \"weight_decay\": WEIGHT_DECAY,\n    \"warmup_epochs\": WARMUP_EPOCHS,\n    \"fraction\": 1.0,\n    \"amp\": True,\n    \"cache\": False,\n    \"single_cls\": True,\n    \"seed\": SEED,\n    \"deterministic\": True,\n    \"augmentations\": {\n        \"hsv_h\": 0.0,\n        \"hsv_s\": 0.0,\n        \"hsv_v\": 0.10,\n        \"degrees\": 2.0,\n        \"translate\": 0.03,\n        \"scale\": 0.15,\n        \"shear\": 0.0,\n        \"perspective\": 0.0,\n        \"flipud\": 0.0,\n        \"fliplr\": 0.5,\n        \"mosaic\": 0.0,\n        \"mixup\": 0.0,\n    },\n}\nTRAINING_CONTRACT = {\n    \"contract_version\": \"1.0\",\n    \"cleaning_version\": EXPECTED_CLEANING_VERSION,\n    \"source_dataset_signature\": EXPECTED_SOURCE_DATASET_SIGNATURE,\n    \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n    \"manifest_sha256\": preflight_report[\"manifest_sha256\"],\n    \"baseline_640_split_sha256\": SPLIT_LOCK_REPORT[\"baseline_split_sha256\"],\n    \"current_1280_split_sha256\": SPLIT_LOCK_REPORT[\"current_1280_split_sha256\"],\n    \"training\": TRAINING_HYPERPARAMETERS,\n}\nTRAINING_CONTRACT_SIGNATURE = hashlib.sha256(\n    json.dumps(\n        TRAINING_CONTRACT, sort_keys=True, ensure_ascii=False\n    ).encode(\"utf-8\")\n).hexdigest()\nTRAINING_CONTRACT[\"training_contract_signature\"] = (\n    TRAINING_CONTRACT_SIGNATURE\n)\n\nBASE_EXPERIMENT_NAME = (\n    \"E1_yolo11n_1280_\" + TRAINING_CONTRACT_SIGNATURE[:12]\n)\n(MODEL_REPORT_ROOT / \"training_contract.json\").write_text(\n    json.dumps(TRAINING_CONTRACT, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\nprint(\"Training contract:\", TRAINING_CONTRACT_SIGNATURE)\nprint(\"Experiment:\", BASE_EXPERIMENT_NAME)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 9. GPU、模型权重与可验证断点恢复\nimport torch\nimport ultralytics\nfrom ultralytics import YOLO\n\ntorch.manual_seed(SEED)\nif torch.cuda.is_available():\n    torch.cuda.manual_seed_all(SEED)\n\nruntime_versions = {\n    \"python\": platform.python_version(),\n    \"torch\": torch.__version__,\n    \"ultralytics\": ultralytics.__version__,\n    \"numpy\": np.__version__,\n    \"pandas\": pd.__version__,\n    \"cuda_available\": bool(torch.cuda.is_available()),\n    \"gpu\": (\n        torch.cuda.get_device_name(0)\n        if torch.cuda.is_available() else None\n    ),\n}\n(MODEL_REPORT_ROOT / \"runtime_versions.json\").write_text(\n    json.dumps(runtime_versions, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\nprint(json.dumps(runtime_versions, indent=2, ensure_ascii=False))\n\nif RUN_MODE in {\"train\", \"full\"} and not torch.cuda.is_available():\n    raise RuntimeError(\n        \"正式训练需要 GPU。请在 Kaggle Settings → Accelerator 选择 GPU。\"\n    )\n\ndef load_contract(path: Path):\n    try:\n        return read_json(path)\n    except Exception:\n        return None\n\ndef run_progress(run_dir: Path):\n    results_csv = run_dir / \"results.csv\"\n    if not results_csv.is_file():\n        return 0\n    try:\n        return int(len(pd.read_csv(results_csv)))\n    except Exception:\n        return 0\n\ndef run_is_contract_compatible(run_dir: Path):\n    contract = load_contract(run_dir / \"training_contract.json\")\n    return (\n        contract is not None\n        and contract.get(\"training_contract_signature\")\n        == TRAINING_CONTRACT_SIGNATURE\n    )\n\nprevious_run_candidates = []\nfor local_run in MODEL_RUNS_ROOT.glob(\n    BASE_EXPERIMENT_NAME + \"*\"\n):\n    if local_run.is_dir() and run_is_contract_compatible(local_run):\n        previous_run_candidates.append(local_run)\n\nif INPUT_ROOT.exists():\n    for contract_path in INPUT_ROOT.rglob(\"training_contract.json\"):\n        run_dir = contract_path.parent\n        contract = load_contract(contract_path)\n        if (\n            contract is not None\n            and contract.get(\"training_contract_signature\")\n            == TRAINING_CONTRACT_SIGNATURE\n            and (\n                (run_dir / \"weights\" / \"last.pt\").is_file()\n                or (run_dir / \"weights\" / \"best.pt\").is_file()\n            )\n        ):\n            previous_run_candidates.append(run_dir)\n\ndef candidate_score(run_dir: Path):\n    complete_path = run_dir / \"TRAINING_COMPLETE.json\"\n    complete = False\n    if complete_path.is_file():\n        status = load_contract(complete_path) or {}\n        complete = (\n            status.get(\"training_contract_signature\")\n            == TRAINING_CONTRACT_SIGNATURE\n            and status.get(\"status\") == \"TRAINING_COMPLETE\"\n        )\n    return (int(complete), run_progress(run_dir), str(run_dir))\n\nprevious_run_candidates = sorted(\n    set(previous_run_candidates), key=candidate_score, reverse=True\n)\n\nif previous_run_candidates:\n    source_run = previous_run_candidates[0]\n    target_run = MODEL_RUNS_ROOT / source_run.name\n    if source_run.resolve() != target_run.resolve():\n        print(\"Importing compatible previous run:\", source_run)\n        shutil.copytree(source_run, target_run, dirs_exist_ok=True)\n    ACTIVE_RUN_DIR = target_run\nelse:\n    ACTIVE_RUN_DIR = MODEL_RUNS_ROOT / BASE_EXPERIMENT_NAME\n\nACTIVE_RUN_DIR.mkdir(parents=True, exist_ok=True)\n(ACTIVE_RUN_DIR / \"training_contract.json\").write_text(\n    json.dumps(TRAINING_CONTRACT, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\n\ncompletion_path = ACTIVE_RUN_DIR / \"TRAINING_COMPLETE.json\"\ncompletion_status = (\n    load_contract(completion_path) if completion_path.is_file() else {}\n) or {}\nTRAINING_ALREADY_COMPLETE = (\n    completion_status.get(\"status\") == \"TRAINING_COMPLETE\"\n    and completion_status.get(\"training_contract_signature\")\n    == TRAINING_CONTRACT_SIGNATURE\n    and (ACTIVE_RUN_DIR / \"weights\" / \"best.pt\").is_file()\n)\nRESUME_CHECKPOINT = ACTIVE_RUN_DIR / \"weights\" / \"last.pt\"\nCAN_RESUME = (\n    not TRAINING_ALREADY_COMPLETE and RESUME_CHECKPOINT.is_file()\n)\n\nweight_candidates = (\n    sorted(\n        path for path in INPUT_ROOT.rglob(MODEL_WEIGHTS)\n        if path.is_file()\n    )\n    if INPUT_ROOT.exists() else []\n)\nPRETRAINED_SOURCE = (\n    str(weight_candidates[0]) if weight_candidates else MODEL_WEIGHTS\n)\n\nif RUN_MODE == \"evaluate\" and not TRAINING_ALREADY_COMPLETE:\n    raise FileNotFoundError(\n        \"RUN_MODE=evaluate，但没有同一训练合同的完整 best.pt。\"\n    )\n\nif torch.cuda.is_available():\n    free_vram, total_vram = torch.cuda.mem_get_info()\n    print(\n        f\"GPU memory: {free_vram / 1024**3:.2f}/\"\n        f\"{total_vram / 1024**3:.2f} GiB free\"\n    )\nprint(\"Active run:\", ACTIVE_RUN_DIR)\nprint(\"Training already complete:\", TRAINING_ALREADY_COMPLETE)\nprint(\"Can resume exact checkpoint:\", CAN_RESUME)\nprint(\"Pretrained source:\", PRETRAINED_SOURCE)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## D. YOLO11n 1280 正式训练\n\n- 训练集、验证集划分必须与 E0 完全一致。\n- 模型、优化器、学习率、增强、epoch、patience 和 seed 与 E0 保持一致。\n- 唯一主要实验变量是分辨率：`640 → 1280`。\n- 默认 `RUN_MODE=\"train\"`，不会在训练结束后立即进行大规模逐图测试。\n- `best.pt` 仍只由验证集选择。\n","metadata":{}},{"cell_type":"code","source":"# 10. 开始新训练或精确恢复 optimizer/scheduler/epoch\ntraining_result = None\nif RUN_MODE in {\"train\", \"full\"} and not TRAINING_ALREADY_COMPLETE:\n    gc.collect()\n    torch.cuda.empty_cache()\n    try:\n        if CAN_RESUME:\n            print(\"Resuming exact checkpoint:\", RESUME_CHECKPOINT)\n            training_model = YOLO(str(RESUME_CHECKPOINT))\n            training_result = training_model.train(resume=True)\n        else:\n            print(\"Starting fresh transfer learning:\", PRETRAINED_SOURCE)\n            training_model = YOLO(PRETRAINED_SOURCE)\n            training_result = training_model.train(\n                data=str(RUNTIME_DATA_YAML),\n                epochs=EPOCHS,\n                patience=PATIENCE,\n                imgsz=TRAIN_IMGSZ,\n                batch=TRAIN_BATCH,\n                nbs=64,\n                fraction=1.0,\n                device=0,\n                workers=WORKERS,\n                project=str(MODEL_RUNS_ROOT),\n                name=ACTIVE_RUN_DIR.name,\n                exist_ok=True,\n                pretrained=True,\n                optimizer=OPTIMIZER,\n                lr0=LR0,\n                lrf=LRF,\n                momentum=0.9,\n                weight_decay=WEIGHT_DECAY,\n                warmup_epochs=WARMUP_EPOCHS,\n                cos_lr=True,\n                seed=SEED,\n                deterministic=True,\n                amp=True,\n                cache=False,\n                single_cls=True,\n                val=True,\n                plots=True,\n                save=True,\n                save_period=5,\n                close_mosaic=0,\n                hsv_h=0.0,\n                hsv_s=0.0,\n                hsv_v=0.10,\n                degrees=2.0,\n                translate=0.03,\n                scale=0.15,\n                shear=0.0,\n                perspective=0.0,\n                flipud=0.0,\n                fliplr=0.5,\n                mosaic=0.0,\n                mixup=0.0,\n                max_det=100,\n                verbose=True,\n            )\n    except RuntimeError as error:\n        error_text = str(error)\n        (MODEL_REPORT_ROOT / \"training_runtime_error.txt\").write_text(\n            traceback.format_exc(), encoding=\"utf-8\"\n        )\n        if \"out of memory\" in error_text.lower():\n            gc.collect()\n            torch.cuda.empty_cache()\n            raise RuntimeError(\n                \"1280 训练发生 CUDA 显存不足。\"\n                \"请先确认没有其他 GPU 任务；若 AutoBatch 仍失败，\"\n                \"把 TRAIN_BATCH 改为 2，再从同一训练合同重新运行。\"\n            ) from error\n        raise\n\n    best_after_train = ACTIVE_RUN_DIR / \"weights\" / \"best.pt\"\n    last_after_train = ACTIVE_RUN_DIR / \"weights\" / \"last.pt\"\n    if not best_after_train.is_file() or not last_after_train.is_file():\n        raise FileNotFoundError(\n            \"训练调用结束，但没有同时生成 best.pt 和 last.pt。\"\n        )\n    trainer_args = getattr(\n        getattr(training_model, \"trainer\", None), \"args\", None\n    )\n    resolved_batch = getattr(trainer_args, \"batch\", TRAIN_BATCH)\n    completion_status = {\n        \"status\": \"TRAINING_COMPLETE\",\n        \"completed_at\": datetime.now(timezone.utc).isoformat(),\n        \"training_contract_signature\": TRAINING_CONTRACT_SIGNATURE,\n        \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n        \"best_pt_sha256\": sha256_file(best_after_train),\n        \"last_pt_sha256\": sha256_file(last_after_train),\n        \"resolved_batch\": resolved_batch,\n        \"completed_epochs\": run_progress(ACTIVE_RUN_DIR),\n    }\n    completion_path.write_text(\n        json.dumps(completion_status, indent=2, ensure_ascii=False),\n        encoding=\"utf-8\",\n    )\n    TRAINING_ALREADY_COMPLETE = True\n    if training_model is not None:\n        training_model.trainer = None\n    del training_model\n    training_result = None\n    gc.collect()\n    torch.cuda.empty_cache()\nelif TRAINING_ALREADY_COMPLETE:\n    print(\"Compatible completed run found; training skipped.\")\nelse:\n    print(\"Training skipped by RUN_MODE.\")\n\nBEST_CHECKPOINT = ACTIVE_RUN_DIR / \"weights\" / \"best.pt\"\nLAST_CHECKPOINT = ACTIVE_RUN_DIR / \"weights\" / \"last.pt\"\nif not BEST_CHECKPOINT.is_file():\n    raise FileNotFoundError(f\"没有找到 best.pt：{BEST_CHECKPOINT}\")\nif completion_path.is_file():\n    print(json.dumps(read_json(completion_path), indent=2, ensure_ascii=False))\nprint(\"Best checkpoint:\", BEST_CHECKPOINT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## E. 独立评估代码（默认跳过）\n\n默认 `RUN_MODE=\"train\"`，以下测试集、阈值、FROC 和错误分析代码不会执行。\n\n保留这些 Cell 仅用于未来独立评估版本；不要在正式训练提交中把 `RUN_MODE` 改为 `full`，以免再次在模型已训练完成后因逐图预测显存不足导致版本显示失败。\n","metadata":{}},{"cell_type":"code","source":"# 11. Ultralytics 独立 test split 检测评估\nRUN_EVALUATION = RUN_MODE in {\"full\", \"evaluate\"}\nbest_model = YOLO(str(BEST_CHECKPOINT))\ntest_detection_result = None\ndetection_metrics = None\n\ndef json_scalar(value):\n    if isinstance(value, (int, float, np.integer, np.floating)):\n        return float(value)\n    if hasattr(value, \"numel\") and value.numel() == 1:\n        return float(value.item())\n    return None\n\nif RUN_EVALUATION:\n    test_detection_result = best_model.val(\n        data=str(RUNTIME_DATA_YAML),\n        split=\"test\",\n        imgsz=DEPLOY_IMGSZ,\n        batch=EVAL_BATCH,\n        device=0 if torch.cuda.is_available() else \"cpu\",\n        workers=WORKERS,\n        plots=True,\n        project=str(MODEL_REPORT_ROOT / \"test_detection\"),\n        name=ACTIVE_RUN_DIR.name,\n        exist_ok=True,\n        max_det=100,\n        verbose=True,\n    )\n    detection_metrics = {}\n    for key, value in test_detection_result.results_dict.items():\n        scalar = json_scalar(value)\n        if scalar is not None:\n            detection_metrics[key] = scalar\n    detection_metrics.update({\n        \"split\": \"test\",\n        \"imgsz\": DEPLOY_IMGSZ,\n        \"checkpoint\": str(BEST_CHECKPOINT),\n        \"checkpoint_sha256\": sha256_file(BEST_CHECKPOINT),\n        \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n        \"training_contract_signature\": TRAINING_CONTRACT_SIGNATURE,\n    })\n    (MODEL_REPORT_ROOT / \"test_detection_metrics.json\").write_text(\n        json.dumps(detection_metrics, indent=2, ensure_ascii=False),\n        encoding=\"utf-8\",\n    )\n    print(json.dumps(detection_metrics, indent=2, ensure_ascii=False))\nelse:\n    print(\"Independent test evaluation skipped: RUN_MODE=train\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 12. 验证集选筛查阈值；测试集图像级指标与预测框\nfrom sklearn.metrics import (\n    average_precision_score,\n    confusion_matrix,\n    precision_recall_curve,\n    roc_auc_score,\n    roc_curve,\n)\n\ndef collect_predictions(split_name: str):\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n    split_frame = clean_manifest.loc[\n        clean_manifest[\"split\"].eq(split_name),\n        [\n            \"image_id\", \"is_positive\",\n            \"runtime_image_path\", \"runtime_label_path\",\n        ],\n    ].copy()\n    path_to_row = {\n        str(Path(row.runtime_image_path).resolve()): row\n        for row in split_frame.itertuples(index=False)\n    }\n    score_rows = []\n    detection_rows = []\n    prediction_stream = best_model.predict(\n        source=split_frame[\"runtime_image_path\"].tolist(),\n        stream=True,\n        imgsz=DEPLOY_IMGSZ,\n        conf=PREDICT_CONF_FLOOR,\n        iou=0.70,\n        max_det=100,\n        batch=1,\n        device=0 if torch.cuda.is_available() else \"cpu\",\n        half=bool(torch.cuda.is_available()),\n        verbose=False,\n    )\n    for result in tqdm(\n        prediction_stream,\n        total=len(split_frame),\n        desc=f\"Predicting {split_name}\",\n    ):\n        resolved = str(Path(result.path).resolve())\n        if resolved not in path_to_row:\n            raise RuntimeError(\n                f\"预测结果路径无法映射回 manifest：{result.path}\"\n            )\n        manifest_row = path_to_row[resolved]\n        boxes = result.boxes\n        if boxes is not None and len(boxes):\n            confidences = boxes.conf.detach().cpu().numpy().astype(float)\n            xyxy = boxes.xyxy.detach().cpu().numpy().astype(float)\n            image_score = float(confidences.max())\n            for confidence, coordinates in zip(confidences, xyxy):\n                detection_rows.append({\n                    \"image_id\": manifest_row.image_id,\n                    \"split\": split_name,\n                    \"confidence\": float(confidence),\n                    \"x1\": float(coordinates[0]),\n                    \"y1\": float(coordinates[1]),\n                    \"x2\": float(coordinates[2]),\n                    \"y2\": float(coordinates[3]),\n                })\n        else:\n            image_score = 0.0\n        score_rows.append({\n            \"image_id\": manifest_row.image_id,\n            \"split\": split_name,\n            \"y_true\": int(manifest_row.is_positive),\n            \"score\": image_score,\n        })\n    scores = pd.DataFrame(score_rows).sort_values(\"image_id\").reset_index(drop=True)\n    detections = pd.DataFrame(\n        detection_rows,\n        columns=[\n            \"image_id\", \"split\", \"confidence\",\n            \"x1\", \"y1\", \"x2\", \"y2\",\n        ],\n    )\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n    return scores, detections\n\ndef binary_metrics(y_true, scores, threshold):\n    y_true = np.asarray(y_true, dtype=int)\n    scores = np.asarray(scores, dtype=float)\n    y_pred = (scores >= threshold).astype(int)\n    tn, fp, fn, tp = confusion_matrix(\n        y_true, y_pred, labels=[0, 1]\n    ).ravel()\n    sensitivity = tp / (tp + fn) if tp + fn else float(\"nan\")\n    specificity = tn / (tn + fp) if tn + fp else float(\"nan\")\n    precision = tp / (tp + fp) if tp + fp else 0.0\n    npv = tn / (tn + fn) if tn + fn else 0.0\n    return {\n        \"threshold\": float(threshold),\n        \"sensitivity\": float(sensitivity),\n        \"specificity\": float(specificity),\n        \"precision\": float(precision),\n        \"negative_predictive_value\": float(npv),\n        \"tp\": int(tp), \"fp\": int(fp), \"tn\": int(tn), \"fn\": int(fn),\n    }\n\ndef select_threshold(y_true, scores, target_sensitivity):\n    y_true = np.asarray(y_true, dtype=int)\n    scores = np.asarray(scores, dtype=float)\n    candidates = np.unique(np.r_[\n        0.0,\n        scores,\n        np.nextafter(float(np.max(scores)), np.inf),\n    ])\n    rows = [\n        binary_metrics(y_true, scores, threshold)\n        for threshold in candidates\n    ]\n    eligible = [\n        row for row in rows\n        if row[\"sensitivity\"] >= target_sensitivity\n    ]\n    if eligible:\n        return max(\n            eligible,\n            key=lambda row: (row[\"specificity\"], row[\"threshold\"]),\n        )\n    return max(\n        rows,\n        key=lambda row: (\n            row[\"sensitivity\"] + row[\"specificity\"] - 1,\n            row[\"threshold\"],\n        ),\n    )\n\ndef stratified_bootstrap_ci(\n    y_true, scores, threshold, repeats=1000, seed=2026\n):\n    y_true = np.asarray(y_true, dtype=int)\n    scores = np.asarray(scores, dtype=float)\n    positive_indices = np.flatnonzero(y_true == 1)\n    negative_indices = np.flatnonzero(y_true == 0)\n    if not len(positive_indices) or not len(negative_indices):\n        raise RuntimeError(\"bootstrap 需要测试集同时包含阳性和阴性。\")\n    rng = np.random.default_rng(seed)\n    samples = {\n        name: [] for name in (\n            \"roc_auc\", \"average_precision\",\n            \"sensitivity\", \"specificity\",\n        )\n    }\n    for _ in tqdm(range(repeats), desc=\"Stratified bootstrap\"):\n        indices = np.r_[\n            rng.choice(\n                positive_indices, len(positive_indices), replace=True\n            ),\n            rng.choice(\n                negative_indices, len(negative_indices), replace=True\n            ),\n        ]\n        sampled_y = y_true[indices]\n        sampled_scores = scores[indices]\n        fixed = binary_metrics(\n            sampled_y, sampled_scores, threshold\n        )\n        samples[\"roc_auc\"].append(\n            roc_auc_score(sampled_y, sampled_scores)\n        )\n        samples[\"average_precision\"].append(\n            average_precision_score(sampled_y, sampled_scores)\n        )\n        samples[\"sensitivity\"].append(fixed[\"sensitivity\"])\n        samples[\"specificity\"].append(fixed[\"specificity\"])\n    return {\n        name: [\n            float(np.percentile(values, 2.5)),\n            float(np.percentile(values, 97.5)),\n        ]\n        for name, values in samples.items()\n    }\n\nif RUN_EVALUATION:\n    validation_image_scores, validation_detections = (\n        collect_predictions(\"val\")\n    )\n    test_image_scores, test_detections = collect_predictions(\"test\")\n    operating_point = select_threshold(\n        validation_image_scores[\"y_true\"],\n        validation_image_scores[\"score\"],\n        TARGET_VALIDATION_SENSITIVITY,\n    )\n    test_fixed = binary_metrics(\n        test_image_scores[\"y_true\"],\n        test_image_scores[\"score\"],\n        operating_point[\"threshold\"],\n    )\n    test_auc = roc_auc_score(\n        test_image_scores[\"y_true\"], test_image_scores[\"score\"]\n    )\n    test_ap = average_precision_score(\n        test_image_scores[\"y_true\"], test_image_scores[\"score\"]\n    )\n    confidence_intervals = stratified_bootstrap_ci(\n        test_image_scores[\"y_true\"].to_numpy(),\n        test_image_scores[\"score\"].to_numpy(),\n        operating_point[\"threshold\"],\n        repeats=BOOTSTRAP_REPEATS,\n        seed=SEED,\n    )\n    clinical_metrics = {\n        \"threshold_selection\": (\n            \"maximum validation specificity subject to target sensitivity\"\n        ),\n        \"target_validation_sensitivity\": (\n            TARGET_VALIDATION_SENSITIVITY\n        ),\n        \"validation_operating_point\": operating_point,\n        \"test\": {\n            \"roc_auc\": float(test_auc),\n            \"average_precision\": float(test_ap),\n            **test_fixed,\n        },\n        \"test_95_percent_stratified_bootstrap_ci\": (\n            confidence_intervals\n        ),\n        \"bootstrap_repeats\": BOOTSTRAP_REPEATS,\n        \"deploy_imgsz\": DEPLOY_IMGSZ,\n    }\n    validation_image_scores.to_csv(\n        MODEL_REPORT_ROOT / \"validation_image_scores.csv\", index=False\n    )\n    test_image_scores.to_csv(\n        MODEL_REPORT_ROOT / \"test_image_scores.csv\", index=False\n    )\n    validation_detections.to_csv(\n        MODEL_REPORT_ROOT / \"validation_detections.csv\", index=False\n    )\n    test_detections.to_csv(\n        MODEL_REPORT_ROOT / \"test_detections.csv\", index=False\n    )\n    (MODEL_REPORT_ROOT / \"clinical_screening_metrics.json\").write_text(\n        json.dumps(clinical_metrics, indent=2, ensure_ascii=False),\n        encoding=\"utf-8\",\n    )\n\n    fpr, tpr, _ = roc_curve(\n        test_image_scores[\"y_true\"], test_image_scores[\"score\"]\n    )\n    precision_curve, recall_curve, _ = precision_recall_curve(\n        test_image_scores[\"y_true\"], test_image_scores[\"score\"]\n    )\n    figure, axes = plt.subplots(1, 2, figsize=(12, 5))\n    axes[0].plot(fpr, tpr, label=f\"AUC={test_auc:.3f}\")\n    axes[0].plot([0, 1], [0, 1], \"--\", color=\"grey\")\n    axes[0].set(\n        xlabel=\"False Positive Rate\",\n        ylabel=\"True Positive Rate\",\n        title=\"Independent Test ROC\",\n    )\n    axes[0].legend()\n    axes[1].plot(\n        recall_curve, precision_curve, label=f\"AP={test_ap:.3f}\"\n    )\n    axes[1].set(\n        xlabel=\"Recall\",\n        ylabel=\"Precision\",\n        title=\"Independent Test Precision–Recall\",\n    )\n    axes[1].legend()\n    figure.tight_layout()\n    figure.savefig(\n        MODEL_REPORT_ROOT / \"test_roc_pr_curves.png\",\n        dpi=180,\n        bbox_inches=\"tight\",\n    )\n    plt.show()\n    plt.close(figure)\n    print(json.dumps(clinical_metrics, indent=2, ensure_ascii=False))\nelse:\n    clinical_metrics = None\n    validation_image_scores = None\n    test_image_scores = None\n    validation_detections = None\n    test_detections = None\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 13. 病灶级 FROC：一对一贪心匹配，IoU 阈值预先锁定为 0.50\ndef read_yolo_ground_truth(label_path: Path, width: int, height: int):\n    boxes = []\n    for line in label_path.read_text(encoding=\"utf-8\").splitlines():\n        if not line.strip():\n            continue\n        class_id, x, y, w, h = line.split()\n        if int(class_id) != 0:\n            raise RuntimeError(f\"FROC 遇到非零类别：{label_path}\")\n        x, y, w, h = map(float, (x, y, w, h))\n        boxes.append([\n            (x - w / 2) * width,\n            (y - h / 2) * height,\n            (x + w / 2) * width,\n            (y + h / 2) * height,\n        ])\n    return np.asarray(boxes, dtype=float).reshape(-1, 4)\n\ndef box_iou_one_to_many(box, boxes):\n    if len(boxes) == 0:\n        return np.empty(0, dtype=float)\n    intersection_x1 = np.maximum(box[0], boxes[:, 0])\n    intersection_y1 = np.maximum(box[1], boxes[:, 1])\n    intersection_x2 = np.minimum(box[2], boxes[:, 2])\n    intersection_y2 = np.minimum(box[3], boxes[:, 3])\n    intersection = (\n        np.maximum(0.0, intersection_x2 - intersection_x1)\n        * np.maximum(0.0, intersection_y2 - intersection_y1)\n    )\n    box_area = max(0.0, box[2] - box[0]) * max(\n        0.0, box[3] - box[1]\n    )\n    boxes_area = np.maximum(0.0, boxes[:, 2] - boxes[:, 0]) * np.maximum(\n        0.0, boxes[:, 3] - boxes[:, 1]\n    )\n    union = box_area + boxes_area - intersection\n    return np.divide(\n        intersection,\n        union,\n        out=np.zeros_like(intersection),\n        where=union > 0,\n    )\n\ndef match_image_detections(\n    ground_truth_boxes, prediction_frame, iou_threshold\n):\n    prediction_frame = prediction_frame.sort_values(\n        \"confidence\", ascending=False\n    )\n    matched_ground_truth = np.zeros(\n        len(ground_truth_boxes), dtype=bool\n    )\n    rows = []\n    for prediction in prediction_frame.itertuples(index=False):\n        prediction_box = np.array([\n            prediction.x1, prediction.y1,\n            prediction.x2, prediction.y2,\n        ])\n        unmatched_indices = np.flatnonzero(~matched_ground_truth)\n        is_true_positive = False\n        matched_index = None\n        matched_iou = 0.0\n        if len(unmatched_indices):\n            ious = box_iou_one_to_many(\n                prediction_box,\n                ground_truth_boxes[unmatched_indices],\n            )\n            local_best = int(np.argmax(ious))\n            matched_iou = float(ious[local_best])\n            if matched_iou >= iou_threshold:\n                matched_index = int(unmatched_indices[local_best])\n                matched_ground_truth[matched_index] = True\n                is_true_positive = True\n        rows.append({\n            \"confidence\": float(prediction.confidence),\n            \"is_true_positive\": int(is_true_positive),\n            \"is_false_positive\": int(not is_true_positive),\n            \"matched_ground_truth_index\": matched_index,\n            \"matched_iou\": matched_iou,\n        })\n    return rows\n\ndef build_froc(\n    test_frame, detections, iou_threshold, fp_targets\n):\n    match_rows = []\n    total_lesions = 0\n    grouped_predictions = {\n        image_id: frame.copy()\n        for image_id, frame in detections.groupby(\"image_id\")\n    } if len(detections) else {}\n\n    for row in tqdm(\n        test_frame.itertuples(index=False),\n        total=len(test_frame),\n        desc=\"FROC matching\",\n    ):\n        width = int(row.output_width)\n        height = int(row.output_height)\n        ground_truth = read_yolo_ground_truth(\n            Path(row.runtime_label_path), width, height\n        )\n        total_lesions += len(ground_truth)\n        predictions = grouped_predictions.get(\n            row.image_id,\n            pd.DataFrame(\n                columns=[\n                    \"confidence\", \"x1\", \"y1\", \"x2\", \"y2\"\n                ]\n            ),\n        )\n        for matched in match_image_detections(\n            ground_truth, predictions, iou_threshold\n        ):\n            match_rows.append({\n                \"image_id\": row.image_id,\n                **matched,\n            })\n\n    matched_frame = pd.DataFrame(\n        match_rows,\n        columns=[\n            \"image_id\", \"confidence\",\n            \"is_true_positive\", \"is_false_positive\",\n            \"matched_ground_truth_index\", \"matched_iou\",\n        ],\n    )\n    if total_lesions <= 0:\n        raise RuntimeError(\"测试集没有病灶，无法计算 FROC。\")\n\n    if len(matched_frame):\n        grouped = (\n            matched_frame.groupby(\"confidence\", as_index=False)\n            .agg(\n                true_positives=(\"is_true_positive\", \"sum\"),\n                false_positives=(\"is_false_positive\", \"sum\"),\n            )\n            .sort_values(\"confidence\", ascending=False)\n        )\n        grouped[\"cumulative_true_positives\"] = (\n            grouped[\"true_positives\"].cumsum()\n        )\n        grouped[\"cumulative_false_positives\"] = (\n            grouped[\"false_positives\"].cumsum()\n        )\n        grouped[\"sensitivity\"] = (\n            grouped[\"cumulative_true_positives\"] / total_lesions\n        )\n        grouped[\"false_positives_per_image\"] = (\n            grouped[\"cumulative_false_positives\"] / len(test_frame)\n        )\n        origin = pd.DataFrame([{\n            \"confidence\": float(\"inf\"),\n            \"true_positives\": 0,\n            \"false_positives\": 0,\n            \"cumulative_true_positives\": 0,\n            \"cumulative_false_positives\": 0,\n            \"sensitivity\": 0.0,\n            \"false_positives_per_image\": 0.0,\n        }])\n        curve = pd.concat([origin, grouped], ignore_index=True)\n    else:\n        curve = pd.DataFrame([{\n            \"confidence\": float(\"inf\"),\n            \"true_positives\": 0,\n            \"false_positives\": 0,\n            \"cumulative_true_positives\": 0,\n            \"cumulative_false_positives\": 0,\n            \"sensitivity\": 0.0,\n            \"false_positives_per_image\": 0.0,\n        }])\n\n    sensitivities_at_targets = {}\n    for target in fp_targets:\n        eligible = curve.loc[\n            curve[\"false_positives_per_image\"] <= float(target),\n            \"sensitivity\",\n        ]\n        sensitivities_at_targets[str(target)] = (\n            float(eligible.max()) if len(eligible) else 0.0\n        )\n    summary = {\n        \"iou_threshold\": float(iou_threshold),\n        \"test_images\": int(len(test_frame)),\n        \"ground_truth_lesions\": int(total_lesions),\n        \"prediction_confidence_floor\": PREDICT_CONF_FLOOR,\n        \"sensitivity_at_false_positives_per_image\": (\n            sensitivities_at_targets\n        ),\n        \"mean_sensitivity_at_targets\": float(\n            np.mean(list(sensitivities_at_targets.values()))\n        ),\n    }\n    return curve, matched_frame, summary\n\nif RUN_EVALUATION:\n    froc_test_frame = clean_manifest.loc[\n        clean_manifest[\"split\"].eq(\"test\")\n    ].copy()\n    froc_curve, froc_matches, froc_summary = build_froc(\n        froc_test_frame,\n        test_detections,\n        FROC_IOU_THRESHOLD,\n        FROC_FP_PER_IMAGE_TARGETS,\n    )\n    froc_curve.to_csv(\n        MODEL_REPORT_ROOT / \"test_froc_curve.csv\", index=False\n    )\n    froc_matches.to_csv(\n        MODEL_REPORT_ROOT / \"test_froc_detection_matches.csv\",\n        index=False,\n    )\n    (MODEL_REPORT_ROOT / \"test_froc_summary.json\").write_text(\n        json.dumps(froc_summary, indent=2, ensure_ascii=False),\n        encoding=\"utf-8\",\n    )\n\n    figure, axis = plt.subplots(figsize=(7, 5))\n    axis.plot(\n        froc_curve[\"false_positives_per_image\"],\n        froc_curve[\"sensitivity\"],\n        linewidth=2,\n    )\n    for target, sensitivity in (\n        froc_summary[\n            \"sensitivity_at_false_positives_per_image\"\n        ].items()\n    ):\n        axis.scatter(float(target), sensitivity, s=28)\n    axis.set_xlim(\n        0, max(FROC_FP_PER_IMAGE_TARGETS)\n    )\n    axis.set_ylim(0, 1)\n    axis.set(\n        xlabel=\"False Positives per Image\",\n        ylabel=\"Lesion Sensitivity\",\n        title=f\"Independent Test FROC (IoU ≥ {FROC_IOU_THRESHOLD:.2f})\",\n    )\n    axis.grid(alpha=0.25)\n    figure.tight_layout()\n    figure.savefig(\n        MODEL_REPORT_ROOT / \"test_froc_curve.png\",\n        dpi=180,\n        bbox_inches=\"tight\",\n    )\n    plt.show()\n    plt.close(figure)\n    print(json.dumps(froc_summary, indent=2, ensure_ascii=False))\nelse:\n    froc_curve = None\n    froc_matches = None\n    froc_summary = None\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 14. 固定测试样例可视化：绿色真值，红色预测\ndef boxes_for_display(image_row, detection_frame, threshold):\n    with Image.open(image_row.runtime_image_path) as opened:\n        width, height = opened.size\n    ground_truth = read_yolo_ground_truth(\n        Path(image_row.runtime_label_path), width, height\n    )\n    predictions = detection_frame.loc[\n        (detection_frame[\"image_id\"] == image_row.image_id)\n        & (detection_frame[\"confidence\"] >= threshold)\n    ].sort_values(\"confidence\", ascending=False).head(10)\n    return ground_truth, predictions\n\nif RUN_EVALUATION:\n    test_display_frame = clean_manifest.loc[\n        clean_manifest[\"split\"].eq(\"test\")\n    ].copy()\n    positive_examples = test_display_frame.loc[\n        test_display_frame[\"is_positive\"].eq(1)\n    ].sort_values(\"image_id\").head(3)\n    negative_examples = test_display_frame.loc[\n        test_display_frame[\"is_positive\"].eq(0)\n    ].sort_values(\"image_id\").head(3)\n    display_examples = pd.concat(\n        [positive_examples, negative_examples], ignore_index=True\n    )\n    figure, axes = plt.subplots(2, 3, figsize=(15, 10))\n    threshold = clinical_metrics[\n        \"validation_operating_point\"\n    ][\"threshold\"]\n    for axis, row in zip(\n        axes.flat, display_examples.itertuples(index=False)\n    ):\n        with Image.open(row.runtime_image_path) as opened:\n            image = opened.convert(\"L\")\n        axis.imshow(image, cmap=\"gray\")\n        ground_truth, predictions = boxes_for_display(\n            row, test_detections, threshold\n        )\n        for x1, y1, x2, y2 in ground_truth:\n            axis.add_patch(plt.Rectangle(\n                (x1, y1), x2 - x1, y2 - y1,\n                fill=False, edgecolor=\"lime\", linewidth=1.5,\n            ))\n        for prediction in predictions.itertuples(index=False):\n            axis.add_patch(plt.Rectangle(\n                (prediction.x1, prediction.y1),\n                prediction.x2 - prediction.x1,\n                prediction.y2 - prediction.y1,\n                fill=False, edgecolor=\"red\", linewidth=1.2,\n            ))\n            axis.text(\n                prediction.x1,\n                max(0, prediction.y1 - 3),\n                f\"{prediction.confidence:.2f}\",\n                color=\"red\",\n                fontsize=7,\n                backgroundcolor=\"white\",\n            )\n        axis.set_title(\n            f\"{row.image_id[:8]} | truth={int(row.is_positive)}\"\n        )\n        axis.axis(\"off\")\n    figure.suptitle(\n        \"Independent Test Examples — GT green / Prediction red\"\n    )\n    figure.tight_layout()\n    figure.savefig(\n        MODEL_REPORT_ROOT / \"test_prediction_examples.png\",\n        dpi=180,\n        bbox_inches=\"tight\",\n    )\n    plt.show()\n    plt.close(figure)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## F. 模型文件与训练实验包\n\n默认训练模式只打包 1280 的 `best.pt`、`last.pt`、训练曲线、参数、数据指纹和划分锁定报告。\n\nONNX、TensorRT、逐图错误分析和测试集评价应在独立 Notebook 中完成。\n","metadata":{}},{"cell_type":"code","source":"# 15. best.pt、ONNX、报告和指纹清单\nonnx_artifact = None\nonnx_export_error = None\nbest_artifact = MODEL_ARTIFACT_ROOT / (\n    \"vindr_nodule_mass_yolo11n_1280_best.pt\"\n)\nlast_artifact = MODEL_ARTIFACT_ROOT / (\n    \"vindr_nodule_mass_yolo11n_1280_last.pt\"\n)\nshutil.copy2(BEST_CHECKPOINT, best_artifact)\nif LAST_CHECKPOINT.is_file():\n    shutil.copy2(LAST_CHECKPOINT, last_artifact)\n\nif RUN_MODE == \"full\":\n    try:\n        exported_path = Path(best_model.export(\n            format=\"onnx\",\n            imgsz=DEPLOY_IMGSZ,\n            batch=1,\n            dynamic=False,\n            simplify=True,\n            opset=12,\n            nms=False,\n            device=\"cpu\",\n        ))\n        onnx_artifact = MODEL_ARTIFACT_ROOT / (\n            \"vindr_nodule_mass_yolo11n_1280_static.onnx\"\n        )\n        shutil.copy2(exported_path, onnx_artifact)\n        import onnx\n        checked_model = onnx.load(str(onnx_artifact))\n        onnx.checker.check_model(checked_model)\n        del checked_model\n        print(\"ONNX validation passed:\", onnx_artifact)\n    except Exception:\n        onnx_export_error = traceback.format_exc()\n        (MODEL_REPORT_ROOT / \"onnx_export_error.txt\").write_text(\n            onnx_export_error, encoding=\"utf-8\"\n        )\n        warnings.warn(\n            \"best.pt 已安全保存，但 ONNX 导出失败；\"\n            \"查看 onnx_export_error.txt。\"\n        )\n\ncopy_report_names = [\n    \"model_data_preflight.json\",\n    \"locked_dataset_split_statistics.csv\",\n    \"training_contract.json\",\n    \"runtime_versions.json\",\n    \"resolution_split_lock.json\",\n    \"test_detection_metrics.json\",\n    \"clinical_screening_metrics.json\",\n    \"test_froc_summary.json\",\n    \"test_roc_pr_curves.png\",\n    \"test_froc_curve.png\",\n    \"test_prediction_examples.png\",\n]\nfor name in copy_report_names:\n    source_path = MODEL_REPORT_ROOT / name\n    if source_path.is_file():\n        shutil.copy2(source_path, MODEL_ARTIFACT_ROOT / name)\n\nfor source_path, output_name in (\n    (clean_identity_path, \"dataset_identity.json\"),\n    (CLEAN_REPORT_ROOT / \"TRAINING_GATE.txt\", \"TRAINING_GATE.txt\"),\n    (\n        CLEAN_REPORT_ROOT / \"cleaning_acceptance_final.json\",\n        \"cleaning_acceptance_final.json\",\n    ),\n    (RUNTIME_DATA_YAML, \"runtime_data.yaml\"),\n    (ACTIVE_RUN_DIR / \"results.csv\", \"training_results.csv\"),\n):\n    if Path(source_path).is_file():\n        shutil.copy2(source_path, MODEL_ARTIFACT_ROOT / output_name)\n\noperating_threshold = (\n    clinical_metrics[\"validation_operating_point\"][\"threshold\"]\n    if clinical_metrics is not None else None\n)\nmodel_card = \"\\n\".join([\n    \"# VinDr-CXR Nodule/Mass YOLO11n 1280 Resolution Experiment E1\",\n    \"\",\n    \"- Task: single-class `Nodule/Mass` object detection\",\n    \"- Research use only: true\",\n    \"- Architecture: YOLO11n\",\n    \"- Experiment: E1 input-resolution ablation\",\n    \"- Reference baseline: E0 YOLO11n 640\",\n    f\"- Training image size: {TRAIN_IMGSZ}\",\n    f\"- Deployment image size: {DEPLOY_IMGSZ}\",\n    f\"- Dataset fingerprint: `{CLEAN_DATASET_FINGERPRINT}`\",\n    f\"- Training contract: `{TRAINING_CONTRACT_SIGNATURE}`\",\n    \"- Operating threshold source: validation set only\",\n    f\"- Operating threshold: {operating_threshold}\",\n    \"- Test set used for training/model selection: false\",\n    f\"- ONNX exported: {onnx_artifact is not None}\",\n    \"- TensorRT engine: build later on the target Jetson environment\",\n    \"\",\n])\n(MODEL_ARTIFACT_ROOT / \"MODEL_CARD.md\").write_text(\n    model_card, encoding=\"utf-8\"\n)\n\nartifact_hashes = {}\nfor path in sorted(MODEL_ARTIFACT_ROOT.iterdir()):\n    if path.is_file():\n        artifact_hashes[path.name] = {\n            \"sha256\": sha256_file(path),\n            \"bytes\": int(path.stat().st_size),\n        }\nartifact_manifest = {\n    \"status\": (\n        \"MODEL_AND_ONNX_READY\"\n        if onnx_artifact is not None\n        else (\n            \"MODEL_READY_TRAIN_ONLY\"\n            if RUN_MODE == \"train\"\n            else \"MODEL_READY_ONNX_WARNING\"\n        )\n    ),\n    \"created_at\": datetime.now(timezone.utc).isoformat(),\n    \"clean_dataset_fingerprint\": CLEAN_DATASET_FINGERPRINT,\n    \"training_contract_signature\": TRAINING_CONTRACT_SIGNATURE,\n    \"best_pt\": best_artifact.name,\n    \"onnx\": onnx_artifact.name if onnx_artifact is not None else None,\n    \"artifacts\": artifact_hashes,\n}\nartifact_manifest_path = MODEL_ARTIFACT_ROOT / \"artifact_manifest.json\"\nartifact_manifest_path.write_text(\n    json.dumps(artifact_manifest, indent=2, ensure_ascii=False),\n    encoding=\"utf-8\",\n)\n\narchive_path = shutil.make_archive(\n    str(MODEL_OUTPUT_ROOT / \"vindr_nodule_mass_yolo11n_1280_E1\"),\n    \"zip\",\n    root_dir=MODEL_ARTIFACT_ROOT,\n)\nprint(json.dumps(artifact_manifest, indent=2, ensure_ascii=False))\nprint(\"Artifact archive:\", archive_path)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 16. 最终结论\nprint(\"=\" * 92)\nprint(\"E1 1280 MODEL STATUS:\", artifact_manifest[\"status\"])\nprint(\"TRAINING READY DATA:\", CLEAN_GATE[\"TRAINING_READY\"])\nprint(\"DATASET FINGERPRINT:\", CLEAN_DATASET_FINGERPRINT)\nprint(\"SPLIT LOCK:\", SPLIT_LOCK_REPORT[\"status\"])\nprint(\"TRAINING CONTRACT:\", TRAINING_CONTRACT_SIGNATURE)\nprint(\"BEST.PT:\", best_artifact)\nprint(\"BEST.PT SHA256:\", sha256_file(best_artifact))\nprint(\"ONNX:\", onnx_artifact if onnx_artifact is not None else \"EXPORT WARNING\")\nprint(\"REPORTS:\", MODEL_REPORT_ROOT)\nprint(\"ARCHIVE:\", archive_path)\nprint(\"=\" * 92)\n\nif RUN_MODE == \"full\" and not RUN_EVALUATION:\n    raise RuntimeError(\"full 模式却没有执行独立测试评估。\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 正确结束时看什么\n\n正式 60 epoch 结束后，最后一个 Cell 应显示：\n\n```text\nE1 1280 MODEL STATUS: MODEL_READY_TRAIN_ONLY\nTRAINING READY DATA: True\nSPLIT LOCK: PASS\nBEST.PT: ...vindr_nodule_mass_yolo11n_1280_best.pt\nARCHIVE: ...vindr_nodule_mass_yolo11n_1280_E1.zip\n```\n\n`MODEL_READY_TRAIN_ONLY` 表示 1280 模型已经训练并保存，但该训练 Notebook 按实验协议没有运行测试集、逐图错误分析或 ONNX 导出。\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\n\nfor p in Path(\"/kaggle\").rglob(\"best.pt\"):\n    print(p)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-10T07:27:09.929702Z","iopub.execute_input":"2026-08-10T07:27:09.929938Z","iopub.status.idle":"2026-08-10T07:28:22.05475Z","shell.execute_reply.started":"2026-08-10T07:27:09.929915Z","shell.execute_reply":"2026-08-10T07:28:22.054064Z"}},"outputs":[],"execution_count":null}]}