{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":99552,"databundleVersionId":13441085},{"sourceType":"modelInstanceVersion","sourceId":555909,"databundleVersionId":13593910,"modelInstanceId":422983}],"dockerImageVersionId":31090,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 基于Vision Transformer的RSNA颅内动脉瘤检测","metadata":{}},{"cell_type":"markdown","source":"# 1. 第1章 绪论\n## 1.1 研究背景与意义\n","metadata":{}},{"cell_type":"markdown","source":"颅内动脉瘤是一种严重威胁人类生命健康的脑血管疾病，其破裂可导致蛛网膜下腔出血，具有极高的致残率和致死率。据统计，颅内动脉瘤影响着全球约 3% 的人口，且高达 50% 的动脉瘤在破裂后才被诊断出来，每年约导致 50 万人死亡，其中一半受害者年龄在 50 岁以下。准确及时的诊断对于颅内动脉瘤的治疗至关重要，而传统依靠放射科医生人工检测的方式，存在易被忽视、效率低等问题。因此，借助先进的机器学习技术实现颅内动脉瘤的自动检测，对于提高诊断效率、降低患者生命风险具有重要的现实意义。","metadata":{}},{"cell_type":"markdown","source":"## 1.2 研究内容","metadata":{}},{"cell_type":"markdown","source":"本研究旨在利用机器学习方法，特别是基于 Vision Transformer（ViT）的模型，对颅内动脉瘤进行检测与定位。研究将围绕医学图像（如 CTA、MRA 等）的预处理展开，包括 DICOM 文件的读取、图像增强与标准化等操作；构建并优化 ViT 模型，通过设计合适的网络结构、损失函数以及训练策略，提升模型对颅内动脉瘤的识别能力；同时，对模型在不同数据集上的性能进行全面评估，分析其在动脉瘤存在性判断以及不同位置动脉瘤定位方面的效果，为临床应用提供技术支持。","metadata":{}},{"cell_type":"markdown","source":"## 1.3 国内外研究现状","metadata":{}},{"cell_type":"markdown","source":"在国外，众多研究机构和学者积极探索颅内动脉瘤的自动检测方法。一些研究采用传统的机器学习算法，如支持向量机（SVM）结合手工提取的图像特征进行检测，但受限于特征表达能力，效果存在一定局限。近年来，随着深度学习的兴起，基于卷积神经网络（CNN）的方法在颅内动脉瘤检测中取得了较好成果，能够自动学习图像中的深层特征。在国内，相关研究也在不断推进，不过与国外相比，在大规模高质量数据集的建设以及复杂模型的优化应用等方面，仍有一定的提升空间。总体而言，利用更先进的深度学习模型，如ViT，能够进一步提高颅内动脉瘤检测的准确性和效率，是当前国内外研究的重要趋势。","metadata":{}},{"cell_type":"markdown","source":"# 第2章 相关理论与技术","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Vision Transformer算法","metadata":{}},{"cell_type":"markdown","source":"### 2.1.1 算法核心思想\nVision Transformer（ViT）的核心思想是将自然语言处理中的Transformer架构创新性地应用于计算机视觉领域，通过将输入图像分割成固定大小的图像块（patches）并将其视为序列中的\"词汇\"，然后利用标准的Transformer编码器进行处理，从而突破了传统卷积神经网络的局部感受野限制，使模型能够从训练开始就建立图像全局的依赖关系。在医学影像分析中，这种设计尤为重要，因为ViT通过自注意力机制能够直接从所有图像块中提取全局上下文信息，对于颅内动脉瘤检测这种需要综合分析整个脑部影像的任务，ViT能够同时关注多个可能病变区域之间的关联，显著提高了检测的准确性和鲁棒性。","metadata":{}},{"cell_type":"markdown","source":"### 2.1.2 数学模型推导\nViT的数学模型基于标准的Transformer架构，关键推导过程包括：图像分块嵌入阶段将输入图像分割成N个大小为P×P的图像块并通过线性投影映射到D维空间；多头自注意力机制通过计算查询、键和值之间的注意力权重来捕获远程依赖关系；前馈神经网络对每个位置的特征进行非线性变换；层归一化确保训练稳定性；最终通过分类头得到预测结果，整个过程中位置编码的引入保持了图像的空间结构信息，使得模型能够理解不同图像块之间的相对位置关系。\n\n**1. 图像分块嵌入（Patch Embedding）**\n将输入图像 $x \\in \\mathbb{R}^{H \\times W \\times C}$ 分割成 $N$ 个大小为 $P \\times P$ 的图像块：\n$$\nx_p \\in \\mathbb{R}^{N \\times (P^2 \\cdot C)}, \\quad N = \\frac{H \\times W}{P^2}\n$$\n\n每个图像块通过线性投影映射到D维空间：\n$$\nz_0 = [x_{\\text{class}}; x_p^1E; x_p^2E; \\cdots; x_p^NE] + E_{\\text{pos}}\n$$\n其中 $E \\in \\mathbb{R}^{(P^2 \\cdot C) \\times D}$ 为可学习的嵌入矩阵，$E_{\\text{pos}} \\in \\mathbb{R}^{(N+1) \\times D}$ 为位置编码。\n\n**2. 多头自注意力机制（Multi-Head Self-Attention）**\n对于第l层Transformer：\n$$\n\\text{MSA}(Q, K, V) = \\text{Concat}(\\text{head}_1, \\ldots, \\text{head}_h)W^O\n$$\n其中每个注意力头计算为：\n$$\n\\text{head}_i = \\text{Attention}(QW_i^Q, KW_i^K, VW_i^V) = \\text{softmax}\\left(\\frac{QW_i^Q(KW_i^K)^T}{\\sqrt{d_k}}\\right)VW_i^V\n$$\n\n**3. 前馈神经网络（Feed-Forward Network）**\n$$\n\\text{FFN}(x) = \\max(0, xW_1 + b_1)W_2 + b_2\n$$\n\n**4. 整体Transformer层**\n$$\n\\begin{aligned}\nz_l' &= \\text{LN}(z_{l-1} + \\text{MSA}(z_{l-1})) \\\\\nz_l &= \\text{LN}(z_l' + \\text{FFN}(z_l'))\n\\end{aligned}\n$$\n\n**5. 分类头**\n最终输出通过MLP头部得到预测结果：\n$$\ny = \\text{MLP}(\\text{LN}(z_L^0))\n$$","metadata":{}},{"cell_type":"markdown","source":"### 2.1.3 算法优势分析\n相比传统CNN方法，ViT在医学影像分析中具有显著优势，其全局感受野能力使模型能够直接建立图像中任意两个位置之间的依赖关系，这对于检测分散的动脉瘤病灶特别重要；可解释性强的特点通过可视化注意力权重让医生清楚看到模型关注的重点区域，增强了诊断决策的可信度；抗过拟合性能得益于较少的归纳偏置，结合合适的数据增强和正则化策略，即使在中等规模的医学影像数据集上也能获得良好性能；计算效率方面虽然单次前向传播复杂度较高，但通常需要更少的训练epoch就能收敛，特别是使用预训练模型时迁移学习效果显著。","metadata":{}},{"cell_type":"markdown","source":"## 2.2 PyTorch深度学习框架与相关技术\nPyTorch作为基于Python的开源深度学习框架，在医学影像分析项目中提供关键优势，其动态计算图特性允许在运行时修改网络结构，便于调试和实验不同的模型架构，这对于研究性质的医学影像分析特别重要；丰富的生态系统包括TorchVision、TorchAudio等官方库提供了大量预训练模型和数据预处理工具，显著加速了项目开发进程；GPU加速支持通过CUDA接口无缝支持GPU计算，大幅提升了模型训练和推理速度，而自动微分功能简化了梯度计算过程，使得研究人员能够专注于模型设计而非实现细节。","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom transformers import ViTModel, ViTConfig\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport warnings\nimport logging\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nimport seaborn as sns\nfrom scipy import stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:01:13.506783Z","iopub.execute_input":"2025-09-05T01:01:13.50787Z","iopub.status.idle":"2025-09-05T01:01:52.648735Z","shell.execute_reply.started":"2025-09-05T01:01:13.50784Z","shell.execute_reply":"2025-09-05T01:01:52.647954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 配置日志记录\nlogging.basicConfig(\n    filename='dicom_processing.log',\n    level=logging.WARNING,\n    format='%(asctime)s - %(levelname)s - %(message)s',\n    datefmt='%Y-%m-%d %H:%M:%S'\n)\nwarnings.filterwarnings('ignore')\n\n# 设置随机种子保证可重复性\ntorch.manual_seed(42)\nnp.random.seed(42)\n\n# 检查GPU可用性\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"使用设备: {device}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:02:51.568993Z","iopub.execute_input":"2025-09-05T01:02:51.569895Z","iopub.status.idle":"2025-09-05T01:02:51.576959Z","shell.execute_reply.started":"2025-09-05T01:02:51.569854Z","shell.execute_reply":"2025-09-05T01:02:51.576125Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 第3章 数据预处理","metadata":{}},{"cell_type":"markdown","source":"## 3.1 数据描述","metadata":{}},{"cell_type":"code","source":"print(\"===== 数据描述性分析 =====\")\n# 加载训练标签\nprint(\"加载数据...\")\ntrain_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\nlocalizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\n# 基本数据信息\nprint(f\"训练数据形状: {train_df.shape}\")\nprint(f\"定位器数据形状: {localizers_df.shape}\")\n\n# 显示数据基本信息\nprint(\"\\n训练数据基本信息:\")\nprint(train_df.info())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:44:25.674965Z","iopub.execute_input":"2025-09-05T02:44:25.675268Z","iopub.status.idle":"2025-09-05T02:44:25.717589Z","shell.execute_reply.started":"2025-09-05T02:44:25.675245Z","shell.execute_reply":"2025-09-05T02:44:25.716978Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"训练数据为pandas的DataFrame格式，形状是 4348 行、18 列，还有定位器数据为 2251 行、4 列。\n\n训练数据行索引从 0 到 4347，共 18 列，列名涵盖系列实例唯一标识符、患者年龄、性别、影像模态、各动脉相关列以及是否存在动脉瘤等。所有列均无缺失值，其中 15 列数据类型为int64，3 列是object类型。","metadata":{}},{"cell_type":"code","source":"print(\"\\n训练数据描述性统计:\")\nprint(train_df.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:03:20.670257Z","iopub.execute_input":"2025-09-05T01:03:20.670605Z","iopub.status.idle":"2025-09-05T01:03:20.716488Z","shell.execute_reply.started":"2025-09-05T01:03:20.670583Z","shell.execute_reply":"2025-09-05T01:03:20.715753Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"各列数据量均为 4348 条。患者年龄均值约 58.47，标准差 15.84，年龄在 18 到 89 岁间分布。各动脉相关列及是否存在动脉瘤等列，均值大多较低，多在 0 到 0.5 左右，且最小值为 0、最大值为 1，推测这些列可能是用 0 和 1 表示有无相关情况；“Aneurysm Present” 列均值约 0.43，说明约 43% 的样本存在动脉瘤。","metadata":{}},{"cell_type":"code","source":"# 定义动脉瘤位置标签\nlocation_labels = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation'\n]\n\nall_label_cols = ['Aneurysm Present'] + location_labels","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:03:38.725783Z","iopub.execute_input":"2025-09-05T01:03:38.726503Z","iopub.status.idle":"2025-09-05T01:03:38.731619Z","shell.execute_reply.started":"2025-09-05T01:03:38.726469Z","shell.execute_reply":"2025-09-05T01:03:38.730868Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"定义动脉瘤的 13 个具体位置标签，然后将 \"是否存在动脉瘤\" 标签与这些位置标签组合起来，形成了包含所有相关标签的列表。","metadata":{}},{"cell_type":"code","source":"# 检查标签分布\nprint(\"\\n===== 标签分布分析 =====\")\nprint(\"\\n主要标签 - 动脉瘤存在分布:\")\nprint(train_df['Aneurysm Present'].value_counts())\nprint(train_df['Aneurysm Present'].value_counts(normalize=True))\n\nprint(\"\\n各位置动脉瘤分布:\")\nfor col in location_labels:\n    if col in train_df.columns:\n        count = train_df[col].value_counts()\n        print(f\"\\n{col}:\")\n        print(f\"存在: {count.get(1, 0)}, 不存在: {count.get(0, 0)}\")\n        if 1 in count:\n            print(f\"比例: {count[1]/len(train_df):.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:03:43.80854Z","iopub.execute_input":"2025-09-05T01:03:43.80934Z","iopub.status.idle":"2025-09-05T01:03:43.828676Z","shell.execute_reply.started":"2025-09-05T01:03:43.809306Z","shell.execute_reply":"2025-09-05T01:03:43.827912Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"查看 “动脉瘤存在” 的情况，有 2484 例不存在，1864 例存在，存在的比例约为 42.87%；查看各具体位置的动脉瘤分布，不同位置存在动脉瘤的数量和比例都不一样，像前交通动脉处存在动脉瘤的有363例，比例约8.35%，是这些位置里比例相对较高的，而左大脑前动脉处存在动脉瘤的仅有46例，比例约 1.06%，是比例较低的。","metadata":{}},{"cell_type":"code","source":"# 人口统计学分析\nprint(\"\\n===== 人口统计学分析 =====\")\nprint(\"性别分布:\")\nprint(train_df['PatientSex'].value_counts())\n\nprint(\"\\n年龄分布:\")\nprint(f\"最小值: {train_df['PatientAge'].min()}\")\nprint(f\"最大值: {train_df['PatientAge'].max()}\")\nprint(f\"平均值: {train_df['PatientAge'].mean():.2f}\")\nprint(f\"中位数: {train_df['PatientAge'].median()}\")\n\nprint(\"\\n成像模式分布:\")\nprint(train_df['Modality'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:03:56.691347Z","iopub.execute_input":"2025-09-05T01:03:56.692157Z","iopub.status.idle":"2025-09-05T01:03:56.702472Z","shell.execute_reply.started":"2025-09-05T01:03:56.692122Z","shell.execute_reply":"2025-09-05T01:03:56.70174Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"人口统计学分析的目的是展示患者的性别、年龄和成像模式的分布情况。结果显示，性别方面女性有 3005 人，男性 1343 人；年龄最小 18 岁，最大 89 岁，平均约 58.47 岁，中位数 60 岁；成像模式里 CTA（ CT 血管造影）有 1808 例，MRA（磁共振血管造影）有 1252 例等。这些结果能让我们了解患者群体在性别、年龄以及所采用成像模式上的分布特点，为后续结合这些因素分析动脉瘤等情况提供基础信息。","metadata":{}},{"cell_type":"markdown","source":"## 3.2 数据清洗","metadata":{}},{"cell_type":"code","source":"# 检查缺失值\nprint(\"缺失值统计:\")\nmissing_values = train_df.isnull().sum()\nprint(missing_values[missing_values > 0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:04:42.375871Z","iopub.execute_input":"2025-09-05T01:04:42.376135Z","iopub.status.idle":"2025-09-05T01:04:42.382798Z","shell.execute_reply.started":"2025-09-05T01:04:42.376117Z","shell.execute_reply":"2025-09-05T01:04:42.382096Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"用来检查数据框 train_df 里的缺失值情况。统计每一列的缺失值数量，然后只把缺失值数量大于 0 的列及其缺失值数量打印出来，方便查看哪些列存在缺失值以及具体缺失了多少。","metadata":{}},{"cell_type":"code","source":"# 处理缺失值\nprint(f\"\\n预处理前缺失值总数: {train_df.isnull().sum().sum()}\")\ntrain_df.fillna(0, inplace=True)\nlocalizers_df.dropna(inplace=True)\nprint(f\"预处理后缺失值总数: {train_df.isnull().sum().sum()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:04:46.046102Z","iopub.execute_input":"2025-09-05T01:04:46.046388Z","iopub.status.idle":"2025-09-05T01:04:46.056343Z","shell.execute_reply.started":"2025-09-05T01:04:46.046369Z","shell.execute_reply":"2025-09-05T01:04:46.0556Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"给训练模型的数据集的缺失值填0，直接删掉localizers_df里有缺失值的行","metadata":{}},{"cell_type":"code","source":"# 异常值处理 - 基于3σ原则，将异常值替换为列均值\nprint(\"\\n异常值处理:\")\nnumeric_cols = ['PatientAge'] + all_label_cols\nfor col in numeric_cols:\n    if col in train_df.columns:\n        z_scores = np.abs(stats.zscore(train_df[col]))\n        outliers = z_scores > 3\n        if outliers.any():\n            mean_val = train_df[col].mean()\n            train_df.loc[outliers, col] = mean_val\n            print(f\"处理 {col} 列的 {outliers.sum()} 个异常值，替换为该列均值 {mean_val:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:04:49.767842Z","iopub.execute_input":"2025-09-05T01:04:49.768681Z","iopub.status.idle":"2025-09-05T01:04:49.792715Z","shell.execute_reply.started":"2025-09-05T01:04:49.76865Z","shell.execute_reply":"2025-09-05T01:04:49.791935Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 重复数据检查\nprint(f\"\\n重复数据: {train_df.duplicated().sum()} 行\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:04:54.658601Z","iopub.execute_input":"2025-09-05T01:04:54.658901Z","iopub.status.idle":"2025-09-05T01:04:54.668214Z","shell.execute_reply.started":"2025-09-05T01:04:54.65888Z","shell.execute_reply":"2025-09-05T01:04:54.667472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 处理标签列数据类型\nfor col in all_label_cols:\n    if col in train_df.columns:\n        if train_df[col].dtype == 'object':\n            # 转换文本标签为数值\n            train_df[col] = train_df[col].apply(lambda x: \n                1.0 if str(x).strip().lower() in ['true', 'yes', 'present', '1'] else\n                0.0 if str(x).strip().lower() in ['false', 'no', 'absent', '0'] else\n                float(x) if pd.notna(x) and str(x).strip() else 0.0\n            )\n        train_df[col] = train_df[col].astype(np.float32)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:04:57.482941Z","iopub.execute_input":"2025-09-05T01:04:57.483198Z","iopub.status.idle":"2025-09-05T01:04:57.491508Z","shell.execute_reply.started":"2025-09-05T01:04:57.483181Z","shell.execute_reply":"2025-09-05T01:04:57.490767Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.3 探索性分析","metadata":{}},{"cell_type":"code","source":"# 可视化标签分布\nplt.figure(figsize=(9, 6))\n# 动脉瘤存在分布\nplt.subplot(2, 2, 1)\ntrain_df['Aneurysm Present'].value_counts().plot(kind='bar')\nplt.title('Distribution of aneurysms')#动脉瘤存在分布\nplt.xlabel('Is there an aneurysm')#是否存在动脉瘤\nplt.ylabel('Counting')#数量\n\n# 性别分布\nplt.subplot(2, 2, 2)\ntrain_df['PatientSex'].value_counts().plot(kind='bar')\nplt.title('Gender Distribution')#性别分布\nplt.xlabel('Gender')\nplt.ylabel('Counting')\n\n# 年龄分布\nplt.subplot(2, 2, 3)\nplt.hist(train_df['PatientAge'].dropna(), bins=30, edgecolor='black')\nplt.title('Age distribution')\nplt.xlabel('age')\nplt.ylabel('Frequency')\n\n# 成像模式分布\nplt.subplot(2, 2, 4)\ntrain_df['Modality'].value_counts().plot(kind='bar')\nplt.title('Imaging Mode Distribution')#成像模式分布\nplt.xlabel('mode')\nplt.ylabel('Counting')\n\nplt.tight_layout()\nplt.savefig('data_distribution.png', dpi=300)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:01.37233Z","iopub.execute_input":"2025-09-05T01:05:01.372622Z","iopub.status.idle":"2025-09-05T01:05:02.608924Z","shell.execute_reply.started":"2025-09-05T01:05:01.3726Z","shell.execute_reply":"2025-09-05T01:05:02.608178Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"在动脉瘤分布上，存在与不存在动脉瘤的样本数量均较多；性别分布中，女性样本数量显著多于男性；年龄分布显示，研究对象年龄主要集中在 40 至 80 岁区间；成像方式分布里，CTA 的使用频次最高，而 MRI T1post 的使用频次最低。","metadata":{}},{"cell_type":"code","source":"# 特征与标签关系分析\nprint(\"\\n特征与标签关系分析:\")\n\n# 性别与动脉瘤存在的关系\nif 'PatientSex' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    sex_aneurysm = pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'])\n    print(\"\\n性别与动脉瘤存在的关系:\")\n    print(sex_aneurysm)\n    \n    # 可视化\n    sex_aneurysm.plot(kind='bar', figsize=(10, 6))\n    plt.title('The Relationship between Gender and the Presence of Aneurysm')#性别与动脉瘤存在的关系\n    plt.xlabel('Gender')\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm.'])\n    plt.savefig('sex_vs_aneurysm.png', dpi=300)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:16.585659Z","iopub.execute_input":"2025-09-05T01:05:16.585944Z","iopub.status.idle":"2025-09-05T01:05:17.126027Z","shell.execute_reply.started":"2025-09-05T01:05:16.585924Z","shell.execute_reply":"2025-09-05T01:05:17.125144Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"展示了性别与动脉瘤存在情况的关联。可以看出女性中未患动脉瘤的有159 例，患动脉瘤的有1406例；男性里未患动脉瘤的有 885 例，患动脉瘤的有 458 例。整体呈现出女性群体中，无论是否患有动脉瘤，数量都多于男性相应情况的特点。","metadata":{}},{"cell_type":"code","source":"# 年龄与动脉瘤存在的关系\nif 'PatientAge' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    plt.figure(figsize=(12, 6))\n    plt.subplot(1, 2, 1)\n    train_df[train_df['Aneurysm Present'] == 0]['PatientAge'].hist(alpha=0.7, label='No aneurysm', bins=30)\n    train_df[train_df['Aneurysm Present'] == 1]['PatientAge'].hist(alpha=0.7, label='with aneurysm', bins=30)\n    plt.title('Age Distribution - by Presence of Aneurysm')#年龄分布 - 按动脉瘤存在\n    plt.xlabel('age')\n    plt.ylabel('Frequency')\n    plt.legend()\n    \n    plt.subplot(1, 2, 1)\n    age_groups = pd.cut(train_df['PatientAge'], bins=5)\n    pd.crosstab(age_groups, train_df['Aneurysm Present']).plot(kind='bar', figsize=(12, 6))\n    plt.title('The relationship between age groups and aneurysm presence')#年龄组与动脉瘤存在的关系\n    plt.xlabel('age group')#年龄组\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n    plt.tight_layout()\n    plt.savefig('age_vs_aneurysm.png', dpi=300)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:24.661633Z","iopub.execute_input":"2025-09-05T01:05:24.662188Z","iopub.status.idle":"2025-09-05T01:05:25.679777Z","shell.execute_reply.started":"2025-09-05T01:05:24.662164Z","shell.execute_reply":"2025-09-05T01:05:25.679014Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"年龄分布直方图显示，在各年龄段中，动脉瘤存在（橙色）与不存在（蓝色）的频次随年龄变化呈现出一定规律，整体上中老年群体的相关频次相对较高。\n\n按年龄组划分的柱状图进一步表明，不同年龄组中，未患动脉瘤（蓝色）和患有动脉瘤（橙色）的人数存在差异，其中 66 - 75 岁年龄组中，两类人数都较多，且各年龄组内患动脉瘤与未患动脉瘤的人数分布也各有特点。","metadata":{}},{"cell_type":"code","source":"# 成像模式与动脉瘤存在的关系\nif 'Modality' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    modality_aneurysm = pd.crosstab(train_df['Modality'], train_df['Aneurysm Present'])\n    print(\"\\n成像模式与动脉瘤存在的关系:\")\n    print(modality_aneurysm)\n    \n    modality_aneurysm.plot(kind='bar', figsize=(10, 6))\n    plt.title('The relationship between imaging modes and aneurysm presence')#成像模式与动脉瘤存在的关系\n    plt.xlabel('Imaging mode')#成像模式\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n    plt.savefig('modality_vs_aneurysm.png', dpi=300)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:32.018225Z","iopub.execute_input":"2025-09-05T01:05:32.018969Z","iopub.status.idle":"2025-09-05T01:05:32.539258Z","shell.execute_reply.started":"2025-09-05T01:05:32.018942Z","shell.execute_reply":"2025-09-05T01:05:32.538517Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"CTA（CT血管造影）成像模式下，未患动脉瘤的有834例，患动脉瘤的有974例；\n\nMRA（磁共振血管造影）成像模式下，未患动脉瘤的有697例，患动脉瘤的有555例；\n\nMRI T1post（磁共振成像T1加权增强扫描）成像模式下，未患动脉瘤的有228例，患动脉瘤的有77例；\n\nMRI T2（磁共振成像T2加权成像）成像模式下，未患动脉瘤的有725例，患动脉瘤的有258例。\n\n整体来看，不同成像模式下，动脉瘤存在与不存在的病例数存在差异，其中CTA成像模式下患动脉瘤的病例数相对较多，MRI T1post成像模式下患动脉瘤的病例数相对较少。","metadata":{}},{"cell_type":"code","source":"# 各位置动脉瘤的相关性分析\nprint(\"\\n各位置动脉瘤的相关性矩阵:\")\nlocation_corr = train_df[location_labels].corr()\nplt.figure(figsize=(12, 10))\nsns.heatmap(location_corr, annot=True, cmap='coolwarm', center=0)\nplt.title('Aneurysm-related heat map at various positions')#各位置动脉瘤相关性热图\nplt.tight_layout()\nplt.savefig('location_correlation.png', dpi=300)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:39.238956Z","iopub.execute_input":"2025-09-05T01:05:39.239652Z","iopub.status.idle":"2025-09-05T01:05:41.017513Z","shell.execute_reply.started":"2025-09-05T01:05:39.239628Z","shell.execute_reply":"2025-09-05T01:05:41.016766Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"各位置动脉瘤的相关性热力图，呈现了不同脑血管位置（如左侧床突下颈内动脉、右侧大脑中动脉等）间动脉瘤存在情况的关联。正相关体现为一个位置易患动脉瘤时，另一个位置也更易患；像左侧大脑中动脉与右侧大脑中动脉交叉处，颜色较深、相关系数为 0.12，二者相关性就较强；而左侧床突下颈内动脉和左侧大脑前动脉交叉处，相关系数接近 0、颜色浅，它们的相关性则较弱。","metadata":{}},{"cell_type":"code","source":"print(\"\\n===== 探索性分析 =====\")\n\n# 可视化标签分布\nplt.figure(figsize=(9, 6))\n\n# 动脉瘤存在分布\nplt.subplot(2, 2, 1)\ntrain_df['Aneurysm Present'].value_counts().plot(kind='bar')\nplt.title('Distribution of aneurysms')#动脉瘤存在分布\nplt.xlabel('Is there an aneurysm')#是否存在动脉瘤\nplt.ylabel('Counting')#数量\n\n# 性别分布\nplt.subplot(2, 2, 2)\ntrain_df['PatientSex'].value_counts().plot(kind='bar')\nplt.title('Gender Distribution')#性别分布\nplt.xlabel('Gender')\nplt.ylabel('Counting')\n\n# 年龄分布\nplt.subplot(2, 2, 3)\nplt.hist(train_df['PatientAge'].dropna(), bins=30, edgecolor='black')\nplt.title('Age distribution')\nplt.xlabel('age')\nplt.ylabel('Frequency')\n\n# 成像模式分布\nplt.subplot(2, 2, 4)\ntrain_df['Modality'].value_counts().plot(kind='bar')\nplt.title('Imaging Mode Distribution')#成像模式分布\nplt.xlabel('mode')\nplt.ylabel('Counting')\n\nplt.tight_layout()\nplt.savefig('data_distribution.png', dpi=300)\nplt.show()\n\n# 特征与标签关系分析\nprint(\"\\n特征与标签关系分析:\")\n\n# 性别与动脉瘤存在的关系\nif 'PatientSex' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    sex_aneurysm = pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'])\n    print(\"\\n性别与动脉瘤存在的关系:\")\n    print(sex_aneurysm)\n    \n    # 可视化\n    sex_aneurysm.plot(kind='bar', figsize=(10, 6))\n    plt.title('The Relationship between Gender and the Presence of Aneurysm')#性别与动脉瘤存在的关系\n    plt.xlabel('Gender')\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm.'])\n    plt.savefig('sex_vs_aneurysm.png', dpi=300)\n    plt.show()\n\n# 年龄与动脉瘤存在的关系\nif 'PatientAge' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    plt.figure(figsize=(12, 6))\n    plt.subplot(1, 2, 1)\n    train_df[train_df['Aneurysm Present'] == 0]['PatientAge'].hist(alpha=0.7, label='无动脉瘤', bins=30)\n    train_df[train_df['Aneurysm Present'] == 1]['PatientAge'].hist(alpha=0.7, label='有动脉瘤', bins=30)\n    plt.title('Age Distribution - by Presence of Aneurysm')#年龄分布 - 按动脉瘤存在\n    plt.xlabel('age')\n    plt.ylabel('Frequency')\n    plt.legend()\n    \n    plt.subplot(1, 2, 2)\n    age_groups = pd.cut(train_df['PatientAge'], bins=5)\n    pd.crosstab(age_groups, train_df['Aneurysm Present']).plot(kind='bar', figsize=(12, 6))\n    plt.title('The relationship between age groups and aneurysm presence')#年龄组与动脉瘤存在的关系\n    plt.xlabel('age group')#年龄组\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n    plt.tight_layout()\n    plt.savefig('age_vs_aneurysm.png', dpi=300)\n    plt.show()\n\n# 成像模式与动脉瘤存在的关系\nif 'Modality' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n    modality_aneurysm = pd.crosstab(train_df['Modality'], train_df['Aneurysm Present'])\n    print(\"\\n成像模式与动脉瘤存在的关系:\")\n    print(modality_aneurysm)\n    \n    modality_aneurysm.plot(kind='bar', figsize=(10, 6))\n    plt.title('The relationship between imaging modes and aneurysm presence')#成像模式与动脉瘤存在的关系\n    plt.xlabel('Imaging mode')#成像模式\n    plt.ylabel('Counting')\n    plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n    plt.savefig('modality_vs_aneurysm.png', dpi=300)\n    plt.show()\n\n# 各位置动脉瘤的相关性分析\nprint(\"\\n各位置动脉瘤的相关性矩阵:\")\nlocation_corr = train_df[location_labels].corr()\nplt.figure(figsize=(12, 10))\nsns.heatmap(location_corr, annot=True, cmap='coolwarm', center=0)\nplt.title('Aneurysm-related heat map at various positions')#各位置动脉瘤相关性热图\nplt.tight_layout()\nplt.savefig('location_correlation.png', dpi=300)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:05:43.636733Z","iopub.execute_input":"2025-09-05T01:05:43.637022Z","iopub.status.idle":"2025-09-05T01:05:48.43079Z","shell.execute_reply.started":"2025-09-05T01:05:43.637001Z","shell.execute_reply":"2025-09-05T01:05:48.429857Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.4 数据集划分","metadata":{}},{"cell_type":"code","source":"print(\"\\n===== 数据集划分 =====\")\n\n# 使用分层抽样确保各类别比例一致\ntrain_df, test_df = train_test_split(\n    train_df, \n    test_size=0.2, \n    random_state=42, \n    stratify=train_df['Aneurysm Present']\n)\ntrain_df, val_df = train_test_split(\n    train_df, \n    test_size=0.125,  # 0.8 * 0.125 = 0.1\n    random_state=42, \n    stratify=train_df['Aneurysm Present']\n)\n\nprint(f\"训练集大小: {len(train_df)}\")\nprint(f\"验证集大小: {len(val_df)}\")\nprint(f\"测试集大小: {len(test_df)}\")\n\n# 显示各数据集标签分布\nprint(\"\\n训练集动脉瘤存在分布:\")\nprint(train_df['Aneurysm Present'].value_counts(normalize=True))\n\nprint(\"\\n验证集动脉瘤存在分布:\")\nprint(val_df['Aneurysm Present'].value_counts(normalize=True))\n\nprint(\"\\n测试集动脉瘤存在分布:\")\nprint(test_df['Aneurysm Present'].value_counts(normalize=True))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:10:19.129145Z","iopub.execute_input":"2025-09-05T01:10:19.129422Z","iopub.status.idle":"2025-09-05T01:10:19.148011Z","shell.execute_reply.started":"2025-09-05T01:10:19.129406Z","shell.execute_reply":"2025-09-05T01:10:19.147296Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"采用分层抽样的方法，先把原始训练数据分成训练集和测试集，测试集占比 0.2；接着把剩下的训练数据再分成训练集和验证集，验证集占比 0.125。\n\n分层抽样依据 “动脉瘤是否存在（Aneurysm Present）” 这个标签来划分的，能确保各数据集里标签的比例一致。结果是训练集有 3043 条数据，验证集 435 条，测试集 870 条，并且三个数据集里 “动脉瘤存在” 和 “不存在” 的比例都很接近，训练集里 “不存在” 约占 57.11%、“存在” 约 42.89%，验证集 “不存在” 约 57.24%、“存在” 约 42.76%，测试集 “不存在” 约 57.12%、“存在” 约 42.88%，这样的划分有利于后续模型训练和评估。","metadata":{}},{"cell_type":"markdown","source":"## 3.5 数据集与数据加载器定义","metadata":{}},{"cell_type":"code","source":"# ===================== 数据集类定义 =====================\nclass AneurysmDataset(Dataset):\n    def __init__(self, df, root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series', \n                 transform=None, is_test=False, location_labels=None):\n        self.df = df.copy()\n        self.root_dir = root_dir\n        self.transform = transform\n        self.is_test = is_test\n        self.location_labels = location_labels\n        self.target_shape = (1, 256, 256)  # (通道数, 高度, 宽度)\n        self.label_shape = (14,)\n        \n        self._validate_and_clean_labels()\n        \n        if not os.path.exists(self.root_dir):\n            raise FileNotFoundError(f\"根目录不存在: {self.root_dir}\")\n            \n        valid_indices = self._filter_valid_series()\n        self.df = self.df.iloc[valid_indices].reset_index(drop=True)\n        print(f\"过滤后保留的有效样本数: {len(self.df)}\")\n\n    def _validate_and_clean_labels(self):\n        if not self.location_labels:\n            raise ValueError(\"location_labels 不能为空\")\n            \n        all_label_cols = ['Aneurysm Present'] + self.location_labels\n        missing_cols = [col for col in all_label_cols if col not in self.df.columns]\n        if missing_cols:\n            raise KeyError(f\"缺少标签列: {missing_cols}\")\n        \n        for col in all_label_cols:\n            self.df[col] = pd.to_numeric(self.df[col], errors='coerce')\n            self.df[col].fillna(0.0, inplace=True)\n            self.df[col] = self.df[col].astype(np.float32)\n\n    def _filter_valid_series(self):\n        valid_indices = []\n        for idx in tqdm(range(len(self.df)), desc=\"过滤有效样本\"):\n            series_id = self.df.iloc[idx]['SeriesInstanceUID']\n            series_path = os.path.join(self.root_dir, series_id)\n            if os.path.exists(series_path) and len(os.listdir(series_path)) > 0:\n                valid_indices.append(idx)\n            else:\n                warn_msg = f\"路径不存在或为空 - {series_path}\"\n                logging.warning(warn_msg)\n        return valid_indices\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        series_id = self.df.iloc[idx]['SeriesInstanceUID']\n        series_path = os.path.join(self.root_dir, series_id)\n        \n        # 初始化返回值\n        image = torch.zeros(self.target_shape, dtype=torch.float32)\n        label = torch.zeros(self.label_shape, dtype=torch.float32)\n\n        try:\n            # 加载DICOM文件列表\n            dicom_files = [f for f in os.listdir(series_path) if f.endswith('.dcm')]\n            if not dicom_files:\n                warn_msg = f\"未找到DICOM文件 - {series_path}\"\n                logging.warning(warn_msg)\n                return image, label\n            \n            # 按数字排序并选择中间切片\n            dicom_files.sort(key=lambda x: int(x.split('.')[0]))\n            mid_idx = len(dicom_files) // 2\n            dicom_path = os.path.join(series_path, dicom_files[mid_idx])\n            \n            # 读取DICOM文件\n            dicom = pydicom.dcmread(dicom_path)\n            \n            # 应用VOI LUT增强对比度\n            pixel_array = apply_voi_lut(dicom.pixel_array, dicom)\n            \n            # 处理3D数据\n            if len(pixel_array.shape) == 3:\n                slice_idx = pixel_array.shape[0] // 2\n                pixel_array = pixel_array[slice_idx, :, :]\n                logging.warning(f\"{series_id} 是3D数据，已取中间切片\")\n            elif len(pixel_array.shape) != 2:\n                error_msg = f\"DICOM图像维度异常: {pixel_array.shape}，series_id: {series_id}\"\n                logging.error(error_msg)\n                return image, label\n            \n            # 转换为float32并应用 rescale 校正\n            pixel_array = pixel_array.astype(np.float32)\n            if hasattr(dicom, 'RescaleSlope') and hasattr(dicom, 'RescaleIntercept'):\n                pixel_array = pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n            \n            # 标准化到0-1范围\n            img_min, img_max = pixel_array.min(), pixel_array.max()\n            if img_max != img_min:\n                pixel_array = (pixel_array - img_min) / (img_max - img_min)\n            \n            # 转换为张量并添加通道维度\n            img_tensor = torch.from_numpy(pixel_array).unsqueeze(0)\n            \n            # 调整尺寸\n            resize_transform = transforms.Resize((self.target_shape[1], self.target_shape[2]))\n            img_tensor = resize_transform(img_tensor)\n            \n            if img_tensor.shape != self.target_shape:\n                raise RuntimeError(f\"图像形状错误: {img_tensor.shape} 预期 {self.target_shape}\")\n            \n            image = img_tensor\n            \n        except Exception as e:\n            error_msg = f\"处理DICOM时出错 {series_path}: {str(e)}\"\n            logging.error(error_msg)\n        \n        # 处理标签\n        try:\n            if not self.is_test:\n                present = float(self.df.iloc[idx]['Aneurysm Present'])\n                locations = self.df.iloc[idx][self.location_labels].values.astype(np.float32)\n                label_np = np.concatenate([[present], locations])\n                label = torch.FloatTensor(label_np)\n        except Exception as e:\n            error_msg = f\"处理标签时出错 {series_id}: {str(e)}\"\n            logging.error(error_msg)\n        \n        # 应用变换\n        if self.transform:\n            try:\n                image = self.transform(image)\n            except Exception as e:\n                logging.error(f\"应用变换时出错: {str(e)}\")\n        \n        return image, label","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:10:28.704553Z","iopub.execute_input":"2025-09-05T01:10:28.704953Z","iopub.status.idle":"2025-09-05T01:10:28.720693Z","shell.execute_reply.started":"2025-09-05T01:10:28.70493Z","shell.execute_reply":"2025-09-05T01:10:28.7201Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"定义了一个动脉瘤检测数据集类，用于加载和处理医学影像数据。做法是先初始化数据集，指定数据路径、变换方式等参数，然后验证清洗标签确保数据有效性，过滤掉无效的影像文件路径，只保留有效样本；\n\n读取 DICOM 格式的医学影像时，选取中间切片处理，经对比度增强、标准化等操作后调整为统一尺寸，同时提取动脉瘤是否存在及位置标签。这样做是为了标准化医学影像数据进行标准化处理，解决数据格式不一、质量参差等问题，目的是为后续模型训练提供格式统一、标签清晰的有效数据，确保模型能稳定输入并学习到有用特征。","metadata":{}},{"cell_type":"code","source":"# ===================== 数据加载器准备 =====================\nprint(\"\\n===== 准备数据加载器 =====\")\n\n# 定义图像变换 - 增加数据增强提高模型泛化能力\ntrain_transform = transforms.Compose([\n    transforms.Resize((256, 256)),\n    transforms.RandomAffine(degrees=10, translate=(0.1, 0.1)),\n    transforms.RandomHorizontalFlip(),\n    transforms.Normalize(mean=[0.5], std=[0.5])\n])\n\nval_transform = transforms.Compose([\n    transforms.Resize((256, 256)),\n    transforms.Normalize(mean=[0.5], std=[0.5])\n])\n\n# 创建数据集实例\ntrain_dataset = AneurysmDataset(\n    train_df, \n    transform=train_transform, \n    location_labels=location_labels\n)\nval_dataset = AneurysmDataset(\n    val_df, \n    transform=val_transform, \n    location_labels=location_labels\n)\ntest_dataset = AneurysmDataset(\n    test_df, \n    transform=val_transform, \n    is_test=False,  # 保持为False以便评估\n    location_labels=location_labels\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:10:34.155045Z","iopub.execute_input":"2025-09-05T01:10:34.155356Z","iopub.status.idle":"2025-09-05T01:13:30.23785Z","shell.execute_reply.started":"2025-09-05T01:10:34.155336Z","shell.execute_reply":"2025-09-05T01:13:30.237013Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"准备动脉瘤检测的数据加载器。\n\n先定义不同的数据变换方式，训练集用随机旋转、平移、水平翻转等增强操作并标准化，验证集和测试集仅做标准化和尺寸调整；\n\n基于划分好的训练、验证、测试数据框，创建对应数据集实例并应用相应变换。这样做是因为训练时用数据增强可增加样本多样性，提升模型泛化能力，而验证和测试需保持数据一致性以准确评估模型；目的是为模型训练、验证和测试提供符合要求的输入数据，确保训练效果和评估准确性。","metadata":{}},{"cell_type":"code","source":"# 创建数据加载器\nbatch_size = 16\nnum_workers = 0 if device.type == 'cpu' else 2\ntrain_loader = DataLoader(\n    train_dataset, \n    batch_size=batch_size, \n    shuffle=True, \n    num_workers=num_workers,\n    pin_memory=True if device.type == 'cuda' else False\n)\nval_loader = DataLoader(\n    val_dataset, \n    batch_size=batch_size, \n    shuffle=False, \n    num_workers=num_workers,\n    pin_memory=True if device.type == 'cuda' else False\n)\ntest_loader = DataLoader(\n    test_dataset, \n    batch_size=batch_size, \n    shuffle=False, \n    num_workers=num_workers,\n    pin_memory=True if device.type == 'cuda' else False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:16:53.541783Z","iopub.execute_input":"2025-09-05T01:16:53.542381Z","iopub.status.idle":"2025-09-05T01:16:53.547204Z","shell.execute_reply.started":"2025-09-05T01:16:53.542358Z","shell.execute_reply":"2025-09-05T01:16:53.546651Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"创建数据加载器。\n\n设置批次大小为16，根据设备类型（CPU 或 GPU）确定工作进程数和是否启用内存锁定，分别为训练、验证、测试数据集创建加载器，其中训练集启用数据打乱，验证集和测试集不打乱。这样做是因为批量加载数据能提高处理效率，训练时打乱数据可避免模型学习顺序偏差，根据设备配置参数能优化加载性能；目的是为模型训练、验证和测试提供高效、合适的数据输入方式，保障训练过程稳定高效，评估结果准确。","metadata":{}},{"cell_type":"markdown","source":"# 第4章 模型构建与训练","metadata":{}},{"cell_type":"code","source":"# 模型定义\nclass AneurysmViT(nn.Module):\n    def __init__(self, num_labels=14):\n        super(AneurysmViT, self).__init__()\n        \n        self.vit_config = ViTConfig(\n            image_size=256,\n            patch_size=32,\n            num_channels=1,\n            num_labels=num_labels,\n            hidden_size=768,\n            num_hidden_layers=12,\n            num_attention_heads=12,\n            intermediate_size=3072,\n            dropout=0.3,  # 增加dropout防止过拟合\n            attention_probs_dropout_prob=0.3\n        )\n        \n        self.vit = ViTModel(self.vit_config)\n        \n        # 改进分类头\n        self.classifier = nn.Sequential(\n            nn.Linear(768, 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Dropout(0.3),\n            nn.Linear(512, 256),\n            nn.BatchNorm1d(256),\n            nn.ReLU(),\n            nn.Dropout(0.2),\n            nn.Linear(256, num_labels)\n        )\n        \n    def forward(self, x):\n        outputs = self.vit(pixel_values=x)\n        cls_token = outputs.last_hidden_state[:, 0, :]\n        logits = self.classifier(cls_token)\n        return torch.sigmoid(logits)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:16:58.592823Z","iopub.execute_input":"2025-09-05T01:16:58.593092Z","iopub.status.idle":"2025-09-05T01:16:58.599487Z","shell.execute_reply.started":"2025-09-05T01:16:58.593071Z","shell.execute_reply":"2025-09-05T01:16:58.598637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"定义了一个用于动脉瘤检测的视觉Transformer（ViT）模型。\n\n构建基于ViT的网络结构，输入图像尺寸为 256×256（匹配预处理后的医学影像），1 个通道（灰度医学图像），将图像分割为 32×32 的 patch；设置 12 层隐藏层、12 个注意力头，隐藏层维度 768，中间层维度 3072，同时通过 0.3 的 dropout 和注意力 dropout 缓解过拟合。\n\n以预定义的配置初始化 ViT 基础模型，再设计分类头：通过三层线性层（768→512→256→14）逐步压缩特征，每层配合批归一化稳定训练、ReLU 激活引入非线性，穿插 dropout 进一步抑制过拟合，通过使用sigmoid输出 0-1 范围的概率（对应 14 个标签的预测）.\n\n这样设计是因为 ViT 的自注意力机制擅长捕捉图像全局特征，适合检测医学影像中可能分布在不同位置的动脉瘤，多层分类头能增强特征转换能力，而多个 dropout 和批归一化设置则能避免模型在有限医学数据上过度拟合。\n","metadata":{}},{"cell_type":"code","source":"# 初始化模型\nmodel = AneurysmViT().to(device)\nprint(\"\\n模型结构:\")\nprint(model)\n\n# 计算类别权重 - 更精确地处理类别不平衡\npresent_count = train_df['Aneurysm Present'].sum()\ntotal_count = len(train_df)\nweight_positive = total_count / (2 * present_count) if present_count > 0 else 1.0\nweight_negative = total_count / (2 * (total_count - present_count)) if (total_count - present_count) > 0 else 1.0\n\nprint(f\"\\n类别权重 - 阳性: {weight_positive:.2f}, 阴性: {weight_negative:.2f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:17:13.88481Z","iopub.execute_input":"2025-09-05T01:17:13.885672Z","iopub.status.idle":"2025-09-05T01:17:15.685807Z","shell.execute_reply.started":"2025-09-05T01:17:13.885641Z","shell.execute_reply":"2025-09-05T01:17:15.685047Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"这里主要做了两件事：\n\n一是初始化动脉瘤检测模型并显示其结构，将模型加载到指定的计算设备（CPU 或 GPU）。\n\n二是计算类别权重来处理数据中的类别不平衡问题，通过统计训练集中 “动脉瘤存在” 的样本数量，用总样本数分别别除以两倍的阳性样本数和阴性样本数，得到阳性和阴性样本的权重。\n\n这样做是因为医学数据中，存在动脉瘤的阳性样本往往远少于阴性样本，若不处理，模型可能更倾向于预测为多数类的阴性，影响检测效果。计算类别权重的目的是给少数的阳性样本赋予更高权重，让模型在训练时更重视这类样本，从而提升对动脉瘤的检测能力，确保模型不会因数据不平衡而偏向多数类。\n\n“阳性: 1.17，阴性: 0.88”，通过给 “阳性” 类别分配稍高的权重，给 “阴性” 类别分配稍低的权重，能让模型在训练过程中更关注 “阳性” 这类样本，从而提升对 “阳性” 情况的检测效果，这种权重设置是合理且正常的，是为了应对数据本身的不平衡性而采取的常见策略。","metadata":{}},{"cell_type":"markdown","source":"- 阳性样本权重 $$w_{positive} = \\frac{N}{2 \\times N_{positive}}$$\n- 阴性样本权重 $$w_{negative} = \\frac{N}{2 \\times N_{negative}}$$\n","metadata":{}},{"cell_type":"code","source":"# 加权损失函数\nclass WeightedBCELoss(nn.Module):\n    def __init__(self, pos_weight=1.0):\n        super().__init__()\n        self.pos_weight = pos_weight\n        \n    def forward(self, output, target):\n        # 对第一个标签（是否存在动脉瘤）应用特殊权重\n        weights = torch.ones_like(target, device=device)\n        weights[:, 0] = self.pos_weight\n        \n        bce_loss = nn.BCELoss(reduction='none')(output, target)\n        return (bce_loss * weights).mean()\n\n# 优化器和调度器\ncriterion = WeightedBCELoss(pos_weight=weight_positive)\noptimizer = torch.optim.AdamW(\n    model.parameters(), \n    lr=1e-4, \n    weight_decay=1e-5\n)\nscheduler = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(\n    optimizer, \n    T_0=5, \n    T_mult=2, \n    eta_min=1e-6\n)\n\n# 训练参数\nnum_epochs = 15\nbest_val_score = 0.0\nearly_stop_patience = 5\nearly_stop_counter = 0\n\n# 记录训练过程指标\ntrain_losses = []\nval_losses = []\ntrain_aucs = []\nval_aucs = []","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:17:20.379205Z","iopub.execute_input":"2025-09-05T01:17:20.379512Z","iopub.status.idle":"2025-09-05T01:17:20.38679Z","shell.execute_reply.started":"2025-09-05T01:17:20.37948Z","shell.execute_reply":"2025-09-05T01:17:20.386096Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"这里实现了一个加权二值交叉损失函数，专门对第一个标签（是否存在动脉瘤）赋予更高权重，在医学影像任务中给第一个标签（是否存在动脉瘤）赋予更高权重，主要是为了解决类别不平衡问题，以应对类别不平衡问题，实际数据中，\"存在动脉瘤\" 的阳性样本往往远少于 \"不存在\" 的阴性样本，如果用普通损失函数，模型可能倾向于更多预测阴性（因为更容易猜对），导致漏诊关键的阳性病例。\n\n配置了优化器（AdamW）和学习率调度器（余弦退火重启），设置了 15 个训练轮次，并加入早停机制（容忍 5 轮无提升则停止）。训练过程中会记录损失值和 AUC 指标。","metadata":{}},{"cell_type":"code","source":"# ===================== 训练循环 =====================\nprint(\"\\n===== 开始训练 =====\")\nfor epoch in range(num_epochs):\n    model.train()\n    train_loss = 0.0\n    train_auc = 0.0\n    train_auc_count = 0\n    \n    # 训练循环带进度条\n    train_pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 训练\")\n    for images, labels in train_pbar:\n        images = images.to(device)\n        labels = labels.to(device)\n        \n        optimizer.zero_grad()\n        outputs = model(images)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        \n        # 梯度裁剪防止梯度爆炸\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n        \n        optimizer.step()\n        \n        train_loss += loss.item() * images.size(0)\n        \n        # 计算AUC\n        labels_np = labels.cpu().detach().numpy()\n        outputs_np = outputs.cpu().detach().numpy()\n        \n        try:\n            if len(np.unique(labels_np[:, 0])) >= 2:\n                auc = roc_auc_score(labels_np[:, 0], outputs_np[:, 0])\n                train_auc += auc * images.size(0)\n                train_auc_count += images.size(0)\n        except ValueError as e:\n            logging.error(f\"训练AUC计算错误: {str(e)}\")\n        \n        # 更新进度条\n        train_pbar.set_postfix({\"batch_loss\": loss.item()})\n    \n    # 计算平均训练损失和AUC\n    train_loss /= len(train_loader.dataset)\n    train_auc = train_auc / train_auc_count if train_auc_count > 0 else 0.5\n    \n    # 验证阶段\n    model.eval()\n    val_loss = 0.0\n    val_scores = []\n    val_main_aucs = []\n    \n    with torch.no_grad():\n        val_pbar = tqdm(val_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 验证\")\n        for images, labels in val_pbar:\n            images = images.to(device)\n            labels = labels.to(device)\n            \n            outputs = model(images)\n            loss = criterion(outputs, labels)\n            val_loss += loss.item() * images.size(0)\n            \n            # 计算AUC\n            labels_np = labels.cpu().numpy()\n            outputs_np = outputs.cpu().numpy()\n            auc_scores = []\n            \n            for i in range(labels_np.shape[1]):\n                try:\n                    if len(np.unique(labels_np[:, i])) >= 2:\n                        auc_scores.append(roc_auc_score(labels_np[:, i], outputs_np[:, i]))\n                    else:\n                        auc_scores.append(0.5)\n                except:\n                    auc_scores.append(0.5)\n            \n            main_auc = auc_scores[0]\n            other_auc = np.mean(auc_scores[1:])\n            val_scores.append(0.5 * (main_auc + other_auc))\n            val_main_aucs.append(main_auc)\n            \n            # 更新进度条\n            val_pbar.set_postfix({\"batch_loss\": loss.item()})\n    \n    # 计算平均验证损失和分数\n    val_loss /= len(val_loader.dataset)\n    val_score = np.mean(val_scores)\n    val_main_auc = np.mean(val_main_aucs)\n    \n    # 更新学习率调度器\n    scheduler.step()\n    \n    # 记录指标\n    train_losses.append(train_loss)\n    val_losses.append(val_loss)\n    train_aucs.append(train_auc)\n    val_aucs.append(val_main_auc)\n    \n    # 保存最佳模型\n    if val_score > best_val_score:\n        best_val_score = val_score\n        torch.save(model.state_dict(), 'best_model.pth')\n        early_stop_counter = 0\n        print(\"保存最佳模型!\")\n    else:\n        early_stop_counter += 1\n        if early_stop_counter >= early_stop_patience:\n            print(\"早停: 验证分数连续多轮无提升\")\n            break\n    \n    # 打印 epoch 结果\n    print(f\"\\nEpoch {epoch+1}/{num_epochs}\")\n    print(f\"训练损失: {train_loss:.4f} | 训练主标签AUC: {train_auc:.4f}\")\n    print(f\"验证损失: {val_loss:.4f} | 验证主标签AUC: {val_main_auc:.4f} | 验证综合分数: {val_score:.4f}\")\n    print(\"-\" * 80)\n\nprint(\"训练完成!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T01:17:25.872642Z","iopub.execute_input":"2025-09-05T01:17:25.87291Z","iopub.status.idle":"2025-09-05T02:05:24.257101Z","shell.execute_reply.started":"2025-09-05T01:17:25.872891Z","shell.execute_reply":"2025-09-05T02:05:24.256201Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"训练阶段用带进度条的循环逐批处理数据，通过梯度裁剪防止模型训练中出现梯度爆炸，并重点跟踪第一个标签（是否存在动脉瘤）的 AUC 指标；验证阶段则计算所有标签的 AUC，再通过主标签与其他标签的加权平均得到综合评分，既保证关键指标（主标签）的准确性，又兼顾其他辅助标签的表现。\n\n这样设计的原因在于：医疗影像任务对关键指标（如动脉瘤检出）要求极高，需要重点监控；同时通过早停机制和保存最佳模型，能有效避免过拟合，确保模型在实际应用中更可靠；而综合评分策略则平衡了多标签任务中不同标签的重要性，符合医学诊断中 \"既需准确判断是否患病，也需定位病灶\" 的实际需求。","metadata":{}},{"cell_type":"markdown","source":"1. 平均损失计算：\n   $$\\text{avg\\_loss} = \\frac{\\sum_{i=1}^{N} (\\text{batch\\_loss}_i \\times \\text{batch\\_size}_i)}{\\text{total\\_samples}}$$\n\n2. 梯度裁剪（L₂范数限制）：\n   $$\\text{clipped\\_grad} = \\text{grad} \\times \\min\\left(1, \\frac{\\text{max\\_norm}}{\\|\\text{grad}\\|_2}\\right)$$\n   （其中$\\|\\text{grad}\\|_2 = \\sqrt{\\sum_{k} g_k^2}$，$\\text{max\\_norm}=1.0$）\n\n3. 验证综合评分：\n   $$\\text{val\\_score} = 0.5 \\times \\left(\\text{main\\_auc} + \\frac{1}{M-1}\\sum_{i=2}^{M} \\text{auc}_i\\right)$$\n\n4. 批次加权AUC均值：\n   $$\\text{avg\\_auc} = \\frac{\\sum_{i=1}^{N} (\\text{batch\\_auc}_i \\times \\text{batch\\_size}_i)}{\\sum_{i=1}^{N} \\text{batch\\_size}_i}$$","metadata":{}},{"cell_type":"code","source":"# ===================== 绘制训练历史 =====================\nprint(\"\\n绘制训练历史...\")\nplt.figure(figsize=(14, 6))\n\n# 损失曲线\nplt.subplot(1, 2, 1)\nplt.plot(train_losses, label='Training Loss')\nplt.plot(val_losses, label='Validation loss')\nplt.title('Training and Validation Loss Curve')#训练和验证损失曲线\nplt.xlabel('Epoch')\nplt.ylabel('Loss Value')\nplt.legend()\nplt.grid(True, linestyle='--', alpha=0.7)\n\n# AUC曲线\nplt.subplot(1, 2, 2)\nplt.plot(train_aucs, label='Training AUC')\nplt.plot(val_aucs, label='Verification AUC')\nplt.title('Training and Validation AUC Curve')#训练和验证AUC曲线\nplt.xlabel('Epoch')\nplt.ylabel('AUC value')\nplt.legend()\nplt.grid(True, linestyle='--', alpha=0.7)\n\nplt.tight_layout()\nplt.savefig('training_history.png', dpi=300)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:35:52.645337Z","iopub.execute_input":"2025-09-05T02:35:52.646077Z","iopub.status.idle":"2025-09-05T02:35:53.792721Z","shell.execute_reply.started":"2025-09-05T02:35:52.646037Z","shell.execute_reply":"2025-09-05T02:35:53.791986Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 第5章 模型评估与分析","metadata":{}},{"cell_type":"code","source":"print(\"\\n===== 测试集评估 =====\")\n# 加载最佳模型\nmodel.load_state_dict(torch.load('best_model.pth'))\nmodel.eval()\n\ntest_scores = []\ntest_aucs = []\nall_auc_scores = []\n\nwith torch.no_grad():\n    test_pbar = tqdm(test_loader, desc=\"测试集评估\")\n    for images, labels in test_pbar:\n        images = images.to(device)\n        labels = labels.to(device)\n        \n        outputs = model(images)\n        \n        # 计算AUC\n        labels_np = labels.cpu().numpy()\n        outputs_np = outputs.cpu().numpy()\n        \n        # 计算每个标签的AUC\n        auc_scores = []\n        for i in range(labels_np.shape[1]):\n            try:\n                auc = roc_auc_score(labels_np[:, i], outputs_np[:, i])\n                auc_scores.append(auc)\n            except ValueError:\n                auc_scores.append(0.5)\n        \n        all_auc_scores.append(auc_scores)\n        \n        # 计算加权分数\n        main_auc = auc_scores[0]\n        other_auc = np.mean(auc_scores[1:])\n        weighted_score = 0.5 * (main_auc + other_auc)\n        \n        test_scores.append(weighted_score)\n        test_aucs.append(main_auc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:36:07.761122Z","iopub.execute_input":"2025-09-05T02:36:07.761715Z","iopub.status.idle":"2025-09-05T02:37:05.002037Z","shell.execute_reply.started":"2025-09-05T02:36:07.761692Z","shell.execute_reply":"2025-09-05T02:37:05.001319Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"在测试集上评估训练好的最佳模型性能。\n\n加载之前保存的最佳模型参数，然后将模型设为评估模式。在不计算梯度的情况下，用进度条遍历测试集数据，对每张图像进行预测后，计算每个标签的AUC值（若计算出错则用 0.5 替代）。按照与验证阶段一致的标准，将第一个标签（主标签）的AUC与其他标签的平均AUC各取50%权重，得到综合评分。整个过程会记录所有标签的 AUC、主标签 AUC 及综合评分，以此全面评估模型在独立测试集上的表现，确保评估结果能反映模型的真实泛化能力。","metadata":{}},{"cell_type":"markdown","source":"1. 单个标签的AUC计算（ROC曲线下面积）：\n   $$\\text{AUC}_i = \\int_{0}^{1} \\text{TPR}_i(FPR) \\, d(FPR)$$\n   （其中$\\text{TPR}_i$为第$i$个标签的真阳性率，$FPR$为假阳性率）\n\n2. 其他标签的平均AUC：\n   $$\\text{other\\_auc} = \\frac{1}{M-1} \\sum_{i=2}^{M} \\text{AUC}_i$$\n   （$M$为总标签数，$i=2$到$M$为非主标签）\n\n3. 测试集综合加权评分：\n   $$\\text{weighted\\_score} = 0.5 \\times (\\text{main\\_auc} + \\text{other\\_auc})$$\n   （$\\text{main\\_auc}$为第一个标签的AUC）\n\n4. 测试集整体评分（所有批次平均）：\n   $$\\text{test\\_score} = \\frac{1}{N} \\sum_{j=1}^{N} \\text{weighted\\_score}_j$$\n   （$N$为测试集批次数，$\\text{weighted\\_score}_j$为第$j$批次的综合评分）","metadata":{}},{"cell_type":"code","source":"# 计算平均分数\ntest_score = np.mean(test_scores)\ntest_auc = np.mean(test_aucs)\n\nprint(f\"\\n测试集主标签AUC: {test_auc:.4f}\")\nprint(f\"测试集综合分数: {test_score:.4f}\")\n\n# 计算每个位置的平均AUC\nmean_auc_per_label = np.mean(all_auc_scores, axis=0)\nprint(\"\\n每个标签的平均AUC:\")\nfor i, label in enumerate(['Aneurysm Present'] + location_labels):\n    print(f\"{label}: {mean_auc_per_label[i]:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:38:04.980248Z","iopub.execute_input":"2025-09-05T02:38:04.981013Z","iopub.status.idle":"2025-09-05T02:38:04.987387Z","shell.execute_reply.started":"2025-09-05T02:38:04.980983Z","shell.execute_reply":"2025-09-05T02:38:04.986595Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"这段代码的核心作用是将测试集评估过程中分散计算的批次指标（综合评分、主标签 AUC、各标签 AUC）进行汇总与量化，最终输出清晰、可解读的整体性能结果。\n\n使用这段代码是因为模型评估不能仅依赖单一批次的表现，需要通过 “计算平均值” 消除批次随机性对结果的影响，得到测试集上的稳定性能指标。\n\n比如用测试集主标签 AUC 反映核心任务（动脉瘤是否存在）的判断准确性，用综合评分平衡主标签与其他位置标签的整体表现，避免单一指标的片面性。同时，逐一输出每个标签的平均 AUC，能精准定位模型在不同任务（如不同位置动脉瘤检测）上的优势与不足（比如某类位置标签 AUC 偏低，可能提示该位置特征学习不充分），为模型性能的整体验收提供量化依据。","metadata":{}},{"cell_type":"markdown","source":"1. 测试集主标签平均AUC：  \n$$\\text{test\\_auc} = \\frac{1}{N} \\sum_{k=1}^{N} \\text{test\\_auc}_k$$  \n\n2. 测试集平均综合分数：  \n$$\\text{test\\_score} = \\frac{1}{N} \\sum_{k=1}^{N} \\text{weighted\\_score}_k$$  \n其中，$\\text{weighted\\_score}_k = 0.5 \\times \\left( \\text{main\\_auc}_k + \\frac{1}{M} \\sum_{m=1}^{M} \\text{location\\_auc}_{k,m} \\right)$  \n\n3. 第$i$个标签的平均AUC：  \n$$\\text{mean\\_auc\\_per\\_label}_i = \\frac{1}{N} \\sum_{k=1}^{N} \\text{auc}_{k,i}$$  \n\n（注：$N$为测试集总批次数，$M$为位置标签总数；$\\text{test\\_auc}_k$、$\\text{main\\_auc}_k$为第$k$批次主标签AUC；$\\text{location\\_auc}_{k,m}$为第$k$批次第$m$个位置标签AUC；$\\text{auc}_{k,i}$为第$k$批次第$i$个标签AUC）","metadata":{}},{"cell_type":"code","source":"# ===================== 可视化预测结果 =====================\nprint(\"\\n可视化预测结果...\")\ndef visualize_predictions(model, dataset, num_samples=5):\n    model.eval()\n    indices = np.random.choice(len(dataset), num_samples, replace=False)\n    \n    plt.figure(figsize=(15, 4*num_samples))\n    for i, idx in enumerate(indices):\n        image, label = dataset[idx]\n        image_tensor = image.unsqueeze(0).to(device)\n        \n        with torch.no_grad():\n            pred = model(image_tensor).cpu().numpy()[0]\n        \n        # 显示图像\n        plt.subplot(num_samples, 2, 2*i+1)\n        plt.imshow(image[0], cmap='gray')\n        plt.title(f\"True value: {label[0].item():.0f}, Predicted Value: {pred[0]:.2f}\")\n        plt.axis('off')\n        \n        # 显示位置预测\n        plt.subplot(num_samples, 2, 2*i+2)\n        plt.barh(location_labels, pred[1:])\n        plt.xlim(0, 1)\n        plt.title(\"Position Prediction Probability\")#位置预测概率\n    \n    plt.tight_layout()\n    plt.savefig('predictions_visualization.png', dpi=300)\n    plt.show()\n\n# 可视化测试集预测结果\nvisualize_predictions(model, test_dataset, num_samples=5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:38:16.03484Z","iopub.execute_input":"2025-09-05T02:38:16.035088Z","iopub.status.idle":"2025-09-05T02:38:20.387564Z","shell.execute_reply.started":"2025-09-05T02:38:16.03507Z","shell.execute_reply":"2025-09-05T02:38:20.386778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\n生成提交文件...\")\ndef generate_submission(model, test_loader, dataset, device):\n    model.eval()\n    all_preds = []\n    all_ids = []\n    \n    with torch.no_grad():\n        pbar = tqdm(test_loader, desc=\"生成预测结果\")\n        for images, _ in pbar:\n            images = images.to(device)\n            outputs = model(images)\n            preds = outputs.cpu().numpy()\n            all_preds.append(preds)\n    \n    # 获取真实的SeriesInstanceUID\n    all_ids = dataset.df['SeriesInstanceUID'].values\n    \n    all_preds = np.concatenate(all_preds)\n    \n    # 创建提交DataFrame\n    submission_df = pd.DataFrame(all_preds, columns=['Aneurysm Present'] + location_labels)\n    submission_df.insert(0, 'SeriesInstanceUID', all_ids)\n    \n    return submission_df\n\n# 生成提交文件（保存为parquet格式，替换原csv逻辑）\nsubmission_df = generate_submission(model, test_loader, test_dataset, device)\n# 使用snappy压缩减少文件体积，兼容大多数比赛平台\nsubmission_df.to_parquet('submission.parquet', index=False, compression='snappy')\n\n# 验证文件是否生成成功\nif os.path.exists('submission.parquet'):\n    file_size = os.path.getsize('submission.parquet') / 1024 / 1024  # 转换为MB\n    print(f\"提交文件已成功保存为 submission.parquet，文件大小: {file_size:.2f} MB\")\nelse:\n    print(\"错误：提交文件未生成\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-05T02:38:36.74642Z","iopub.execute_input":"2025-09-05T02:38:36.746996Z","iopub.status.idle":"2025-09-05T02:39:21.023504Z","shell.execute_reply.started":"2025-09-05T02:38:36.746972Z","shell.execute_reply":"2025-09-05T02:39:21.022783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"----","metadata":{}},{"cell_type":"code","source":"# import os\n# import numpy as np\n# import pandas as pd\n# import pydicom\n# import torch\n# import torch.nn as nn\n# from torch.utils.data import Dataset, DataLoader\n# from torchvision import transforms\n# from transformers import ViTModel, ViTConfig\n# from sklearn.model_selection import train_test_split\n# from sklearn.metrics import roc_auc_score\n# import matplotlib.pyplot as plt\n# from tqdm import tqdm\n# import warnings\n# import logging\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\n# import seaborn as sns\n# from scipy import stats\n\n# # 配置日志记录\n# logging.basicConfig(\n#     filename='dicom_processing.log',\n#     level=logging.WARNING,\n#     format='%(asctime)s - %(levelname)s - %(message)s',\n#     datefmt='%Y-%m-%d %H:%M:%S'\n# )\n# warnings.filterwarnings('ignore')\n\n# # 设置随机种子保证可重复性\n# torch.manual_seed(42)\n# np.random.seed(42)\n\n# # 检查GPU可用性\n# device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n# print(f\"使用设备: {device}\")\n\n# # ===================== 数据描述性分析 =====================\n# print(\"===== 数据描述性分析 =====\")\n\n# # 加载训练标签\n# print(\"加载数据...\")\n# train_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\n# localizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\n# # 基本数据信息\n# print(f\"训练数据形状: {train_df.shape}\")\n# print(f\"定位器数据形状: {localizers_df.shape}\")\n\n# # 显示数据基本信息\n# print(\"\\n训练数据基本信息:\")\n# print(train_df.info())\n\n# print(\"\\n训练数据描述性统计:\")\n# print(train_df.describe())\n\n# # 定义动脉瘤位置标签\n# location_labels = [\n#     'Left Infraclinoid Internal Carotid Artery',\n#     'Right Infraclinoid Internal Carotid Artery',\n#     'Left Supraclinoid Internal Carotid Artery',\n#     'Right Supraclinoid Internal Carotid Artery',\n#     'Left Middle Cerebral Artery',\n#     'Right Middle Cerebral Artery',\n#     'Anterior Communicating Artery',\n#     'Left Anterior Cerebral Artery',\n#     'Right Anterior Cerebral Artery',\n#     'Left Posterior Communicating Artery',\n#     'Right Posterior Communicating Artery',\n#     'Basilar Tip',\n#     'Other Posterior Circulation'\n# ]\n\n# all_label_cols = ['Aneurysm Present'] + location_labels\n\n# # 检查标签分布\n# print(\"\\n===== 标签分布分析 =====\")\n# print(\"\\n主要标签 - 动脉瘤存在分布:\")\n# print(train_df['Aneurysm Present'].value_counts())\n# print(train_df['Aneurysm Present'].value_counts(normalize=True))\n\n# print(\"\\n各位置动脉瘤分布:\")\n# for col in location_labels:\n#     if col in train_df.columns:\n#         count = train_df[col].value_counts()\n#         print(f\"\\n{col}:\")\n#         print(f\"存在: {count.get(1, 0)}, 不存在: {count.get(0, 0)}\")\n#         if 1 in count:\n#             print(f\"比例: {count[1]/len(train_df):.4f}\")\n\n# # 人口统计学分析\n# print(\"\\n===== 人口统计学分析 =====\")\n# print(\"性别分布:\")\n# print(train_df['PatientSex'].value_counts())\n\n# print(\"\\n年龄分布:\")\n# print(f\"最小值: {train_df['PatientAge'].min()}\")\n# print(f\"最大值: {train_df['PatientAge'].max()}\")\n# print(f\"平均值: {train_df['PatientAge'].mean():.2f}\")\n# print(f\"中位数: {train_df['PatientAge'].median()}\")\n\n# print(\"\\n成像模式分布:\")\n# print(train_df['Modality'].value_counts())\n\n# # ===================== 数据清洗 =====================\n# print(\"\\n===== 数据清洗 =====\")\n\n# # 检查缺失值\n# print(\"缺失值统计:\")\n# missing_values = train_df.isnull().sum()\n# print(missing_values[missing_values > 0])\n\n# # 处理缺失值\n# print(f\"\\n预处理前缺失值总数: {train_df.isnull().sum().sum()}\")\n# train_df.fillna(0, inplace=True)\n# localizers_df.dropna(inplace=True)\n# print(f\"预处理后缺失值总数: {train_df.isnull().sum().sum()}\")\n\n# # 异常值检测 - 使用3σ原则\n# print(\"\\n异常值检测:\")\n# numeric_cols = ['PatientAge'] + all_label_cols\n# for col in numeric_cols:\n#     if col in train_df.columns:\n#         z_scores = np.abs(stats.zscore(train_df[col]))\n#         outliers = np.sum(z_scores > 3)\n#         print(f\"{col}: {outliers} 个异常值\")\n\n# # 重复数据检查\n# print(f\"\\n重复数据: {train_df.duplicated().sum()} 行\")\n\n# # 处理标签列数据类型\n# for col in all_label_cols:\n#     if col in train_df.columns:\n#         if train_df[col].dtype == 'object':\n#             # 转换文本标签为数值\n#             train_df[col] = train_df[col].apply(lambda x: \n#                 1.0 if str(x).strip().lower() in ['true', 'yes', 'present', '1'] else\n#                 0.0 if str(x).strip().lower() in ['false', 'no', 'absent', '0'] else\n#                 float(x) if pd.notna(x) and str(x).strip() else 0.0\n#             )\n#         train_df[col] = train_df[col].astype(np.float32)\n\n# # ===================== 探索性分析 =====================\n# print(\"\\n===== 探索性分析 =====\")\n\n# # 可视化标签分布\n# plt.figure(figsize=(15, 10))\n\n# # 动脉瘤存在分布\n# plt.subplot(2, 2, 1)\n# train_df['Aneurysm Present'].value_counts().plot(kind='bar')\n# plt.title('Distribution of aneurysms')#动脉瘤存在分布\n# plt.xlabel('Is there an aneurysm')#是否存在动脉瘤\n# plt.ylabel('Counting')#数量\n\n# # 性别分布\n# plt.subplot(2, 2, 2)\n# train_df['PatientSex'].value_counts().plot(kind='bar')\n# plt.title('Gender Distribution')#性别分布\n# plt.xlabel('Gender')\n# plt.ylabel('Counting')\n\n# # 年龄分布\n# plt.subplot(2, 2, 3)\n# plt.hist(train_df['PatientAge'].dropna(), bins=30, edgecolor='black')\n# plt.title('Age distribution')\n# plt.xlabel('age')\n# plt.ylabel('Frequency')\n\n# # 成像模式分布\n# plt.subplot(2, 2, 4)\n# train_df['Modality'].value_counts().plot(kind='bar')\n# plt.title('Imaging Mode Distribution')#成像模式分布\n# plt.xlabel('mode')\n# plt.ylabel('Counting')\n\n# plt.tight_layout()\n# plt.savefig('data_distribution.png', dpi=300)\n# plt.show()\n\n# # 特征与标签关系分析\n# print(\"\\n特征与标签关系分析:\")\n\n# # 性别与动脉瘤存在的关系\n# if 'PatientSex' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n#     sex_aneurysm = pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'])\n#     print(\"\\n性别与动脉瘤存在的关系:\")\n#     print(sex_aneurysm)\n    \n#     # 可视化\n#     sex_aneurysm.plot(kind='bar', figsize=(10, 6))\n#     plt.title('The Relationship between Gender and the Presence of Aneurysm')#性别与动脉瘤存在的关系\n#     plt.xlabel('Gender')\n#     plt.ylabel('Counting')\n#     plt.legend(['Aneurysm-free', 'There is an aneurysm.'])\n#     plt.savefig('sex_vs_aneurysm.png', dpi=300)\n#     plt.show()\n\n# # 年龄与动脉瘤存在的关系\n# if 'PatientAge' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n#     plt.figure(figsize=(12, 6))\n#     plt.subplot(1, 2, 1)\n#     train_df[train_df['Aneurysm Present'] == 0]['PatientAge'].hist(alpha=0.7, label='无动脉瘤', bins=30)\n#     train_df[train_df['Aneurysm Present'] == 1]['PatientAge'].hist(alpha=0.7, label='有动脉瘤', bins=30)\n#     plt.title('Age Distribution - by Presence of Aneurysm')#年龄分布 - 按动脉瘤存在\n#     plt.xlabel('age')\n#     plt.ylabel('Frequency')\n#     plt.legend()\n    \n#     plt.subplot(1, 2, 2)\n#     age_groups = pd.cut(train_df['PatientAge'], bins=5)\n#     pd.crosstab(age_groups, train_df['Aneurysm Present']).plot(kind='bar', figsize=(12, 6))\n#     plt.title('The relationship between age groups and aneurysm presence')#年龄组与动脉瘤存在的关系\n#     plt.xlabel('age group')#年龄组\n#     plt.ylabel('Counting')\n#     plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n#     plt.tight_layout()\n#     plt.savefig('age_vs_aneurysm.png', dpi=300)\n#     plt.show()\n\n# # 成像模式与动脉瘤存在的关系\n# if 'Modality' in train_df.columns and 'Aneurysm Present' in train_df.columns:\n#     modality_aneurysm = pd.crosstab(train_df['Modality'], train_df['Aneurysm Present'])\n#     print(\"\\n成像模式与动脉瘤存在的关系:\")\n#     print(modality_aneurysm)\n    \n#     modality_aneurysm.plot(kind='bar', figsize=(10, 6))\n#     plt.title('The relationship between imaging modes and aneurysm presence')#成像模式与动脉瘤存在的关系\n#     plt.xlabel('Imaging mode')#成像模式\n#     plt.ylabel('Counting')\n#     plt.legend(['Aneurysm-free', 'There is an aneurysm'])\n#     plt.savefig('modality_vs_aneurysm.png', dpi=300)\n#     plt.show()\n\n# # 各位置动脉瘤的相关性分析\n# print(\"\\n各位置动脉瘤的相关性矩阵:\")\n# location_corr = train_df[location_labels].corr()\n# plt.figure(figsize=(12, 10))\n# sns.heatmap(location_corr, annot=True, cmap='coolwarm', center=0)\n# plt.title('Aneurysm-related heat map at various positions')#各位置动脉瘤相关性热图\n# plt.tight_layout()\n# plt.savefig('location_correlation.png', dpi=300)\n# plt.show()\n\n# # ===================== 数据集划分 =====================\n# print(\"\\n===== 数据集划分 =====\")\n\n# # 使用分层抽样确保各类别比例一致\n# train_df, test_df = train_test_split(\n#     train_df, \n#     test_size=0.2, \n#     random_state=42, \n#     stratify=train_df['Aneurysm Present']\n# )\n# train_df, val_df = train_test_split(\n#     train_df, \n#     test_size=0.125,  # 0.8 * 0.125 = 0.1\n#     random_state=42, \n#     stratify=train_df['Aneurysm Present']\n# )\n\n# print(f\"训练集大小: {len(train_df)}\")\n# print(f\"验证集大小: {len(val_df)}\")\n# print(f\"测试集大小: {len(test_df)}\")\n\n# # 显示各数据集标签分布\n# print(\"\\n训练集动脉瘤存在分布:\")\n# print(train_df['Aneurysm Present'].value_counts(normalize=True))\n\n# print(\"\\n验证集动脉瘤存在分布:\")\n# print(val_df['Aneurysm Present'].value_counts(normalize=True))\n\n# print(\"\\n测试集动脉瘤存在分布:\")\n# print(test_df['Aneurysm Present'].value_counts(normalize=True))\n\n# # ===================== 数据集类定义 =====================\n# class AneurysmDataset(Dataset):\n#     def __init__(self, df, root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series', \n#                  transform=None, is_test=False, location_labels=None):\n#         self.df = df.copy()\n#         self.root_dir = root_dir\n#         self.transform = transform\n#         self.is_test = is_test\n#         self.location_labels = location_labels\n#         self.target_shape = (1, 256, 256)  # (通道数, 高度, 宽度)\n#         self.label_shape = (14,)\n        \n#         self._validate_and_clean_labels()\n        \n#         if not os.path.exists(self.root_dir):\n#             raise FileNotFoundError(f\"根目录不存在: {self.root_dir}\")\n            \n#         valid_indices = self._filter_valid_series()\n#         self.df = self.df.iloc[valid_indices].reset_index(drop=True)\n#         print(f\"过滤后保留的有效样本数: {len(self.df)}\")\n\n#     def _validate_and_clean_labels(self):\n#         if not self.location_labels:\n#             raise ValueError(\"location_labels 不能为空\")\n            \n#         all_label_cols = ['Aneurysm Present'] + self.location_labels\n#         missing_cols = [col for col in all_label_cols if col not in self.df.columns]\n#         if missing_cols:\n#             raise KeyError(f\"缺少标签列: {missing_cols}\")\n        \n#         for col in all_label_cols:\n#             self.df[col] = pd.to_numeric(self.df[col], errors='coerce')\n#             self.df[col].fillna(0.0, inplace=True)\n#             self.df[col] = self.df[col].astype(np.float32)\n\n#     def _filter_valid_series(self):\n#         valid_indices = []\n#         for idx in tqdm(range(len(self.df)), desc=\"过滤有效样本\"):\n#             series_id = self.df.iloc[idx]['SeriesInstanceUID']\n#             series_path = os.path.join(self.root_dir, series_id)\n#             if os.path.exists(series_path) and len(os.listdir(series_path)) > 0:\n#                 valid_indices.append(idx)\n#             else:\n#                 warn_msg = f\"路径不存在或为空 - {series_path}\"\n#                 logging.warning(warn_msg)\n#         return valid_indices\n\n#     def __len__(self):\n#         return len(self.df)\n\n#     def __getitem__(self, idx):\n#         series_id = self.df.iloc[idx]['SeriesInstanceUID']\n#         series_path = os.path.join(self.root_dir, series_id)\n        \n#         # 初始化返回值\n#         image = torch.zeros(self.target_shape, dtype=torch.float32)\n#         label = torch.zeros(self.label_shape, dtype=torch.float32)\n\n#         try:\n#             # 加载DICOM文件列表\n#             dicom_files = [f for f in os.listdir(series_path) if f.endswith('.dcm')]\n#             if not dicom_files:\n#                 warn_msg = f\"未找到DICOM文件 - {series_path}\"\n#                 logging.warning(warn_msg)\n#                 return image, label\n            \n#             # 按数字排序并选择中间切片\n#             dicom_files.sort(key=lambda x: int(x.split('.')[0]))\n#             mid_idx = len(dicom_files) // 2\n#             dicom_path = os.path.join(series_path, dicom_files[mid_idx])\n            \n#             # 读取DICOM文件\n#             dicom = pydicom.dcmread(dicom_path)\n            \n#             # 应用VOI LUT增强对比度\n#             pixel_array = apply_voi_lut(dicom.pixel_array, dicom)\n            \n#             # 处理3D数据\n#             if len(pixel_array.shape) == 3:\n#                 slice_idx = pixel_array.shape[0] // 2\n#                 pixel_array = pixel_array[slice_idx, :, :]\n#                 logging.warning(f\"{series_id} 是3D数据，已取中间切片\")\n#             elif len(pixel_array.shape) != 2:\n#                 error_msg = f\"DICOM图像维度异常: {pixel_array.shape}，series_id: {series_id}\"\n#                 logging.error(error_msg)\n#                 return image, label\n            \n#             # 转换为float32并应用 rescale 校正\n#             pixel_array = pixel_array.astype(np.float32)\n#             if hasattr(dicom, 'RescaleSlope') and hasattr(dicom, 'RescaleIntercept'):\n#                 pixel_array = pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n            \n#             # 标准化到0-1范围\n#             img_min, img_max = pixel_array.min(), pixel_array.max()\n#             if img_max != img_min:\n#                 pixel_array = (pixel_array - img_min) / (img_max - img_min)\n            \n#             # 转换为张量并添加通道维度\n#             img_tensor = torch.from_numpy(pixel_array).unsqueeze(0)\n            \n#             # 调整尺寸\n#             resize_transform = transforms.Resize((self.target_shape[1], self.target_shape[2]))\n#             img_tensor = resize_transform(img_tensor)\n            \n#             if img_tensor.shape != self.target_shape:\n#                 raise RuntimeError(f\"图像形状错误: {img_tensor.shape} 预期 {self.target_shape}\")\n            \n#             image = img_tensor\n            \n#         except Exception as e:\n#             error_msg = f\"处理DICOM时出错 {series_path}: {str(e)}\"\n#             logging.error(error_msg)\n        \n#         # 处理标签\n#         try:\n#             if not self.is_test:\n#                 present = float(self.df.iloc[idx]['Aneurysm Present'])\n#                 locations = self.df.iloc[idx][self.location_labels].values.astype(np.float32)\n#                 label_np = np.concatenate([[present], locations])\n#                 label = torch.FloatTensor(label_np)\n#         except Exception as e:\n#             error_msg = f\"处理标签时出错 {series_id}: {str(e)}\"\n#             logging.error(error_msg)\n        \n#         # 应用变换\n#         if self.transform:\n#             try:\n#                 image = self.transform(image)\n#             except Exception as e:\n#                 logging.error(f\"应用变换时出错: {str(e)}\")\n        \n#         return image, label\n\n# # ===================== 数据加载器准备 =====================\n# print(\"\\n===== 准备数据加载器 =====\")\n\n# # 定义图像变换 - 增加数据增强提高模型泛化能力\n# train_transform = transforms.Compose([\n#     transforms.Resize((256, 256)),\n#     transforms.RandomAffine(degrees=10, translate=(0.1, 0.1)),\n#     transforms.RandomHorizontalFlip(),\n#     transforms.Normalize(mean=[0.5], std=[0.5])\n# ])\n\n# val_transform = transforms.Compose([\n#     transforms.Resize((256, 256)),\n#     transforms.Normalize(mean=[0.5], std=[0.5])\n# ])\n\n# # 创建数据集实例\n# train_dataset = AneurysmDataset(\n#     train_df, \n#     transform=train_transform, \n#     location_labels=location_labels\n# )\n# val_dataset = AneurysmDataset(\n#     val_df, \n#     transform=val_transform, \n#     location_labels=location_labels\n# )\n# test_dataset = AneurysmDataset(\n#     test_df, \n#     transform=val_transform, \n#     is_test=False,  # 保持为False以便评估\n#     location_labels=location_labels\n# )\n\n# # 创建数据加载器\n# batch_size = 16\n# num_workers = 0 if device.type == 'cpu' else 2\n# train_loader = DataLoader(\n#     train_dataset, \n#     batch_size=batch_size, \n#     shuffle=True, \n#     num_workers=num_workers,\n#     pin_memory=True if device.type == 'cuda' else False\n# )\n# val_loader = DataLoader(\n#     val_dataset, \n#     batch_size=batch_size, \n#     shuffle=False, \n#     num_workers=num_workers,\n#     pin_memory=True if device.type == 'cuda' else False\n# )\n# test_loader = DataLoader(\n#     test_dataset, \n#     batch_size=batch_size, \n#     shuffle=False, \n#     num_workers=num_workers,\n#     pin_memory=True if device.type == 'cuda' else False\n# )\n\n# # ===================== 模型定义 =====================\n# class AneurysmViT(nn.Module):\n#     def __init__(self, num_labels=14):\n#         super(AneurysmViT, self).__init__()\n        \n#         self.vit_config = ViTConfig(\n#             image_size=256,\n#             patch_size=32,\n#             num_channels=1,\n#             num_labels=num_labels,\n#             hidden_size=768,\n#             num_hidden_layers=12,\n#             num_attention_heads=12,\n#             intermediate_size=3072,\n#             dropout=0.3,  # 增加dropout防止过拟合\n#             attention_probs_dropout_prob=0.3\n#         )\n        \n#         self.vit = ViTModel(self.vit_config)\n        \n#         # 改进分类头\n#         self.classifier = nn.Sequential(\n#             nn.Linear(768, 512),\n#             nn.BatchNorm1d(512),\n#             nn.ReLU(),\n#             nn.Dropout(0.3),\n#             nn.Linear(512, 256),\n#             nn.BatchNorm1d(256),\n#             nn.ReLU(),\n#             nn.Dropout(0.2),\n#             nn.Linear(256, num_labels)\n#         )\n        \n#     def forward(self, x):\n#         outputs = self.vit(pixel_values=x)\n#         cls_token = outputs.last_hidden_state[:, 0, :]\n#         logits = self.classifier(cls_token)\n#         return torch.sigmoid(logits)\n\n# # 初始化模型\n# model = AneurysmViT().to(device)\n# print(\"\\n模型结构:\")\n# print(model)\n\n# # ===================== 训练配置 =====================\n# # 计算类别权重 - 更精确地处理类别不平衡\n# present_count = train_df['Aneurysm Present'].sum()\n# total_count = len(train_df)\n# weight_positive = total_count / (2 * present_count) if present_count > 0 else 1.0\n# weight_negative = total_count / (2 * (total_count - present_count)) if (total_count - present_count) > 0 else 1.0\n\n# print(f\"\\n类别权重 - 阳性: {weight_positive:.2f}, 阴性: {weight_negative:.2f}\")\n\n# # 加权损失函数\n# class WeightedBCELoss(nn.Module):\n#     def __init__(self, pos_weight=1.0):\n#         super().__init__()\n#         self.pos_weight = pos_weight\n        \n#     def forward(self, output, target):\n#         # 对第一个标签（是否存在动脉瘤）应用特殊权重\n#         weights = torch.ones_like(target, device=device)\n#         weights[:, 0] = self.pos_weight\n        \n#         bce_loss = nn.BCELoss(reduction='none')(output, target)\n#         return (bce_loss * weights).mean()\n\n# # 优化器和调度器\n# criterion = WeightedBCELoss(pos_weight=weight_positive)\n# optimizer = torch.optim.AdamW(\n#     model.parameters(), \n#     lr=1e-4, \n#     weight_decay=1e-5\n# )\n# scheduler = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(\n#     optimizer, \n#     T_0=5, \n#     T_mult=2, \n#     eta_min=1e-6\n# )\n\n# # 训练参数\n# num_epochs = 15\n# best_val_score = 0.0\n# early_stop_patience = 5\n# early_stop_counter = 0\n\n# # 记录训练过程指标\n# train_losses = []\n# val_losses = []\n# train_aucs = []\n# val_aucs = []\n\n# # ===================== 训练循环 =====================\n# print(\"\\n===== 开始训练 =====\")\n# for epoch in range(num_epochs):\n#     model.train()\n#     train_loss = 0.0\n#     train_auc = 0.0\n#     train_auc_count = 0\n    \n#     # 训练循环带进度条\n#     train_pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 训练\")\n#     for images, labels in train_pbar:\n#         images = images.to(device)\n#         labels = labels.to(device)\n        \n#         optimizer.zero_grad()\n#         outputs = model(images)\n#         loss = criterion(outputs, labels)\n#         loss.backward()\n        \n#         # 梯度裁剪防止梯度爆炸\n#         torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n        \n#         optimizer.step()\n        \n#         train_loss += loss.item() * images.size(0)\n        \n#         # 计算AUC\n#         labels_np = labels.cpu().detach().numpy()\n#         outputs_np = outputs.cpu().detach().numpy()\n        \n#         try:\n#             if len(np.unique(labels_np[:, 0])) >= 2:\n#                 auc = roc_auc_score(labels_np[:, 0], outputs_np[:, 0])\n#                 train_auc += auc * images.size(0)\n#                 train_auc_count += images.size(0)\n#         except ValueError as e:\n#             logging.error(f\"训练AUC计算错误: {str(e)}\")\n        \n#         # 更新进度条\n#         train_pbar.set_postfix({\"batch_loss\": loss.item()})\n    \n#     # 计算平均训练损失和AUC\n#     train_loss /= len(train_loader.dataset)\n#     train_auc = train_auc / train_auc_count if train_auc_count > 0 else 0.5\n    \n#     # 验证阶段\n#     model.eval()\n#     val_loss = 0.0\n#     val_scores = []\n#     val_main_aucs = []\n    \n#     with torch.no_grad():\n#         val_pbar = tqdm(val_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 验证\")\n#         for images, labels in val_pbar:\n#             images = images.to(device)\n#             labels = labels.to(device)\n            \n#             outputs = model(images)\n#             loss = criterion(outputs, labels)\n#             val_loss += loss.item() * images.size(0)\n            \n#             # 计算AUC\n#             labels_np = labels.cpu().numpy()\n#             outputs_np = outputs.cpu().numpy()\n#             auc_scores = []\n            \n#             for i in range(labels_np.shape[1]):\n#                 try:\n#                     if len(np.unique(labels_np[:, i])) >= 2:\n#                         auc_scores.append(roc_auc_score(labels_np[:, i], outputs_np[:, i]))\n#                     else:\n#                         auc_scores.append(0.5)\n#                 except:\n#                     auc_scores.append(0.5)\n            \n#             main_auc = auc_scores[0]\n#             other_auc = np.mean(auc_scores[1:])\n#             val_scores.append(0.5 * (main_auc + other_auc))\n#             val_main_aucs.append(main_auc)\n            \n#             # 更新进度条\n#             val_pbar.set_postfix({\"batch_loss\": loss.item()})\n    \n#     # 计算平均验证损失和分数\n#     val_loss /= len(val_loader.dataset)\n#     val_score = np.mean(val_scores)\n#     val_main_auc = np.mean(val_main_aucs)\n    \n#     # 更新学习率调度器\n#     scheduler.step()\n    \n#     # 记录指标\n#     train_losses.append(train_loss)\n#     val_losses.append(val_loss)\n#     train_aucs.append(train_auc)\n#     val_aucs.append(val_main_auc)\n    \n#     # 保存最佳模型\n#     if val_score > best_val_score:\n#         best_val_score = val_score\n#         torch.save(model.state_dict(), 'best_model.pth')\n#         early_stop_counter = 0\n#         print(\"保存最佳模型!\")\n#     else:\n#         early_stop_counter += 1\n#         if early_stop_counter >= early_stop_patience:\n#             print(\"早停: 验证分数连续多轮无提升\")\n#             break\n    \n#     # 打印 epoch 结果\n#     print(f\"\\nEpoch {epoch+1}/{num_epochs}\")\n#     print(f\"训练损失: {train_loss:.4f} | 训练主标签AUC: {train_auc:.4f}\")\n#     print(f\"验证损失: {val_loss:.4f} | 验证主标签AUC: {val_main_auc:.4f} | 验证综合分数: {val_score:.4f}\")\n#     print(\"-\" * 80)\n\n# print(\"训练完成!\")\n\n# # ===================== 绘制训练历史 =====================\n# print(\"\\n绘制训练历史...\")\n# plt.figure(figsize=(14, 6))\n\n# # 损失曲线\n# plt.subplot(1, 2, 1)\n# plt.plot(train_losses, label='训练损失')\n# plt.plot(val_losses, label='验证损失')\n# plt.title('Training and Validation Loss Curve')#训练和验证损失曲线\n# plt.xlabel('Epoch')\n# plt.ylabel('Loss Value')\n# plt.legend()\n# plt.grid(True, linestyle='--', alpha=0.7)\n\n# # AUC曲线\n# plt.subplot(1, 2, 2)\n# plt.plot(train_aucs, label='训练AUC')\n# plt.plot(val_aucs, label='验证AUC')\n# plt.title('Training and Validation AUC Curve')#训练和验证AUC曲线\n# plt.xlabel('Epoch')\n# plt.ylabel('AUC value')\n# plt.legend()\n# plt.grid(True, linestyle='--', alpha=0.7)\n\n# plt.tight_layout()\n# plt.savefig('training_history.png', dpi=300)\n# plt.show()\n\n# # ===================== 测试集评估 =====================\n# print(\"\\n===== 测试集评估 =====\")\n# # 加载最佳模型\n# model.load_state_dict(torch.load('best_model.pth'))\n# model.eval()\n\n# test_scores = []\n# test_aucs = []\n# all_auc_scores = []\n\n# with torch.no_grad():\n#     test_pbar = tqdm(test_loader, desc=\"测试集评估\")\n#     for images, labels in test_pbar:\n#         images = images.to(device)\n#         labels = labels.to(device)\n        \n#         outputs = model(images)\n        \n#         # 计算AUC\n#         labels_np = labels.cpu().numpy()\n#         outputs_np = outputs.cpu().numpy()\n        \n#         # 计算每个标签的AUC\n#         auc_scores = []\n#         for i in range(labels_np.shape[1]):\n#             try:\n#                 auc = roc_auc_score(labels_np[:, i], outputs_np[:, i])\n#                 auc_scores.append(auc)\n#             except ValueError:\n#                 auc_scores.append(0.5)\n        \n#         all_auc_scores.append(auc_scores)\n        \n#         # 计算加权分数\n#         main_auc = auc_scores[0]\n#         other_auc = np.mean(auc_scores[1:])\n#         weighted_score = 0.5 * (main_auc + other_auc)\n        \n#         test_scores.append(weighted_score)\n#         test_aucs.append(main_auc)\n\n# # 计算平均分数\n# test_score = np.mean(test_scores)\n# test_auc = np.mean(test_aucs)\n\n# print(f\"\\n测试集主标签AUC: {test_auc:.4f}\")\n# print(f\"测试集综合分数: {test_score:.4f}\")\n\n# # 计算每个位置的平均AUC\n# mean_auc_per_label = np.mean(all_auc_scores, axis=0)\n# print(\"\\n每个标签的平均AUC:\")\n# for i, label in enumerate(['Aneurysm Present'] + location_labels):\n#     print(f\"{label}: {mean_auc_per_label[i]:.4f}\")\n\n# # ===================== 可视化预测结果 =====================\n# print(\"\\n可视化预测结果...\")\n# def visualize_predictions(model, dataset, num_samples=5):\n#     model.eval()\n#     indices = np.random.choice(len(dataset), num_samples, replace=False)\n    \n#     plt.figure(figsize=(15, 4*num_samples))\n#     for i, idx in enumerate(indices):\n#         image, label = dataset[idx]\n#         image_tensor = image.unsqueeze(0).to(device)\n        \n#         with torch.no_grad():\n#             pred = model(image_tensor).cpu().numpy()[0]\n        \n#         # 显示图像\n#         plt.subplot(num_samples, 2, 2*i+1)\n#         plt.imshow(image[0], cmap='gray')\n#         plt.title(f\"真实值: {label[0].item():.0f}, 预测值: {pred[0]:.2f}\")\n#         plt.axis('off')\n        \n#         # 显示位置预测\n#         plt.subplot(num_samples, 2, 2*i+2)\n#         plt.barh(location_labels, pred[1:])\n#         plt.xlim(0, 1)\n#         plt.title(\"Position Prediction Probability\")#位置预测概率\n    \n#     plt.tight_layout()\n#     plt.savefig('predictions_visualization.png', dpi=300)\n#     plt.show()\n\n# # 可视化测试集预测结果\n# visualize_predictions(model, test_dataset, num_samples=5)\n\n# # ===================== 生成提交文件 =====================\n# # print(\"\\n生成提交文件...\")\n# # def generate_submission(model, test_loader, dataset, device):\n# #     model.eval()\n# #     all_preds = []\n# #     all_ids = []\n    \n# #     with torch.no_grad():\n# #         pbar = tqdm(test_loader, desc=\"生成预测结果\")\n# #         for images, _ in pbar:\n# #             images = images.to(device)\n# #             outputs = model(images)\n# #             preds = outputs.cpu().numpy()\n# #             all_preds.append(preds)\n    \n# #     # 获取真实的SeriesInstanceUID\n# #     all_ids = dataset.df['SeriesInstanceUID'].values\n    \n# #     all_preds = np.concatenate(all_preds)\n    \n# #     # 创建提交DataFrame\n# #     submission_df = pd.DataFrame(all_preds, columns=['Aneurysm Present'] + location_labels)\n# #     submission_df.insert(0, 'SeriesInstanceUID', all_ids)\n    \n# #     return submission_df\n\n# # # 生成提交文件\n# # submission_df = generate_submission(model, test_loader, test_dataset, device)\n# # submission_df.to_csv('submission.csv', index=False)\n# # print(\"提交文件已保存为 submission.csv\")\n\n# # print(\"\\n===== 代码执行完成 =====\")\n# print(\"\\n生成提交文件...\")\n# def generate_submission(model, test_loader, dataset, device):\n#     model.eval()\n#     all_preds = []\n#     all_ids = []\n    \n#     with torch.no_grad():\n#         pbar = tqdm(test_loader, desc=\"生成预测结果\")\n#         for images, _ in pbar:\n#             images = images.to(device)\n#             outputs = model(images)\n#             preds = outputs.cpu().numpy()\n#             all_preds.append(preds)\n    \n#     # 获取真实的SeriesInstanceUID\n#     all_ids = dataset.df['SeriesInstanceUID'].values\n    \n#     all_preds = np.concatenate(all_preds)\n    \n#     # 创建提交DataFrame\n#     submission_df = pd.DataFrame(all_preds, columns=['Aneurysm Present'] + location_labels)\n#     submission_df.insert(0, 'SeriesInstanceUID', all_ids)\n    \n#     return submission_df\n\n# # 生成提交文件（保存为parquet格式，替换原csv逻辑）\n# submission_df = generate_submission(model, test_loader, test_dataset, device)\n# # 使用snappy压缩减少文件体积，兼容大多数比赛平台\n# submission_df.to_parquet('submission.parquet', index=False, compression='snappy')\n\n# # 验证文件是否生成成功\n# if os.path.exists('submission.parquet'):\n#     file_size = os.path.getsize('submission.parquet') / 1024 / 1024  # 转换为MB\n#     print(f\"提交文件已成功保存为 submission.parquet，文件大小: {file_size:.2f} MB\")\n# else:\n#     print(\"错误：提交文件未生成\")\n\n# print(\"\\n===== 代码执行完成 =====\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-31T00:06:04.3941Z","iopub.execute_input":"2025-08-31T00:06:04.394799Z","iopub.status.idle":"2025-08-31T00:56:16.449104Z","shell.execute_reply.started":"2025-08-31T00:06:04.394774Z","shell.execute_reply":"2025-08-31T00:56:16.448365Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---------","metadata":{}},{"cell_type":"code","source":"# import os\n# import numpy as np\n# import pandas as pd\n# import pydicom\n# import torch\n# import torch.nn as nn\n# from torch.utils.data import Dataset, DataLoader\n# from torchvision import transforms\n# from transformers import ViTModel, ViTConfig\n# from safetensors.torch import load_file\n# from sklearn.model_selection import train_test_split\n# from sklearn.metrics import roc_auc_score\n# import matplotlib.pyplot as plt\n# from tqdm import tqdm\n# import warnings\n# import logging\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\n\n# # 配置日志记录\n# logging.basicConfig(\n#     filename='dicom_processing.log',\n#     level=logging.WARNING,\n#     format='%(asctime)s - %(levelname)s - %(message)s',\n#     datefmt='%Y-%m-%d %H:%M:%S'\n# )\n# warnings.filterwarnings('ignore')\n\n# # 设置随机种子保证可重复性\n# torch.manual_seed(42)\n# np.random.seed(42)\n\n# # 检查GPU可用性\n# device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n# print(f\"使用设备: {device}\")\n\n# # 定义动脉瘤位置标签\n# location_labels = [\n#     'Left Infraclinoid Internal Carotid Artery',\n#     'Right Infraclinoid Internal Carotid Artery',\n#     'Left Supraclinoid Internal Carotid Artery',\n#     'Right Supraclinoid Internal Carotid Artery',\n#     'Left Middle Cerebral Artery',\n#     'Right Middle Cerebral Artery',\n#     'Anterior Communicating Artery',\n#     'Left Anterior Cerebral Artery',\n#     'Right Anterior Cerebral Artery',\n#     'Left Posterior Communicating Artery',\n#     'Right Posterior Communicating Artery',\n#     'Basilar Tip',\n#     'Other Posterior Circulation'\n# ]\n# all_label_cols = ['Aneurysm Present'] + location_labels\n\n# # ===================== 数据集类定义 =====================\n# class AneurysmDataset(Dataset):\n#     def __init__(self, df, root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series', \n#                  transform=None, is_test=False, location_labels=None):\n#         self.df = df.copy()\n#         self.root_dir = root_dir\n#         self.transform = transform\n#         self.is_test = is_test\n#         self.location_labels = location_labels\n#         self.target_shape = (1, 224, 224)  # 匹配模型输入尺寸224x224\n#         self.label_shape = (14,)\n        \n#         self._validate_and_clean_labels()\n        \n#         if not os.path.exists(self.root_dir):\n#             raise FileNotFoundError(f\"根目录不存在: {self.root_dir}\")\n            \n#         valid_indices = self._filter_valid_series()\n#         self.df = self.df.iloc[valid_indices].reset_index(drop=True)\n#         print(f\"过滤后保留的有效样本数: {len(self.df)}\")\n\n#     def _validate_and_clean_labels(self):\n#         if not self.location_labels:\n#             raise ValueError(\"location_labels 不能为空\")\n            \n#         all_label_cols = ['Aneurysm Present'] + self.location_labels\n#         missing_cols = [col for col in all_label_cols if col not in self.df.columns]\n#         if missing_cols:\n#             raise KeyError(f\"缺少标签列: {missing_cols}\")\n        \n#         for col in all_label_cols:\n#             self.df[col] = pd.to_numeric(self.df[col], errors='coerce')\n#             self.df[col].fillna(0.0, inplace=True)\n#             self.df[col] = self.df[col].astype(np.float32)\n\n#     def _filter_valid_series(self):\n#         valid_indices = []\n#         for idx in tqdm(range(len(self.df)), desc=\"过滤有效样本\"):\n#             series_id = self.df.iloc[idx]['SeriesInstanceUID']\n#             series_path = os.path.join(self.root_dir, series_id)\n#             if os.path.exists(series_path) and len(os.listdir(series_path)) > 0:\n#                 valid_indices.append(idx)\n#             else:\n#                 warn_msg = f\"路径不存在或为空 - {series_path}\"\n#                 logging.warning(warn_msg)\n#         return valid_indices\n\n#     def __len__(self):\n#         return len(self.df)\n\n#     def __getitem__(self, idx):\n#         series_id = self.df.iloc[idx]['SeriesInstanceUID']\n#         series_path = os.path.join(self.root_dir, series_id)\n        \n#         # 初始化返回值\n#         image = torch.zeros(self.target_shape, dtype=torch.float32)\n#         label = torch.zeros(self.label_shape, dtype=torch.float32)\n\n#         try:\n#             # 加载DICOM文件列表\n#             dicom_files = [f for f in os.listdir(series_path) if f.endswith('.dcm')]\n#             if not dicom_files:\n#                 warn_msg = f\"未找到DICOM文件 - {series_path}\"\n#                 logging.warning(warn_msg)\n#                 return image, label\n            \n#             # 按数字排序并选择中间切片\n#             dicom_files.sort(key=lambda x: int(x.split('.')[0]))\n#             mid_idx = len(dicom_files) // 2\n#             dicom_path = os.path.join(series_path, dicom_files[mid_idx])\n            \n#             # 读取DICOM文件\n#             dicom = pydicom.dcmread(dicom_path)\n            \n#             # 应用VOI LUT增强对比度\n#             pixel_array = apply_voi_lut(dicom.pixel_array, dicom)\n            \n#             # 处理3D数据\n#             if len(pixel_array.shape) == 3:\n#                 slice_idx = pixel_array.shape[0] // 2\n#                 pixel_array = pixel_array[slice_idx, :, :]\n#                 logging.warning(f\"{series_id} 是3D数据，已取中间切片\")\n#             elif len(pixel_array.shape) != 2:\n#                 error_msg = f\"DICOM图像维度异常: {pixel_array.shape}，series_id: {series_id}\"\n#                 logging.error(error_msg)\n#                 return image, label\n            \n#             # 转换为float32并应用 rescale 校正\n#             pixel_array = pixel_array.astype(np.float32)\n#             if hasattr(dicom, 'RescaleSlope') and hasattr(dicom, 'RescaleIntercept'):\n#                 pixel_array = pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n            \n#             # 标准化到0-1范围\n#             img_min, img_max = pixel_array.min(), pixel_array.max()\n#             if img_max != img_min:\n#                 pixel_array = (pixel_array - img_min) / (img_max - img_min)\n            \n#             # 转换为张量并添加通道维度\n#             img_tensor = torch.from_numpy(pixel_array).unsqueeze(0)\n            \n#             # 调整尺寸到224x224\n#             resize_transform = transforms.Resize((self.target_shape[1], self.target_shape[2]))\n#             img_tensor = resize_transform(img_tensor)\n            \n#             if img_tensor.shape != self.target_shape:\n#                 raise RuntimeError(f\"图像形状错误: {img_tensor.shape} 预期 {self.target_shape}\")\n            \n#             image = img_tensor\n            \n#         except Exception as e:\n#             error_msg = f\"处理DICOM时出错 {series_path}: {str(e)}\"\n#             logging.error(error_msg)\n        \n#         # 处理标签\n#         try:\n#             if not self.is_test:\n#                 present = float(self.df.iloc[idx]['Aneurysm Present'])\n#                 locations = self.df.iloc[idx][self.location_labels].values.astype(np.float32)\n#                 label_np = np.concatenate([[present], locations])\n#                 label = torch.FloatTensor(label_np)\n#         except Exception as e:\n#             error_msg = f\"处理标签时出错 {series_id}: {str(e)}\"\n#             logging.error(error_msg)\n        \n#         # 应用变换\n#         if self.transform:\n#             try:\n#                 image = self.transform(image)\n#             except Exception as e:\n#                 logging.error(f\"应用变换时出错: {str(e)}\")\n        \n#         return image, label\n\n# # ===================== 模型定义 =====================\n# class AneurysmViT(nn.Module):\n#     def __init__(self, num_labels=14):\n#         super(AneurysmViT, self).__init__()\n        \n#         # 配置与预训练权重完全匹配（224x224图像、16x16 patch、3通道）\n#         self.vit_config = ViTConfig(\n#             image_size=224,\n#             patch_size=16,\n#             num_channels=3,\n#             hidden_size=768,\n#             num_hidden_layers=12,\n#             num_attention_heads=12,\n#             intermediate_size=3072,\n#             dropout=0.3,\n#             attention_probs_dropout_prob=0.3\n#         )\n        \n#         # 直接用配置初始化ViT模型\n#         self.vit = ViTModel(self.vit_config)\n        \n#         # 加载外部预训练权重\n#         model_path = '/kaggle/input/vit/transformers/default/1/vit_model/model.safetensors'\n#         if os.path.exists(model_path):\n#             print(f\"正在加载预训练模型: {model_path}\")\n#             state_dict = load_file(model_path)\n            \n#             # 清理权重键名（移除可能的\"vit.\"前缀）\n#             cleaned_state_dict = {}\n#             for param_name, param_value in state_dict.items():\n#                 if param_name.startswith('vit.'):\n#                     cleaned_name = param_name[4:]\n#                 else:\n#                     cleaned_name = param_name\n#                 cleaned_state_dict[cleaned_name] = param_value\n            \n#             # 加载权重（strict=True，确保完全匹配）\n#             self.vit.load_state_dict(cleaned_state_dict, strict=True)\n#             print(\"预训练权重加载成功（结构完全匹配）\")\n#         else:\n#             raise FileNotFoundError(f\"预训练模型文件不存在: {model_path}\")\n        \n#         # 冻结部分预训练层\n#         freeze_layers = 8  # 冻结前8层\n#         for i, (name, param) in enumerate(self.vit.named_parameters()):\n#             if i < freeze_layers * 12:  # 每层约12个参数组\n#                 param.requires_grad = False\n#             else:\n#                 param.requires_grad = True\n        \n#         # 分类头\n#         self.classifier = nn.Sequential(\n#             nn.Linear(768, 512),\n#             nn.BatchNorm1d(512),\n#             nn.ReLU(),\n#             nn.Dropout(0.3),\n#             nn.Linear(512, 256),\n#             nn.BatchNorm1d(256),\n#             nn.ReLU(),\n#             nn.Dropout(0.2),\n#             nn.Linear(256, num_labels)\n#         )\n        \n#     def forward(self, x):\n#         # 1通道灰度图 → 3通道（复制3次匹配预训练输入）\n#         x = x.repeat(1, 3, 1, 1)  # 形状: [batch, 1, 224, 224] → [batch, 3, 224, 224]\n        \n#         # ViT前向传播\n#         vit_outputs = self.vit(pixel_values=x)\n#         cls_token = vit_outputs.last_hidden_state[:, 0, :]\n        \n#         # 分类头预测\n#         logits = self.classifier(cls_token)\n#         return torch.sigmoid(logits)\n\n# # ===================== 主函数 =====================\n# def main():\n#     # 1. 数据加载与清洗\n#     print(\"===== 1. 数据加载与清洗 =====\")\n#     train_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\n#     localizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n    \n#     print(f\"原始训练数据形状: {train_df.shape}\")\n#     print(f\"定位器数据形状: {localizers_df.shape}\")\n    \n#     # 处理缺失值\n#     print(f\"\\n预处理前缺失值总数: {train_df.isnull().sum().sum()}\")\n#     train_df.fillna(0, inplace=True)\n#     localizers_df.dropna(inplace=True)\n#     print(f\"预处理后缺失值总数: {train_df.isnull().sum().sum()}\")\n    \n#     # 统一标签类型\n#     for col in all_label_cols:\n#         if col in train_df.columns:\n#             if train_df[col].dtype == 'object':\n#                 train_df[col] = train_df[col].apply(lambda x: \n#                     1.0 if str(x).strip().lower() in ['true', 'yes', 'present', '1'] else\n#                     0.0 if str(x).strip().lower() in ['false', 'no', 'absent', '0'] else\n#                     float(x) if pd.notna(x) and str(x).strip() else 0.0\n#                 )\n#             train_df[col] = train_df[col].astype(np.float32)\n\n#     # 2. 数据集划分\n#     print(\"\\n===== 2. 数据集划分 =====\")\n#     train_df, test_df = train_test_split(\n#         train_df, \n#         test_size=0.2, \n#         random_state=42, \n#         stratify=train_df['Aneurysm Present']\n#     )\n#     train_df, val_df = train_test_split(\n#         train_df, \n#         test_size=0.125,  # 0.8×0.125=0.1\n#         random_state=42, \n#         stratify=train_df['Aneurysm Present']\n#     )\n    \n#     print(f\"训练集大小: {len(train_df)}\")\n#     print(f\"验证集大小: {len(val_df)}\")\n#     print(f\"测试集大小: {len(test_df)}\")\n    \n#     # 验证分层效果\n#     print(\"\\n各数据集动脉瘤存在比例:\")\n#     print(f\"训练集: {train_df['Aneurysm Present'].mean():.4f}\")\n#     print(f\"验证集: {val_df['Aneurysm Present'].mean():.4f}\")\n#     print(f\"测试集: {test_df['Aneurysm Present'].mean():.4f}\")\n\n#     # 3. 数据加载器准备\n#     print(\"\\n===== 3. 数据加载器准备 =====\")\n#     # 训练集数据增强\n#     train_transform = transforms.Compose([\n#         transforms.Resize((224, 224)),\n#         transforms.RandomAffine(degrees=10, translate=(0.1, 0.1)),\n#         transforms.RandomHorizontalFlip(p=0.5),\n#         transforms.Normalize(mean=[0.5], std=[0.5])\n#     ])\n    \n#     # 验证集/测试集无增强\n#     val_transform = transforms.Compose([\n#         transforms.Resize((224, 224)),\n#         transforms.Normalize(mean=[0.5], std=[0.5])\n#     ])\n    \n#     # 创建数据集实例\n#     train_dataset = AneurysmDataset(\n#         train_df, \n#         transform=train_transform, \n#         location_labels=location_labels\n#     )\n#     val_dataset = AneurysmDataset(\n#         val_df, \n#         transform=val_transform, \n#         location_labels=location_labels\n#     )\n#     test_dataset = AneurysmDataset(\n#         test_df, \n#         transform=val_transform, \n#         is_test=False,\n#         location_labels=location_labels\n#     )\n    \n#     # 创建数据加载器\n#     batch_size = 16\n#     num_workers = 2 if device.type == 'cuda' else 0\n#     train_loader = DataLoader(\n#         train_dataset, \n#         batch_size=batch_size, \n#         shuffle=True,\n#         num_workers=num_workers,\n#         pin_memory=True if device.type == 'cuda' else False\n#     )\n#     val_loader = DataLoader(\n#         val_dataset, \n#         batch_size=batch_size, \n#         shuffle=False,\n#         num_workers=num_workers,\n#         pin_memory=True if device.type == 'cuda' else False\n#     )\n#     test_loader = DataLoader(\n#         test_dataset, \n#         batch_size=batch_size, \n#         shuffle=False,\n#         num_workers=num_workers,\n#         pin_memory=True if device.type == 'cuda' else False\n#     )\n\n#     # 4. 模型初始化与配置\n#     print(\"\\n===== 4. 模型初始化 =====\")\n#     model = AneurysmViT().to(device)\n#     print(\"模型结构概览:\")\n#     print(model)\n    \n#     # 计算模型参数\n#     total_params = sum(p.numel() for p in model.parameters())\n#     trainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)\n#     print(f\"总参数: {total_params:,} | 可训练参数: {trainable_params:,}\")\n\n#     # 5. 训练配置\n#     print(\"\\n===== 5. 训练配置 =====\")\n#     # 类别权重计算\n#     present_count = train_df['Aneurysm Present'].sum()\n#     total_count = len(train_df)\n#     weight_positive = total_count / (2 * present_count) if present_count > 0 else 1.0\n#     print(f\"动脉瘤存在样本占比: {present_count/total_count:.4f}\")\n#     print(f\"类别权重（阳性样本）: {weight_positive:.2f}\")\n    \n#     # 自定义加权BCE损失\n#     class WeightedBCELoss(nn.Module):\n#         def __init__(self, pos_weight=1.0):\n#             super().__init__()\n#             self.pos_weight = pos_weight\n            \n#         def forward(self, output, target):\n#             weights = torch.ones_like(target, device=device)\n#             weights[:, 0] = self.pos_weight  # 对主标签加权\n#             bce_loss = nn.BCELoss(reduction='none')(output, target)\n#             return (bce_loss * weights).mean()\n    \n#     # 优化器和调度器\n#     optimizer = torch.optim.AdamW(\n#         model.parameters(), \n#         lr=1e-4, \n#         weight_decay=1e-5\n#     )\n#     scheduler = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(\n#         optimizer, \n#         T_0=5, \n#         T_mult=2, \n#         eta_min=1e-6\n#     )\n\n#     # 6. 模型训练与验证\n#     print(\"\\n===== 6. 开始训练 =====\")\n#     num_epochs = 15\n#     best_val_score = 0.0\n#     early_stop_patience = 5\n#     early_stop_counter = 0\n    \n#     # 记录训练指标\n#     train_losses = []\n#     val_losses = []\n#     train_aucs = []\n#     val_aucs = []\n    \n#     for epoch in range(num_epochs):\n#         # 训练阶段\n#         model.train()\n#         train_loss = 0.0\n#         train_auc = 0.0\n#         train_auc_count = 0\n        \n#         train_pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 训练\")\n#         for images, labels in train_pbar:\n#             images = images.to(device)\n#             labels = labels.to(device)\n            \n#             optimizer.zero_grad()\n#             outputs = model(images)\n#             loss = WeightedBCELoss(pos_weight=weight_positive)(outputs, labels)\n#             loss.backward()\n            \n#             # 梯度裁剪\n#             torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n            \n#             optimizer.step()\n            \n#             train_loss += loss.item() * images.size(0)\n            \n#             # 计算AUC\n#             labels_np = labels.cpu().detach().numpy()\n#             outputs_np = outputs.cpu().detach().numpy()\n            \n#             try:\n#                 if len(np.unique(labels_np[:, 0])) >= 2:\n#                     auc = roc_auc_score(labels_np[:, 0], outputs_np[:, 0])\n#                     train_auc += auc * images.size(0)\n#                     train_auc_count += images.size(0)\n#             except ValueError as e:\n#                 logging.error(f\"训练AUC计算错误: {str(e)}\")\n            \n#             train_pbar.set_postfix({\"batch_loss\": loss.item()})\n        \n#         # 计算训练集平均指标\n#         train_loss /= len(train_loader.dataset)\n#         train_auc = train_auc / train_auc_count if train_auc_count > 0 else 0.5\n        \n#         # 验证阶段\n#         model.eval()\n#         val_loss = 0.0\n#         val_scores = []\n#         val_main_aucs = []\n        \n#         with torch.no_grad():\n#             val_pbar = tqdm(val_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 验证\")\n#             for images, labels in val_pbar:\n#                 images = images.to(device)\n#                 labels = labels.to(device)\n                \n#                 outputs = model(images)\n#                 loss = WeightedBCELoss(pos_weight=weight_positive)(outputs, labels)\n#                 val_loss += loss.item() * images.size(0)\n                \n#                 # 计算AUC\n#                 labels_np = labels.cpu().numpy()\n#                 outputs_np = outputs.cpu().numpy()\n#                 auc_scores = []\n                \n#                 for i in range(labels_np.shape[1]):\n#                     try:\n#                         if len(np.unique(labels_np[:, i])) >= 2:\n#                             auc_scores.append(roc_auc_score(labels_np[:, i], outputs_np[:, i]))\n#                         else:\n#                             auc_scores.append(0.5)\n#                     except:\n#                         auc_scores.append(0.5)\n                \n#                 main_auc = auc_scores[0]\n#                 other_auc = np.mean(auc_scores[1:])\n#                 val_scores.append(0.5 * (main_auc + other_auc))\n#                 val_main_aucs.append(main_auc)\n                \n#                 val_pbar.set_postfix({\"batch_loss\": loss.item()})\n        \n#         # 计算验证集平均指标\n#         val_loss /= len(val_loader.dataset)\n#         val_score = np.mean(val_scores)\n#         val_main_auc = np.mean(val_main_aucs)\n        \n#         # 更新调度器\n#         scheduler.step()\n        \n#         # 记录指标\n#         train_losses.append(train_loss)\n#         val_losses.append(val_loss)\n#         train_aucs.append(train_auc)\n#         val_aucs.append(val_main_auc)\n        \n#         # 保存最佳模型\n#         if val_score > best_val_score:\n#             best_val_score = val_score\n#             torch.save(model.state_dict(), 'best_model.pth')\n#             early_stop_counter = 0\n#             print(f\"保存最佳模型（验证综合评分: {val_score:.4f}）\")\n#         else:\n#             early_stop_counter += 1\n#             if early_stop_counter >= early_stop_patience:\n#                 print(f\"早停触发：连续{early_stop_patience}轮验证分数无提升\")\n#                 break\n        \n#         # 打印本轮结果\n#         print(f\"\\nEpoch {epoch+1}/{num_epochs} 结果:\")\n#         print(f\"训练损失: {train_loss:.4f} | 训练主标签AUC: {train_auc:.4f}\")\n#         print(f\"验证损失: {val_loss:.4f} | 验证主标签AUC: {val_main_auc:.4f} | 验证综合评分: {val_score:.4f}\")\n#         print(\"-\" * 80)\n    \n#     print(\"训练完成!\")\n\n#     # 7. 训练历史可视化\n#     print(\"\\n===== 7. 训练历史可视化 =====\")\n#     plt.figure(figsize=(14, 6))\n    \n#     # 损失曲线\n#     plt.subplot(1, 2, 1)\n#     plt.plot(train_losses, label='训练损失')\n#     plt.plot(val_losses, label='验证损失')\n#     plt.title('训练和验证损失曲线')\n#     plt.xlabel('Epoch')\n#     plt.ylabel('损失值')\n#     plt.legend()\n#     plt.grid(True, linestyle='--', alpha=0.7)\n    \n#     # AUC曲线\n#     plt.subplot(1, 2, 2)\n#     plt.plot(train_aucs, label='训练AUC')\n#     plt.plot(val_aucs, label='验证AUC')\n#     plt.title('训练和验证AUC曲线')\n#     plt.xlabel('Epoch')\n#     plt.ylabel('AUC值')\n#     plt.legend()\n#     plt.grid(True, linestyle='--', alpha=0.7)\n    \n#     plt.tight_layout()\n#     plt.savefig('training_history.png', dpi=300)\n#     plt.show()\n\n#     # 8. 测试集评估\n#     print(\"\\n===== 8. 测试集评估 =====\")\n#     # 加载最佳模型\n#     model.load_state_dict(torch.load('best_model.pth'))\n#     model.eval()\n    \n#     test_scores = []\n#     test_aucs = []\n#     all_auc_scores = []\n    \n#     with torch.no_grad():\n#         test_pbar = tqdm(test_loader, desc=\"测试集评估\")\n#         for images, labels in test_pbar:\n#             images = images.to(device)\n#             labels = labels.to(device)\n            \n#             outputs = model(images)\n            \n#             # 计算AUC\n#             labels_np = labels.cpu().numpy()\n#             outputs_np = outputs.cpu().numpy()\n#             auc_scores = []\n            \n#             for i in range(labels_np.shape[1]):\n#                 try:\n#                     auc = roc_auc_score(labels_np[:, i], outputs_np[:, i])\n#                     auc_scores.append(auc)\n#                 except ValueError:\n#                     auc_scores.append(0.5)\n            \n#             all_auc_scores.append(auc_scores)\n            \n#             # 计算综合评分\n#             main_auc = auc_scores[0]\n#             other_auc = np.mean(auc_scores[1:])\n#             weighted_score = 0.5 * (main_auc + other_auc)\n            \n#             test_scores.append(weighted_score)\n#             test_aucs.append(main_auc)\n    \n#     # 计算测试集最终指标\n#     test_score = np.mean(test_scores)\n#     test_auc = np.mean(test_aucs)\n    \n#     print(f\"\\n测试集结果:\")\n#     print(f\"主标签（是否存在动脉瘤）AUC: {test_auc:.4f}\")\n#     print(f\"综合评分（主标签+位置标签）: {test_score:.4f}\")\n    \n#     # 打印各位置标签的AUC\n#     print(\"\\n各标签AUC值:\")\n#     mean_auc_per_label = np.mean(all_auc_scores, axis=0)\n#     for i, label in enumerate(['Aneurysm Present'] + location_labels):\n#         print(f\"{label}: {mean_auc_per_label[i]:.4f}\")\n\n#     # 9. 预测结果可视化\n#     print(\"\\n===== 9. 预测结果可视化 =====\")\n#     def visualize_predictions(model, dataset, num_samples=5):\n#         model.eval()\n#         indices = np.random.choice(len(dataset), num_samples, replace=False)\n        \n#         plt.figure(figsize=(15, 4*num_samples))\n#         for i, idx in enumerate(indices):\n#             image, label = dataset[idx]\n#             image_tensor = image.unsqueeze(0).to(device)\n            \n#             with torch.no_grad():\n#                 pred = model(image_tensor).cpu().numpy()[0]\n            \n#             # 显示医学图像\n#             plt.subplot(num_samples, 2, 2*i+1)\n#             plt.imshow(image[0], cmap='gray')\n#             plt.title(f\"真实值: {label[0].item():.0f}, 预测值: {pred[0]:.2f}\")\n#             plt.axis('off')\n            \n#             # 显示位置预测概率\n#             plt.subplot(num_samples, 2, 2*i+2)\n#             plt.barh(location_labels, pred[1:])\n#             plt.xlim(0, 1)\n#             plt.title(\"位置预测概率\")\n        \n#         plt.tight_layout()\n#         plt.savefig('predictions_visualization.png', dpi=300)\n#         plt.show()\n    \n#     # 可视化测试集预测结果\n#     visualize_predictions(model, test_dataset, num_samples=5)\n\n#     # 10. 生成提交文件\n#     print(\"\\n===== 10. 生成提交文件 =====\")\n#     def generate_submission(model, test_loader, dataset, device):\n#         model.eval()\n#         all_preds = []\n#         all_ids = dataset.df['SeriesInstanceUID'].values\n        \n#         with torch.no_grad():\n#             pbar = tqdm(test_loader, desc=\"生成提交结果\")\n#             for images, _ in pbar:\n#                 images = images.to(device)\n#                 outputs = model(images)\n#                 preds = outputs.cpu().numpy()\n#                 all_preds.append(preds)\n        \n#         all_preds = np.concatenate(all_preds)\n        \n#         # 创建提交DataFrame\n#         submission_df = pd.DataFrame(\n#             all_preds, \n#             columns=['Aneurysm Present'] + location_labels\n#         )\n#         submission_df.insert(0, 'SeriesInstanceUID', all_ids)\n        \n#         return submission_df\n    \n#     # 生成并保存提交文件\n#     submission_df = generate_submission(model, test_loader, test_dataset, device)\n#     submission_df.to_parquet('submission.parquet', index=False, compression='snappy')\n    \n#     # 验证提交文件\n#     if os.path.exists('submission.parquet'):\n#         file_size = os.path.getsize('submission.parquet') / 1024 / 1024\n#         print(f\"提交文件已保存: submission.parquet ({file_size:.2f} MB)\")\n#         print(\"前5行预览:\")\n#         print(submission_df.head())\n#     else:\n#         print(\"错误：提交文件未生成\")\n    \n#     print(\"\\n===== 所有流程执行完成 =====\")\n\n# if __name__ == \"__main__\":\n#     main()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-03T00:49:48.199097Z","iopub.execute_input":"2025-09-03T00:49:48.199721Z","iopub.status.idle":"2025-09-03T01:23:02.783337Z","shell.execute_reply.started":"2025-09-03T00:49:48.199688Z","shell.execute_reply":"2025-09-03T01:23:02.782346Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 动脉瘤检测完整代码（无函数嵌套版）\n# 所有模块合并为单一文件，保留独立性，可按顺序执行\n\n# ==================================================\n# 1. 导入依赖库与基础配置\n# ==================================================\nimport os\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom transformers import ViTModel, ViTConfig\nfrom safetensors.torch import load_file\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport warnings\nimport logging\nimport json\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n# 配置日志记录\nlogging.basicConfig(\n    filename='dicom_processing.log',\n    level=logging.WARNING,\n    format='%(asctime)s - %(levelname)s - %(message)s',\n    datefmt='%Y-%m-%d %H:%M:%S'\n)\nwarnings.filterwarnings('ignore')\n\n# 设置随机种子保证可重复性\ntorch.manual_seed(42)\nnp.random.seed(42)\n\n# 检查GPU可用性\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"当前使用设备: {device}\")\n\n# 定义动脉瘤位置标签\nlocation_labels = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation'\n]\nall_label_cols = ['Aneurysm Present'] + location_labels\nprint(f\"标签列总数: {len(all_label_cols)} (1个存在标签 + 13个位置标签)\")\n\n\n# ==================================================\n# 2. 数据加载与预处理\n# ==================================================\nprint(\"\\n===== 开始数据加载与清洗 =====\")\n# 加载原始数据（Kaggle数据集路径示例）\ntrain_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\nlocalizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\nprint(f\"原始训练数据形状: {train_df.shape}\")\nprint(f\"定位器数据形状: {localizers_df.shape}\")\n\n# 处理缺失值\nprint(f\"\\n预处理前缺失值总数: {train_df.isnull().sum().sum()}\")\ntrain_df.fillna(0, inplace=True)  # 标签缺失填充为0（无动脉瘤）\nlocalizers_df.dropna(inplace=True)  # 定位器数据直接删除缺失行\nprint(f\"预处理后缺失值总数: {train_df.isnull().sum().sum()}\")\n\n# 统一标签类型\nfor col in all_label_cols:\n    if col in train_df.columns:\n        if train_df[col].dtype == 'object':\n            train_df[col] = train_df[col].apply(lambda x: \n                1.0 if str(x).strip().lower() in ['true', 'yes', 'present', '1'] else\n                0.0 if str(x).strip().lower() in ['false', 'no', 'absent', '0'] else\n                float(x) if pd.notna(x) and str(x).strip() else 0.0\n            )\n        train_df[col] = train_df[col].astype(np.float32)\n\n# 数据集划分（7:1:2）\nprint(\"\\n===== 开始数据集划分 =====\")\ntrain_df_split, test_df = train_test_split(\n    train_df, \n    test_size=0.2, \n    random_state=42, \n    stratify=train_df['Aneurysm Present']\n)\ntrain_df_final, val_df = train_test_split(\n    train_df_split, \n    test_size=0.125, \n    random_state=42, \n    stratify=train_df_split['Aneurysm Present']\n)\n\n# 输出划分结果\nprint(f\"训练集大小: {len(train_df_final)}\")\nprint(f\"验证集大小: {len(val_df)}\")\nprint(f\"测试集大小: {len(test_df)}\")\nprint(\"\\n各数据集动脉瘤存在比例:\")\nprint(f\"训练集: {train_df_final['Aneurysm Present'].mean():.4f}\")\nprint(f\"验证集: {val_df['Aneurysm Present'].mean():.4f}\")\nprint(f\"测试集: {test_df['Aneurysm Present'].mean():.4f}\")\n\n# 保存清洗后的数据集\ntrain_df_final.to_csv('cleaned_train.csv', index=False)\nval_df.to_csv('cleaned_val.csv', index=False)\ntest_df.to_csv('cleaned_test.csv', index=False)\nprint(\"\\n清洗后的数据集已保存为: cleaned_train.csv / cleaned_val.csv / cleaned_test.csv\")\n\n\n# ==================================================\n# 3. 自定义数据集类实现\n# ==================================================\n# 重新加载清洗后的数据集\ntrain_df = pd.read_csv('cleaned_train.csv')\nval_df = pd.read_csv('cleaned_val.csv')\ntest_df = pd.read_csv('cleaned_test.csv')\n\n# 定义数据集类\nclass AneurysmDataset(Dataset):\n    def __init__(self, df, root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series', \n                 transform=None, is_test=False, location_labels=None):\n        self.df = df.copy()\n        self.root_dir = root_dir\n        self.transform = transform\n        self.is_test = is_test\n        self.location_labels = location_labels\n        self.target_shape = (1, 224, 224)\n        self.label_shape = (14,)\n        \n        self._validate_and_clean_labels()\n        \n        if not os.path.exists(self.root_dir):\n            raise FileNotFoundError(f\"DICOM根目录不存在: {self.root_dir}\")\n            \n        valid_indices = self._filter_valid_series()\n        self.df = self.df.iloc[valid_indices].reset_index(drop=True)\n        print(f\"有效样本数（过滤后）: {len(self.df)}\")\n\n    def _validate_and_clean_labels(self):\n        if not self.location_labels:\n            raise ValueError(\"location_labels 不能为空，请传入位置标签列表\")\n            \n        all_label_cols = ['Aneurysm Present'] + self.location_labels\n        missing_cols = [col for col in all_label_cols if col not in self.df.columns]\n        if missing_cols:\n            raise KeyError(f\"数据框缺少标签列: {missing_cols}\")\n        \n        for col in all_label_cols:\n            self.df[col] = pd.to_numeric(self.df[col], errors='coerce')\n            self.df[col].fillna(0.0, inplace=True)\n            self.df[col] = self.df[col].astype(np.float32)\n\n    def _filter_valid_series(self):\n        valid_indices = []\n        for idx in tqdm(range(len(self.df)), desc=\"过滤无效样本\"):\n            series_id = self.df.iloc[idx]['SeriesInstanceUID']\n            series_path = os.path.join(self.root_dir, series_id)\n            if os.path.exists(series_path) and len(os.listdir(series_path)) > 0:\n                valid_indices.append(idx)\n            else:\n                warn_msg = f\"样本{idx}路径不存在或为空: {series_path}\"\n                logging.warning(warn_msg)\n        return valid_indices\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        series_id = self.df.iloc[idx]['SeriesInstanceUID']\n        series_path = os.path.join(self.root_dir, series_id)\n        \n        image = torch.zeros(self.target_shape, dtype=torch.float32)\n        label = torch.zeros(self.label_shape, dtype=torch.float32)\n\n        try:\n            dicom_files = [f for f in os.listdir(series_path) if f.endswith('.dcm')]\n            if not dicom_files:\n                warn_msg = f\"样本{idx}未找到DICOM文件\"\n                logging.warning(warn_msg)\n                return image, label\n            \n            dicom_files.sort(key=lambda x: int(x.split('.')[0]))\n            mid_idx = len(dicom_files) // 2\n            dicom_path = os.path.join(series_path, dicom_files[mid_idx])\n            \n            dicom = pydicom.dcmread(dicom_path)\n            pixel_array = apply_voi_lut(dicom.pixel_array, dicom)\n            \n            if len(pixel_array.shape) == 3:\n                slice_idx = pixel_array.shape[0] // 2\n                pixel_array = pixel_array[slice_idx, :, :]\n                logging.warning(f\"样本{idx}是3D数据，已取中间切片\")\n            elif len(pixel_array.shape) != 2:\n                error_msg = f\"样本{idx}DICOM维度异常: {pixel_array.shape}\"\n                logging.error(error_msg)\n                return image, label\n            \n            pixel_array = pixel_array.astype(np.float32)\n            if hasattr(dicom, 'RescaleSlope') and hasattr(dicom, 'RescaleIntercept'):\n                pixel_array = pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n            \n            img_min, img_max = pixel_array.min(), pixel_array.max()\n            if img_max != img_min:\n                pixel_array = (pixel_array - img_min) / (img_max - img_min)\n            \n            img_tensor = torch.from_numpy(pixel_array).unsqueeze(0)\n            resize_transform = transforms.Resize((self.target_shape[1], self.target_shape[2]))\n            img_tensor = resize_transform(img_tensor)\n            \n            if img_tensor.shape != self.target_shape:\n                raise RuntimeError(f\"样本{idx}图像形状错误: {img_tensor.shape}\")\n            \n            image = img_tensor\n            \n        except Exception as e:\n            error_msg = f\"样本{idx}处理DICOM时出错: {str(e)}\"\n            logging.error(error_msg)\n        \n        try:\n            if not self.is_test:\n                present = float(self.df.iloc[idx]['Aneurysm Present'])\n                locations = self.df.iloc[idx][self.location_labels].values.astype(np.float32)\n                label_np = np.concatenate([[present], locations])\n                label = torch.FloatTensor(label_np)\n        except Exception as e:\n            error_msg = f\"样本{idx}处理标签时出错: {str(e)}\"\n            logging.error(error_msg)\n        \n        if self.transform:\n            try:\n                image = self.transform(image)\n            except Exception as e:\n                logging.error(f\"样本{idx}应用变换时出错: {str(e)}\")\n        \n        return image, label\n\n# 测试数据集加载\nprint(\"\\n===== 测试数据集加载 =====\")\ntrain_transform = transforms.Compose([\n    transforms.Resize((224, 224)),\n    transforms.RandomAffine(degrees=10, translate=(0.1, 0.1)),\n    transforms.RandomHorizontalFlip(p=0.5),\n    transforms.Normalize(mean=[0.5], std=[0.5])\n])\nval_test_transform = transforms.Compose([\n    transforms.Resize((224, 224)),\n    transforms.Normalize(mean=[0.5], std=[0.5])\n])\n\n# 创建数据集实例\ntrain_dataset = AneurysmDataset(\n    train_df, \n    transform=train_transform, \n    location_labels=location_labels\n)\nval_dataset = AneurysmDataset(\n    val_df, \n    transform=val_test_transform, \n    location_labels=location_labels\n)\ntest_dataset = AneurysmDataset(\n    test_df, \n    transform=val_test_transform, \n    is_test=False,\n    location_labels=location_labels\n)\n\n# 输出数据集信息\nprint(f\"\\n训练集样本数: {len(train_dataset)}\")\nprint(f\"验证集样本数: {len(val_dataset)}\")\nprint(f\"测试集样本数: {len(test_dataset)}\")\n\n# 测试单样本加载\nsample_img, sample_label = train_dataset[0]\nprint(f\"\\n单样本图像形状: {sample_img.shape}\")\nprint(f\"单样本标签形状: {sample_label.shape}\")\nprint(f\"单样本标签值: {sample_label.numpy()}\")\n\n\n# ==================================================\n# 4. 模型定义与初始化\n# ==================================================\nclass AneurysmViT(nn.Module):\n    def __init__(self, num_labels=14):\n        super(AneurysmViT, self).__init__()\n        \n        self.vit_config = ViTConfig(\n            image_size=224,\n            patch_size=16,\n            num_channels=3,\n            hidden_size=768,\n            num_hidden_layers=12,\n            num_attention_heads=12,\n            intermediate_size=3072,\n            dropout=0.3,\n            attention_probs_dropout_prob=0.3\n        )\n        \n        self.vit = ViTModel(self.vit_config)\n        \n        model_path = '/kaggle/input/vit/transformers/default/1/vit_model/model.safetensors'\n        if os.path.exists(model_path):\n            print(f\"正在加载预训练模型: {model_path}\")\n            state_dict = load_file(model_path)\n            \n            cleaned_state_dict = {}\n            for param_name, param_value in state_dict.items():\n                if param_name.startswith('vit.'):\n                    cleaned_name = param_name[4:]\n                else:\n                    cleaned_name = param_name\n                cleaned_state_dict[cleaned_name] = param_value\n            \n            self.vit.load_state_dict(cleaned_state_dict, strict=True)\n            print(\"预训练权重加载成功\")\n        else:\n            raise FileNotFoundError(f\"预训练模型文件不存在: {model_path}\")\n        \n        # 冻结部分预训练层\n        freeze_layers = 8\n        for i, (name, param) in enumerate(self.vit.named_parameters()):\n            if i < freeze_layers * 12:\n                param.requires_grad = False\n            else:\n                param.requires_grad = True\n        \n        # 分类头\n        self.classifier = nn.Sequential(\n            nn.Linear(768, 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(),\n            nn.Dropout(0.3),\n            nn.Linear(512, 256),\n            nn.BatchNorm1d(256),\n            nn.ReLU(),\n            nn.Dropout(0.2),\n            nn.Linear(256, num_labels)\n        )\n        \n    def forward(self, x):\n        x = x.repeat(1, 3, 1, 1)  # 1通道转3通道\n        vit_outputs = self.vit(pixel_values=x)\n        cls_token = vit_outputs.last_hidden_state[:, 0, :]\n        logits = self.classifier(cls_token)\n        return torch.sigmoid(logits)\n\n# 初始化模型并验证\nprint(\"\\n===== 模型初始化与结构验证 =====\")\nmodel = AneurysmViT(num_labels=14).to(device)\nprint(f\"模型已移动到 {device} 设备\")\n\n# 打印模型结构概览\nprint(\"\\n模型结构概览:\")\nprint(model)\n\n# 计算模型参数\ntotal_params = sum(p.numel() for p in model.parameters())\ntrainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)\nprint(f\"\\n模型参数统计:\")\nprint(f\"总参数数量: {total_params:,}\")\nprint(f\"可训练参数数量: {trainable_params:,}\")\nprint(f\"冻结参数数量: {total_params - trainable_params:,}\")\n\n# 测试模型前向传播\ntest_input = torch.randn(2, 1, 224, 224).to(device)\ntest_output = model(test_input)\nprint(f\"\\n前向传播测试:\")\nprint(f\"输入形状: {test_input.shape}\")\nprint(f\"输出形状: {test_output.shape}\")\nprint(f\"输出值范围: [{test_output.min():.4f}, {test_output.max():.4f}]\")\n\n\n# ==================================================\n# 5. 数据加载器创建\n# ==================================================\nprint(\"\\n===== 重新创建Dataset实例 =====\")\ntrain_dataset = AneurysmDataset(\n    train_df,\n    root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series',\n    transform=train_transform,\n    location_labels=location_labels\n)\nval_dataset = AneurysmDataset(\n    val_df,\n    root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series',\n    transform=val_test_transform,\n    location_labels=location_labels\n)\ntest_dataset = AneurysmDataset(\n    test_df,\n    root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series',\n    transform=val_test_transform,\n    is_test=False,\n    location_labels=location_labels\n)\n\n# 创建DataLoader\nprint(\"\\n===== 创建DataLoader =====\")\nbatch_size = 16\nnum_workers = 2 if device.type == 'cuda' else 0\npin_memory = True if device.type == 'cuda' else False\n\ntrain_loader = DataLoader(\n    train_dataset,\n    batch_size=batch_size,\n    shuffle=True,\n    num_workers=num_workers,\n    pin_memory=pin_memory,\n    drop_last=False\n)\n\nval_loader = DataLoader(\n    val_dataset,\n    batch_size=batch_size,\n    shuffle=False,\n    num_workers=num_workers,\n    pin_memory=pin_memory,\n    drop_last=False\n)\n\ntest_loader = DataLoader(\n    test_dataset,\n    batch_size=batch_size,\n    shuffle=False,\n    num_workers=num_workers,\n    pin_memory=pin_memory,\n    drop_last=False\n)\n\n# 验证Loader有效性\nprint(f\"\\nDataLoader验证:\")\nprint(f\"训练集Loader批次数: {len(train_loader)} (每批{batch_size}样本)\")\nprint(f\"验证集Loader批次数: {len(val_loader)}\")\nprint(f\"测试集Loader批次数: {len(test_loader)}\")\n\n# 测试批量加载数据\ntrain_batch = next(iter(train_loader))\nval_batch = next(iter(val_loader))\nprint(f\"\\n训练集批量数据形状:\")\nprint(f\"图像形状: {train_batch[0].shape}\")\nprint(f\"标签形状: {train_batch[1].shape}\")\nprint(f\"验证集批量数据形状:\")\nprint(f\"图像形状: {val_batch[0].shape}\")\nprint(f\"标签形状: {val_batch[1].shape}\")\n\n# 保存Loader配置\nloader_config = {\n    'batch_size': batch_size,\n    'num_workers': num_workers,\n    'pin_memory': pin_memory,\n    'train_loader_len': len(train_loader),\n    'val_loader_len': len(val_loader),\n    'test_loader_len': len(test_loader)\n}\nwith open('loader_config.json', 'w') as f:\n    json.dump(loader_config, f, indent=2)\nprint(\"\\nLoader配置已保存为: loader_config.json\")\n\n\n# ==================================================\n# 6. 训练配置与模型训练\n# ==================================================\n# 计算类别权重\nprint(\"\\n===== 计算类别权重 =====\")\npresent_count = train_df['Aneurysm Present'].sum()\ntotal_count = len(train_df)\nweight_positive = total_count / (2 * present_count) if present_count > 0 else 1.0\nprint(f\"训练集动脉瘤阳性样本占比: {present_count/total_count:.4f}\")\nprint(f\"阳性样本权重: {weight_positive:.2f}\")\n\n# 定义加权BCE损失\nclass WeightedBCELoss(nn.Module):\n    def __init__(self, pos_weight=1.0):\n        super().__init__()\n        self.pos_weight = pos_weight\n        \n    def forward(self, output, target):\n        weights = torch.ones_like(target, device=device)\n        weights[:, 0] = self.pos_weight\n        bce_loss = nn.BCELoss(reduction='none')(output, target)\n        return (bce_loss * weights).mean()\n\n# 配置优化器与调度器\nprint(\"\\n===== 配置优化器与调度器 =====\")\noptimizer = torch.optim.AdamW(\n    model.parameters(),\n    lr=1e-4,\n    weight_decay=1e-5\n)\nscheduler = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(\n    optimizer,\n    T_0=5,\n    T_mult=2,\n    eta_min=1e-6\n)\n\n# 训练参数初始化\nnum_epochs = 15\nbest_val_score = 0.0\nearly_stop_patience = 5\nearly_stop_counter = 0\n\n# 记录训练指标\ntrain_losses = []\nval_losses = []\ntrain_aucs = []\nval_aucs = []\n\n# 开始训练\nprint(\"\\n===== 开始模型训练 =====\")\nfor epoch in range(num_epochs):\n    # 训练阶段\n    model.train()\n    train_loss = 0.0\n    train_auc = 0.0\n    train_auc_count = 0\n    \n    train_pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 训练\")\n    for images, labels in train_pbar:\n        images = images.to(device)\n        labels = labels.to(device)\n        \n        optimizer.zero_grad()\n        outputs = model(images)\n        loss = WeightedBCELoss(pos_weight=weight_positive)(outputs, labels)\n        loss.backward()\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n        optimizer.step()\n        \n        train_loss += loss.item() * images.size(0)\n        \n        # 计算AUC\n        labels_np = labels.cpu().detach().numpy()\n        outputs_np = outputs.cpu().detach().numpy()\n        if len(np.unique(labels_np[:, 0])) >= 2:\n            auc = roc_auc_score(labels_np[:, 0], outputs_np[:, 0])\n            train_auc += auc * images.size(0)\n            train_auc_count += images.size(0)\n        \n        train_pbar.set_postfix({\"batch_loss\": f\"{loss.item():.4f}\"})\n    \n    # 计算训练集平均指标\n    train_loss_avg = train_loss / len(train_loader.dataset)\n    train_auc_avg = train_auc / train_auc_count if train_auc_count > 0 else 0.5\n    train_losses.append(train_loss_avg)\n    train_aucs.append(train_auc_avg)\n    \n    # 验证阶段\n    model.eval()\n    val_loss = 0.0\n    val_scores = []\n    val_main_aucs = []\n    \n    with torch.no_grad():\n        val_pbar = tqdm(val_loader, desc=f\"Epoch {epoch+1}/{num_epochs} - 验证\")\n        for images, labels in val_pbar:\n            images = images.to(device)\n            labels = labels.to(device)\n            \n            outputs = model(images)\n            loss = WeightedBCELoss(pos_weight=weight_positive)(outputs, labels)\n            val_loss += loss.item() * images.size(0)\n            \n            # 计算各标签AUC\n            labels_np = labels.cpu().numpy()\n            outputs_np = outputs.cpu().numpy()\n            auc_scores = []\n            \n            for i in range(labels_np.shape[1]):\n                try:\n                    if len(np.unique(labels_np[:, i])) >= 2:\n                        auc_scores.append(roc_auc_score(labels_np[:, i], outputs_np[:, i]))\n                    else:\n                        auc_scores.append(0.5)\n                except:\n                    auc_scores.append(0.5)\n            \n            main_auc = auc_scores[0]\n            other_auc = np.mean(auc_scores[1:])\n            val_scores.append(0.5 * (main_auc + other_auc))\n            val_main_aucs.append(main_auc)\n            \n            val_pbar.set_postfix({\"batch_loss\": f\"{loss.item():.4f}\"})\n    \n    # 计算验证集平均指标\n    val_loss_avg = val_loss / len(val_loader.dataset)\n    val_score_avg = np.mean(val_scores)\n    val_main_auc_avg = np.mean(val_main_aucs)\n    val_losses.append(val_loss_avg)\n    val_aucs.append(val_main_auc_avg)\n    \n    # 更新调度器\n    scheduler.step()\n    \n    # 保存最佳模型\n    if val_score_avg > best_val_score:\n        best_val_score = val_score_avg\n        torch.save(model.state_dict(), 'best_model.pth')\n        early_stop_counter = 0\n        print(f\"✅ 保存最佳模型（验证综合评分: {val_score_avg:.4f}）\")\n    else:\n        early_stop_counter += 1\n        if early_stop_counter >= early_stop_patience:\n            print(f\"❌ 早停触发：连续{early_stop_patience}轮验证分数无提升\")\n            break\n    \n    # 打印本轮结果\n    print(f\"\\nEpoch {epoch+1}/{num_epochs} 结果:\")\n    print(f\"训练损失: {train_loss_avg:.4f} | 训练主标签AUC: {train_auc_avg:.4f}\")\n    print(f\"验证损失: {val_loss_avg:.4f} | 验证主标签AUC: {val_main_auc_avg:.4f} | 验证综合评分: {val_score_avg:.4f}\")\n    print(\"-\" * 80)\n\nprint(\"训练完成!\")\n\n# 保存训练指标\nnp.savez('training_metrics.npz', \n         train_losses=train_losses, \n         val_losses=val_losses,\n         train_aucs=train_aucs,\n         val_aucs=val_aucs)\nprint(\"训练指标已保存为: training_metrics.npz\")\n\n\n# ==================================================\n# 7. 训练历史可视化\n# ==================================================\nprint(\"\\n===== 生成训练历史可视化 =====\")\ntry:\n    metrics = np.load('training_metrics.npz')\n    train_losses = metrics['train_losses']\n    val_losses = metrics['val_losses']\n    train_aucs = metrics['train_aucs']\n    val_aucs = metrics['val_aucs']\n    num_epochs = len(train_losses)\n    print(f\"成功加载训练指标，共{num_epochs}个训练轮次\")\nexcept FileNotFoundError:\n    print(\"错误：未找到training_metrics.npz，请先运行训练模块\")\n\n# 创建可视化图表\nplt.figure(figsize=(14, 6))\n\n# 损失曲线\nplt.subplot(1, 2, 1)\nplt.plot(range(1, num_epochs+1), train_losses, label='Training Loss', marker='o', markersize=4)\nplt.plot(range(1, num_epochs+1), val_losses, label='Validation loss', marker='s', markersize=4)\nplt.title('Training and Validation Loss Curve', fontsize=12)\nplt.xlabel('Epoch', fontsize=10)\nplt.ylabel('Loss Value', fontsize=10)\nplt.legend(fontsize=10)\nplt.grid(True, linestyle='--', alpha=0.7)\n\n# AUC曲线\nplt.subplot(1, 2, 2)\nplt.plot(range(1, num_epochs+1), train_aucs, label='TrainingAUC', marker='o', markersize=4)\nplt.plot(range(1, num_epochs+1), val_aucs, label='VerificationAUC', marker='s', markersize=4)\nplt.title('Training and Validation AUC Curve (Primary Label)', fontsize=12)\nplt.xlabel('Epoch', fontsize=10)\nplt.ylabel('AUC value', fontsize=10)\nplt.ylim(0.5, 1.0)\nplt.legend(fontsize=10)\nplt.grid(True, linestyle='--', alpha=0.7)\n\n# 保存与显示\nplt.tight_layout()\nplt.savefig('training_history.png', dpi=300, bbox_inches='tight')\nprint(\"训练历史图表已保存为: training_history.png\")\nplt.show()\n\n\n# ==================================================\n# 8. 测试集评估\n# ==================================================\nprint(\"\\n===== 开始测试集评估 =====\")\n# 重新创建测试集Loader\ntest_dataset = AneurysmDataset(\n    test_df,\n    root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series',\n    transform=val_test_transform,\n    is_test=False,\n    location_labels=location_labels\n)\ntest_loader = DataLoader(\n    test_dataset,\n    batch_size=16,\n    shuffle=False,\n    num_workers=2 if device.type == 'cuda' else 0,\n    pin_memory=True if device.type == 'cuda' else False\n)\n\n# 加载最佳模型\nmodel = AneurysmViT(num_labels=14).to(device)\ntry:\n    model.load_state_dict(torch.load('best_model.pth', map_location=device))\n    print(\"成功加载最佳模型: best_model.pth\")\nexcept FileNotFoundError:\n    print(\"错误：未找到best_model.pth，请先运行训练模块\")\n\nmodel.eval()\n\n# 测试集评估\ntest_scores = []\ntest_aucs = []\nall_auc_scores = []\n\nwith torch.no_grad():\n    test_pbar = tqdm(test_loader, desc=\"测试集评估\")\n    for images, labels in test_pbar:\n        images = images.to(device)\n        labels = labels.to(device)\n        \n        outputs = model(images)\n        \n        # 计算各标签AUC\n        labels_np = labels.cpu().numpy()\n        outputs_np = outputs.cpu().numpy()\n        auc_scores = []\n        \n        for i in range(labels_np.shape[1]):\n            try:\n                auc = roc_auc_score(labels_np[:, i], outputs_np[:, i])\n                auc_scores.append(auc)\n            except ValueError:\n                auc_scores.append(0.5)\n        \n        all_auc_scores.append(auc_scores)\n        \n        # 计算综合评分\n        main_auc = auc_scores[0]\n        other_auc = np.mean(auc_scores[1:])\n        weighted_score = 0.5 * (main_auc + other_auc)\n        \n        test_scores.append(weighted_score)\n        test_aucs.append(main_auc)\n\n# 计算最终指标\ntest_score = np.mean(test_scores)\ntest_auc = np.mean(test_aucs)\nmean_auc_per_label = np.mean(all_auc_scores, axis=0)\n\n# 输出评估结果\nprint(f\"\\n测试集最终结果:\")\nprint(f\"主标签（是否存在动脉瘤）AUC: {test_auc:.4f}\")\nprint(f\"综合评分（主标签+位置标签）: {test_score:.4f}\")\n\nprint(\"\\n各标签AUC值:\")\nfor i, label in enumerate(['Aneurysm Present'] + location_labels):\n    print(f\"{label}: {mean_auc_per_label[i]:.4f}\")\n\n# 保存评估结果\neval_results = {\n    'main_auc': test_auc,\n    'overall_score': test_score,\n    'per_label_auc': dict(zip(['Aneurysm Present'] + location_labels, mean_auc_per_label))\n}\nwith open('evaluation_results.json', 'w') as f:\n    json.dump(eval_results, f, indent=2)\nprint(\"\\n评估结果已保存为: evaluation_results.json\")\n\n\n# ==================================================\n# 9. 预测结果可视化\n# ==================================================\nprint(\"\\n===== 生成预测结果可视化 =====\")\nnum_samples = 5\nindices = np.random.choice(len(test_dataset), num_samples, replace=False)\n\nplt.figure(figsize=(15, 4*num_samples))\nfor i, idx in enumerate(indices):\n    image, label = test_dataset[idx]\n    image_tensor = image.unsqueeze(0).to(device)\n    \n    with torch.no_grad():\n        pred = model(image_tensor).cpu().numpy()[0]\n    \n    # 显示医学图像\n    plt.subplot(num_samples, 2, 2*i+1)\n    plt.imshow(image[0], cmap='gray')\n    plt.title(f\"样本 {idx}\\n true value: {label[0].item():.0f} | Predicted Value: {pred[0]:.2f}\", fontsize=10)\n    plt.axis('off')\n    \n    # 显示位置预测概率\n    plt.subplot(num_samples, 2, 2*i+2)\n    plt.barh(location_labels, pred[1:])\n    plt.xlim(0, 1)\n    plt.title(\"Aneurysm location prediction probability\", fontsize=10)\n    plt.yticks(fontsize=8)\n\n# 保存与显示\nplt.tight_layout()\nplt.savefig('predictions_visualization.png', dpi=300, bbox_inches='tight')\nprint(\"预测结果可视化已保存为: predictions_visualization.png\")\nplt.show()\n\n\n# ==================================================\n# 10. 生成提交文件\n# ==================================================\nprint(\"\\n===== 生成提交文件 =====\")\n# 重新创建测试集Loader（测试模式）\ntest_dataset = AneurysmDataset(\n    test_df,\n    root_dir='/kaggle/input/rsna-intracranial-aneurysm-detection/series',\n    transform=val_test_transform,\n    is_test=True,\n    location_labels=location_labels\n)\ntest_loader = DataLoader(\n    test_dataset,\n    batch_size=16,\n    shuffle=False,\n    num_workers=2 if device.type == 'cuda' else 0,\n    pin_memory=True if device.type == 'cuda' else False\n)\n\n# 加载最佳模型\nmodel = AneurysmViT(num_labels=14).to(device)\nmodel.load_state_dict(torch.load('best_model.pth', map_location=device))\nmodel.eval()\n\n# 生成预测结果\nall_preds = []\nall_ids = test_dataset.df['SeriesInstanceUID'].values\n\nwith torch.no_grad():\n    pbar = tqdm(test_loader, desc=\"生成预测结果\")\n    for images, _ in pbar:\n        images = images.to(device)\n        outputs = model(images)\n        preds = outputs.cpu().numpy()\n        all_preds.append(preds)\n\n# 合并所有预测结果\nall_preds = np.concatenate(all_preds)\n\n# 创建提交DataFrame\nsubmission_df = pd.DataFrame(\n    all_preds,\n    columns=['Aneurysm Present'] + location_labels\n)\nsubmission_df.insert(0, 'SeriesInstanceUID', all_ids)\n\n# 保存提交文件\nsubmission_df.to_parquet('submission.parquet', index=False, compression='snappy')\n\n# 验证提交文件\nif os.path.exists('submission.parquet'):\n    file_size = os.path.getsize('submission.parquet') / 1024 / 1024\n    print(f\"提交文件已保存: submission.parquet ({file_size:.2f} MB)\")\n    print(\"前5行预览:\")\n    print(submission_df.head())\nelse:\n    print(\"错误：提交文件未生成\")\n\nprint(\"\\n===== 所有流程执行完成 =====\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-03T01:41:01.869105Z","iopub.execute_input":"2025-09-03T01:41:01.869778Z","iopub.status.idle":"2025-09-03T02:23:36.987247Z","shell.execute_reply.started":"2025-09-03T01:41:01.869743Z","shell.execute_reply":"2025-09-03T02:23:36.986223Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\n\n# 创建一个简单的DataFrame作为示例数据\ndata = {\n    'id': [1, 2, 3, 4, 5],\n    'name': ['Alice', 'Bob', 'Charlie', 'David', 'Eve'],\n    'value': [10.5, 20.3, 15.7, 25.1, 30.0]\n}\ndf = pd.DataFrame(data)\n\n# 将DataFrame保存为parquet文件\ndf.to_parquet('submission.parquet', index=False)\n\nprint(\"文件已成功保存为submission.parquet\")\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-04T06:28:20.28271Z","iopub.execute_input":"2025-09-04T06:28:20.283133Z","iopub.status.idle":"2025-09-04T06:28:20.294363Z","shell.execute_reply.started":"2025-09-04T06:28:20.283096Z","shell.execute_reply":"2025-09-04T06:28:20.293349Z"}},"outputs":[],"execution_count":null}]}