{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"A simple first look at the [**RSNA Knee Abnormality Detection** dataset](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/data).\n\nThe goal of this notebook is only to understand the data before touching the MRI pixels or building a model.\n\nWe will look at:\n\n- the available tables\n- studies, series, and DICOM organization\n- MRI series characteristics\n- the 12 target abnormalities\n- label availability\n- target prevalence and co-occurrence\n- radiology reports\n- the sample submission\n\n> **Important:** missing target values are not treated as negative labels in this notebook.","metadata":{}},{"cell_type":"markdown","source":"## 1. Imports","metadata":{}},{"cell_type":"code","source":"import kagglehub\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom IPython.display import display","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:45:46.812317Z","iopub.execute_input":"2026-08-25T20:45:46.812717Z","iopub.status.idle":"2026-08-25T20:45:47.632916Z","shell.execute_reply.started":"2026-08-25T20:45:46.812683Z","shell.execute_reply":"2026-08-25T20:45:47.632048Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Locate the competition data","metadata":{}},{"cell_type":"code","source":"DATA_DIR = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\n\nif DATA_DIR.exists():\n    print(\"Dataset found!\")\n\nelse:\n    print(\"Dataset not found. Downloading...\")\n\n    DATA_DIR = Path(\n        kagglehub.competition_download(\n            \"rsna-knee-abnormality-detection\"\n        )\n    )\n    print(\"Dataset downloaded!\")\nprint(\"\\nAvailable files/folders:\\n\")\n\nfor path in sorted(DATA_DIR.iterdir()):\n    print(\"-\", path.name)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:46:35.795584Z","iopub.execute_input":"2026-08-25T20:46:35.796167Z","iopub.status.idle":"2026-08-25T20:46:35.804354Z","shell.execute_reply.started":"2026-08-25T20:46:35.796131Z","shell.execute_reply":"2026-08-25T20:46:35.803337Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The competition provides tabular information together with MRI data stored as DICOM files.","metadata":{}},{"cell_type":"markdown","source":"## 3. Load the tables","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(DATA_DIR / \"train.csv\")\ntest = pd.read_csv(DATA_DIR / \"test.csv\")\n\ntrain_series = pd.read_csv(DATA_DIR / \"train_series.csv\")\ntest_series = pd.read_csv(DATA_DIR / \"test_series.csv\")\n\nsample_submission = pd.read_csv(DATA_DIR / \"sample_submission.csv\")\n\nprint(f\"train.csv:        {train.shape}\")\nprint(f\"test.csv:         {test.shape}\")\nprint(f\"train_series.csv: {train_series.shape}\")\nprint(f\"test_series.csv:  {test_series.shape}\")\nprint(f\"sample_submission:{sample_submission.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:55.600022Z","iopub.execute_input":"2026-08-25T20:44:55.600395Z","iopub.status.idle":"2026-08-25T20:44:55.932264Z","shell.execute_reply.started":"2026-08-25T20:44:55.600342Z","shell.execute_reply":"2026-08-25T20:44:55.931136Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Study-level data","metadata":{}},{"cell_type":"code","source":"display(train.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:55.934638Z","iopub.execute_input":"2026-08-25T20:44:55.935089Z","iopub.status.idle":"2026-08-25T20:44:55.97688Z","shell.execute_reply.started":"2026-08-25T20:44:55.93503Z","shell.execute_reply":"2026-08-25T20:44:55.976029Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`train.csv` contains one row per MRI study.\n\nEach study is identified by a `StudyInstanceUID`.\n\nThe training table also contains:\n\n- a radiology `Report`\n- 12 structured target columns","metadata":{}},{"cell_type":"markdown","source":"## 5. Series-level data","metadata":{}},{"cell_type":"code","source":"display(train_series.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:55.978142Z","iopub.execute_input":"2026-08-25T20:44:55.978519Z","iopub.status.idle":"2026-08-25T20:44:55.989062Z","shell.execute_reply.started":"2026-08-25T20:44:55.97848Z","shell.execute_reply":"2026-08-25T20:44:55.987748Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`train_series.csv` contains one row per MRI series.\n\nA single study can contain several series, each identified by a `SeriesInstanceUID`.","metadata":{}},{"cell_type":"markdown","source":"## 6. How are the MRI files organized?\n\nThe basic hierarchy is:\n\n**Study -> Series -> DICOM slices**\n\n- **Study:** one knee MRI examination\n- **Series:** one MRI acquisition within that examination\n- **Slice:** one 2D DICOM image belonging to a series\n\nSo one study can contain several MRI series, and one series contains multiple DICOM slices.","metadata":{}},{"cell_type":"markdown","source":"## 7. Quick DataFrame overview","metadata":{}},{"cell_type":"code","source":"def summarize_dataframe(df):\n    return pd.DataFrame({\n        \"dtype\": df.dtypes,\n        \"missing\": df.isna().sum(),\n        \"missing_%\": (df.isna().mean() * 100).round(2),\n        \"unique\": df.nunique(dropna=True)\n    })","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:55.990326Z","iopub.execute_input":"2026-08-25T20:44:55.990674Z","iopub.status.idle":"2026-08-25T20:44:56.005393Z","shell.execute_reply.started":"2026-08-25T20:44:55.990637Z","shell.execute_reply":"2026-08-25T20:44:56.004252Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"train.csv\")\ndisplay(summarize_dataframe(train))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.006955Z","iopub.execute_input":"2026-08-25T20:44:56.007666Z","iopub.status.idle":"2026-08-25T20:44:56.070704Z","shell.execute_reply.started":"2026-08-25T20:44:56.007611Z","shell.execute_reply":"2026-08-25T20:44:56.06986Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"train_series.csv\")\ndisplay(summarize_dataframe(train_series))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.072101Z","iopub.execute_input":"2026-08-25T20:44:56.072377Z","iopub.status.idle":"2026-08-25T20:44:56.112593Z","shell.execute_reply.started":"2026-08-25T20:44:56.072352Z","shell.execute_reply":"2026-08-25T20:44:56.111669Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Train vs visible test tables","metadata":{}},{"cell_type":"code","source":"train_only = sorted(set(train.columns) - set(test.columns))\ntest_only = sorted(set(test.columns) - set(train.columns))\nshared = sorted(set(train.columns) & set(test.columns))\n\nprint(\"Train-only columns:\")\nprint(train_only)\n\nprint(\"\\nTest-only columns:\")\nprint(test_only)\n\nprint(\"\\nShared columns:\")\nprint(shared)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.113668Z","iopub.execute_input":"2026-08-25T20:44:56.114005Z","iopub.status.idle":"2026-08-25T20:44:56.120688Z","shell.execute_reply.started":"2026-08-25T20:44:56.113968Z","shell.execute_reply":"2026-08-25T20:44:56.119429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Training studies             :\", train[\"StudyInstanceUID\"].nunique())\nprint(\"Visible test studies         :\", test[\"StudyInstanceUID\"].nunique())\n\nprint()\nprint(\"Training series              :\", train_series[\"SeriesInstanceUID\"].nunique())\nprint(\"Visible test series          :\", test_series[\"SeriesInstanceUID\"].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.124132Z","iopub.execute_input":"2026-08-25T20:44:56.124541Z","iopub.status.idle":"2026-08-25T20:44:56.152924Z","shell.execute_reply.started":"2026-08-25T20:44:56.124508Z","shell.execute_reply":"2026-08-25T20:44:56.151864Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. How many MRI series are in one study?","metadata":{}},{"cell_type":"code","source":"series_per_study = (\n    train_series\n    .groupby(\"StudyInstanceUID\")\n    .size()\n)\n\ndisplay(series_per_study.describe().to_frame(\"series_per_study\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.15399Z","iopub.execute_input":"2026-08-25T20:44:56.154261Z","iopub.status.idle":"2026-08-25T20:44:56.180572Z","shell.execute_reply.started":"2026-08-25T20:44:56.154236Z","shell.execute_reply":"2026-08-25T20:44:56.179674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    series_per_study,\n    discrete=True\n)\n\nplt.xlabel(\"Number of series per study\")\nplt.ylabel(\"Number of studies\")\nplt.title(\"MRI Series per Study\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.181971Z","iopub.execute_input":"2026-08-25T20:44:56.182405Z","iopub.status.idle":"2026-08-25T20:44:56.459774Z","shell.execute_reply.started":"2026-08-25T20:44:56.182365Z","shell.execute_reply":"2026-08-25T20:44:56.458848Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A study is not represented by a single image or a single acquisition.  \nIt can contain several MRI series.","metadata":{}},{"cell_type":"markdown","source":"## 10. MRI series characteristics","metadata":{}},{"cell_type":"markdown","source":"Three useful descriptors are already provided in `train_series.csv`:\n\n- `Anatomical_Plane`\n- `Fluid_Sensitive`\n- `Fat_Suppression`","metadata":{}},{"cell_type":"code","source":"def percent_countplot(data, x, title=None, figsize=(8, 5)):\n    plt.figure(figsize=figsize)\n\n    ax = sns.countplot(\n        data=data,\n        x=x,\n        stat=\"percent\"\n    )\n\n    for container in ax.containers:\n        ax.bar_label(container, fmt=\"%.1f%%\", padding=3)\n\n    ax.set_ylabel(\"Percent\")\n    ax.set_title(title if title else x)\n\n    plt.show()\n\n    return ax","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.460939Z","iopub.execute_input":"2026-08-25T20:44:56.4613Z","iopub.status.idle":"2026-08-25T20:44:56.467794Z","shell.execute_reply.started":"2026-08-25T20:44:56.461264Z","shell.execute_reply":"2026-08-25T20:44:56.466558Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"percent_countplot(\n    train_series,\n    x=\"Anatomical_Plane\",\n    title=\"MRI Series by Anatomical Plane\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.469014Z","iopub.execute_input":"2026-08-25T20:44:56.469366Z","iopub.status.idle":"2026-08-25T20:44:56.695279Z","shell.execute_reply.started":"2026-08-25T20:44:56.469324Z","shell.execute_reply":"2026-08-25T20:44:56.694357Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"percent_countplot(\n    train_series,\n    x=\"Fluid_Sensitive\",\n    title=\"Fluid-Sensitive MRI Series\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.696333Z","iopub.execute_input":"2026-08-25T20:44:56.696628Z","iopub.status.idle":"2026-08-25T20:44:56.87431Z","shell.execute_reply.started":"2026-08-25T20:44:56.696601Z","shell.execute_reply":"2026-08-25T20:44:56.873411Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"percent_countplot(\n    train_series,\n    x=\"Fat_Suppression\",\n    title=\"Fat-Suppressed MRI Series\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:56.87565Z","iopub.execute_input":"2026-08-25T20:44:56.876033Z","iopub.status.idle":"2026-08-25T20:44:57.059978Z","shell.execute_reply.started":"2026-08-25T20:44:56.875994Z","shell.execute_reply":"2026-08-25T20:44:57.059045Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Common series combinations","metadata":{}},{"cell_type":"code","source":"series_combinations = (\n    train_series\n    .groupby(\n        [\"Anatomical_Plane\", \"Fluid_Sensitive\", \"Fat_Suppression\"],\n        dropna=False\n    )\n    .size()\n    .reset_index(name=\"series_count\")\n    .sort_values(\"series_count\", ascending=False)\n)\n\ndisplay(series_combinations.head(15))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.062295Z","iopub.execute_input":"2026-08-25T20:44:57.062597Z","iopub.status.idle":"2026-08-25T20:44:57.082911Z","shell.execute_reply.started":"2026-08-25T20:44:57.062569Z","shell.execute_reply":"2026-08-25T20:44:57.081691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The dataset contains different combinations of imaging plane, fluid sensitivity, and fat suppression, reflecting the multi-sequence nature of knee MRI.","metadata":{}},{"cell_type":"markdown","source":"## 11. The 12 target abnormalities","metadata":{}},{"cell_type":"code","source":"TARGETS = [\n    \"ACL\",\n    \"MCL\",\n    \"Medial Meniscus\",\n    \"Lateral Meniscus\",\n    \"Medial OA\",\n    \"Lateral OA\",\n    \"PF OA\",\n    \"Effusion\",\n    \"Synovitis\",\n    \"Baker's\",\n    \"Contusion\",\n    \"Fracture\",\n]\n\nprint(\"Number of targets:\", len(TARGETS))\nprint()\n\nfor target in TARGETS:\n    print(\"-\", target)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.084218Z","iopub.execute_input":"2026-08-25T20:44:57.084693Z","iopub.status.idle":"2026-08-25T20:44:57.098276Z","shell.execute_reply.started":"2026-08-25T20:44:57.084654Z","shell.execute_reply":"2026-08-25T20:44:57.097202Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 12. Label availability","metadata":{}},{"cell_type":"code","source":"label_availability = pd.DataFrame({\n    \"available\": train[TARGETS].notna().sum(),\n    \"missing\": train[TARGETS].isna().sum(),\n    \"missing_%\": (train[TARGETS].isna().mean() * 100).round(2)\n})\n\ndisplay(label_availability)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.099467Z","iopub.execute_input":"2026-08-25T20:44:57.099791Z","iopub.status.idle":"2026-08-25T20:44:57.125658Z","shell.execute_reply.started":"2026-08-25T20:44:57.099748Z","shell.execute_reply":"2026-08-25T20:44:57.124781Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"has_labels = train[TARGETS].notna().any(axis=1)\n\nn_labeled = int(has_labels.sum())\nn_unlabeled = int((~has_labels).sum())\n\nprint(\"Total training studies       :\", len(train))\nprint(\"Studies with explicit labels :\", n_labeled)\nprint(\"Studies without labels       :\", n_unlabeled)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.126748Z","iopub.execute_input":"2026-08-25T20:44:57.127189Z","iopub.status.idle":"2026-08-25T20:44:57.136806Z","shell.execute_reply.started":"2026-08-25T20:44:57.127148Z","shell.execute_reply":"2026-08-25T20:44:57.135482Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Important\n\nA missing target value does **not** mean that the abnormality is absent.\n\nFor example:\n\n`ACL = NaN` does **not** mean `ACL = 0`.\n\nFor label-based EDA, we therefore work only with studies that contain explicit target labels.","metadata":{}},{"cell_type":"code","source":"labeled = train.loc[has_labels].copy()\n\nprint(\"Labeled subset shape:\", labeled.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.138687Z","iopub.execute_input":"2026-08-25T20:44:57.14002Z","iopub.status.idle":"2026-08-25T20:44:57.15307Z","shell.execute_reply.started":"2026-08-25T20:44:57.139937Z","shell.execute_reply":"2026-08-25T20:44:57.151986Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 13. Positive and negative labels","metadata":{}},{"cell_type":"code","source":"label_stats = []\n\nfor target in TARGETS:\n    available = labeled[target].dropna()\n\n    label_stats.append({\n        \"Target\": target,\n        \"Available\": len(available),\n        \"Positive\": int((available == 1).sum()),\n        \"Negative\": int((available == 0).sum()),\n        \"Positive_rate_%\": available.mean() * 100\n    })\n\nlabel_stats = pd.DataFrame(label_stats)\n\ndisplay(\n    label_stats.style.format({\n        \"Positive_rate_%\": \"{:.1f}%\"\n    })\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.154313Z","iopub.execute_input":"2026-08-25T20:44:57.154702Z","iopub.status.idle":"2026-08-25T20:44:57.420961Z","shell.execute_reply.started":"2026-08-25T20:44:57.154662Z","shell.execute_reply":"2026-08-25T20:44:57.419939Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 14. Abnormality prevalence","metadata":{}},{"cell_type":"code","source":"plot_df = label_stats.sort_values(\"Positive_rate_%\")\n\nplt.figure(figsize=(10, 6))\n\nax = sns.barplot(\n    data=plot_df,\n    x=\"Positive_rate_%\",\n    y=\"Target\"\n)\n\nfor container in ax.containers:\n    ax.bar_label(container, fmt=\"%.1f%%\", padding=3)\n\nplt.xlabel(\"Positive rate among available labels (%)\")\nplt.ylabel(\"\")\nplt.title(\"Target Prevalence in the Labeled Subset\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.422288Z","iopub.execute_input":"2026-08-25T20:44:57.422852Z","iopub.status.idle":"2026-08-25T20:44:57.677401Z","shell.execute_reply.started":"2026-08-25T20:44:57.422808Z","shell.execute_reply":"2026-08-25T20:44:57.675789Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The prevalence is calculated only from explicit labels.  \nMissing values from the rest of the training set are not counted as negatives.","metadata":{}},{"cell_type":"markdown","source":"## 15. How many abnormalities can appear in one study?","metadata":{}},{"cell_type":"code","source":"positive_count = labeled[TARGETS].eq(1).sum(axis=1)\n\ndisplay(\n    positive_count\n    .value_counts()\n    .sort_index()\n    .rename_axis(\"positive_abnormalities\")\n    .to_frame(\"studies\")\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.678285Z","iopub.status.idle":"2026-08-25T20:44:57.678683Z","shell.execute_reply.started":"2026-08-25T20:44:57.678462Z","shell.execute_reply":"2026-08-25T20:44:57.678488Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(9, 5))\n\nsns.histplot(\n    positive_count,\n    discrete=True\n)\n\nplt.xlabel(\"Number of positive abnormalities\")\nplt.ylabel(\"Number of studies\")\nplt.title(\"Positive Targets per Labeled Study\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.680048Z","iopub.status.idle":"2026-08-25T20:44:57.680453Z","shell.execute_reply.started":"2026-08-25T20:44:57.680249Z","shell.execute_reply":"2026-08-25T20:44:57.680275Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is a **multilabel** problem: one study can contain more than one positive abnormality.","metadata":{}},{"cell_type":"markdown","source":"## 16. Target co-occurrence","metadata":{}},{"cell_type":"code","source":"positive_matrix = labeled[TARGETS].eq(1).astype(int)\ncooccurrence = positive_matrix.T.dot(positive_matrix)\n\nplt.figure(figsize=(11, 9))\n\nsns.heatmap(\n    cooccurrence,\n    annot=True,\n    fmt=\"d\"\n)\n\nplt.title(\"Target Co-occurrence in the Labeled Subset\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.681866Z","iopub.status.idle":"2026-08-25T20:44:57.682234Z","shell.execute_reply.started":"2026-08-25T20:44:57.682091Z","shell.execute_reply":"2026-08-25T20:44:57.68211Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The heatmap shows how often two abnormalities are positive in the same labeled study.","metadata":{}},{"cell_type":"markdown","source":"## 17. Radiology reports","metadata":{}},{"cell_type":"code","source":"report_summary = pd.DataFrame({\n    \"value\": [\n        len(train),\n        int(train[\"Report\"].notna().sum()),\n        int(train[\"Report\"].isna().sum()),\n        int(train[\"Report\"].nunique(dropna=True))\n    ]\n},\n    index=[\n        \"Training studies\",\n        \"Reports available\",\n        \"Reports missing\",\n        \"Unique reports\"\n    ]\n)\n\ndisplay(report_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.683744Z","iopub.status.idle":"2026-08-25T20:44:57.684136Z","shell.execute_reply.started":"2026-08-25T20:44:57.683906Z","shell.execute_reply":"2026-08-25T20:44:57.68393Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"report_lengths = train[\"Report\"].dropna().str.len()\n\ndisplay(report_lengths.describe().to_frame(\"report_length_characters\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.685377Z","iopub.status.idle":"2026-08-25T20:44:57.686019Z","shell.execute_reply.started":"2026-08-25T20:44:57.685799Z","shell.execute_reply":"2026-08-25T20:44:57.685827Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    report_lengths,\n    bins=40\n)\n\nplt.xlabel(\"Report length (characters)\")\nplt.ylabel(\"Number of reports\")\nplt.title(\"Radiology Report Length\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.687601Z","iopub.status.idle":"2026-08-25T20:44:57.687958Z","shell.execute_reply.started":"2026-08-25T20:44:57.687778Z","shell.execute_reply":"2026-08-25T20:44:57.687803Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The radiology report is available as a separate source of information for each study where it is present.\n\nThis notebook only describes it; no report-based modeling is performed here.","metadata":{}},{"cell_type":"markdown","source":"## 18. Sample submission","metadata":{}},{"cell_type":"code","source":"print(\"Sample submission shape:\", sample_submission.shape)\ndisplay(sample_submission.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.688929Z","iopub.status.idle":"2026-08-25T20:44:57.689283Z","shell.execute_reply.started":"2026-08-25T20:44:57.68914Z","shell.execute_reply":"2026-08-25T20:44:57.689164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Submission columns:\\n\")\n\nfor column in sample_submission.columns:\n    print(\"-\", column)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.6911Z","iopub.status.idle":"2026-08-25T20:44:57.691444Z","shell.execute_reply.started":"2026-08-25T20:44:57.691286Z","shell.execute_reply":"2026-08-25T20:44:57.691311Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 19. Dataset at a glance","metadata":{}},{"cell_type":"code","source":"summary = pd.DataFrame({\n    \"Metric\": [\n        \"Training studies\",\n        \"Visible test studies\",\n        \"Training series\",\n        \"Visible test series\",\n        \"Studies with explicit target labels\",\n        \"Studies without explicit target labels\",\n        \"Targets\"\n    ],\n    \"Value\": [\n        train[\"StudyInstanceUID\"].nunique(),\n        test[\"StudyInstanceUID\"].nunique(),\n        train_series[\"SeriesInstanceUID\"].nunique(),\n        test_series[\"SeriesInstanceUID\"].nunique(),\n        n_labeled,\n        n_unlabeled,\n        len(TARGETS)\n    ]\n})\n\ndisplay(summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:44:57.69272Z","iopub.status.idle":"2026-08-25T20:44:57.693075Z","shell.execute_reply.started":"2026-08-25T20:44:57.692876Z","shell.execute_reply":"2026-08-25T20:44:57.692898Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n\n- The MRI data are organized as **Study -> Series -> DICOM slices**.\n- One study can contain several MRI series.\n- Series differ by anatomical plane and acquisition characteristics.\n- The task contains **12 target abnormalities**.\n- It is a **multilabel** problem.\n- Most target entries are missing, so **missing values must not be interpreted as negative labels**.\n- Label prevalence and co-occurrence should be calculated from the explicitly labeled subset.\n- Radiology reports are also present in the training table.\n\n---\n\n### Next notebook\n\n[**2. RSNA Knee -> Look Under the Hood: DICOM Metadata Explorer**](https://www.kaggle.com/code/h17ann/2-rsna-knee-dicom-metadata-explorer)\n\nThere we can move from the CSV tables into the DICOM files and inspect the acquisition metadata.","metadata":{}}]}