{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"1b9d39ff","cell_type":"markdown","source":"**Look Under the Hood: DICOM Metadata Explorer**\n\nIn the first notebook [here](https://www.kaggle.com/code/h17ann/1-rsna-knee-simple-starting-point-eda), we looked at the CSV files and the general dataset structure.\n\nNow let's open the DICOM headers and see what information is stored there.\n\nI will keep this notebook focused on metadata only. The actual MRI pixels will be explored in the next notebook.\n\nWe will look at:\n\n- manufacturer\n- patient sex and body part\n- scanning sequence\n- acquisition type\n- magnetic field strength\n- image size\n- slice thickness\n- spacing between slices\n- pixel spacing","metadata":{}},{"id":"9c27af34","cell_type":"markdown","source":"## 1. Imports","metadata":{}},{"id":"267ea29b","cell_type":"code","source":"import kagglehub\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pydicom as pdc\nfrom pydicom.multival import MultiValue\n\nfrom tqdm.auto import tqdm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:07.936152Z","iopub.execute_input":"2026-08-25T12:14:07.936446Z","iopub.status.idle":"2026-08-25T12:14:07.941576Z","shell.execute_reply.started":"2026-08-25T12:14:07.936419Z","shell.execute_reply":"2026-08-25T12:14:07.940556Z"}},"outputs":[],"execution_count":null},{"id":"d9a23d92","cell_type":"markdown","source":"## 2. Data","metadata":{}},{"id":"bcbf9a7b","cell_type":"code","source":"DATA_DIR = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\n# Download data if it is not already attached\nif not DATA_DIR.exists():\n    DATA_DIR = Path(\n        kagglehub.competition_download(\n            \"rsna-knee-abnormality-detection\"\n        )\n    )\nTRAIN_DIR = DATA_DIR / \"train_series\"\n\ntrain_series = pd.read_csv(DATA_DIR / \"train_series.csv\")\n\ndisplay(train_series.head())\nprint(\"Studies:\", train_series[\"StudyInstanceUID\"].nunique())\nprint(\"Series :\", train_series[\"SeriesInstanceUID\"].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:07.942729Z","iopub.execute_input":"2026-08-25T12:14:07.943146Z","iopub.status.idle":"2026-08-25T12:14:08.049036Z","shell.execute_reply.started":"2026-08-25T12:14:07.94311Z","shell.execute_reply":"2026-08-25T12:14:08.047885Z"}},"outputs":[],"execution_count":null},{"id":"8f5fb877","cell_type":"markdown","source":"The image files are organized as:\n\n`StudyInstanceUID / SeriesInstanceUID / *.dcm`\n\nFor this notebook I will use a sample of series so the metadata extraction stays reasonably fast.","metadata":{}},{"id":"56e18bb4","cell_type":"markdown","source":"## 3. Pick a sample of series","metadata":{}},{"id":"62847175","cell_type":"code","source":"N_SERIES = 300\n\nsample_series = train_series.sample(\n    n=min(N_SERIES, len(train_series)),\n    random_state=42\n)\n\nsample_series.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:08.050112Z","iopub.execute_input":"2026-08-25T12:14:08.050362Z","iopub.status.idle":"2026-08-25T12:14:08.064211Z","shell.execute_reply.started":"2026-08-25T12:14:08.050338Z","shell.execute_reply":"2026-08-25T12:14:08.063252Z"}},"outputs":[],"execution_count":null},{"id":"a3599386","cell_type":"code","source":"dicom_paths = []\n\nfor row in sample_series.itertuples(index=False):\n    series_dir = (\n        TRAIN_DIR\n        / str(row.StudyInstanceUID)\n        / str(row.SeriesInstanceUID)\n    )\n\n    dicom_paths.extend(series_dir.glob(\"*.dcm\"))\n\nprint(\"DICOM files found:\", len(dicom_paths))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:08.065363Z","iopub.execute_input":"2026-08-25T12:14:08.065679Z","iopub.status.idle":"2026-08-25T12:14:08.341099Z","shell.execute_reply.started":"2026-08-25T12:14:08.065639Z","shell.execute_reply":"2026-08-25T12:14:08.340119Z"}},"outputs":[],"execution_count":null},{"id":"7b4ce7c0","cell_type":"markdown","source":"## 4. Read one DICOM header","metadata":{}},{"id":"85eb11ec","cell_type":"code","source":"dcm_path = dicom_paths[0]\n\nds = pdc.dcmread(\n    dcm_path,\n    stop_before_pixels=True\n)\n\nprint(ds)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:08.342089Z","iopub.execute_input":"2026-08-25T12:14:08.342352Z","iopub.status.idle":"2026-08-25T12:14:08.352656Z","shell.execute_reply.started":"2026-08-25T12:14:08.342327Z","shell.execute_reply":"2026-08-25T12:14:08.351624Z"}},"outputs":[],"execution_count":null},{"id":"3f01c997","cell_type":"markdown","source":"`stop_before_pixels=True` means that we read the DICOM header without loading the image pixels.\n\nThat is enough for this notebook and makes the reading faster.","metadata":{}},{"id":"a8f199b5","cell_type":"markdown","source":"## 5. Metadata we want to extract","metadata":{}},{"id":"55f714e4","cell_type":"code","source":"elements_list = [\n    \"study_uid\",\n    \"series_uid\",\n    \"modality\",\n    \"manufacturer\",\n    \"patient_sex\",\n    \"body_part\",\n    \"scanning_sequence\",\n    \"acquisition_type\",\n    \"magnetic_field_strength\",\n    \"spacing_between_slices\",\n    \"slice_thickness\",\n    \"rows\",\n    \"columns\",\n    \"pixel_spacing\"\n]\n\ntags_list = [\n    (0x0020, 0x000D),  # StudyInstanceUID\n    (0x0020, 0x000E),  # SeriesInstanceUID\n    (0x0008, 0x0060),  # Modality\n    (0x0008, 0x0070),  # Manufacturer\n    (0x0010, 0x0040),  # PatientSex\n    (0x0018, 0x0015),  # BodyPartExamined\n    (0x0018, 0x0020),  # ScanningSequence\n    (0x0018, 0x0023),  # MRAcquisitionType\n    (0x0018, 0x0087),  # MagneticFieldStrength\n    (0x0018, 0x0088),  # SpacingBetweenSlices\n    (0x0018, 0x0050),  # SliceThickness\n    (0x0028, 0x0010),  # Rows\n    (0x0028, 0x0011),  # Columns\n    (0x0028, 0x0030)   # PixelSpacing\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:08.353795Z","iopub.execute_input":"2026-08-25T12:14:08.354212Z","iopub.status.idle":"2026-08-25T12:14:08.367546Z","shell.execute_reply.started":"2026-08-25T12:14:08.354182Z","shell.execute_reply":"2026-08-25T12:14:08.366565Z"}},"outputs":[],"execution_count":null},{"id":"d134dd2e","cell_type":"markdown","source":"## 6. Extract the metadata","metadata":{}},{"id":"cd7be2ad","cell_type":"code","source":"metadata_ = {\n    \"path\": [],\n    **{name: [] for name in elements_list}\n}\n\nfor dcm_path in tqdm(dicom_paths):\n\n    ds = pdc.dcmread(\n        dcm_path,\n        stop_before_pixels=True\n    )\n\n    metadata_[\"path\"].append(str(dcm_path))\n\n    for name, tag in zip(elements_list, tags_list):\n\n        content = ds.get(tag)\n\n        if content is not None:\n            value = content.value\n\n            if isinstance(value, MultiValue):\n\n                if name == \"pixel_spacing\":\n                    value = [float(v) for v in value]\n                else:\n                    value = \"\\\\\".join(str(v) for v in value)\n\n        else:\n            value = None\n\n        metadata_[name].append(value)\n\nmetadata_df = pd.DataFrame(metadata_)\n\ndisplay(metadata_df.head())\nprint(\"Shape:\", metadata_df.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:08.368916Z","iopub.execute_input":"2026-08-25T12:14:08.369316Z","iopub.status.idle":"2026-08-25T12:14:27.504915Z","shell.execute_reply.started":"2026-08-25T12:14:08.369289Z","shell.execute_reply":"2026-08-25T12:14:27.50383Z"}},"outputs":[],"execution_count":null},{"id":"11626088","cell_type":"markdown","source":"Each row of `metadata_df` corresponds to one DICOM slice.\n\nA lot of metadata are repeated across all slices of the same MRI series.","metadata":{}},{"id":"91275111","cell_type":"markdown","source":"## 7. Save it once","metadata":{}},{"id":"555d692f","cell_type":"markdown","source":"Before saving, multi-valued DICOM fields are converted to regular Python values.\n\n- `PixelSpacing` becomes a normal list of two floats.\n- Other `MultiValue` fields are stored as simple text.\n\nThis keeps the Parquet export happy.","metadata":{}},{"id":"ba25c85d","cell_type":"code","source":"metadata_df.to_parquet(\n    \"/kaggle/working/dicom_metadata_sample.parquet\",\n    index=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.506028Z","iopub.execute_input":"2026-08-25T12:14:27.506387Z","iopub.status.idle":"2026-08-25T12:14:27.55749Z","shell.execute_reply.started":"2026-08-25T12:14:27.506347Z","shell.execute_reply":"2026-08-25T12:14:27.55652Z"}},"outputs":[],"execution_count":null},{"id":"b8e3c475","cell_type":"markdown","source":"This is useful because extracting DICOM metadata can take some time.  \nWe can save the result and reuse it later instead of reading every DICOM file again.","metadata":{}},{"id":"e333ea4f","cell_type":"markdown","source":"## 8. Slice level vs series level","metadata":{}},{"id":"740f62d8","cell_type":"code","source":"print(\"DICOM slices :\", len(metadata_df))\nprint(\"Unique series:\", metadata_df[\"series_uid\"].nunique())\nprint(\"Unique studies:\", metadata_df[\"study_uid\"].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.558501Z","iopub.execute_input":"2026-08-25T12:14:27.558788Z","iopub.status.idle":"2026-08-25T12:14:27.568405Z","shell.execute_reply.started":"2026-08-25T12:14:27.558762Z","shell.execute_reply":"2026-08-25T12:14:27.567402Z"}},"outputs":[],"execution_count":null},{"id":"62db86bd","cell_type":"markdown","source":"For acquisition metadata, I prefer to count each MRI series once.\n\nOtherwise, a series with 60 slices would have twice the weight of a series with 30 slices.","metadata":{}},{"id":"6abd5fb9","cell_type":"code","source":"series_df = metadata_df.drop_duplicates(\"series_uid\").copy()\n\nprint(\"Series-level table:\", series_df.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.569416Z","iopub.execute_input":"2026-08-25T12:14:27.569683Z","iopub.status.idle":"2026-08-25T12:14:27.584683Z","shell.execute_reply.started":"2026-08-25T12:14:27.569657Z","shell.execute_reply":"2026-08-25T12:14:27.583743Z"}},"outputs":[],"execution_count":null},{"id":"246afd1a","cell_type":"markdown","source":"## 9. Missing metadata","metadata":{}},{"id":"c3e19898","cell_type":"code","source":"missing = (\n    series_df\n    .isna()\n    .mean()\n    .mul(100)\n    .sort_values(ascending=False)\n)\n\ndisplay(missing.round(2).to_frame(\"missing_%\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.585778Z","iopub.execute_input":"2026-08-25T12:14:27.586164Z","iopub.status.idle":"2026-08-25T12:14:27.607994Z","shell.execute_reply.started":"2026-08-25T12:14:27.586135Z","shell.execute_reply":"2026-08-25T12:14:27.607095Z"}},"outputs":[],"execution_count":null},{"id":"3d1492b8","cell_type":"code","source":"plt.figure(figsize=(10, 6))\n\nax = sns.barplot(\n    x=missing.values,\n    y=missing.index\n)\n\nfor container in ax.containers:\n    ax.bar_label(container, fmt=\"%.1f%%\", padding=3)\n    \nplt.xlabel(\"Missing values (%)\")\nplt.ylabel(\"\")\nplt.title(\"Missing DICOM metadata\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:19:07.746923Z","iopub.execute_input":"2026-08-25T12:19:07.74731Z","iopub.status.idle":"2026-08-25T12:19:08.016872Z","shell.execute_reply.started":"2026-08-25T12:19:07.747274Z","shell.execute_reply":"2026-08-25T12:19:08.016055Z"}},"outputs":[],"execution_count":null},{"id":"89482027","cell_type":"markdown","source":"Some DICOM tags are optional, so missing values are expected for some fields. If a missing field is required for a specific analysis, the corresponding series may need to be excluded from that analysis.","metadata":{}},{"id":"708d9474","cell_type":"markdown","source":"## 10. Modality","metadata":{}},{"id":"3e2a69f1","cell_type":"code","source":"plt.figure(figsize=(6, 4))\n\nsns.countplot(\n    data=series_df,\n    x=\"modality\",\n    stat=\"percent\"\n)\n\nplt.ylabel(\"Percent\")\nplt.title(\"Modality\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.851103Z","iopub.execute_input":"2026-08-25T12:14:27.851398Z","iopub.status.idle":"2026-08-25T12:14:27.96773Z","shell.execute_reply.started":"2026-08-25T12:14:27.851369Z","shell.execute_reply":"2026-08-25T12:14:27.966915Z"}},"outputs":[],"execution_count":null},{"id":"e8f755c1","cell_type":"markdown","source":"## 11. Manufacturer","metadata":{}},{"id":"02de0fbd","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nax = sns.countplot(\n    data=series_df,\n    x=\"manufacturer\",\n    stat=\"percent\"\n)\n\nplt.xticks(rotation=45)\nplt.ylabel(\"Percent\")\nplt.title(\"MRI manufacturer\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:27.968811Z","iopub.execute_input":"2026-08-25T12:14:27.969113Z","iopub.status.idle":"2026-08-25T12:14:28.196175Z","shell.execute_reply.started":"2026-08-25T12:14:27.969084Z","shell.execute_reply":"2026-08-25T12:14:28.195087Z"}},"outputs":[],"execution_count":null},{"id":"eb282745","cell_type":"markdown","source":"## 12. Patient sex","metadata":{}},{"id":"b9f02160","cell_type":"code","source":"study_df = metadata_df.drop_duplicates(\"study_uid\").copy()\n\nplt.figure(figsize=(6, 4))\n\nsns.countplot(\n    data=study_df,\n    x=\"patient_sex\",\n    stat=\"percent\"\n)\n\nplt.ylabel(\"Percent\")\nplt.title(\"Patient sex\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.197212Z","iopub.execute_input":"2026-08-25T12:14:28.19746Z","iopub.status.idle":"2026-08-25T12:14:28.319705Z","shell.execute_reply.started":"2026-08-25T12:14:28.197438Z","shell.execute_reply":"2026-08-25T12:14:28.318889Z"}},"outputs":[],"execution_count":null},{"id":"75f0b5e6","cell_type":"markdown","source":"For patient sex I use one row per study instead of one row per slice.","metadata":{}},{"id":"d7f34278","cell_type":"markdown","source":"## 13. Body part","metadata":{}},{"id":"d294967c-3864-4e16-9706-4fd6a97fddd7","cell_type":"code","source":"series_df[\"body_part\"].unique()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:20:59.652277Z","iopub.execute_input":"2026-08-25T12:20:59.652659Z","iopub.status.idle":"2026-08-25T12:20:59.66416Z","shell.execute_reply.started":"2026-08-25T12:20:59.652627Z","shell.execute_reply":"2026-08-25T12:20:59.663152Z"}},"outputs":[],"execution_count":null},{"id":"3e82533e","cell_type":"code","source":"plt.figure(figsize=(8, 4))\n\nsns.countplot(\n    data=series_df,\n    x=\"body_part\",\n    stat=\"percent\"\n)\n\nplt.xticks(rotation=45)\nplt.ylabel(\"Percent\")\nplt.title(\"Body part examined\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.320842Z","iopub.execute_input":"2026-08-25T12:14:28.321192Z","iopub.status.idle":"2026-08-25T12:14:28.480318Z","shell.execute_reply.started":"2026-08-25T12:14:28.321163Z","shell.execute_reply":"2026-08-25T12:14:28.479204Z"}},"outputs":[],"execution_count":null},{"id":"23c44f31","cell_type":"markdown","source":"## 14. Scanning sequence","metadata":{}},{"id":"9fce838d","cell_type":"code","source":"series_df[\"scanning_sequence\"] = series_df[\"scanning_sequence\"].astype(str)\n\nprint(\"Unique values:\", series_df[\"scanning_sequence\"].nunique())\n\ndisplay(\n    series_df[\"scanning_sequence\"]\n    .value_counts()\n    .head(15)\n    .to_frame(\"series\")\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.481516Z","iopub.execute_input":"2026-08-25T12:14:28.481873Z","iopub.status.idle":"2026-08-25T12:14:28.494243Z","shell.execute_reply.started":"2026-08-25T12:14:28.481836Z","shell.execute_reply":"2026-08-25T12:14:28.493422Z"}},"outputs":[],"execution_count":null},{"id":"95de850e","cell_type":"code","source":"top_sequences = (\n    series_df[\"scanning_sequence\"]\n    .value_counts()\n    .head(12)\n    .index\n)\n\nplot_df = series_df[\n    series_df[\"scanning_sequence\"].isin(top_sequences)\n]\n\nplt.figure(figsize=(11, 6))\n\nsns.countplot(\n    data=plot_df,\n    y=\"scanning_sequence\",\n    order=top_sequences\n)\n\nplt.xlabel(\"Number of series\")\nplt.ylabel(\"\")\nplt.title(\"Most common scanning sequences\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.495405Z","iopub.execute_input":"2026-08-25T12:14:28.495806Z","iopub.status.idle":"2026-08-25T12:14:28.654785Z","shell.execute_reply.started":"2026-08-25T12:14:28.495759Z","shell.execute_reply":"2026-08-25T12:14:28.653767Z"}},"outputs":[],"execution_count":null},{"id":"1471ff23","cell_type":"markdown","source":"## 15. Acquisition type","metadata":{}},{"id":"90631617","cell_type":"code","source":"plt.figure(figsize=(7, 4))\n\nsns.countplot(\n    data=series_df,\n    x=\"acquisition_type\",\n    stat=\"percent\"\n)\n\nplt.ylabel(\"Percent\")\nplt.title(\"MR acquisition type\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.655859Z","iopub.execute_input":"2026-08-25T12:14:28.656178Z","iopub.status.idle":"2026-08-25T12:14:28.776674Z","shell.execute_reply.started":"2026-08-25T12:14:28.656148Z","shell.execute_reply":"2026-08-25T12:14:28.775832Z"}},"outputs":[],"execution_count":null},{"id":"5e1867c7","cell_type":"markdown","source":"## 16. Magnetic field strength","metadata":{}},{"id":"44ba9cc4","cell_type":"code","source":"print(series_df[\"magnetic_field_strength\"].value_counts(dropna=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.777727Z","iopub.execute_input":"2026-08-25T12:14:28.778082Z","iopub.status.idle":"2026-08-25T12:14:28.784628Z","shell.execute_reply.started":"2026-08-25T12:14:28.778051Z","shell.execute_reply":"2026-08-25T12:14:28.783689Z"}},"outputs":[],"execution_count":null},{"id":"fe7b45ed","cell_type":"code","source":"plt.figure(figsize=(7, 4))\n\nsns.countplot(\n    data=series_df,\n    x=\"magnetic_field_strength\",\n    stat=\"percent\"\n)\n\nplt.xlabel(\"Magnetic field strength (T)\")\nplt.ylabel(\"Percent\")\nplt.title(\"Magnetic field strength\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.785565Z","iopub.execute_input":"2026-08-25T12:14:28.785844Z","iopub.status.idle":"2026-08-25T12:14:28.925497Z","shell.execute_reply.started":"2026-08-25T12:14:28.785819Z","shell.execute_reply":"2026-08-25T12:14:28.924505Z"}},"outputs":[],"execution_count":null},{"id":"1419ac32","cell_type":"markdown","source":"`MagneticFieldStrength` is expressed in Tesla.","metadata":{}},{"id":"968a821a","cell_type":"markdown","source":"## 17. Image matrix","metadata":{}},{"id":"a72947ea","cell_type":"code","source":"print(\"Unique row values   :\", series_df[\"rows\"].nunique())\nprint(\"Unique column values:\", series_df[\"columns\"].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.926711Z","iopub.execute_input":"2026-08-25T12:14:28.927106Z","iopub.status.idle":"2026-08-25T12:14:28.932656Z","shell.execute_reply.started":"2026-08-25T12:14:28.92707Z","shell.execute_reply":"2026-08-25T12:14:28.931725Z"}},"outputs":[],"execution_count":null},{"id":"c4d71d61","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    data=series_df,\n    x=\"rows\",\n    bins=20,\n    stat=\"percent\"\n)\n\nplt.xlabel(\"Rows\")\nplt.ylabel(\"Percent\")\nplt.title(\"Image rows\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:28.934129Z","iopub.execute_input":"2026-08-25T12:14:28.934405Z","iopub.status.idle":"2026-08-25T12:14:29.136833Z","shell.execute_reply.started":"2026-08-25T12:14:28.93438Z","shell.execute_reply":"2026-08-25T12:14:29.135838Z"}},"outputs":[],"execution_count":null},{"id":"97b9fac7","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    data=series_df,\n    x=\"columns\",\n    bins=20,\n    stat=\"percent\"\n)\n\nplt.xlabel(\"Columns\")\nplt.ylabel(\"Percent\")\nplt.title(\"Image columns\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.13809Z","iopub.execute_input":"2026-08-25T12:14:29.138393Z","iopub.status.idle":"2026-08-25T12:14:29.307302Z","shell.execute_reply.started":"2026-08-25T12:14:29.138362Z","shell.execute_reply":"2026-08-25T12:14:29.306504Z"}},"outputs":[],"execution_count":null},{"id":"24cc0271","cell_type":"markdown","source":"`Rows` and `Columns` are the image dimensions in pixels.\n\nThey do not tell us the physical size of the pixels.","metadata":{}},{"id":"6838c631","cell_type":"markdown","source":"## 18. Slice thickness","metadata":{}},{"id":"c50803ca","cell_type":"code","source":"series_df[\"slice_thickness\"] = pd.to_numeric(\n    series_df[\"slice_thickness\"],\n    errors=\"coerce\"\n)\n\ndisplay(series_df[\"slice_thickness\"].describe().to_frame())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.3083Z","iopub.execute_input":"2026-08-25T12:14:29.308597Z","iopub.status.idle":"2026-08-25T12:14:29.320242Z","shell.execute_reply.started":"2026-08-25T12:14:29.308568Z","shell.execute_reply":"2026-08-25T12:14:29.31912Z"}},"outputs":[],"execution_count":null},{"id":"0abc482b","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    data=series_df,\n    x=\"slice_thickness\",\n    bins=20\n)\n\nplt.xlabel(\"Slice thickness (mm)\")\nplt.title(\"Slice thickness\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.321444Z","iopub.execute_input":"2026-08-25T12:14:29.321726Z","iopub.status.idle":"2026-08-25T12:14:29.50808Z","shell.execute_reply.started":"2026-08-25T12:14:29.321698Z","shell.execute_reply":"2026-08-25T12:14:29.507232Z"}},"outputs":[],"execution_count":null},{"id":"f6ce829c","cell_type":"code","source":"plt.figure(figsize=(10, 2.5))\n\nsns.boxplot(\n    data=series_df,\n    x=\"slice_thickness\"\n)\n\nplt.xlabel(\"Slice thickness (mm)\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.509169Z","iopub.execute_input":"2026-08-25T12:14:29.50955Z","iopub.status.idle":"2026-08-25T12:14:29.605255Z","shell.execute_reply.started":"2026-08-25T12:14:29.509501Z","shell.execute_reply":"2026-08-25T12:14:29.604071Z"}},"outputs":[],"execution_count":null},{"id":"05b24508","cell_type":"markdown","source":"## 19. Spacing between slices","metadata":{}},{"id":"d3350808","cell_type":"code","source":"series_df[\"spacing_between_slices\"] = pd.to_numeric(\n    series_df[\"spacing_between_slices\"],\n    errors=\"coerce\"\n)\n\ndisplay(\n    series_df[\"spacing_between_slices\"]\n    .describe()\n    .to_frame()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.606456Z","iopub.execute_input":"2026-08-25T12:14:29.606713Z","iopub.status.idle":"2026-08-25T12:14:29.619288Z","shell.execute_reply.started":"2026-08-25T12:14:29.606688Z","shell.execute_reply":"2026-08-25T12:14:29.618011Z"}},"outputs":[],"execution_count":null},{"id":"8b468959","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    data=series_df,\n    x=\"spacing_between_slices\",\n    bins=20\n)\n\nplt.xlabel(\"Spacing between slices (mm)\")\nplt.title(\"Spacing between slices\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.620344Z","iopub.execute_input":"2026-08-25T12:14:29.620662Z","iopub.status.idle":"2026-08-25T12:14:29.813978Z","shell.execute_reply.started":"2026-08-25T12:14:29.620628Z","shell.execute_reply":"2026-08-25T12:14:29.812992Z"}},"outputs":[],"execution_count":null},{"id":"a57ed01c","cell_type":"markdown","source":"`SliceThickness` and `SpacingBetweenSlices` are not exactly the same thing.\n\nAlso, `SpacingBetweenSlices` is not always present in the DICOM header.","metadata":{}},{"id":"2c767a61","cell_type":"markdown","source":"## 20. Pixel spacing","metadata":{}},{"id":"6a16eea8","cell_type":"markdown","source":"`PixelSpacing` contains two values:\n\n`[row spacing, column spacing]`\n\nThe unit is millimeters per pixel.","metadata":{}},{"id":"b1c40fcc","cell_type":"code","source":"series_df[\"pixel_spacing_row\"] = series_df[\"pixel_spacing\"].apply(\n    lambda x: float(x[0]) if x is not None else np.nan\n)\n\nseries_df[\"pixel_spacing_col\"] = series_df[\"pixel_spacing\"].apply(\n    lambda x: float(x[1]) if x is not None else np.nan\n)\n\ndisplay(\n    series_df[\n        [\"pixel_spacing_row\", \"pixel_spacing_col\"]\n    ].describe()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.815247Z","iopub.execute_input":"2026-08-25T12:14:29.815683Z","iopub.status.idle":"2026-08-25T12:14:29.834784Z","shell.execute_reply.started":"2026-08-25T12:14:29.815642Z","shell.execute_reply":"2026-08-25T12:14:29.833962Z"}},"outputs":[],"execution_count":null},{"id":"bff08dff","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    data=series_df,\n    x=\"pixel_spacing_row\",\n    bins=25\n)\n\nplt.xlabel(\"Row pixel spacing (mm/pixel)\")\nplt.title(\"Pixel spacing\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:29.836045Z","iopub.execute_input":"2026-08-25T12:14:29.836414Z","iopub.status.idle":"2026-08-25T12:14:30.042014Z","shell.execute_reply.started":"2026-08-25T12:14:29.836376Z","shell.execute_reply":"2026-08-25T12:14:30.040927Z"}},"outputs":[],"execution_count":null},{"id":"0247c829","cell_type":"markdown","source":"## 21. Are the in-plane pixels isotropic?","metadata":{}},{"id":"af43c3fe","cell_type":"code","source":"spacing_df = series_df.dropna(\n    subset=[\"pixel_spacing_row\", \"pixel_spacing_col\"]\n).copy()\n\nsame_spacing = np.isclose(\n    spacing_df[\"pixel_spacing_row\"],\n    spacing_df[\"pixel_spacing_col\"]\n)\n\nprint(\n    \"Same row and column spacing:\",\n    f\"{same_spacing.mean() * 100:.2f}%\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:30.043126Z","iopub.execute_input":"2026-08-25T12:14:30.043401Z","iopub.status.idle":"2026-08-25T12:14:30.052688Z","shell.execute_reply.started":"2026-08-25T12:14:30.043374Z","shell.execute_reply":"2026-08-25T12:14:30.051647Z"}},"outputs":[],"execution_count":null},{"id":"fdc1bf9a","cell_type":"markdown","source":"If row and column spacing are equal, the pixels are isotropic **inside the 2D slice**.\n\nThat does not automatically mean that the complete 3D volume is isotropic, because the spacing between slices can be different.","metadata":{}},{"id":"42a55c28","cell_type":"markdown","source":"## 22. Link the DICOM metadata back to the CSV","metadata":{}},{"id":"bebca1f0","cell_type":"code","source":"series_df = series_df.merge(\n    train_series[\n        [\n            \"StudyInstanceUID\",\n            \"SeriesInstanceUID\",\n            \"Anatomical_Plane\",\n            \"Fluid_Sensitive\",\n            \"Fat_Suppression\"\n        ]\n    ],\n    left_on=[\"study_uid\", \"series_uid\"],\n    right_on=[\"StudyInstanceUID\", \"SeriesInstanceUID\"],\n    how=\"left\"\n)\n\ndisplay(\n    series_df[\n        [\n            \"manufacturer\",\n            \"magnetic_field_strength\",\n            \"rows\",\n            \"columns\",\n            \"pixel_spacing_row\",\n            \"pixel_spacing_col\",\n            \"Anatomical_Plane\",\n            \"Fluid_Sensitive\",\n            \"Fat_Suppression\"\n        ]\n    ].head()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T12:14:30.053959Z","iopub.execute_input":"2026-08-25T12:14:30.054347Z","iopub.status.idle":"2026-08-25T12:14:30.098874Z","shell.execute_reply.started":"2026-08-25T12:14:30.054304Z","shell.execute_reply":"2026-08-25T12:14:30.098045Z"}},"outputs":[],"execution_count":null},{"id":"3927c940","cell_type":"markdown","source":"The DICOM identifiers make it easy to connect the header information back to `train_series.csv`.","metadata":{}},{"id":"7c4f365e","cell_type":"markdown","source":"## Takeaways","metadata":{}},{"id":"e7eb30a8","cell_type":"markdown","source":"- DICOM files contain much more than the MRI pixel values.\n- We can inspect the acquisition metadata without loading the image itself.\n- A study contains several series, and each series contains several DICOM slices.\n- For acquisition metadata, series-level statistics are more meaningful than counting every slice.\n- Manufacturer, field strength, matrix size, slice thickness and spatial resolution vary across the data.\n- `PixelSpacing` gives the physical pixel size inside the image plane.\n- Some DICOM tags are missing because not every field is mandatory.\n\n### Next notebook\n\n[**3. RSNA Knee -> See the MRI: From DICOM to Slices & Studies**](https://www.kaggle.com/code/h17ann/3-rsna-knee-see-the-mri-ax-cor-sag)\n\nNext, we will load the pixel data and visualize complete MRI series.","metadata":{}}]}