{
  "id": 608767,
  "title": "Clarification on the Meaning of PatientID and StudyUID in the DICOM Files",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/608767",
  "author_name": "morixx",
  "post_date": "2025-09-22T05:46:38.936000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have a question regarding the identifiers in the provided DICOM files.<br>\nIn the dataset description, it is stated:</p>\n<p>“The DICOM files contain the identifier tags: PatientID, StudyUID, and SeriesUID. All three tags are included in the DICOMs for conformity to the DICOM standard, but only SeriesUID is meaningful. The PatientID and StudyUID tags do not identify anything.”</p>\n<p>Could you please confirm if this means that:<br>\n    1.  The PatientID and StudyUID tags are anonymized placeholders and do not correspond to actual patient-level grouping, and<br>\n    2.  For data splitting, we should only treat each SeriesInstanceUID as an independent sample, without considering patient-level leakage?</p>",
  "messages": [
    {
      "id": 3292854,
      "postDate": "2025-09-22T17:00:55.483Z",
      "content": "<h1>1 is correct. The goal is of the challenge create an algorithm that can evaluate a single image series at a time, which could be one of the four different modalities.</h1>\n<p>As far as #2, given that #1 is true I think the most reasonable approach is to treat all training cases as independent samples. I can confirm there is no patient-level leakage between training and test sets and your models will only have access to one test series at a time. </p>",
      "rawMarkdown": "#1 is correct. The goal is of the challenge create an algorithm that can evaluate a single image series at a time, which could be one of the four different modalities. \nAs far as #2, given that #1 is true I think the most reasonable approach is to treat all training cases as independent samples. I can confirm there is no patient-level leakage between training and test sets and your models will only have access to one test series at a time. ",
      "votes": 3
    },
    {
      "id": 3292551,
      "postDate": "2025-09-22T05:46:38.937Z",
      "content": "<p>I have a question regarding the identifiers in the provided DICOM files.<br>\nIn the dataset description, it is stated:</p>\n<p>“The DICOM files contain the identifier tags: PatientID, StudyUID, and SeriesUID. All three tags are included in the DICOMs for conformity to the DICOM standard, but only SeriesUID is meaningful. The PatientID and StudyUID tags do not identify anything.”</p>\n<p>Could you please confirm if this means that:<br>\n    1.  The PatientID and StudyUID tags are anonymized placeholders and do not correspond to actual patient-level grouping, and<br>\n    2.  For data splitting, we should only treat each SeriesInstanceUID as an independent sample, without considering patient-level leakage?</p>",
      "rawMarkdown": "I have a question regarding the identifiers in the provided DICOM files.\nIn the dataset description, it is stated:\n\n“The DICOM files contain the identifier tags: PatientID, StudyUID, and SeriesUID. All three tags are included in the DICOMs for conformity to the DICOM standard, but only SeriesUID is meaningful. The PatientID and StudyUID tags do not identify anything.”\n\nCould you please confirm if this means that:\n\t1.\tThe PatientID and StudyUID tags are anonymized placeholders and do not correspond to actual patient-level grouping, and\n\t2.\tFor data splitting, we should only treat each SeriesInstanceUID as an independent sample, without considering patient-level leakage?\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3292854,
      "author_name": "JeffRudie",
      "author_url": "",
      "post_date": "2025-09-22T17:00:55.483000",
      "content": "<h1>1 is correct. The goal is of the challenge create an algorithm that can evaluate a single image series at a time, which could be one of the four different modalities.</h1>\n<p>As far as #2, given that #1 is true I think the most reasonable approach is to treat all training cases as independent samples. I can confirm there is no patient-level leakage between training and test sets and your models will only have access to one test series at a time. </p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3292854": "#1 is correct. The goal is of the challenge create an algorithm that can evaluate a single image series at a time, which could be one of the four different modalities. \nAs far as #2, given that #1 is true I think the most reasonable approach is to treat all training cases as independent samples. I can confirm there is no patient-level leakage between training and test sets and your models will only have access to one test series at a time. ",
    "3292551": "I have a question regarding the identifiers in the provided DICOM files.\nIn the dataset description, it is stated:\n\n“The DICOM files contain the identifier tags: PatientID, StudyUID, and SeriesUID. All three tags are included in the DICOMs for conformity to the DICOM standard, but only SeriesUID is meaningful. The PatientID and StudyUID tags do not identify anything.”\n\nCould you please confirm if this means that:\n\t1.\tThe PatientID and StudyUID tags are anonymized placeholders and do not correspond to actual patient-level grouping, and\n\t2.\tFor data splitting, we should only treat each SeriesInstanceUID as an independent sample, without considering patient-level leakage?\n"
  }
}