{
  "id": 185187,
  "title": "Do I understand Study/Series/SOP correctly?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/185187",
  "author_name": "CoreyJamesLevinson",
  "post_date": "2020-09-19T18:36:21.649000",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Do I understand that…</p>\n<ol>\n<li>StudyInstanceUID = an ID uniquely identifying a single human being, who might come to the doctor's office multiple times</li>\n<li>SeriesInstanceUID = an ID uniquely identifying which visit to the doctor's office you're at. So if you are in a study and have to come in 5 times, then you have 5 different SeriesInstances. And do I understand correctly that this is on different days?</li>\n<li>SOPInstanceUID = an ID uniquely identifying the scan</li>\n</ol>\n<p>Furthermore, do I understand correctly that to prevent leaks, we should split by StudyInstanceUID, i.e. split by human? Ty</p>",
  "messages": [
    {
      "id": 1018613,
      "postDate": "2020-09-19T19:20:29.963Z",
      "content": "<p>Not quite.</p>\n<p>StudyInstanceUID - a single examination. Such as a CT scan, Xray, MRI.</p>\n<p>SeriesInstanceUID - a group of one or more images within the StudyInstanceUID. A real CT study would usually have 2-5 sequences. Either separate acquisitions of data or post-processing of images.</p>\n<p>SOPInstanceUID - a single DICOM image.</p>\n<p>In this competition, we have no patient groupings. We don't know if more than one StudyInstanceUID comes from the same patient. The DICOM data is anonymized. I don't think they left anything in the DICOM that would identify the patient.</p>\n<p>Unless the organizer states it somewhere, we don't know if any patient has more than one StudyInstanceUID. One could try and identify duplicates based on the images (as was done in the Melanoma Competition).</p>\n<p>In this competition (as far as I have seen) we have only one SeriesInstanceUID within each Study. In the real world you would have more. But for this competition you only have one.</p>\n<p>So, your final statement, that we should split by StudyInstanceUID is correct, but we are splitting by Study, not human.</p>",
      "rawMarkdown": "Not quite.\n\nStudyInstanceUID - a single examination. Such as a CT scan, Xray, MRI.\n\nSeriesInstanceUID - a group of one or more images within the StudyInstanceUID. A real CT study would usually have 2-5 sequences. Either separate acquisitions of data or post-processing of images.\n\nSOPInstanceUID - a single DICOM image.\n\nIn this competition, we have no patient groupings. We don't know if more than one StudyInstanceUID comes from the same patient. The DICOM data is anonymized. I don't think they left anything in the DICOM that would identify the patient.\n\nUnless the organizer states it somewhere, we don't know if any patient has more than one StudyInstanceUID. One could try and identify duplicates based on the images (as was done in the Melanoma Competition).\n\nIn this competition (as far as I have seen) we have only one SeriesInstanceUID within each Study. In the real world you would have more. But for this competition you only have one.\n\nSo, your final statement, that we should split by StudyInstanceUID is correct, but we are splitting by Study, not human.",
      "votes": 9,
      "replies": [
        {
          "id": 1018628,
          "postDate": "2020-09-19T19:29:03.300Z",
          "content": "<p>Ohhh, Okay I think I understand. So SeriesInstanceUID is the 3d scan, and SOPInstanceUID is the slice from that 3d image.</p>",
          "rawMarkdown": "Ohhh, Okay I think I understand. So SeriesInstanceUID is the 3d scan, and SOPInstanceUID is the slice from that 3d image.",
          "votes": 3,
          "replies": [
            {
              "id": 1018629,
              "postDate": "2020-09-19T19:33:04.623Z",
              "content": "<p>correct</p>\n<p>-Rich</p>",
              "rawMarkdown": "correct\n\n-Rich",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 1018559,
      "postDate": "2020-09-19T18:36:21.650Z",
      "content": "<p>Do I understand that…</p>\n<ol>\n<li>StudyInstanceUID = an ID uniquely identifying a single human being, who might come to the doctor's office multiple times</li>\n<li>SeriesInstanceUID = an ID uniquely identifying which visit to the doctor's office you're at. So if you are in a study and have to come in 5 times, then you have 5 different SeriesInstances. And do I understand correctly that this is on different days?</li>\n<li>SOPInstanceUID = an ID uniquely identifying the scan</li>\n</ol>\n<p>Furthermore, do I understand correctly that to prevent leaks, we should split by StudyInstanceUID, i.e. split by human? Ty</p>",
      "rawMarkdown": "Do I understand that...\n\n1. StudyInstanceUID = an ID uniquely identifying a single human being, who might come to the doctor's office multiple times\n2. SeriesInstanceUID = an ID uniquely identifying which visit to the doctor's office you're at. So if you are in a study and have to come in 5 times, then you have 5 different SeriesInstances. And do I understand correctly that this is on different days?\n3. SOPInstanceUID = an ID uniquely identifying the scan\n\nFurthermore, do I understand correctly that to prevent leaks, we should split by StudyInstanceUID, i.e. split by human? Ty",
      "votes": 9
    },
    {
      "id": 1021247,
      "postDate": "2020-09-21T17:57:37.630Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true,
      "replies": [
        {
          "id": 1021307,
          "postDate": "2020-09-21T18:37:34.893Z",
          "content": "<p><a href=\"https://www.kaggle.com/vijaysimhareddyp\" target=\"_blank\">@vijaysimhareddyp</a> Please stop spamming your kernel links around the forums. This is irrelevant to this post and is not welcome on Kaggle.</p>",
          "rawMarkdown": "@vijaysimhareddyp Please stop spamming your kernel links around the forums. This is irrelevant to this post and is not welcome on Kaggle.",
          "votes": 8
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1018613,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-19T19:20:29.963000",
      "content": "<p>Not quite.</p>\n<p>StudyInstanceUID - a single examination. Such as a CT scan, Xray, MRI.</p>\n<p>SeriesInstanceUID - a group of one or more images within the StudyInstanceUID. A real CT study would usually have 2-5 sequences. Either separate acquisitions of data or post-processing of images.</p>\n<p>SOPInstanceUID - a single DICOM image.</p>\n<p>In this competition, we have no patient groupings. We don't know if more than one StudyInstanceUID comes from the same patient. The DICOM data is anonymized. I don't think they left anything in the DICOM that would identify the patient.</p>\n<p>Unless the organizer states it somewhere, we don't know if any patient has more than one StudyInstanceUID. One could try and identify duplicates based on the images (as was done in the Melanoma Competition).</p>\n<p>In this competition (as far as I have seen) we have only one SeriesInstanceUID within each Study. In the real world you would have more. But for this competition you only have one.</p>\n<p>So, your final statement, that we should split by StudyInstanceUID is correct, but we are splitting by Study, not human.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1018628,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-09-19T19:29:03.300000",
          "content": "<p>Ohhh, Okay I think I understand. So SeriesInstanceUID is the 3d scan, and SOPInstanceUID is the slice from that 3d image.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 1018629,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-19T19:33:04.623000",
              "content": "<p>correct</p>\n<p>-Rich</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1021247,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-21T17:57:37.630000",
      "content": "",
      "votes": -4,
      "replies": [
        {
          "id": 1021307,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-09-21T18:37:34.893000",
          "content": "<p><a href=\"https://www.kaggle.com/vijaysimhareddyp\" target=\"_blank\">@vijaysimhareddyp</a> Please stop spamming your kernel links around the forums. This is irrelevant to this post and is not welcome on Kaggle.</p>",
          "votes": 8,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1018613": "Not quite.\n\nStudyInstanceUID - a single examination. Such as a CT scan, Xray, MRI.\n\nSeriesInstanceUID - a group of one or more images within the StudyInstanceUID. A real CT study would usually have 2-5 sequences. Either separate acquisitions of data or post-processing of images.\n\nSOPInstanceUID - a single DICOM image.\n\nIn this competition, we have no patient groupings. We don't know if more than one StudyInstanceUID comes from the same patient. The DICOM data is anonymized. I don't think they left anything in the DICOM that would identify the patient.\n\nUnless the organizer states it somewhere, we don't know if any patient has more than one StudyInstanceUID. One could try and identify duplicates based on the images (as was done in the Melanoma Competition).\n\nIn this competition (as far as I have seen) we have only one SeriesInstanceUID within each Study. In the real world you would have more. But for this competition you only have one.\n\nSo, your final statement, that we should split by StudyInstanceUID is correct, but we are splitting by Study, not human.",
    "1018559": "Do I understand that...\n\n1. StudyInstanceUID = an ID uniquely identifying a single human being, who might come to the doctor's office multiple times\n2. SeriesInstanceUID = an ID uniquely identifying which visit to the doctor's office you're at. So if you are in a study and have to come in 5 times, then you have 5 different SeriesInstances. And do I understand correctly that this is on different days?\n3. SOPInstanceUID = an ID uniquely identifying the scan\n\nFurthermore, do I understand correctly that to prevent leaks, we should split by StudyInstanceUID, i.e. split by human? Ty",
    "1021247": ""
  }
}