{
  "id": 188019,
  "title": "Will there ever be multiple series per study?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/188019",
  "author_name": "Louka Ewington-Pitsos",
  "post_date": "2020-10-01T09:50:09.965000",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Has anyone looked into this yet? </p>\n<p>There is exactly one series per study/exam in train.csv and test.csv.</p>\n<p>There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.</p>\n<p>Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.</p>",
  "messages": [
    {
      "id": 1033832,
      "postDate": "2020-10-01T09:50:09.967Z",
      "content": "<p>Has anyone looked into this yet? </p>\n<p>There is exactly one series per study/exam in train.csv and test.csv.</p>\n<p>There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.</p>\n<p>Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.</p>",
      "rawMarkdown": "Has anyone looked into this yet? \n\nThere is exactly one series per study/exam in train.csv and test.csv.\n\nThere are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.\n\nIs this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.",
      "votes": 8
    },
    {
      "id": 1043610,
      "postDate": "2020-10-09T05:22:46.493Z",
      "content": "<p>Hello all, </p>\n<p>I created a notebook to show how to handle an edgecase where there are two series in a study. I hope this helps someone! cc: <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/lewington\" target=\"_blank\">@lewington</a> </p>\n<p><a href=\"https://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort\" target=\"_blank\">https://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort</a></p>",
      "rawMarkdown": "Hello all, \n\nI created a notebook to show how to handle an edgecase where there are two series in a study. I hope this helps someone! cc: @cdeotte @lewington \n\nhttps://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort",
      "votes": 1
    },
    {
      "id": 1043318,
      "postDate": "2020-10-08T21:55:18.607Z",
      "content": "<blockquote>\n  <p>Has anyone looked into this yet? </p>\n  <p>There is exactly one series per study/exam in train.csv and test.csv.</p>\n  <p>There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.</p>\n  <p>Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.</p>\n</blockquote>\n<p>We have cases with two series per exam in training! Search for StudyInstanceUID 759a5963508b</p>",
      "rawMarkdown": "> Has anyone looked into this yet? \n> \n> There is exactly one series per study/exam in train.csv and test.csv.\n> \n> There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.\n> \n> Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.\n\nWe have cases with two series per exam in training! Search for StudyInstanceUID 759a5963508b\n",
      "replies": [
        {
          "id": 1043491,
          "postDate": "2020-10-09T03:14:39.053Z",
          "content": "<p>Weird. When i search, I only see 1 series in that study. What am I doing wrong?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fd3672c30c2807ea29510e86f59c34778%2Fex1.png?generation=1602213270678774&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Weird. When i search, I only see 1 series in that study. What am I doing wrong?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fd3672c30c2807ea29510e86f59c34778%2Fex1.png?generation=1602213270678774&alt=media)"
        },
        {
          "id": 1043505,
          "postDate": "2020-10-09T03:32:08.947Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, if you open up that data in open source imaging software you will see that there are two series (ImageJ, Horos, ITK-snap, 3D-slicer). The Study Instance UID doesn't separate them our (perhaps these are edge cases in the dataset which were missed when organizing the data).</p>\n<p>In pydicom you can also sort by InstanceNumber or ImagePositionPatient and separate them out.</p>\n<p>One series has 96 Images (CT ABD/PEL) and the other has 284 Images (CTA PE).</p>",
          "rawMarkdown": "@cdeotte, if you open up that data in open source imaging software you will see that there are two series (ImageJ, Horos, ITK-snap, 3D-slicer). The Study Instance UID doesn't separate them our (perhaps these are edge cases in the dataset which were missed when organizing the data).\n\nIn pydicom you can also sort by InstanceNumber or ImagePositionPatient and separate them out.\n\nOne series has 96 Images (CT ABD/PEL) and the other has 284 Images (CTA PE).",
          "votes": 2
        }
      ]
    },
    {
      "id": 1043314,
      "postDate": "2020-10-08T21:52:24.903Z",
      "content": "<p>Hello, </p>\n<p>I have looked at the training data and there are many examples where there is more than one series in a study. I attached an example that I isolated from the train.csv.  In <a href=\"https://www.dropbox.com/s/sv0s0x1gvvtikll/additional_series.mp4?dl=0\" target=\"_blank\">this example</a> the additional series is NOT a PE study and includes the abdomen, pelvis and lower-chest. I'm curious if there are examples with PE identified on an image level in these studies or if these are simply ignored and treated as negative.</p>\n<p>Any clarifications here? cc: <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>\n<p>CSV of this example is also attached.</p>",
      "rawMarkdown": "Hello, \n\nI have looked at the training data and there are many examples where there is more than one series in a study. I attached an example that I isolated from the train.csv.  In [this example](https://www.dropbox.com/s/sv0s0x1gvvtikll/additional_series.mp4?dl=0) the additional series is NOT a PE study and includes the abdomen, pelvis and lower-chest. I'm curious if there are examples with PE identified on an image level in these studies or if these are simply ignored and treated as negative.\n\nAny clarifications here? cc: @philculliton\n\nCSV of this example is also attached.\n"
    },
    {
      "id": 1036620,
      "postDate": "2020-10-04T04:40:37.507Z",
      "content": "<p>any update on this . its an important thing to know </p>",
      "rawMarkdown": "any update on this . its an important thing to know "
    },
    {
      "id": 1034368,
      "postDate": "2020-10-01T17:11:13.680Z",
      "content": "<p>We have to wait the host' respond I suppose</p>",
      "rawMarkdown": "We have to wait the host' respond I suppose"
    }
  ],
  "comments": [
    {
      "id": 1043610,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-09T05:22:46.493000",
      "content": "<p>Hello all, </p>\n<p>I created a notebook to show how to handle an edgecase where there are two series in a study. I hope this helps someone! cc: <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/lewington\" target=\"_blank\">@lewington</a> </p>\n<p><a href=\"https://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort\" target=\"_blank\">https://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1043318,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-08T21:55:18.607000",
      "content": "<blockquote>\n  <p>Has anyone looked into this yet? </p>\n  <p>There is exactly one series per study/exam in train.csv and test.csv.</p>\n  <p>There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.</p>\n  <p>Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.</p>\n</blockquote>\n<p>We have cases with two series per exam in training! Search for StudyInstanceUID 759a5963508b</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1043491,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-09T03:14:39.053000",
          "content": "<p>Weird. When i search, I only see 1 series in that study. What am I doing wrong?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fd3672c30c2807ea29510e86f59c34778%2Fex1.png?generation=1602213270678774&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1043505,
          "author_name": "aksg87",
          "author_url": "",
          "post_date": "2020-10-09T03:32:08.947000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, if you open up that data in open source imaging software you will see that there are two series (ImageJ, Horos, ITK-snap, 3D-slicer). The Study Instance UID doesn't separate them our (perhaps these are edge cases in the dataset which were missed when organizing the data).</p>\n<p>In pydicom you can also sort by InstanceNumber or ImagePositionPatient and separate them out.</p>\n<p>One series has 96 Images (CT ABD/PEL) and the other has 284 Images (CTA PE).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1043314,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-08T21:52:24.903000",
      "content": "<p>Hello, </p>\n<p>I have looked at the training data and there are many examples where there is more than one series in a study. I attached an example that I isolated from the train.csv.  In <a href=\"https://www.dropbox.com/s/sv0s0x1gvvtikll/additional_series.mp4?dl=0\" target=\"_blank\">this example</a> the additional series is NOT a PE study and includes the abdomen, pelvis and lower-chest. I'm curious if there are examples with PE identified on an image level in these studies or if these are simply ignored and treated as negative.</p>\n<p>Any clarifications here? cc: <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>\n<p>CSV of this example is also attached.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1036620,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-10-04T04:40:37.507000",
      "content": "<p>any update on this . its an important thing to know </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1034368,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-10-01T17:11:13.680000",
      "content": "<p>We have to wait the host' respond I suppose</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1033832": "Has anyone looked into this yet? \n\nThere is exactly one series per study/exam in train.csv and test.csv.\n\nThere are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.\n\nIs this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.",
    "1043610": "Hello all, \n\nI created a notebook to show how to handle an edgecase where there are two series in a study. I hope this helps someone! cc: @cdeotte @lewington \n\nhttps://www.kaggle.com/aksg87/dicom-sequence-selection-and-sort",
    "1043318": "> Has anyone looked into this yet? \n> \n> There is exactly one series per study/exam in train.csv and test.csv.\n> \n> There are a further 1517 exams in the private test set. We will never have public access to these files so in theory these could contain arbitrarily many series per exam.\n> \n> Is this something we should take into account? Predicting on multiple series per exam seems like it would be difficult given that we are only ever exposed to one per exam during training.\n\nWe have cases with two series per exam in training! Search for StudyInstanceUID 759a5963508b\n",
    "1043314": "Hello, \n\nI have looked at the training data and there are many examples where there is more than one series in a study. I attached an example that I isolated from the train.csv.  In [this example](https://www.dropbox.com/s/sv0s0x1gvvtikll/additional_series.mp4?dl=0) the additional series is NOT a PE study and includes the abdomen, pelvis and lower-chest. I'm curious if there are examples with PE identified on an image level in these studies or if these are simply ignored and treated as negative.\n\nAny clarifications here? cc: @philculliton\n\nCSV of this example is also attached.\n",
    "1036620": "any update on this . its an important thing to know ",
    "1034368": "We have to wait the host' respond I suppose"
  }
}