{
  "id": 218496,
  "title": "DataLoader for computation efficiency. Also, using Radiologist ID data?",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/218496",
  "author_name": "Daniel Hagan",
  "post_date": "2021-02-10T21:55:20.161000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First, I have created a data loader to combine all findings for an image and then separate them out by radiologist <a href=\"url\" target=\"_blank\">https://github.com/soccerguy282/Labels-Loader-for-Kaggle-Contest/blob/main/DataLoader.py</a> or you can look at it in my notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/soccerguy282/yolov5-dataloader</a>. This will save computation because the network will not have to load an image multiple times if there are more than one finding on that image. The data loader simply separates out the findings from the three radiologists, but should I include the radiologist ID as an input to the network? I don't see how this would help much but perhaps someone has tried this or something else already?</p>",
  "messages": [
    {
      "id": 1195533,
      "postDate": "2021-02-10T21:55:20.160Z",
      "content": "<p>First, I have created a data loader to combine all findings for an image and then separate them out by radiologist <a href=\"url\" target=\"_blank\">https://github.com/soccerguy282/Labels-Loader-for-Kaggle-Contest/blob/main/DataLoader.py</a> or you can look at it in my notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/soccerguy282/yolov5-dataloader</a>. This will save computation because the network will not have to load an image multiple times if there are more than one finding on that image. The data loader simply separates out the findings from the three radiologists, but should I include the radiologist ID as an input to the network? I don't see how this would help much but perhaps someone has tried this or something else already?</p>",
      "rawMarkdown": "First, I have created a data loader to combine all findings for an image and then separate them out by radiologist [https://github.com/soccerguy282/Labels-Loader-for-Kaggle-Contest/blob/main/DataLoader.py](url) or you can look at it in my notebook [https://www.kaggle.com/soccerguy282/yolov5-dataloader](url). This will save computation because the network will not have to load an image multiple times if there are more than one finding on that image. The data loader simply separates out the findings from the three radiologists, but should I include the radiologist ID as an input to the network? I don't see how this would help much but perhaps someone has tried this or something else already?",
      "votes": 6
    },
    {
      "id": 1195601,
      "postDate": "2021-02-11T00:18:59.400Z",
      "content": "<p>Regarding what to do with the multiple labels, you could aggregate the annotations (e.g. like <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">here</a>) or use the different annotations as data augmentations e.g. by using different ones in different epochs. </p>\n<p>My suspicion is that processing the label information is unlikely to be a major bottle neck in efficiency, but doing any pre-processing up-front definitely makes sense. Especially so, if you are using the limited free GPU notebook time on Kaggle. For the same reason, it's worth avoiding to read the .dicom files and to pre-process them before any model fitting. Another interesting idea is to store them in a single file together with the labels etc. That way, you can load a whole batch of images and their labels at once, which saves a good bit of time vs. opening files one by one. E.g. the <code>shelve</code> package provides a pretty convenient format for saving files and being able to access them e.g. by <code>image_id</code> or index (and parallelized access is possible). I wrote a PyTorch <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed\" target=\"_blank\">DataLoader that implements that</a>. Speed-wise that beats a PyTorch image dataloader (after converting dicom to images) by a factor of 3 or so (see <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211010\" target=\"_blank\">this previous discussion</a> on that topic). Either of the two beat reading dicom files on the fly by miles (I don't think you'd ever want to repeatedly read the .dicom files, it is rather slow).</p>",
      "rawMarkdown": "Regarding what to do with the multiple labels, you could aggregate the annotations (e.g. like [here](https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset)) or use the different annotations as data augmentations e.g. by using different ones in different epochs. \n\nMy suspicion is that processing the label information is unlikely to be a major bottle neck in efficiency, but doing any pre-processing up-front definitely makes sense. Especially so, if you are using the limited free GPU notebook time on Kaggle. For the same reason, it's worth avoiding to read the .dicom files and to pre-process them before any model fitting. Another interesting idea is to store them in a single file together with the labels etc. That way, you can load a whole batch of images and their labels at once, which saves a good bit of time vs. opening files one by one. E.g. the `shelve` package provides a pretty convenient format for saving files and being able to access them e.g. by `image_id` or index (and parallelized access is possible). I wrote a PyTorch [DataLoader that implements that](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed). Speed-wise that beats a PyTorch image dataloader (after converting dicom to images) by a factor of 3 or so (see [this previous discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211010) on that topic). Either of the two beat reading dicom files on the fly by miles (I don't think you'd ever want to repeatedly read the .dicom files, it is rather slow).",
      "votes": 1,
      "replies": [
        {
          "id": 1196716,
          "postDate": "2021-02-11T15:30:05.247Z",
          "content": "<p>Thanks for your input, I will check out your dataloader! it seems we had a similar idea to group the images and findings together so we don't have to reload the dicom file. We thought it might be better to separate them out further by radiologist so that each image has all findings from one radiologist, that way it is trained to output what one radiologist would find instead of a group of 3</p>",
          "rawMarkdown": "Thanks for your input, I will check out your dataloader! it seems we had a similar idea to group the images and findings together so we don't have to reload the dicom file. We thought it might be better to separate them out further by radiologist so that each image has all findings from one radiologist, that way it is trained to output what one radiologist would find instead of a group of 3",
          "votes": 1
        }
      ]
    },
    {
      "id": 1195576,
      "postDate": "2021-02-10T23:39:32.227Z",
      "content": "<p>Another question we were investigating is whether to feed an image with all radiologists finings combined, or to read them in as separate inputs to the network.</p>",
      "rawMarkdown": "Another question we were investigating is whether to feed an image with all radiologists finings combined, or to read them in as separate inputs to the network.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1195601,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-11T00:18:59.400000",
      "content": "<p>Regarding what to do with the multiple labels, you could aggregate the annotations (e.g. like <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">here</a>) or use the different annotations as data augmentations e.g. by using different ones in different epochs. </p>\n<p>My suspicion is that processing the label information is unlikely to be a major bottle neck in efficiency, but doing any pre-processing up-front definitely makes sense. Especially so, if you are using the limited free GPU notebook time on Kaggle. For the same reason, it's worth avoiding to read the .dicom files and to pre-process them before any model fitting. Another interesting idea is to store them in a single file together with the labels etc. That way, you can load a whole batch of images and their labels at once, which saves a good bit of time vs. opening files one by one. E.g. the <code>shelve</code> package provides a pretty convenient format for saving files and being able to access them e.g. by <code>image_id</code> or index (and parallelized access is possible). I wrote a PyTorch <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed\" target=\"_blank\">DataLoader that implements that</a>. Speed-wise that beats a PyTorch image dataloader (after converting dicom to images) by a factor of 3 or so (see <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211010\" target=\"_blank\">this previous discussion</a> on that topic). Either of the two beat reading dicom files on the fly by miles (I don't think you'd ever want to repeatedly read the .dicom files, it is rather slow).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1196716,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-02-11T15:30:05.247000",
          "content": "<p>Thanks for your input, I will check out your dataloader! it seems we had a similar idea to group the images and findings together so we don't have to reload the dicom file. We thought it might be better to separate them out further by radiologist so that each image has all findings from one radiologist, that way it is trained to output what one radiologist would find instead of a group of 3</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1195576,
      "author_name": "Daniel Hagan",
      "author_url": "",
      "post_date": "2021-02-10T23:39:32.227000",
      "content": "<p>Another question we were investigating is whether to feed an image with all radiologists finings combined, or to read them in as separate inputs to the network.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1195533": "First, I have created a data loader to combine all findings for an image and then separate them out by radiologist [https://github.com/soccerguy282/Labels-Loader-for-Kaggle-Contest/blob/main/DataLoader.py](url) or you can look at it in my notebook [https://www.kaggle.com/soccerguy282/yolov5-dataloader](url). This will save computation because the network will not have to load an image multiple times if there are more than one finding on that image. The data loader simply separates out the findings from the three radiologists, but should I include the radiologist ID as an input to the network? I don't see how this would help much but perhaps someone has tried this or something else already?",
    "1195601": "Regarding what to do with the multiple labels, you could aggregate the annotations (e.g. like [here](https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset)) or use the different annotations as data augmentations e.g. by using different ones in different epochs. \n\nMy suspicion is that processing the label information is unlikely to be a major bottle neck in efficiency, but doing any pre-processing up-front definitely makes sense. Especially so, if you are using the limited free GPU notebook time on Kaggle. For the same reason, it's worth avoiding to read the .dicom files and to pre-process them before any model fitting. Another interesting idea is to store them in a single file together with the labels etc. That way, you can load a whole batch of images and their labels at once, which saves a good bit of time vs. opening files one by one. E.g. the `shelve` package provides a pretty convenient format for saving files and being able to access them e.g. by `image_id` or index (and parallelized access is possible). I wrote a PyTorch [DataLoader that implements that](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed). Speed-wise that beats a PyTorch image dataloader (after converting dicom to images) by a factor of 3 or so (see [this previous discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211010) on that topic). Either of the two beat reading dicom files on the fly by miles (I don't think you'd ever want to repeatedly read the .dicom files, it is rather slow).",
    "1195576": "Another question we were investigating is whether to feed an image with all radiologists finings combined, or to read them in as separate inputs to the network."
  }
}