{
  "id": 208478,
  "title": "Faster Data loading with numpy files",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/208478",
  "author_name": "black_raven",
  "post_date": "2021-01-03T16:47:33.848000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I noticed that a data loader that implements the DICOM reading pipeline by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> is quite slow. Therefore I created a dataset with numpy files corresponding to the 18k DICOM files in the competition dataset. The images have been downscaled to 256X256. </p>\n<p>The dataset can be found here: <a href=\"https://www.kaggle.com/bibhash123/xraynumpy\" target=\"_blank\">https://www.kaggle.com/bibhash123/xraynumpy</a>. </p>\n<p>An end to end pipeline with this dataset can be found here <a href=\"https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras\" target=\"_blank\">https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras</a></p>",
  "messages": [
    {
      "id": 1137071,
      "postDate": "2021-01-03T16:47:33.850Z",
      "content": "<p>I noticed that a data loader that implements the DICOM reading pipeline by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> is quite slow. Therefore I created a dataset with numpy files corresponding to the 18k DICOM files in the competition dataset. The images have been downscaled to 256X256. </p>\n<p>The dataset can be found here: <a href=\"https://www.kaggle.com/bibhash123/xraynumpy\" target=\"_blank\">https://www.kaggle.com/bibhash123/xraynumpy</a>. </p>\n<p>An end to end pipeline with this dataset can be found here <a href=\"https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras\" target=\"_blank\">https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras</a></p>",
      "rawMarkdown": "I noticed that a data loader that implements the DICOM reading pipeline by @raddar is quite slow. Therefore I created a dataset with numpy files corresponding to the 18k DICOM files in the competition dataset. The images have been downscaled to 256X256. \n\nThe dataset can be found here: [https://www.kaggle.com/bibhash123/xraynumpy](https://www.kaggle.com/bibhash123/xraynumpy). \n\nAn end to end pipeline with this dataset can be found here [https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras](https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras)",
      "votes": 4
    },
    {
      "id": 1193067,
      "postDate": "2021-02-09T12:44:04.580Z",
      "content": "<p>Thnx man, i was searching for this.</p>",
      "rawMarkdown": "Thnx man, i was searching for this."
    },
    {
      "id": 1192132,
      "postDate": "2021-02-09T01:48:15.323Z",
      "content": "<p>Wow! This runs a lot faster now that I'm reading smaller npy files directly instead of the larger pixel arrays in dicom format. Before it was estimated around 5 hrs per iteration through the dataset compared to around 7 min with this new format (though I might have been doing something wrong).</p>\n<p>I was looking at this dataset you made and was wondering how exactly the downscaled images were generated? The image arrays seem to be quite different from the original and although I was able to get similar values when using scikit-learn's rescaling and downsampling library, I wasn't able to get the exact arrays back.</p>",
      "rawMarkdown": "Wow! This runs a lot faster now that I'm reading smaller npy files directly instead of the larger pixel arrays in dicom format. Before it was estimated around 5 hrs per iteration through the dataset compared to around 7 min with this new format (though I might have been doing something wrong).\n\nI was looking at this dataset you made and was wondering how exactly the downscaled images were generated? The image arrays seem to be quite different from the original and although I was able to get similar values when using scikit-learn's rescaling and downsampling library, I wasn't able to get the exact arrays back.",
      "replies": [
        {
          "id": 1192604,
          "postDate": "2021-02-09T07:57:16.297Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/shreyasagarwal\" target=\"_blank\">@shreyasagarwal</a>. I have made the dataset preparation notebook public for your reference. You can find it in this link: <a href=\"https://www.kaggle.com/bibhash123/dataset-preparation\" target=\"_blank\">https://www.kaggle.com/bibhash123/dataset-preparation</a>.<br>\n Please let me know if you find any error. You'll notice that the images have been down-sampled to 256X256. I chose this size arbitrarily and it is  possible that some other size might give better results.</p>",
          "rawMarkdown": "Hi @shreyasagarwal. I have made the dataset preparation notebook public for your reference. You can find it in this link: [https://www.kaggle.com/bibhash123/dataset-preparation](https://www.kaggle.com/bibhash123/dataset-preparation).\n Please let me know if you find any error. You'll notice that the images have been down-sampled to 256X256. I chose this size arbitrarily and it is  possible that some other size might give better results."
        }
      ]
    },
    {
      "id": 1138187,
      "postDate": "2021-01-04T13:29:43.490Z",
      "content": "<p>Though downscaling seems to be the only way to do so, but what should be a safe downscaling factor? Nice work <a href=\"https://www.kaggle.com/bibhash123\" target=\"_blank\">@bibhash123</a> </p>",
      "rawMarkdown": "Though downscaling seems to be the only way to do so, but what should be a safe downscaling factor? Nice work @bibhash123 "
    }
  ],
  "comments": [
    {
      "id": 1193067,
      "author_name": "Dr.Bilal",
      "author_url": "",
      "post_date": "2021-02-09T12:44:04.580000",
      "content": "<p>Thnx man, i was searching for this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1192132,
      "author_name": "Shreyas Agarwal",
      "author_url": "",
      "post_date": "2021-02-09T01:48:15.323000",
      "content": "<p>Wow! This runs a lot faster now that I'm reading smaller npy files directly instead of the larger pixel arrays in dicom format. Before it was estimated around 5 hrs per iteration through the dataset compared to around 7 min with this new format (though I might have been doing something wrong).</p>\n<p>I was looking at this dataset you made and was wondering how exactly the downscaled images were generated? The image arrays seem to be quite different from the original and although I was able to get similar values when using scikit-learn's rescaling and downsampling library, I wasn't able to get the exact arrays back.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1192604,
          "author_name": "black_raven",
          "author_url": "",
          "post_date": "2021-02-09T07:57:16.297000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/shreyasagarwal\" target=\"_blank\">@shreyasagarwal</a>. I have made the dataset preparation notebook public for your reference. You can find it in this link: <a href=\"https://www.kaggle.com/bibhash123/dataset-preparation\" target=\"_blank\">https://www.kaggle.com/bibhash123/dataset-preparation</a>.<br>\n Please let me know if you find any error. You'll notice that the images have been down-sampled to 256X256. I chose this size arbitrarily and it is  possible that some other size might give better results.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1138187,
      "author_name": "Ultron",
      "author_url": "",
      "post_date": "2021-01-04T13:29:43.490000",
      "content": "<p>Though downscaling seems to be the only way to do so, but what should be a safe downscaling factor? Nice work <a href=\"https://www.kaggle.com/bibhash123\" target=\"_blank\">@bibhash123</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1137071": "I noticed that a data loader that implements the DICOM reading pipeline by @raddar is quite slow. Therefore I created a dataset with numpy files corresponding to the 18k DICOM files in the competition dataset. The images have been downscaled to 256X256. \n\nThe dataset can be found here: [https://www.kaggle.com/bibhash123/xraynumpy](https://www.kaggle.com/bibhash123/xraynumpy). \n\nAn end to end pipeline with this dataset can be found here [https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras](https://www.kaggle.com/bibhash123/chest-x-ray-abnormalities-baseline-tf-keras)",
    "1193067": "Thnx man, i was searching for this.",
    "1192132": "Wow! This runs a lot faster now that I'm reading smaller npy files directly instead of the larger pixel arrays in dicom format. Before it was estimated around 5 hrs per iteration through the dataset compared to around 7 min with this new format (though I might have been doing something wrong).\n\nI was looking at this dataset you made and was wondering how exactly the downscaled images were generated? The image arrays seem to be quite different from the original and although I was able to get similar values when using scikit-learn's rescaling and downsampling library, I wasn't able to get the exact arrays back.",
    "1138187": "Though downscaling seems to be the only way to do so, but what should be a safe downscaling factor? Nice work @bibhash123 "
  }
}