{
  "id": 369692,
  "title": "How to speed up converting dcm files to images when submitting？",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369692",
  "author_name": "yanqiangmiffy",
  "post_date": "2022-12-01T03:30:55.836000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<blockquote>\n  <p>Files<br>\n  [train/test]_images/[patient_id]/[image_id].dcm The mammograms, in dicom format. You can expect roughly 8,000 patients in the hidden test set. There are usually but not always 4 images per patient. Note that many of the images use the jpeg 2000 format which may you may need special libraries to load.</p>\n</blockquote>\n<p>use code：<a href=\"https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\" target=\"_blank\">https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg</a></p>\n<pre><code>_ = Parallel(n_jobs=)(\n    delayed(process)(uid, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n     uid  tqdm(train_images[:])\n)\n</code></pre>\n<p>Because the test set has nearly 32,000 files, the submission is very slow. Is there any other acceleration solution?</p>\n<p>Or save the hidden processed result to somewhere</p>",
  "messages": [
    {
      "id": 2050853,
      "postDate": "2022-12-01T03:30:55.837Z",
      "content": "<blockquote>\n  <p>Files<br>\n  [train/test]_images/[patient_id]/[image_id].dcm The mammograms, in dicom format. You can expect roughly 8,000 patients in the hidden test set. There are usually but not always 4 images per patient. Note that many of the images use the jpeg 2000 format which may you may need special libraries to load.</p>\n</blockquote>\n<p>use code：<a href=\"https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\" target=\"_blank\">https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg</a></p>\n<pre><code>_ = Parallel(n_jobs=)(\n    delayed(process)(uid, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n     uid  tqdm(train_images[:])\n)\n</code></pre>\n<p>Because the test set has nearly 32,000 files, the submission is very slow. Is there any other acceleration solution?</p>\n<p>Or save the hidden processed result to somewhere</p>",
      "rawMarkdown": "> Files\n[train/test]_images/[patient_id]/[image_id].dcm The mammograms, in dicom format. You can expect roughly 8,000 patients in the hidden test set. There are usually but not always 4 images per patient. Note that many of the images use the jpeg 2000 format which may you may need special libraries to load.\n\nuse code：https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\n\n```python\n_ = Parallel(n_jobs=4)(\n    delayed(process)(uid, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for uid in tqdm(train_images[:10])\n)\n```\n\nBecause the test set has nearly 32,000 files, the submission is very slow. Is there any other acceleration solution?\n\nOr save the hidden processed result to somewhere",
      "votes": 4
    },
    {
      "id": 2050915,
      "postDate": "2022-12-01T04:50:41.787Z",
      "content": "<p>how about:<br>\n<a href=\"https://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/\" target=\"_blank\">https://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/</a><br>\n<a href=\"https://docs.rapids.ai/api/cucim/stable/\" target=\"_blank\">https://docs.rapids.ai/api/cucim/stable/</a></p>",
      "rawMarkdown": "how about:\nhttps://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/\nhttps://docs.rapids.ai/api/cucim/stable/\n ",
      "votes": 1,
      "replies": [
        {
          "id": 2050989,
          "postDate": "2022-12-01T06:21:19.730Z",
          "content": "<p>Thanks，I will try on the kaggle platform。Another,I use Amd EPYC 7773X 64-Core Processor，it processed the trainingset images（54000+） just use 22 minutes,but it took 9hours in kaggle😂</p>",
          "rawMarkdown": "Thanks，I will try on the kaggle platform。Another,I use Amd EPYC 7773X 64-Core Processor，it processed the trainingset images（54000+） just use 22 minutes,but it took 9hours in kaggle😂"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2050915,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-01T04:50:41.787000",
      "content": "<p>how about:<br>\n<a href=\"https://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/\" target=\"_blank\">https://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/</a><br>\n<a href=\"https://docs.rapids.ai/api/cucim/stable/\" target=\"_blank\">https://docs.rapids.ai/api/cucim/stable/</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2050989,
          "author_name": "yanqiangmiffy",
          "author_url": "",
          "post_date": "2022-12-01T06:21:19.730000",
          "content": "<p>Thanks，I will try on the kaggle platform。Another,I use Amd EPYC 7773X 64-Core Processor，it processed the trainingset images（54000+） just use 22 minutes,but it took 9hours in kaggle😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2050853": "> Files\n[train/test]_images/[patient_id]/[image_id].dcm The mammograms, in dicom format. You can expect roughly 8,000 patients in the hidden test set. There are usually but not always 4 images per patient. Note that many of the images use the jpeg 2000 format which may you may need special libraries to load.\n\nuse code：https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\n\n```python\n_ = Parallel(n_jobs=4)(\n    delayed(process)(uid, size=SIZE, save_folder=SAVE_FOLDER, extension=EXTENSION)\n    for uid in tqdm(train_images[:10])\n)\n```\n\nBecause the test set has nearly 32,000 files, the submission is very slow. Is there any other acceleration solution?\n\nOr save the hidden processed result to somewhere",
    "2050915": "how about:\nhttps://developer.nvidia.com/blog/cucim-rapid-n-dimensional-image-processing-and-i-o-on-gpus/\nhttps://docs.rapids.ai/api/cucim/stable/\n "
  }
}