{
  "id": 384231,
  "title": "newest DALI nighly includes decoding for lossless format",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/384231",
  "author_name": "Dieter",
  "post_date": "2023-02-07T06:26:50.969000",
  "votes": 70,
  "comment_count": 19,
  "views": 0,
  "content": "<p>It was discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">here</a> and here that DALI can be used for speedy decoding of the JPEG2000 encoded dicoms. But thats only half of the dicoms. The latest release also includes decoding of the other half (JPEG-lossless). You can install via </p>\n<p><code>pip install --extra-index-url https://developer.download.nvidia.com/compute/redist/nightly nvidia-dali-nightly-cuda110</code></p>\n<p>I created a dataset with the pip wheel <a href=\"https://www.kaggle.com/datasets/christofhenkel/nvidia-dali-nightly-cuda110-1230dev/\" target=\"_blank\">nvidia_dali_nightly_cuda110_1230dev</a> and a kernel showing how to use it here <a href=\"https://www.kaggle.com/code/christofhenkel/se-resnext50-full-gpu-decoding\" target=\"_blank\">SE-ResNeXt50 full GPU decoding</a></p>\n<p>WARNING: Allthough the GPU decoding works for all train images, a few of the JPEG-lossless formated DICOMS (TransferSyntaxUID == '1.2.840.10008.1.2.4.70') of the hidden test set cannot be decoded. So its crucial to have a CPU fallback in place for the few dicoms so the notebook wont throw an exception in the submission re-run</p>",
  "messages": [
    {
      "id": 2132944,
      "postDate": "2023-02-07T06:26:50.970Z",
      "content": "<p>It was discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">here</a> and here that DALI can be used for speedy decoding of the JPEG2000 encoded dicoms. But thats only half of the dicoms. The latest release also includes decoding of the other half (JPEG-lossless). You can install via </p>\n<p><code>pip install --extra-index-url https://developer.download.nvidia.com/compute/redist/nightly nvidia-dali-nightly-cuda110</code></p>\n<p>I created a dataset with the pip wheel <a href=\"https://www.kaggle.com/datasets/christofhenkel/nvidia-dali-nightly-cuda110-1230dev/\" target=\"_blank\">nvidia_dali_nightly_cuda110_1230dev</a> and a kernel showing how to use it here <a href=\"https://www.kaggle.com/code/christofhenkel/se-resnext50-full-gpu-decoding\" target=\"_blank\">SE-ResNeXt50 full GPU decoding</a></p>\n<p>WARNING: Allthough the GPU decoding works for all train images, a few of the JPEG-lossless formated DICOMS (TransferSyntaxUID == '1.2.840.10008.1.2.4.70') of the hidden test set cannot be decoded. So its crucial to have a CPU fallback in place for the few dicoms so the notebook wont throw an exception in the submission re-run</p>",
      "rawMarkdown": "It was discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534) and here that DALI can be used for speedy decoding of the JPEG2000 encoded dicoms. But thats only half of the dicoms. The latest release also includes decoding of the other half (JPEG-lossless). You can install via \n\n`pip install --extra-index-url https://developer.download.nvidia.com/compute/redist/nightly nvidia-dali-nightly-cuda110`\n \nI created a dataset with the pip wheel [nvidia_dali_nightly_cuda110_1230dev](https://www.kaggle.com/datasets/christofhenkel/nvidia-dali-nightly-cuda110-1230dev/) and a kernel showing how to use it here [SE-ResNeXt50 full GPU decoding](https://www.kaggle.com/code/christofhenkel/se-resnext50-full-gpu-decoding)\n\nWARNING: Allthough the GPU decoding works for all train images, a few of the JPEG-lossless formated DICOMS (TransferSyntaxUID == '1.2.840.10008.1.2.4.70') of the hidden test set cannot be decoded. So its crucial to have a CPU fallback in place for the few dicoms so the notebook wont throw an exception in the submission re-run",
      "votes": 70
    },
    {
      "id": 2145373,
      "postDate": "2023-02-15T02:00:54.083Z",
      "content": "<p>I've added some additional GPU optimizations to the notebook from  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.  It achieves the same LB score and runs about 30% faster.  You can check it out <a href=\"https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I've added some additional GPU optimizations to the notebook from  @christofhenkel.  It achieves the same LB score and runs about 30% faster.  You can check it out [here](https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized)",
      "votes": 5,
      "replies": [
        {
          "id": 2145385,
          "postDate": "2023-02-15T02:29:19.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> </p>\n<p>thanks for the notebook.</p>\n<p>you have timing for decoding all test images only?<br>\nor an approximation on how how it will take? based on your results i am guessing it is around 1.8hr?</p>",
          "rawMarkdown": "@tivfrvqhs5 \n\nthanks for the notebook.\n\nyou have timing for decoding all test images only?\nor an approximation on how how it will take? based on your results i am guessing it is around 1.8hr?",
          "replies": [
            {
              "id": 2145387,
              "postDate": "2023-02-15T02:32:55.903Z",
              "content": "<p>Yes it's ~2hrs for decode and ~1.5 hours for inference (5 fold  se-resnext50).</p>",
              "rawMarkdown": "Yes it's ~2hrs for decode and ~1.5 hours for inference (5 fold  se-resnext50).",
              "votes": 1
            }
          ]
        },
        {
          "id": 2154181,
          "postDate": "2023-02-21T21:19:40.597Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>, can you share the steps to convert the <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/datasets/christofhenkel/rsna-seresnext50-5fold\" target=\"_blank\">Pytorch models</a> to the <a href=\"https://www.kaggle.com/datasets/tivfrvqhs5/dietertensorrt2\" target=\"_blank\">TensorRT models</a> you are using in your <a href=\"https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized\" target=\"_blank\">optimized notebook</a>? I am not able to load them in a GeForce GTX 1080Ti GPU, which supports TensorRT with FP32 precision according to the <a href=\"https://docs.nvidia.com/deeplearning/tensorrt/support-matrix/index.html\" target=\"_blank\">TensorRT support matrix</a> since it has CUDA compute capability 6.1, so I believe I will need to re-convert them.</p>\n<p>Thank you!</p>",
          "rawMarkdown": "Hi @tivfrvqhs5, can you share the steps to convert the @christofhenkel [Pytorch models](https://www.kaggle.com/datasets/christofhenkel/rsna-seresnext50-5fold) to the [TensorRT models](https://www.kaggle.com/datasets/tivfrvqhs5/dietertensorrt2) you are using in your [optimized notebook](https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized)? I am not able to load them in a GeForce GTX 1080Ti GPU, which supports TensorRT with FP32 precision according to the [TensorRT support matrix](https://docs.nvidia.com/deeplearning/tensorrt/support-matrix/index.html) since it has CUDA compute capability 6.1, so I believe I will need to re-convert them.\n\nThank you!",
          "votes": 2,
          "replies": [
            {
              "id": 2159618,
              "postDate": "2023-02-25T21:52:52.527Z",
              "content": "<p>The steps to convert pytorch models to nvidia torch_tensorrt are in <a href=\"https://github.com/pytorch/TensorRT/tree/main/notebooks\" target=\"_blank\">these</a> example notebooks from Pytorch or in <a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">this</a> notebook from <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>, Thanks.</p>",
              "rawMarkdown": "The steps to convert pytorch models to nvidia torch_tensorrt are in [these](https://github.com/pytorch/TensorRT/tree/main/notebooks) example notebooks from Pytorch or in [this](https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook) notebook from @tivfrvqhs5, Thanks."
            }
          ]
        }
      ]
    },
    {
      "id": 2132963,
      "postDate": "2023-02-07T06:38:58.480Z",
      "content": "<p>waht is the total speed for decoding all test images with dali?<br>\nis it winthin 2hrs?</p>",
      "rawMarkdown": "waht is the total speed for decoding all test images with dali?\nis it winthin 2hrs?",
      "votes": 3
    },
    {
      "id": 2132956,
      "postDate": "2023-02-07T06:37:11.327Z",
      "content": "<p>oh, they hear our voices!👍</p>",
      "rawMarkdown": "oh, they hear our voices!👍",
      "votes": 1
    },
    {
      "id": 2142295,
      "postDate": "2023-02-13T12:36:36.087Z",
      "content": "<p>Your instructions for installing the latest version of DALI and the link to the kernel demonstrating its usage will be helpful for those who are looking to use DALI for their projects. It's also important to note the warning about the hidden test set and the need to have a CPU fallback in place to avoid exceptions during the submission re-run. Thanks for sharing <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 🔥🔥</p>",
      "rawMarkdown": "Your instructions for installing the latest version of DALI and the link to the kernel demonstrating its usage will be helpful for those who are looking to use DALI for their projects. It's also important to note the warning about the hidden test set and the need to have a CPU fallback in place to avoid exceptions during the submission re-run. Thanks for sharing @christofhenkel 🔥🔥",
      "votes": -1
    },
    {
      "id": 2143620,
      "postDate": "2023-02-14T12:14:15.620Z",
      "content": "<p>oh, they hear our voices!👍</p>",
      "rawMarkdown": "oh, they hear our voices!👍"
    },
    {
      "id": 2142605,
      "postDate": "2023-02-13T16:41:14.907Z",
      "content": "<p>Your amazing share makes me improve my understanding of this field. Thanks.</p>",
      "rawMarkdown": "Your amazing share makes me improve my understanding of this field. Thanks."
    },
    {
      "id": 2142380,
      "postDate": "2023-02-13T13:41:44.697Z",
      "content": "<p>Hey looks likes somebody has copy pasted your code everything and has got the same score like you… <a href=\"https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding\" target=\"_blank\">https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding</a></p>",
      "rawMarkdown": "Hey looks likes somebody has copy pasted your code everything and has got the same score like you... https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding"
    },
    {
      "id": 2135233,
      "postDate": "2023-02-08T14:19:36.427Z",
      "content": "<p>That's awesome, thanks for sharing! Regarding the fallback, I guess it isn't implemented in the notebook you have shared? </p>",
      "rawMarkdown": "That's awesome, thanks for sharing! Regarding the fallback, I guess it isn't implemented in the notebook you have shared? ",
      "replies": [
        {
          "id": 2135327,
          "postDate": "2023-02-08T15:05:28.363Z",
          "content": "<p>cell 13 - 15</p>",
          "rawMarkdown": "cell 13 - 15",
          "votes": 1,
          "replies": [
            {
              "id": 2135340,
              "postDate": "2023-02-08T15:13:44.803Z",
              "content": "<p>Thanks for the quick answer! 🙏</p>",
              "rawMarkdown": "Thanks for the quick answer! 🙏"
            }
          ]
        }
      ]
    },
    {
      "id": 2133134,
      "postDate": "2023-02-07T09:07:04.450Z",
      "content": "<p>Thanks! What's the gain in CV and LB vs the previous processing? Good score for a seresnext architecture!</p>",
      "rawMarkdown": "Thanks! What's the gain in CV and LB vs the previous processing? Good score for a seresnext architecture!"
    },
    {
      "id": 2152790,
      "postDate": "2023-02-21T02:59:53.650Z",
      "content": "<p>Learn a lot,thanks!</p>",
      "rawMarkdown": "Learn a lot,thanks!"
    },
    {
      "id": 2152275,
      "postDate": "2023-02-20T17:34:13.293Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 2134987,
      "postDate": "2023-02-08T11:31:37.003Z",
      "content": "<p>Thanks for update!</p>",
      "rawMarkdown": "Thanks for update!"
    },
    {
      "id": 2133810,
      "postDate": "2023-02-07T16:03:17.520Z",
      "content": "<p>Thank you for sharing! 👍</p>",
      "rawMarkdown": "Thank you for sharing! 👍"
    }
  ],
  "comments": [
    {
      "id": 2145373,
      "author_name": "David Austin",
      "author_url": "",
      "post_date": "2023-02-15T02:00:54.083000",
      "content": "<p>I've added some additional GPU optimizations to the notebook from  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.  It achieves the same LB score and runs about 30% faster.  You can check it out <a href=\"https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized\" target=\"_blank\">here</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2145385,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-02-15T02:29:19.543000",
          "content": "<p><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> </p>\n<p>thanks for the notebook.</p>\n<p>you have timing for decoding all test images only?<br>\nor an approximation on how how it will take? based on your results i am guessing it is around 1.8hr?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2145387,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-02-15T02:32:55.903000",
              "content": "<p>Yes it's ~2hrs for decode and ~1.5 hours for inference (5 fold  se-resnext50).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2154181,
          "author_name": "Pablo Rios",
          "author_url": "",
          "post_date": "2023-02-21T21:19:40.597000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>, can you share the steps to convert the <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/datasets/christofhenkel/rsna-seresnext50-5fold\" target=\"_blank\">Pytorch models</a> to the <a href=\"https://www.kaggle.com/datasets/tivfrvqhs5/dietertensorrt2\" target=\"_blank\">TensorRT models</a> you are using in your <a href=\"https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized\" target=\"_blank\">optimized notebook</a>? I am not able to load them in a GeForce GTX 1080Ti GPU, which supports TensorRT with FP32 precision according to the <a href=\"https://docs.nvidia.com/deeplearning/tensorrt/support-matrix/index.html\" target=\"_blank\">TensorRT support matrix</a> since it has CUDA compute capability 6.1, so I believe I will need to re-convert them.</p>\n<p>Thank you!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2159618,
              "author_name": "Pablo Rios",
              "author_url": "",
              "post_date": "2023-02-25T21:52:52.527000",
              "content": "<p>The steps to convert pytorch models to nvidia torch_tensorrt are in <a href=\"https://github.com/pytorch/TensorRT/tree/main/notebooks\" target=\"_blank\">these</a> example notebooks from Pytorch or in <a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">this</a> notebook from <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>, Thanks.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2132963,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-07T06:38:58.480000",
      "content": "<p>waht is the total speed for decoding all test images with dali?<br>\nis it winthin 2hrs?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2132956,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-07T06:37:11.327000",
      "content": "<p>oh, they hear our voices!👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2142295,
      "author_name": "Ryan",
      "author_url": "",
      "post_date": "2023-02-13T12:36:36.087000",
      "content": "<p>Your instructions for installing the latest version of DALI and the link to the kernel demonstrating its usage will be helpful for those who are looking to use DALI for their projects. It's also important to note the warning about the hidden test set and the need to have a CPU fallback in place to avoid exceptions during the submission re-run. Thanks for sharing <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 🔥🔥</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2143620,
      "author_name": "liang zhu111",
      "author_url": "",
      "post_date": "2023-02-14T12:14:15.620000",
      "content": "<p>oh, they hear our voices!👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2142605,
      "author_name": "Hyunsoo Lee 1010",
      "author_url": "",
      "post_date": "2023-02-13T16:41:14.907000",
      "content": "<p>Your amazing share makes me improve my understanding of this field. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2142380,
      "author_name": "Denis Muriungi",
      "author_url": "",
      "post_date": "2023-02-13T13:41:44.697000",
      "content": "<p>Hey looks likes somebody has copy pasted your code everything and has got the same score like you… <a href=\"https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding\" target=\"_blank\">https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135233,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-02-08T14:19:36.427000",
      "content": "<p>That's awesome, thanks for sharing! Regarding the fallback, I guess it isn't implemented in the notebook you have shared? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2135327,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2023-02-08T15:05:28.363000",
          "content": "<p>cell 13 - 15</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2135340,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-02-08T15:13:44.803000",
              "content": "<p>Thanks for the quick answer! 🙏</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2133134,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2023-02-07T09:07:04.450000",
      "content": "<p>Thanks! What's the gain in CV and LB vs the previous processing? Good score for a seresnext architecture!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2152790,
      "author_name": "Andy",
      "author_url": "",
      "post_date": "2023-02-21T02:59:53.650000",
      "content": "<p>Learn a lot,thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2152275,
      "author_name": "*1.01",
      "author_url": "",
      "post_date": "2023-02-20T17:34:13.293000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134987,
      "author_name": "ToriGlori",
      "author_url": "",
      "post_date": "2023-02-08T11:31:37.003000",
      "content": "<p>Thanks for update!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2133810,
      "author_name": "Dheeraj Pandey",
      "author_url": "",
      "post_date": "2023-02-07T16:03:17.520000",
      "content": "<p>Thank you for sharing! 👍</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2132944": "It was discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534) and here that DALI can be used for speedy decoding of the JPEG2000 encoded dicoms. But thats only half of the dicoms. The latest release also includes decoding of the other half (JPEG-lossless). You can install via \n\n`pip install --extra-index-url https://developer.download.nvidia.com/compute/redist/nightly nvidia-dali-nightly-cuda110`\n \nI created a dataset with the pip wheel [nvidia_dali_nightly_cuda110_1230dev](https://www.kaggle.com/datasets/christofhenkel/nvidia-dali-nightly-cuda110-1230dev/) and a kernel showing how to use it here [SE-ResNeXt50 full GPU decoding](https://www.kaggle.com/code/christofhenkel/se-resnext50-full-gpu-decoding)\n\nWARNING: Allthough the GPU decoding works for all train images, a few of the JPEG-lossless formated DICOMS (TransferSyntaxUID == '1.2.840.10008.1.2.4.70') of the hidden test set cannot be decoded. So its crucial to have a CPU fallback in place for the few dicoms so the notebook wont throw an exception in the submission re-run",
    "2145373": "I've added some additional GPU optimizations to the notebook from  @christofhenkel.  It achieves the same LB score and runs about 30% faster.  You can check it out [here](https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized)",
    "2132963": "waht is the total speed for decoding all test images with dali?\nis it winthin 2hrs?",
    "2132956": "oh, they hear our voices!👍",
    "2142295": "Your instructions for installing the latest version of DALI and the link to the kernel demonstrating its usage will be helpful for those who are looking to use DALI for their projects. It's also important to note the warning about the hidden test set and the need to have a CPU fallback in place to avoid exceptions during the submission re-run. Thanks for sharing @christofhenkel 🔥🔥",
    "2143620": "oh, they hear our voices!👍",
    "2142605": "Your amazing share makes me improve my understanding of this field. Thanks.",
    "2142380": "Hey looks likes somebody has copy pasted your code everything and has got the same score like you... https://www.kaggle.com/code/ashokkumargarain/se-resnext50-full-gpu-decoding",
    "2135233": "That's awesome, thanks for sharing! Regarding the fallback, I guess it isn't implemented in the notebook you have shared? ",
    "2133134": "Thanks! What's the gain in CV and LB vs the previous processing? Good score for a seresnext architecture!",
    "2152790": "Learn a lot,thanks!",
    "2152275": "Thank you for sharing!",
    "2134987": "Thanks for update!",
    "2133810": "Thank you for sharing! 👍"
  }
}