{
  "id": 390785,
  "title": "Scoring taking a long time",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/390785",
  "author_name": "gsharpminor",
  "post_date": "2023-02-27T08:51:45.097000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've gotten my code running successfully, and have made two submissions. The code itself runs through without error. However, after that it appears to be stuck in \"scoring\" limbo for the last 5-6 hours:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F930277%2F0e1e167508c764cf38ac093e747a862e%2FScreenshot%20from%202023-02-27%2001-49-35.png?generation=1677487862034282&amp;alt=media\" alt=\"\"></p>\n<p>Anyone else having this happen? Is it just normal for scoring the results to take a really, really long time?</p>",
  "messages": [
    {
      "id": 2161635,
      "postDate": "2023-02-27T16:13:05.740Z",
      "content": "<p>This happened to me as well, if you are using the regular pyDICOM package in python. Mine first several submission failed after it was running for too long.</p>\n<p>After I switch to DALI package, the process will take 4-6 hours to finish. </p>\n<p>Take a look at <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a> this is proposal to use <a href=\"https://docs.nvidia.com/deeplearning/dali/user-guide/docs\" target=\"_blank\">https://docs.nvidia.com/deeplearning/dali/user-guide/docs</a> which is faster then pydicom based pipeline you've tried. </p>",
      "rawMarkdown": "This happened to me as well, if you are using the regular pyDICOM package in python. Mine first several submission failed after it was running for too long.\n\nAfter I switch to DALI package, the process will take 4-6 hours to finish. \n\nTake a look at https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali this is proposal to use https://docs.nvidia.com/deeplearning/dali/user-guide/docs which is faster then pydicom based pipeline you've tried. ",
      "replies": [
        {
          "id": 2161775,
          "postDate": "2023-02-27T18:10:26.080Z",
          "rawMarkdown": "",
          "votes": -4,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2161155,
      "postDate": "2023-02-27T09:28:42.433Z",
      "content": "<p>For me it takes 6-7 hours to finish the pre-processing and inference of the hidden evaluation data--I am reading the DICOM while I am inferring one examination at a time. </p>",
      "rawMarkdown": "For me it takes 6-7 hours to finish the pre-processing and inference of the hidden evaluation data--I am reading the DICOM while I am inferring one examination at a time. ",
      "replies": [
        {
          "id": 2161769,
          "postDate": "2023-02-27T18:05:47.843Z",
          "content": "<p>Yes so I just woke up to find that both my submissions timed out with no score. So with only 6 hours left I learn I still have no submissions. </p>\n<p>The pipeline runs without error and yields a classifier that (on a validation set) makes predictions well above-chance. And I cannot even get Kaggle to score the thing. Because reasons.</p>",
          "rawMarkdown": "Yes so I just woke up to find that both my submissions timed out with no score. So with only 6 hours left I learn I still have no submissions. \n\nThe pipeline runs without error and yields a classifier that (on a validation set) makes predictions well above-chance. And I cannot even get Kaggle to score the thing. Because reasons.",
          "replies": [
            {
              "id": 2161794,
              "postDate": "2023-02-27T18:32:43.227Z",
              "content": "<p>Are you able to put try-except structures (if you did not already) here and there where you assume some error might occur?</p>\n<p>I initialized the whole submission.csv prior to starting to go over the images so that I would have something to submit at the end. Like so:</p>\n<pre><code>csv_file = Path()\ntest_meta = pd.read_csv(csv_file)\ntest_meta[] = test_meta[].astype() +  + test_meta[].astype()\n</code></pre>\n<pre><code>submission = test_meta[[,]].copy()\nsubmission = submission.drop_duplicates(subset=[], keep=, ignore_index=)\n</code></pre>\n<pre><code>pred = \nsubmission.insert(, , pred)  \n</code></pre>\n<p>And then I update the <code>submission</code> DataFrame:</p>\n<pre><code>submission.loc[submission.prediction_id==curr_prediction_id_right,] = curr_pred_right\nsubmission.loc[submission.prediction_id==curr_prediction_id_left,] = curr_pred_left\n</code></pre>\n<p>Hope it helps!</p>",
              "rawMarkdown": "Are you able to put try-except structures (if you did not already) here and there where you assume some error might occur?\n\nI initialized the whole submission.csv prior to starting to go over the images so that I would have something to submit at the end. Like so:\n\n```python\ncsv_file = Path('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest_meta = pd.read_csv(csv_file)\ntest_meta['prediction_id'] = test_meta['patient_id'].astype(str) + \"_\" + test_meta['laterality'].astype(str)\n```\n\n```python\nsubmission = test_meta[['prediction_id',]].copy()\nsubmission = submission.drop_duplicates(subset=['prediction_id'], keep='first', ignore_index=True)\n```\n\n```python\npred = 0.0\nsubmission.insert(0, 'cancer', pred)  # this should be in a separate cell\n```\n\nAnd then I update the `submission` DataFrame:\n\n```python\nsubmission.loc[submission.prediction_id==curr_prediction_id_right,'cancer'] = curr_pred_right\nsubmission.loc[submission.prediction_id==curr_prediction_id_left,'cancer'] = curr_pred_left\n```\nHope it helps!"
            }
          ]
        }
      ]
    },
    {
      "id": 2161139,
      "postDate": "2023-02-27T08:51:45.097Z",
      "content": "<p>I've gotten my code running successfully, and have made two submissions. The code itself runs through without error. However, after that it appears to be stuck in \"scoring\" limbo for the last 5-6 hours:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F930277%2F0e1e167508c764cf38ac093e747a862e%2FScreenshot%20from%202023-02-27%2001-49-35.png?generation=1677487862034282&amp;alt=media\" alt=\"\"></p>\n<p>Anyone else having this happen? Is it just normal for scoring the results to take a really, really long time?</p>",
      "rawMarkdown": "I've gotten my code running successfully, and have made two submissions. The code itself runs through without error. However, after that it appears to be stuck in \"scoring\" limbo for the last 5-6 hours:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F930277%2F0e1e167508c764cf38ac093e747a862e%2FScreenshot%20from%202023-02-27%2001-49-35.png?generation=1677487862034282&alt=media)\n\nAnyone else having this happen? Is it just normal for scoring the results to take a really, really long time?\n"
    }
  ],
  "comments": [
    {
      "id": 2161635,
      "author_name": "Xiao-Su (Frank) Hu",
      "author_url": "",
      "post_date": "2023-02-27T16:13:05.740000",
      "content": "<p>This happened to me as well, if you are using the regular pyDICOM package in python. Mine first several submission failed after it was running for too long.</p>\n<p>After I switch to DALI package, the process will take 4-6 hours to finish. </p>\n<p>Take a look at <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a> this is proposal to use <a href=\"https://docs.nvidia.com/deeplearning/dali/user-guide/docs\" target=\"_blank\">https://docs.nvidia.com/deeplearning/dali/user-guide/docs</a> which is faster then pydicom based pipeline you've tried. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2161775,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-02-27T18:10:26.080000",
          "content": "",
          "votes": -4,
          "replies": []
        }
      ]
    },
    {
      "id": 2161155,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-27T09:28:42.433000",
      "content": "<p>For me it takes 6-7 hours to finish the pre-processing and inference of the hidden evaluation data--I am reading the DICOM while I am inferring one examination at a time. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2161769,
          "author_name": "gsharpminor",
          "author_url": "",
          "post_date": "2023-02-27T18:05:47.843000",
          "content": "<p>Yes so I just woke up to find that both my submissions timed out with no score. So with only 6 hours left I learn I still have no submissions. </p>\n<p>The pipeline runs without error and yields a classifier that (on a validation set) makes predictions well above-chance. And I cannot even get Kaggle to score the thing. Because reasons.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2161794,
              "author_name": "Antti Isosalo",
              "author_url": "",
              "post_date": "2023-02-27T18:32:43.227000",
              "content": "<p>Are you able to put try-except structures (if you did not already) here and there where you assume some error might occur?</p>\n<p>I initialized the whole submission.csv prior to starting to go over the images so that I would have something to submit at the end. Like so:</p>\n<pre><code>csv_file = Path()\ntest_meta = pd.read_csv(csv_file)\ntest_meta[] = test_meta[].astype() +  + test_meta[].astype()\n</code></pre>\n<pre><code>submission = test_meta[[,]].copy()\nsubmission = submission.drop_duplicates(subset=[], keep=, ignore_index=)\n</code></pre>\n<pre><code>pred = \nsubmission.insert(, , pred)  \n</code></pre>\n<p>And then I update the <code>submission</code> DataFrame:</p>\n<pre><code>submission.loc[submission.prediction_id==curr_prediction_id_right,] = curr_pred_right\nsubmission.loc[submission.prediction_id==curr_prediction_id_left,] = curr_pred_left\n</code></pre>\n<p>Hope it helps!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2161635": "This happened to me as well, if you are using the regular pyDICOM package in python. Mine first several submission failed after it was running for too long.\n\nAfter I switch to DALI package, the process will take 4-6 hours to finish. \n\nTake a look at https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali this is proposal to use https://docs.nvidia.com/deeplearning/dali/user-guide/docs which is faster then pydicom based pipeline you've tried. ",
    "2161155": "For me it takes 6-7 hours to finish the pre-processing and inference of the hidden evaluation data--I am reading the DICOM while I am inferring one examination at a time. ",
    "2161139": "I've gotten my code running successfully, and have made two submissions. The code itself runs through without error. However, after that it appears to be stuck in \"scoring\" limbo for the last 5-6 hours:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F930277%2F0e1e167508c764cf38ac093e747a862e%2FScreenshot%20from%202023-02-27%2001-49-35.png?generation=1677487862034282&alt=media)\n\nAnyone else having this happen? Is it just normal for scoring the results to take a really, really long time?\n"
  }
}