{
  "id": 422566,
  "title": "Random submission scoring error after inference loop?",
  "url": "/competitions/asl-fingerspelling/discussion/422566",
  "author_name": "gezi",
  "post_date": "2023-07-10T14:57:12.138000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have two models (modifed from <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference)\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference)</a>. <br>\nThe only difference is one trained 30 epochs, and another trained 100 epochs.<br>\nI submitted them at similar time, and all after 4 hours online infer run(I would assume they all just finished inference loop), the one with 100 epochs trained finish and get score but the one with 30 epochs trained failed with Submission Scoring Error.     </p>\n<p>They are using same training pipeline and all tflite models passed local and kaggle notebook local test.   <br>\nWhat might be the root cause to have submission scoring error after infering loop? <br>\nHow could I debug or reproduce the problem locally?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F96c65173c52953b3c9cd8d866d5549f2%2Fxx.png?generation=1689003848968283&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2338212,
      "postDate": "2023-07-10T16:04:33Z",
      "content": "<ol>\n<li>I think there is some randomness in scoring, particularly when the model approaches the evaluation limit, based on my experience.</li>\n<li>You're using a transformer model, which stops decoding after predicting an end token. Thus,  a bad model requires longer inference time if it makes longer predictions. This can also be the reason because your 30-epoch model exceeded the time limit, while the 100-epoch model passed.</li>\n</ol>",
      "rawMarkdown": "1. I think there is some randomness in scoring, particularly when the model approaches the evaluation limit, based on my experience.\n2. You're using a transformer model, which stops decoding after predicting an end token. Thus,  a bad model requires longer inference time if it makes longer predictions. This can also be the reason because your 30-epoch model exceeded the time limit, while the 100-epoch model passed.",
      "votes": 3,
      "replies": [
        {
          "id": 2339657,
          "postDate": "2023-07-11T01:57:04.827Z",
          "content": "<p>Thanks Yu. Actually they all run 3.5hours, I do not think this is time limit issue. I have models running infer 6 hours which got finanl score. So here they have all finished inference loop, but fail at last a bit strange, did they do edit distance at last not during inference loop? Anyway good thing is fail rate is not high I only faced scoring error 2 times  so far.</p>",
          "rawMarkdown": "Thanks Yu. Actually they all run 3.5hours, I do not think this is time limit issue. I have models running infer 6 hours which got finanl score. So here they have all finished inference loop, but fail at last a bit strange, did they do edit distance at last not during inference loop? Anyway good thing is fail rate is not high I only faced scoring error 2 times  so far.",
          "votes": 2,
          "replies": [
            {
              "id": 2339691,
              "postDate": "2023-07-11T03:09:44.687Z",
              "content": "<p>This seems strange…🤯</p>",
              "rawMarkdown": "This seems strange...🤯"
            }
          ]
        },
        {
          "id": 2340312,
          "postDate": "2023-07-11T12:13:28.380Z",
          "content": "<p>Point 2 might be the cause of the problem.<br>\nFor each prediction the encoder is called once, which is a constant computational cost.<br>\nThe decoder however, is called <code>n+1</code> times where <code>n</code> is the number of predicted characters (the <code>+1</code> being caused by the End Of Sentence token)<br>\nIf you have a worse model which predicts longer incorrect strings the inference process will take longer.</p>",
          "rawMarkdown": "Point 2 might be the cause of the problem.\nFor each prediction the encoder is called once, which is a constant computational cost.\nThe decoder however, is called `n+1` times where `n` is the number of predicted characters (the `+1` being caused by the End Of Sentence token)\nIf you have a worse model which predicts longer incorrect strings the inference process will take longer.",
          "votes": 2,
          "replies": [
            {
              "id": 2340576,
              "postDate": "2023-07-11T15:02:19.450Z",
              "content": "<p>This is reasonable but still I think here it fail just after 3.5 hours will not due to time limit.</p>",
              "rawMarkdown": "This is reasonable but still I think here it fail just after 3.5 hours will not due to time limit.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2338136,
      "postDate": "2023-07-10T14:57:12.140Z",
      "content": "<p>I have two models (modifed from <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference)\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference)</a>. <br>\nThe only difference is one trained 30 epochs, and another trained 100 epochs.<br>\nI submitted them at similar time, and all after 4 hours online infer run(I would assume they all just finished inference loop), the one with 100 epochs trained finish and get score but the one with 30 epochs trained failed with Submission Scoring Error.     </p>\n<p>They are using same training pipeline and all tflite models passed local and kaggle notebook local test.   <br>\nWhat might be the root cause to have submission scoring error after infering loop? <br>\nHow could I debug or reproduce the problem locally?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F96c65173c52953b3c9cd8d866d5549f2%2Fxx.png?generation=1689003848968283&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have two models (modifed from https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference). \nThe only difference is one trained 30 epochs, and another trained 100 epochs.\nI submitted them at similar time, and all after 4 hours online infer run(I would assume they all just finished inference loop), the one with 100 epochs trained finish and get score but the one with 30 epochs trained failed with Submission Scoring Error.     \n\nThey are using same training pipeline and all tflite models passed local and kaggle notebook local test.   \nWhat might be the root cause to have submission scoring error after infering loop? \nHow could I debug or reproduce the problem locally?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F96c65173c52953b3c9cd8d866d5549f2%2Fxx.png?generation=1689003848968283&alt=media)\n",
      "votes": 1
    },
    {
      "id": 2344764,
      "postDate": "2023-07-14T18:43:12.790Z",
      "content": "<p>Oh, I only see this now, I also made a similar post (<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503)\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503)</a>. For me, it only ran 2 hours and resubmitting solved the issue, but this seems a bit strange. Did you resolve it?</p>",
      "rawMarkdown": "Oh, I only see this now, I also made a similar post (https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503). For me, it only ran 2 hours and resubmitting solved the issue, but this seems a bit strange. Did you resolve it?",
      "replies": [
        {
          "id": 2344815,
          "postDate": "2023-07-14T19:30:38.347Z",
          "content": "<p>No, it might happen again, but as the chance is small, I would not investigate currently.</p>",
          "rawMarkdown": "No, it might happen again, but as the chance is small, I would not investigate currently.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2339540,
      "postDate": "2023-07-10T21:25:48.317Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2338212,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-07-10T16:04:33",
      "content": "<ol>\n<li>I think there is some randomness in scoring, particularly when the model approaches the evaluation limit, based on my experience.</li>\n<li>You're using a transformer model, which stops decoding after predicting an end token. Thus,  a bad model requires longer inference time if it makes longer predictions. This can also be the reason because your 30-epoch model exceeded the time limit, while the 100-epoch model passed.</li>\n</ol>",
      "votes": 3,
      "replies": [
        {
          "id": 2339657,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-07-11T01:57:04.827000",
          "content": "<p>Thanks Yu. Actually they all run 3.5hours, I do not think this is time limit issue. I have models running infer 6 hours which got finanl score. So here they have all finished inference loop, but fail at last a bit strange, did they do edit distance at last not during inference loop? Anyway good thing is fail rate is not high I only faced scoring error 2 times  so far.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2339691,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-07-11T03:09:44.687000",
              "content": "<p>This seems strange…🤯</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2340312,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-07-11T12:13:28.380000",
          "content": "<p>Point 2 might be the cause of the problem.<br>\nFor each prediction the encoder is called once, which is a constant computational cost.<br>\nThe decoder however, is called <code>n+1</code> times where <code>n</code> is the number of predicted characters (the <code>+1</code> being caused by the End Of Sentence token)<br>\nIf you have a worse model which predicts longer incorrect strings the inference process will take longer.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2340576,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-07-11T15:02:19.450000",
              "content": "<p>This is reasonable but still I think here it fail just after 3.5 hours will not due to time limit.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2344764,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2023-07-14T18:43:12.790000",
      "content": "<p>Oh, I only see this now, I also made a similar post (<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503)\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503)</a>. For me, it only ran 2 hours and resubmitting solved the issue, but this seems a bit strange. Did you resolve it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2344815,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-07-14T19:30:38.347000",
          "content": "<p>No, it might happen again, but as the chance is small, I would not investigate currently.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2339540,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-10T21:25:48.317000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2338212": "1. I think there is some randomness in scoring, particularly when the model approaches the evaluation limit, based on my experience.\n2. You're using a transformer model, which stops decoding after predicting an end token. Thus,  a bad model requires longer inference time if it makes longer predictions. This can also be the reason because your 30-epoch model exceeded the time limit, while the 100-epoch model passed.",
    "2338136": "I have two models (modifed from https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference). \nThe only difference is one trained 30 epochs, and another trained 100 epochs.\nI submitted them at similar time, and all after 4 hours online infer run(I would assume they all just finished inference loop), the one with 100 epochs trained finish and get score but the one with 30 epochs trained failed with Submission Scoring Error.     \n\nThey are using same training pipeline and all tflite models passed local and kaggle notebook local test.   \nWhat might be the root cause to have submission scoring error after infering loop? \nHow could I debug or reproduce the problem locally?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F42245%2F96c65173c52953b3c9cd8d866d5549f2%2Fxx.png?generation=1689003848968283&alt=media)\n",
    "2344764": "Oh, I only see this now, I also made a similar post (https://www.kaggle.com/competitions/asl-fingerspelling/discussion/424503). For me, it only ran 2 hours and resubmitting solved the issue, but this seems a bit strange. Did you resolve it?",
    "2339540": ""
  }
}