{
  "id": 424346,
  "title": "The difference between the validation and test results of the transformer is very large.",
  "url": "/competitions/asl-fingerspelling/discussion/424346",
  "author_name": "canlion",
  "post_date": "2023-07-13T14:38:11.618000",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello, I hope you are having a funny competition. </p>\n<p>I am currently using 1D conv and transformer.<br>\nHowever, sometimes the difference between validation (predicting the next token for each token in the target sequence) and test (inputting StartOfSequence token and repeating decoding to complete the sentence) can be very large. I have not yet found any rules.</p>\n<p>For example, in the case of using three layers of conv layers and one encoder and decoder each, the difference between validation and test results is small. (In terms of Levenshtein distance, the test is about twice the validation. I take this for granted.) However, in the case of using one conv layer and two encoders and decoders each, <strong>the train and validation loss are lower than in previous experiments, but the Levenshtein distance between validation and test is about eight times.</strong></p>\n<table>\n<thead>\n<tr>\n<th>3conv-1enc-1dec</th>\n<th>1conv-2enc-2dec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>test Levenshtein = 2 * validation Levenshtein</td>\n<td>test Levenshtein = 8 * validation Levenshtein</td>\n</tr>\n</tbody>\n</table>\n<p>I don't think it's a code problem. 😩 This problem does not always occur. there is no problem with a large model that uses 6 conv layers, 2 decoders, and 2 encoders.</p>\n<p>If you have experienced this problem, please give me some advice.</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": 2343285,
      "postDate": "2023-07-13T14:38:11.620Z",
      "content": "<p>Hello, I hope you are having a funny competition. </p>\n<p>I am currently using 1D conv and transformer.<br>\nHowever, sometimes the difference between validation (predicting the next token for each token in the target sequence) and test (inputting StartOfSequence token and repeating decoding to complete the sentence) can be very large. I have not yet found any rules.</p>\n<p>For example, in the case of using three layers of conv layers and one encoder and decoder each, the difference between validation and test results is small. (In terms of Levenshtein distance, the test is about twice the validation. I take this for granted.) However, in the case of using one conv layer and two encoders and decoders each, <strong>the train and validation loss are lower than in previous experiments, but the Levenshtein distance between validation and test is about eight times.</strong></p>\n<table>\n<thead>\n<tr>\n<th>3conv-1enc-1dec</th>\n<th>1conv-2enc-2dec</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>test Levenshtein = 2 * validation Levenshtein</td>\n<td>test Levenshtein = 8 * validation Levenshtein</td>\n</tr>\n</tbody>\n</table>\n<p>I don't think it's a code problem. 😩 This problem does not always occur. there is no problem with a large model that uses 6 conv layers, 2 decoders, and 2 encoders.</p>\n<p>If you have experienced this problem, please give me some advice.</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hello, I hope you are having a funny competition. \n\nI am currently using 1D conv and transformer.\nHowever, sometimes the difference between validation (predicting the next token for each token in the target sequence) and test (inputting StartOfSequence token and repeating decoding to complete the sentence) can be very large. I have not yet found any rules.\n\nFor example, in the case of using three layers of conv layers and one encoder and decoder each, the difference between validation and test results is small. (In terms of Levenshtein distance, the test is about twice the validation. I take this for granted.) However, in the case of using one conv layer and two encoders and decoders each, **the train and validation loss are lower than in previous experiments, but the Levenshtein distance between validation and test is about eight times.**\n\n| 3conv-1enc-1dec |1conv-2enc-2dec  |\n| --- | --- |\n| test Levenshtein = 2 * validation Levenshtein | test Levenshtein = 8 * validation Levenshtein |\n\nI don't think it's a code problem. 😩 This problem does not always occur. there is no problem with a large model that uses 6 conv layers, 2 decoders, and 2 encoders.\n\nIf you have experienced this problem, please give me some advice.\n\nThank you.",
      "votes": 10
    },
    {
      "id": 2355986,
      "postDate": "2023-07-23T18:43:27.603Z",
      "content": "<p>It seems that this issue is being experienced by more people than I originally thought. It's clear that it's not a matter that can be fully explained by the overfitting of the model. <strong>Since I myself don't have much information on CV and LB, it might be beneficial to listen to others regarding this issue. In particular, I'm interested in hearing the opinions of Kaggler's who have had significant experience with competitions that have presented similar issue.</strong></p>\n<p><a href=\"https://www.kaggle.com/canlion\" target=\"_blank\">@canlion</a> suggested that there might be a problem with the dataset, which is something I can comment on as I've also observed similar issues. Firstly, this dataset inherently contains a certain level of noise - in other words, there exists a discrepancy between the actual landmark data and the ground truth phrase to some extent. This can be understood when looking at how the data was collected. For example, factors such as the clarity of the sign language by each participant, the lighting and camera angles, and the subsequent application of the MediaPipe model can all contribute to the noise. (This was also the case with the dataset in the previous competition, where many people were puzzled by the large difference in scores between the public and private sets.). However, this noise is obviously not significant enough to pose a problem to the dataset, and I believe that part of ML is to create models that are robust to an appropriate level of noise.</p>\n<p>The question that arises here is whether to remove data that is clearly noise and then train the model. Due to the characteristics of the dataset, training the model after removing or modifying a certain portion of the data can significantly change the model's performance on some portion of the dataset. In other words, if you remove(or modify) a part of the dataset, it is possible to cause a large fluctuation in your CV or LB scores. However, it is not easy to predict how this score change will ultimately affect the private LB (since we essentially do not know how much noise is in the private set). Ironically, on the flip side, if you have great confidence, modifying the dataset might greatly benefit the private LB.</p>\n<p>Personally, I believe that if you have a reliable CV strategy and can observe consistent correlation between CV and LB, it might be best not to worry too deeply or seriously about this issue. Thank you.</p>",
      "rawMarkdown": "It seems that this issue is being experienced by more people than I originally thought. It's clear that it's not a matter that can be fully explained by the overfitting of the model. **Since I myself don't have much information on CV and LB, it might be beneficial to listen to others regarding this issue. In particular, I'm interested in hearing the opinions of Kaggler's who have had significant experience with competitions that have presented similar issue.**\n\n@canlion suggested that there might be a problem with the dataset, which is something I can comment on as I've also observed similar issues. Firstly, this dataset inherently contains a certain level of noise - in other words, there exists a discrepancy between the actual landmark data and the ground truth phrase to some extent. This can be understood when looking at how the data was collected. For example, factors such as the clarity of the sign language by each participant, the lighting and camera angles, and the subsequent application of the MediaPipe model can all contribute to the noise. (This was also the case with the dataset in the previous competition, where many people were puzzled by the large difference in scores between the public and private sets.). However, this noise is obviously not significant enough to pose a problem to the dataset, and I believe that part of ML is to create models that are robust to an appropriate level of noise.\n\nThe question that arises here is whether to remove data that is clearly noise and then train the model. Due to the characteristics of the dataset, training the model after removing or modifying a certain portion of the data can significantly change the model's performance on some portion of the dataset. In other words, if you remove(or modify) a part of the dataset, it is possible to cause a large fluctuation in your CV or LB scores. However, it is not easy to predict how this score change will ultimately affect the private LB (since we essentially do not know how much noise is in the private set). Ironically, on the flip side, if you have great confidence, modifying the dataset might greatly benefit the private LB.\n\nPersonally, I believe that if you have a reliable CV strategy and can observe consistent correlation between CV and LB, it might be best not to worry too deeply or seriously about this issue. Thank you.",
      "votes": 7,
      "replies": [
        {
          "id": 2362467,
          "postDate": "2023-07-28T05:38:47.987Z",
          "content": "<p>Thank you for your reply.<br>\nI have excluded the data samples whose input sequences are shorter than the length of the labels and I am starting the experiment again from the beginning. When I use one encoder and decoder, there is no problem. However, there is a problem when I increase the number of encoders and decoders to two. The train loss and validation loss are still surprisingly low but low low low score. large model, low score… When I apply regularization techniques to the encoder-decoder learning, the train loss and validation loss increase, but the score improves. </p>\n<p>I think that [your first advice: the overfitting of the encoder-decoder (can only perfectly predict the next token if all of the tokens in the input sequence are correct)] may be the cause of the problem.</p>\n<p>Of course, it is also possible that I have not noticed any significant errors of my notebook because most top score kaggler, including you, do not seem to be experiencing these problems.</p>",
          "rawMarkdown": "Thank you for your reply.\nI have excluded the data samples whose input sequences are shorter than the length of the labels and I am starting the experiment again from the beginning. When I use one encoder and decoder, there is no problem. However, there is a problem when I increase the number of encoders and decoders to two. The train loss and validation loss are still surprisingly low but low low low score. large model, low score... When I apply regularization techniques to the encoder-decoder learning, the train loss and validation loss increase, but the score improves. \n\nI think that [your first advice: the overfitting of the encoder-decoder (can only perfectly predict the next token if all of the tokens in the input sequence are correct)] may be the cause of the problem.\n\nOf course, it is also possible that I have not noticed any significant errors of my notebook because most top score kaggler, including you, do not seem to be experiencing these problems.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2354233,
      "postDate": "2023-07-22T09:56:07.570Z",
      "content": "<p>In my personal opinion, I believe this issue most likely caused by overfitting of the decoder. One certaitn thing is that there's no guarantee of correlation between the accuracy of a sentence generated from the ground truth's next token and the accuracy of the sentence generated from only the sos token.</p>\n<p>There are two types of decoders: an extremely overfitted decoder and a well-trained, robust decoder that generalizes well. The overfitted decoder can accurately predict the next token when given the ground truth, but its accuracy drastically drops when wrong tokens begin to mix into the input sentence during the generation process. On the other hand, a robust decoder with good generalization can, overall, accurately predict the entire sentence even if it mispredicts a few tokens during the prediction process.</p>\n<p>I believe the solution to this problem may involves reducing the size of the decoder or applying stronger regularization to the decoder. <br>\nHope this helps :)</p>",
      "rawMarkdown": "In my personal opinion, I believe this issue most likely caused by overfitting of the decoder. One certaitn thing is that there's no guarantee of correlation between the accuracy of a sentence generated from the ground truth's next token and the accuracy of the sentence generated from only the sos token.\n\nThere are two types of decoders: an extremely overfitted decoder and a well-trained, robust decoder that generalizes well. The overfitted decoder can accurately predict the next token when given the ground truth, but its accuracy drastically drops when wrong tokens begin to mix into the input sentence during the generation process. On the other hand, a robust decoder with good generalization can, overall, accurately predict the entire sentence even if it mispredicts a few tokens during the prediction process.\n\nI believe the solution to this problem may involves reducing the size of the decoder or applying stronger regularization to the decoder. \nHope this helps :)",
      "votes": 4,
      "replies": [
        {
          "id": 2354587,
          "postDate": "2023-07-22T15:34:35.020Z",
          "content": "<p>Thank you for your answer. However, I think this problem is caused by the dataset rather than the overfitting of the decoder. Today, I repeated the experiment and found the following. </p>\n<ul>\n<li>I divided the data according to the ratio of frames where the landmarks of both hands are null out of the total frames.</li>\n<li>All experiments use the same validation set.</li>\n<li>If I use only the data with a null hand frame ratio of 0.1 or less, the problem mentioned occurs and the LB score is significantly reduced.<ul>\n<li># of data samples: 100,000</li>\n<li>train loss: 0.23~ / validation loss: 0.19~ / validation levenshtein distance: 6.~ / LB score: 0.666</li></ul></li>\n<li>If I use only the data with a null hand frame ratio of 0.25 or less, the above problem is much mitigated.<ul>\n<li># of data samples: 85,000</li>\n<li>train loss: 0.32~ / validation loss: 0.28~ / validation levenshtein distance: 3.~ / LB score: 0.698</li></ul></li>\n<li>In addition, if I remove the data with a loss greater than a certain value by inference with the existing trained model from the data with a null hand frame ratio of 0.1 or less, the problem of decreasing levenshtein distance is resolved, but the LB score decreases.<ul>\n<li># of data samples: 95,000</li>\n<li>train loss: 0.33~ / validation loss: 0.27~ / validation levenshtein distance: 3.~ / LB score: 0.667</li></ul></li>\n</ul>\n<p>summary:</p>\n<ul>\n<li>data(null hand frame ratio &lt; .1) - low train, validation loss but high levenshtein dist., low LB score</li>\n<li>data(null hand frame ratio &lt; .25) - high train, validation loss but low levenshtein dist. high LB score</li>\n<li>data(null hand frame ratio &lt; .1 &amp; remove high loss samples) - high train, validation loss, low levenshtein dist. but low LB score</li>\n</ul>\n<p>Based on the experimental results, I think the additional 15,000 data (added to the 85,000 data) is causing problems. However, I do not know if the problem caused by the 15,000 data is also occurring in the 85,000 data that I am currently using. I have no experience in dealing with sequence data and transformers, so I do not know how to solve this problem.<br>\nIf you have time, please give me some advice.<br>\nthanks you.</p>",
          "rawMarkdown": "Thank you for your answer. However, I think this problem is caused by the dataset rather than the overfitting of the decoder. Today, I repeated the experiment and found the following. \n\n- I divided the data according to the ratio of frames where the landmarks of both hands are null out of the total frames.\n- All experiments use the same validation set.\n- If I use only the data with a null hand frame ratio of 0.1 or less, the problem mentioned occurs and the LB score is significantly reduced.\n  - # of data samples: 100,000\n  - train loss: 0.23~ / validation loss: 0.19~ / validation levenshtein distance: 6.~ / LB score: 0.666\n- If I use only the data with a null hand frame ratio of 0.25 or less, the above problem is much mitigated.\n  - # of data samples: 85,000\n  - train loss: 0.32~ / validation loss: 0.28~ / validation levenshtein distance: 3.~ / LB score: 0.698\n- In addition, if I remove the data with a loss greater than a certain value by inference with the existing trained model from the data with a null hand frame ratio of 0.1 or less, the problem of decreasing levenshtein distance is resolved, but the LB score decreases.\n  - # of data samples: 95,000\n  - train loss: 0.33~ / validation loss: 0.27~ / validation levenshtein distance: 3.~ / LB score: 0.667\n\nsummary:\n- data(null hand frame ratio < .1) - low train, validation loss but high levenshtein dist., low LB score\n- data(null hand frame ratio < .25) - high train, validation loss but low levenshtein dist. high LB score\n- data(null hand frame ratio < .1 & remove high loss samples) - high train, validation loss, low levenshtein dist. but low LB score\n\nBased on the experimental results, I think the additional 15,000 data (added to the 85,000 data) is causing problems. However, I do not know if the problem caused by the 15,000 data is also occurring in the 85,000 data that I am currently using. I have no experience in dealing with sequence data and transformers, so I do not know how to solve this problem.\nIf you have time, please give me some advice.\nthanks you.",
          "votes": 2,
          "replies": [
            {
              "id": 2354911,
              "postDate": "2023-07-22T23:35:43.300Z",
              "content": "<p>I see. Thank you for providing your interesting findings. While I cannot tell you anything in details since I don't know your exact settings, It seems like in this competition, similar to the previous one, the rule of thumb is to trust your CV, if you have a good, reliable CV strategy(in your case as you seem to be using a seq2seq transformer model,  I would only refer to levenshtein distance calculated from sentences generated from only the sos token). I know that dealing with uncorrelated CV, LB metrics is indeed very painful. hope to see you resolve the problem soon.</p>",
              "rawMarkdown": "I see. Thank you for providing your interesting findings. While I cannot tell you anything in details since I don't know your exact settings, It seems like in this competition, similar to the previous one, the rule of thumb is to trust your CV, if you have a good, reliable CV strategy(in your case as you seem to be using a seq2seq transformer model,  I would only refer to levenshtein distance calculated from sentences generated from only the sos token). I know that dealing with uncorrelated CV, LB metrics is indeed very painful. hope to see you resolve the problem soon.",
              "votes": 2
            },
            {
              "id": 2355477,
              "postDate": "2023-07-23T11:23:39.727Z",
              "content": "<p>It took me almost a month and still can't find a way to split my CV to match the LB. The CV predicted by sos and my LB are 10% different. Can you tell me how you split your CV. I tried 5/10/20 split but all of them didn't work :( 😭</p>",
              "rawMarkdown": "It took me almost a month and still can't find a way to split my CV to match the LB. The CV predicted by sos and my LB are 10% different. Can you tell me how you split your CV. I tried 5/10/20 split but all of them didn't work :( 😭"
            },
            {
              "id": 2355819,
              "postDate": "2023-07-23T16:28:19.053Z",
              "content": "<p>I have the same issue that my local validation score is higher than the leaderboard score. I am currently only using the training data provided and splitting it into unique participants, then I take 10% of those and use them for validation. Some of the approaches I used to drop some of the data include: </p>\n<ul>\n<li>drop nothing</li>\n<li>dropping if 2* the length of frames is &lt; phrase</li>\n<li>dropping if the length of frames is &lt; phrase</li>\n<li>dropping if len frames with hands &lt; 1</li>\n<li>as suggested above dropping if null hand frame ratio of 0.1</li>\n<li>as suggested above dropping if null hand frame ratio of 0.25<br>\nFor me all of these give a better local score than the leaderboard, although I did apply all of those to both the train/validation datasets equally. Perhaps I should keep all of the validation data, but drop from the training data; or maybe it is that my model overfits.</li>\n</ul>",
              "rawMarkdown": "I have the same issue that my local validation score is higher than the leaderboard score. I am currently only using the training data provided and splitting it into unique participants, then I take 10% of those and use them for validation. Some of the approaches I used to drop some of the data include: \n- drop nothing\n- dropping if 2* the length of frames is < phrase\n- dropping if the length of frames is < phrase\n- dropping if len frames with hands < 1\n- as suggested above dropping if null hand frame ratio of 0.1\n- as suggested above dropping if null hand frame ratio of 0.25\nFor me all of these give a better local score than the leaderboard, although I did apply all of those to both the train/validation datasets equally. Perhaps I should keep all of the validation data, but drop from the training data; or maybe it is that my model overfits."
            },
            {
              "id": 2355984,
              "postDate": "2023-07-23T18:39:09.900Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2356088,
              "postDate": "2023-07-23T21:48:55.090Z",
              "content": "<p>Did you apply data augmentation/data cleaning on your validation set? My LB scores and CV scores correlate very well with a very simple 95%/5% splitting (although the gap between my CV score and LB score gradually increases as my CV score increases. For instance, I got <strong>LB 0.761</strong> for <strong>CV 0.771</strong>, and <strong>LB 0.771</strong> for <strong>CV 0.796</strong>). More specifically, I leave out 3 files of the original parquet files as a validation set and compute the score on an exported tflite model.</p>",
              "rawMarkdown": "Did you apply data augmentation/data cleaning on your validation set? My LB scores and CV scores correlate very well with a very simple 95%/5% splitting (although the gap between my CV score and LB score gradually increases as my CV score increases. For instance, I got **LB 0.761** for **CV 0.771**, and **LB 0.771** for **CV 0.796**). More specifically, I leave out 3 files of the original parquet files as a validation set and compute the score on an exported tflite model."
            },
            {
              "id": 2356632,
              "postDate": "2023-07-24T09:53:35.440Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2356635,
              "postDate": "2023-07-24T09:54:03.893Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2356638,
              "postDate": "2023-07-24T09:55:29.913Z",
              "content": "<p><a href=\"https://www.kaggle.com/proptiter\" target=\"_blank\">@proptiter</a> I'm currently using typical 5fold split by id.</p>",
              "rawMarkdown": "@proptiter I'm currently using typical 5fold split by id.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2365217,
      "postDate": "2023-07-30T04:50:18.383Z",
      "content": "<p>hahahahaha… I don't know what to do. Having two or more decoder layers definitely causes performance degradation. I don't know how to analyze the cause.</p>",
      "rawMarkdown": "hahahahaha... I don't know what to do. Having two or more decoder layers definitely causes performance degradation. I don't know how to analyze the cause."
    },
    {
      "id": 2348774,
      "postDate": "2023-07-18T00:44:57.633Z",
      "content": "<p>I met this too. Very strange.</p>",
      "rawMarkdown": "I met this too. Very strange."
    }
  ],
  "comments": [
    {
      "id": 2355986,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-07-23T18:43:27.603000",
      "content": "<p>It seems that this issue is being experienced by more people than I originally thought. It's clear that it's not a matter that can be fully explained by the overfitting of the model. <strong>Since I myself don't have much information on CV and LB, it might be beneficial to listen to others regarding this issue. In particular, I'm interested in hearing the opinions of Kaggler's who have had significant experience with competitions that have presented similar issue.</strong></p>\n<p><a href=\"https://www.kaggle.com/canlion\" target=\"_blank\">@canlion</a> suggested that there might be a problem with the dataset, which is something I can comment on as I've also observed similar issues. Firstly, this dataset inherently contains a certain level of noise - in other words, there exists a discrepancy between the actual landmark data and the ground truth phrase to some extent. This can be understood when looking at how the data was collected. For example, factors such as the clarity of the sign language by each participant, the lighting and camera angles, and the subsequent application of the MediaPipe model can all contribute to the noise. (This was also the case with the dataset in the previous competition, where many people were puzzled by the large difference in scores between the public and private sets.). However, this noise is obviously not significant enough to pose a problem to the dataset, and I believe that part of ML is to create models that are robust to an appropriate level of noise.</p>\n<p>The question that arises here is whether to remove data that is clearly noise and then train the model. Due to the characteristics of the dataset, training the model after removing or modifying a certain portion of the data can significantly change the model's performance on some portion of the dataset. In other words, if you remove(or modify) a part of the dataset, it is possible to cause a large fluctuation in your CV or LB scores. However, it is not easy to predict how this score change will ultimately affect the private LB (since we essentially do not know how much noise is in the private set). Ironically, on the flip side, if you have great confidence, modifying the dataset might greatly benefit the private LB.</p>\n<p>Personally, I believe that if you have a reliable CV strategy and can observe consistent correlation between CV and LB, it might be best not to worry too deeply or seriously about this issue. Thank you.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2362467,
          "author_name": "canlion",
          "author_url": "",
          "post_date": "2023-07-28T05:38:47.987000",
          "content": "<p>Thank you for your reply.<br>\nI have excluded the data samples whose input sequences are shorter than the length of the labels and I am starting the experiment again from the beginning. When I use one encoder and decoder, there is no problem. However, there is a problem when I increase the number of encoders and decoders to two. The train loss and validation loss are still surprisingly low but low low low score. large model, low score… When I apply regularization techniques to the encoder-decoder learning, the train loss and validation loss increase, but the score improves. </p>\n<p>I think that [your first advice: the overfitting of the encoder-decoder (can only perfectly predict the next token if all of the tokens in the input sequence are correct)] may be the cause of the problem.</p>\n<p>Of course, it is also possible that I have not noticed any significant errors of my notebook because most top score kaggler, including you, do not seem to be experiencing these problems.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2354233,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-07-22T09:56:07.570000",
      "content": "<p>In my personal opinion, I believe this issue most likely caused by overfitting of the decoder. One certaitn thing is that there's no guarantee of correlation between the accuracy of a sentence generated from the ground truth's next token and the accuracy of the sentence generated from only the sos token.</p>\n<p>There are two types of decoders: an extremely overfitted decoder and a well-trained, robust decoder that generalizes well. The overfitted decoder can accurately predict the next token when given the ground truth, but its accuracy drastically drops when wrong tokens begin to mix into the input sentence during the generation process. On the other hand, a robust decoder with good generalization can, overall, accurately predict the entire sentence even if it mispredicts a few tokens during the prediction process.</p>\n<p>I believe the solution to this problem may involves reducing the size of the decoder or applying stronger regularization to the decoder. <br>\nHope this helps :)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2354587,
          "author_name": "canlion",
          "author_url": "",
          "post_date": "2023-07-22T15:34:35.020000",
          "content": "<p>Thank you for your answer. However, I think this problem is caused by the dataset rather than the overfitting of the decoder. Today, I repeated the experiment and found the following. </p>\n<ul>\n<li>I divided the data according to the ratio of frames where the landmarks of both hands are null out of the total frames.</li>\n<li>All experiments use the same validation set.</li>\n<li>If I use only the data with a null hand frame ratio of 0.1 or less, the problem mentioned occurs and the LB score is significantly reduced.<ul>\n<li># of data samples: 100,000</li>\n<li>train loss: 0.23~ / validation loss: 0.19~ / validation levenshtein distance: 6.~ / LB score: 0.666</li></ul></li>\n<li>If I use only the data with a null hand frame ratio of 0.25 or less, the above problem is much mitigated.<ul>\n<li># of data samples: 85,000</li>\n<li>train loss: 0.32~ / validation loss: 0.28~ / validation levenshtein distance: 3.~ / LB score: 0.698</li></ul></li>\n<li>In addition, if I remove the data with a loss greater than a certain value by inference with the existing trained model from the data with a null hand frame ratio of 0.1 or less, the problem of decreasing levenshtein distance is resolved, but the LB score decreases.<ul>\n<li># of data samples: 95,000</li>\n<li>train loss: 0.33~ / validation loss: 0.27~ / validation levenshtein distance: 3.~ / LB score: 0.667</li></ul></li>\n</ul>\n<p>summary:</p>\n<ul>\n<li>data(null hand frame ratio &lt; .1) - low train, validation loss but high levenshtein dist., low LB score</li>\n<li>data(null hand frame ratio &lt; .25) - high train, validation loss but low levenshtein dist. high LB score</li>\n<li>data(null hand frame ratio &lt; .1 &amp; remove high loss samples) - high train, validation loss, low levenshtein dist. but low LB score</li>\n</ul>\n<p>Based on the experimental results, I think the additional 15,000 data (added to the 85,000 data) is causing problems. However, I do not know if the problem caused by the 15,000 data is also occurring in the 85,000 data that I am currently using. I have no experience in dealing with sequence data and transformers, so I do not know how to solve this problem.<br>\nIf you have time, please give me some advice.<br>\nthanks you.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2354911,
              "author_name": "hoyso48",
              "author_url": "",
              "post_date": "2023-07-22T23:35:43.300000",
              "content": "<p>I see. Thank you for providing your interesting findings. While I cannot tell you anything in details since I don't know your exact settings, It seems like in this competition, similar to the previous one, the rule of thumb is to trust your CV, if you have a good, reliable CV strategy(in your case as you seem to be using a seq2seq transformer model,  I would only refer to levenshtein distance calculated from sentences generated from only the sos token). I know that dealing with uncorrelated CV, LB metrics is indeed very painful. hope to see you resolve the problem soon.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2355477,
              "author_name": "Just A game on your lips",
              "author_url": "",
              "post_date": "2023-07-23T11:23:39.727000",
              "content": "<p>It took me almost a month and still can't find a way to split my CV to match the LB. The CV predicted by sos and my LB are 10% different. Can you tell me how you split your CV. I tried 5/10/20 split but all of them didn't work :( 😭</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2355819,
              "author_name": "wildrunner1",
              "author_url": "",
              "post_date": "2023-07-23T16:28:19.053000",
              "content": "<p>I have the same issue that my local validation score is higher than the leaderboard score. I am currently only using the training data provided and splitting it into unique participants, then I take 10% of those and use them for validation. Some of the approaches I used to drop some of the data include: </p>\n<ul>\n<li>drop nothing</li>\n<li>dropping if 2* the length of frames is &lt; phrase</li>\n<li>dropping if the length of frames is &lt; phrase</li>\n<li>dropping if len frames with hands &lt; 1</li>\n<li>as suggested above dropping if null hand frame ratio of 0.1</li>\n<li>as suggested above dropping if null hand frame ratio of 0.25<br>\nFor me all of these give a better local score than the leaderboard, although I did apply all of those to both the train/validation datasets equally. Perhaps I should keep all of the validation data, but drop from the training data; or maybe it is that my model overfits.</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2355984,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-07-23T18:39:09.900000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2356088,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-07-23T21:48:55.090000",
              "content": "<p>Did you apply data augmentation/data cleaning on your validation set? My LB scores and CV scores correlate very well with a very simple 95%/5% splitting (although the gap between my CV score and LB score gradually increases as my CV score increases. For instance, I got <strong>LB 0.761</strong> for <strong>CV 0.771</strong>, and <strong>LB 0.771</strong> for <strong>CV 0.796</strong>). More specifically, I leave out 3 files of the original parquet files as a validation set and compute the score on an exported tflite model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2356632,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-07-24T09:53:35.440000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2356635,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-07-24T09:54:03.893000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2356638,
              "author_name": "hoyso48",
              "author_url": "",
              "post_date": "2023-07-24T09:55:29.913000",
              "content": "<p><a href=\"https://www.kaggle.com/proptiter\" target=\"_blank\">@proptiter</a> I'm currently using typical 5fold split by id.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2365217,
      "author_name": "canlion",
      "author_url": "",
      "post_date": "2023-07-30T04:50:18.383000",
      "content": "<p>hahahahaha… I don't know what to do. Having two or more decoder layers definitely causes performance degradation. I don't know how to analyze the cause.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2348774,
      "author_name": "zmchen",
      "author_url": "",
      "post_date": "2023-07-18T00:44:57.633000",
      "content": "<p>I met this too. Very strange.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2343285": "Hello, I hope you are having a funny competition. \n\nI am currently using 1D conv and transformer.\nHowever, sometimes the difference between validation (predicting the next token for each token in the target sequence) and test (inputting StartOfSequence token and repeating decoding to complete the sentence) can be very large. I have not yet found any rules.\n\nFor example, in the case of using three layers of conv layers and one encoder and decoder each, the difference between validation and test results is small. (In terms of Levenshtein distance, the test is about twice the validation. I take this for granted.) However, in the case of using one conv layer and two encoders and decoders each, **the train and validation loss are lower than in previous experiments, but the Levenshtein distance between validation and test is about eight times.**\n\n| 3conv-1enc-1dec |1conv-2enc-2dec  |\n| --- | --- |\n| test Levenshtein = 2 * validation Levenshtein | test Levenshtein = 8 * validation Levenshtein |\n\nI don't think it's a code problem. 😩 This problem does not always occur. there is no problem with a large model that uses 6 conv layers, 2 decoders, and 2 encoders.\n\nIf you have experienced this problem, please give me some advice.\n\nThank you.",
    "2355986": "It seems that this issue is being experienced by more people than I originally thought. It's clear that it's not a matter that can be fully explained by the overfitting of the model. **Since I myself don't have much information on CV and LB, it might be beneficial to listen to others regarding this issue. In particular, I'm interested in hearing the opinions of Kaggler's who have had significant experience with competitions that have presented similar issue.**\n\n@canlion suggested that there might be a problem with the dataset, which is something I can comment on as I've also observed similar issues. Firstly, this dataset inherently contains a certain level of noise - in other words, there exists a discrepancy between the actual landmark data and the ground truth phrase to some extent. This can be understood when looking at how the data was collected. For example, factors such as the clarity of the sign language by each participant, the lighting and camera angles, and the subsequent application of the MediaPipe model can all contribute to the noise. (This was also the case with the dataset in the previous competition, where many people were puzzled by the large difference in scores between the public and private sets.). However, this noise is obviously not significant enough to pose a problem to the dataset, and I believe that part of ML is to create models that are robust to an appropriate level of noise.\n\nThe question that arises here is whether to remove data that is clearly noise and then train the model. Due to the characteristics of the dataset, training the model after removing or modifying a certain portion of the data can significantly change the model's performance on some portion of the dataset. In other words, if you remove(or modify) a part of the dataset, it is possible to cause a large fluctuation in your CV or LB scores. However, it is not easy to predict how this score change will ultimately affect the private LB (since we essentially do not know how much noise is in the private set). Ironically, on the flip side, if you have great confidence, modifying the dataset might greatly benefit the private LB.\n\nPersonally, I believe that if you have a reliable CV strategy and can observe consistent correlation between CV and LB, it might be best not to worry too deeply or seriously about this issue. Thank you.",
    "2354233": "In my personal opinion, I believe this issue most likely caused by overfitting of the decoder. One certaitn thing is that there's no guarantee of correlation between the accuracy of a sentence generated from the ground truth's next token and the accuracy of the sentence generated from only the sos token.\n\nThere are two types of decoders: an extremely overfitted decoder and a well-trained, robust decoder that generalizes well. The overfitted decoder can accurately predict the next token when given the ground truth, but its accuracy drastically drops when wrong tokens begin to mix into the input sentence during the generation process. On the other hand, a robust decoder with good generalization can, overall, accurately predict the entire sentence even if it mispredicts a few tokens during the prediction process.\n\nI believe the solution to this problem may involves reducing the size of the decoder or applying stronger regularization to the decoder. \nHope this helps :)",
    "2365217": "hahahahaha... I don't know what to do. Having two or more decoder layers definitely causes performance degradation. I don't know how to analyze the cause.",
    "2348774": "I met this too. Very strange."
  }
}