{
  "id": 432784,
  "title": "Public score is 0.035 lower than my holdout (N-D)/N, Solved!",
  "url": "/competitions/asl-fingerspelling/discussion/432784",
  "author_name": "Andy Atkinson",
  "post_date": "2023-08-18T23:52:02.483000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm not sure what's causing it: leakage, difference in train/test distributions, my still using the old competition data pre update, error in tflite submission pipeline preprocessing (though I ran 1000 validation samples through my tflite model and they all match my pre-tflite predictions), overfitting (seems unlikely as multiple model configurations yield the same bias), maybe I randomly picked easy samples (not probable, though I haven't rotated my data to check this).</p>\n<p>I'm doing a global sum across observations before applying the formula, so like this: (sum(N) - sum(D))/sum(N)</p>\n<p>Has anyone seen a similar delta between holdout and submission score?</p>\n<p><strong>EDIT: I'm using this for my data processing <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset</a>.   It was set to remove videos with less than 4 frames and that paints a rosy picture making my CV higher than LB, setting this MIN_NUM_FRAMES_PER_CHARACTER = 0 resolves the disconnect.</strong></p>",
  "messages": [
    {
      "id": 2397371,
      "postDate": "2023-08-19T01:42:20.620Z",
      "content": "<p>It depends on if you split train/valid considering participatant id or not.</p>",
      "rawMarkdown": "It depends on if you split train/valid considering participatant id or not.",
      "replies": [
        {
          "id": 2397465,
          "postDate": "2023-08-19T03:49:48.550Z",
          "content": "<p>I split grouping by participant id, using this code: <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset</a></p>\n<p>I don't actually know what this groups = participant_ids is doing in the above, whether it's ensuring a given participant is represented in both train and valid, or if it's forcing there to be no participant in both. :D</p>\n<p>I did see another discussion that said that participant_ids related splits could cause CV and LB divergence? What do you think?</p>",
          "rawMarkdown": "I split grouping by participant id, using this code: https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\n\nI don't actually know what this groups = participant_ids is doing in the above, whether it's ensuring a given participant is represented in both train and valid, or if it's forcing there to be no participant in both. :D\n\nI did see another discussion that said that participant_ids related splits could cause CV and LB divergence? What do you think?",
          "replies": [
            {
              "id": 2397501,
              "postDate": "2023-08-19T04:50:58.770Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2397506,
              "postDate": "2023-08-19T04:53:14.947Z",
              "content": "<p>Normaly if you use groups=participant_ids  for cv split then your local cv is lower then online, if you do not use then your local score should be higher then online. <br>\nI have a post asking the official to verify if there are duplicate participant_ids between train and online test. But no reply yet. Personally I think there are new participant_ids and also old(in train) participant_ids in the test, if so the final results will be intesting, for you could not test well locally:)</p>",
              "rawMarkdown": "Normaly if you use groups=participant_ids  for cv split then your local cv is lower then online, if you do not use then your local score should be higher then online. \nI have a post asking the official to verify if there are duplicate participant_ids between train and online test. But no reply yet. Personally I think there are new participant_ids and also old(in train) participant_ids in the test, if so the final results will be intesting, for you could not test well locally:)\n\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2397335,
      "postDate": "2023-08-18T23:52:02.483Z",
      "content": "<p>I'm not sure what's causing it: leakage, difference in train/test distributions, my still using the old competition data pre update, error in tflite submission pipeline preprocessing (though I ran 1000 validation samples through my tflite model and they all match my pre-tflite predictions), overfitting (seems unlikely as multiple model configurations yield the same bias), maybe I randomly picked easy samples (not probable, though I haven't rotated my data to check this).</p>\n<p>I'm doing a global sum across observations before applying the formula, so like this: (sum(N) - sum(D))/sum(N)</p>\n<p>Has anyone seen a similar delta between holdout and submission score?</p>\n<p><strong>EDIT: I'm using this for my data processing <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset</a>.   It was set to remove videos with less than 4 frames and that paints a rosy picture making my CV higher than LB, setting this MIN_NUM_FRAMES_PER_CHARACTER = 0 resolves the disconnect.</strong></p>",
      "rawMarkdown": "I'm not sure what's causing it: leakage, difference in train/test distributions, my still using the old competition data pre update, error in tflite submission pipeline preprocessing (though I ran 1000 validation samples through my tflite model and they all match my pre-tflite predictions), overfitting (seems unlikely as multiple model configurations yield the same bias), maybe I randomly picked easy samples (not probable, though I haven't rotated my data to check this).\n\nI'm doing a global sum across observations before applying the formula, so like this: (sum(N) - sum(D))/sum(N)\n\nHas anyone seen a similar delta between holdout and submission score?\n\n**EDIT: I'm using this for my data processing https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset.   It was set to remove videos with less than 4 frames and that paints a rosy picture making my CV higher than LB, setting this MIN_NUM_FRAMES_PER_CHARACTER = 0 resolves the disconnect.**"
    }
  ],
  "comments": [
    {
      "id": 2397371,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-19T01:42:20.620000",
      "content": "<p>It depends on if you split train/valid considering participatant id or not.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2397465,
          "author_name": "Andy Atkinson",
          "author_url": "",
          "post_date": "2023-08-19T03:49:48.550000",
          "content": "<p>I split grouping by participant id, using this code: <a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset</a></p>\n<p>I don't actually know what this groups = participant_ids is doing in the above, whether it's ensuring a given participant is represented in both train and valid, or if it's forcing there to be no participant in both. :D</p>\n<p>I did see another discussion that said that participant_ids related splits could cause CV and LB divergence? What do you think?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2397501,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-19T04:50:58.770000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2397506,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-19T04:53:14.947000",
              "content": "<p>Normaly if you use groups=participant_ids  for cv split then your local cv is lower then online, if you do not use then your local score should be higher then online. <br>\nI have a post asking the official to verify if there are duplicate participant_ids between train and online test. But no reply yet. Personally I think there are new participant_ids and also old(in train) participant_ids in the test, if so the final results will be intesting, for you could not test well locally:)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2397371": "It depends on if you split train/valid considering participatant id or not.",
    "2397335": "I'm not sure what's causing it: leakage, difference in train/test distributions, my still using the old competition data pre update, error in tflite submission pipeline preprocessing (though I ran 1000 validation samples through my tflite model and they all match my pre-tflite predictions), overfitting (seems unlikely as multiple model configurations yield the same bias), maybe I randomly picked easy samples (not probable, though I haven't rotated my data to check this).\n\nI'm doing a global sum across observations before applying the formula, so like this: (sum(N) - sum(D))/sum(N)\n\nHas anyone seen a similar delta between holdout and submission score?\n\n**EDIT: I'm using this for my data processing https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset.   It was set to remove videos with less than 4 frames and that paints a rosy picture making my CV higher than LB, setting this MIN_NUM_FRAMES_PER_CHARACTER = 0 resolves the disconnect.**"
  }
}