{
  "id": 540663,
  "title": "Train/Public scores aka CV/LB",
  "url": "/competitions/ariel-data-challenge-2024/discussion/540663",
  "author_name": "Sergei Fironov",
  "post_date": "2024-10-15T15:21:19.148000",
  "votes": 10,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Mine 0.660/0.685 and 0.781/0.679. Still have no idea what's wrong with the last one.</p>",
  "messages": [
    {
      "id": 3018191,
      "postDate": "2024-10-15T15:21:19.150Z",
      "content": "<p>Mine 0.660/0.685 and 0.781/0.679. Still have no idea what's wrong with the last one.</p>",
      "rawMarkdown": "Mine 0.660/0.685 and 0.781/0.679. Still have no idea what's wrong with the last one.",
      "votes": 10
    },
    {
      "id": 3018504,
      "postDate": "2024-10-15T19:50:44.697Z",
      "content": "<p>0.641/0.675 (my 0.678 is LB overfitting; let's not consider it). Funny story: I accidentally used sigma_true=1e-6 instead of 1e-5 from the beginning, so all my scores are ~0.4, and my intuition is built around this range. I used the correct sigma only for this post.</p>",
      "rawMarkdown": "0.641/0.675 (my 0.678 is LB overfitting; let's not consider it). Funny story: I accidentally used sigma_true=1e-6 instead of 1e-5 from the beginning, so all my scores are ~0.4, and my intuition is built around this range. I used the correct sigma only for this post.",
      "votes": 5,
      "replies": [
        {
          "id": 3018507,
          "postDate": "2024-10-15T19:56:03.903Z",
          "content": "<p>I noticed that when you start chatting on the forum, you reveal your own mistakes. Just don't say too much. Thanks for sharing! It looks similar. Edit: I'm not sure if this works the same way in English as in Russian. I don't mean you specifically, but I mean in general. Someday I'll learn English too.</p>",
          "rawMarkdown": "I noticed that when you start chatting on the forum, you reveal your own mistakes. Just don't say too much. Thanks for sharing! It looks similar. Edit: I'm not sure if this works the same way in English as in Russian. I don't mean you specifically, but I mean in general. Someday I'll learn English too.",
          "votes": 1,
          "replies": [
            {
              "id": 3018515,
              "postDate": "2024-10-15T20:10:07.647Z",
              "content": "<p>Hehe, it can be interpreted in both ways, but yeah, it sounds more like the first option. Well, English is not my first language, either, but I would probably go with, 'I noticed that when one starts chatting on the forum, one reveals his own mistakes.'</p>",
              "rawMarkdown": "Hehe, it can be interpreted in both ways, but yeah, it sounds more like the first option. Well, English is not my first language, either, but I would probably go with, 'I noticed that when one starts chatting on the forum, one reveals his own mistakes.'",
              "votes": 3
            },
            {
              "id": 3018519,
              "postDate": "2024-10-15T20:14:21.863Z",
              "content": "<p>Yes, it sounds good. Thanks)</p>",
              "rawMarkdown": "Yes, it sounds good. Thanks)"
            }
          ]
        }
      ]
    },
    {
      "id": 3020126,
      "postDate": "2024-10-17T07:19:03.830Z",
      "content": "<p>Mine is 0.679/0.624. I've recently made several \"improvements\" which turned out to be pure overfitting, decreasing public scores 😪<br>\nMy assumption: train and test sets have very different spectra distributions (and by that I mean the \"shape\" of y for each planet). Whatever you do that improves specifically fitting the train distributions, is an overfit, and decreases test score.</p>",
      "rawMarkdown": "Mine is 0.679/0.624. I've recently made several \"improvements\" which turned out to be pure overfitting, decreasing public scores 😪\nMy assumption: train and test sets have very different spectra distributions (and by that I mean the \"shape\" of y for each planet). Whatever you do that improves specifically fitting the train distributions, is an overfit, and decreases test score.",
      "votes": 3,
      "replies": [
        {
          "id": 3020132,
          "postDate": "2024-10-17T07:28:06.323Z",
          "content": "<p>It seems that fitting raw signal will not generalize to LB?</p>",
          "rawMarkdown": "It seems that fitting raw signal will not generalize to LB?",
          "votes": 1,
          "replies": [
            {
              "id": 3020308,
              "postDate": "2024-10-17T11:42:32.263Z",
              "content": "<p>What do you mean?</p>",
              "rawMarkdown": "What do you mean?"
            }
          ]
        }
      ]
    },
    {
      "id": 3019773,
      "postDate": "2024-10-16T20:47:39.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> what is your CV mse? Mine sits at around 4e-9 - 5e-9</p>",
      "rawMarkdown": "@sergeifironov what is your CV mse? Mine sits at around 4e-9 - 5e-9",
      "votes": 3,
      "replies": [
        {
          "id": 3019784,
          "postDate": "2024-10-16T21:07:54.980Z",
          "content": "<p>both about 2e-9. It's overfitting for sure.</p>",
          "rawMarkdown": "both about 2e-9. It's overfitting for sure.",
          "votes": 1,
          "replies": [
            {
              "id": 3019833,
              "postDate": "2024-10-16T22:50:47.417Z",
              "content": "<p>Wow okay, so it's not just my sigma estimation that is bad 😄</p>",
              "rawMarkdown": "Wow okay, so it's not just my sigma estimation that is bad 😄"
            },
            {
              "id": 3024340,
              "postDate": "2024-10-21T14:34:16.780Z",
              "content": "<p>That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)</p>",
              "rawMarkdown": "That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)",
              "votes": 1
            },
            {
              "id": 3024372,
              "postDate": "2024-10-21T15:10:14.563Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3024375,
              "postDate": "2024-10-21T15:15:23.327Z",
              "content": "<blockquote>\n  <p>That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)</p>\n</blockquote>\n<p>It seems to me that you are talking about Rmse, not MSE. My RMSE is approximately the same as yours.</p>",
              "rawMarkdown": "> That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)\n\nIt seems to me that you are talking about Rmse, not MSE. My RMSE is approximately the same as yours.",
              "votes": 2
            },
            {
              "id": 3024942,
              "postDate": "2024-10-22T07:15:39.140Z",
              "content": "<p>Oh yeah, right, my mistake</p>",
              "rawMarkdown": "Oh yeah, right, my mistake"
            }
          ]
        }
      ]
    },
    {
      "id": 3019336,
      "postDate": "2024-10-16T13:27:08.687Z",
      "content": "<p>How do you estimate training score? I use the following function and it returns my training score near 0.9 which has no correlation with lb.</p>\n<p><em>gll_score = competition_score(        \n            train_labels.copy(),\n            submission.copy().reset_index(), \n            train_labels.values.mean(),\n            train_labels.values.std(),\n            sigma_true=1e-5).</em><br>\nDo you use lower sigma_true?</p>",
      "rawMarkdown": "How do you estimate training score? I use the following function and it returns my training score near 0.9 which has no correlation with lb.\n\n*gll_score = competition_score(        \n            train_labels.copy(),\n            submission.copy().reset_index(), \n            train_labels.values.mean(),\n            train_labels.values.std(),\n            sigma_true=1e-5).*\nDo you use lower sigma_true?",
      "replies": [
        {
          "id": 3019376,
          "postDate": "2024-10-16T14:23:15.680Z",
          "content": "<p>Your code is correct. 0.9 is an unrealistically high score. There is probably a leakage in the training process or model is overfitted </p>",
          "rawMarkdown": "Your code is correct. 0.9 is an unrealistically high score. There is probably a leakage in the training process or model is overfitted ",
          "votes": 2,
          "replies": [
            {
              "id": 3019441,
              "postDate": "2024-10-16T15:19:40.490Z",
              "content": "<p>Thanks, that seems a bit weird. I'm currently trying to develop your unsupervised search by polynomial approximation with more careful and stable transit zone handling. As a baseline I forked your great ariel_only_correlation notebook and got local score 0.897. Did you observe the same score when testing this idea? Looking at this discussion, still can't understand if this method may overfit so much somehow or I can't find a bug in my scoring :(</p>",
              "rawMarkdown": "Thanks, that seems a bit weird. I'm currently trying to develop your unsupervised search by polynomial approximation with more careful and stable transit zone handling. As a baseline I forked your great ariel_only_correlation notebook and got local score 0.897. Did you observe the same score when testing this idea? Looking at this discussion, still can't understand if this method may overfit so much somehow or I can't find a bug in my scoring :("
            },
            {
              "id": 3019464,
              "postDate": "2024-10-16T15:49:09.913Z",
              "content": "<p>As I can remember, score was 0.48+ with 1.6e-4 sigma. </p>",
              "rawMarkdown": "As I can remember, score was 0.48+ with 1.6e-4 sigma. ",
              "votes": 2
            },
            {
              "id": 3019667,
              "postDate": "2024-10-16T18:33:30.603Z",
              "content": "<p>Chiming in as well, I believe there's a mistake in the way you're applying your scoring formula. </p>\n<p>Assuming you make a perfect prediction of the mean and copy that across the 283 wavelengths (and Sergei's baseline does that, however without a perfect prediction of the mean, so you'd implicitly expect a lower score), your maximum score could be around 0.53 with a sigma of ~9e-5. Getting to around 0.9 would assume an almost perfect prediction of not only the mean Rp/Rs squared, but also of the distribution across the 283 wavelengths. </p>\n<p>A model that simply predicts the mean value for a planet and copies that across the other wavelengths essentially can't reach a score of 0.9.</p>\n<pre><code>train_target = pd.read_csv('/kaggle/input/ariel-data-challenge-/train_labels.csv')\ny_true = train_target.loc[:, train_target. != 'planet_id'].\nnaive_mean=y_true.(),\nnaive_sigma=y_true.(),\nsigma_true=\nsigma_pred=\n = []\n i  (len(train_target)):\n    .(.repeat(y_true[i].(),))\ny_pred = .()\nGLL_pred = .(scipy.stats.norm.logpdf(y_true, loc=y_pred, =sigma_pred))\nGLL_true = .(scipy.stats.norm.logpdf(y_true, loc=y_true, =sigma_true * .ones_like(y_true)))\nGLL_mean = .(scipy.stats.norm.logpdf(y_true, loc=naive_mean * .ones_like(y_true), =naive_sigma * .ones_like(y_true)))\n((GLL_pred-GLL_mean)/(GLL_true-GLL_mean))\n</code></pre>",
              "rawMarkdown": "Chiming in as well, I believe there's a mistake in the way you're applying your scoring formula. \n\nAssuming you make a perfect prediction of the mean and copy that across the 283 wavelengths (and Sergei's baseline does that, however without a perfect prediction of the mean, so you'd implicitly expect a lower score), your maximum score could be around 0.53 with a sigma of ~9e-5. Getting to around 0.9 would assume an almost perfect prediction of not only the mean Rp/Rs squared, but also of the distribution across the 283 wavelengths. \n\nA model that simply predicts the mean value for a planet and copies that across the other wavelengths essentially can't reach a score of 0.9.\n\n```\ntrain_target = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv')\ny_true = train_target.loc[:, train_target.columns != 'planet_id'].values\nnaive_mean=y_true.mean(),\nnaive_sigma=y_true.std(),\nsigma_true=1e-5\nsigma_pred=9e-5\nlabels = []\nfor i in range(len(train_target)):\n    labels.append(np.repeat(y_true[i].mean(),283))\ny_pred = np.array(labels)\nGLL_pred = np.sum(scipy.stats.norm.logpdf(y_true, loc=y_pred, scale=sigma_pred))\nGLL_true = np.sum(scipy.stats.norm.logpdf(y_true, loc=y_true, scale=sigma_true * np.ones_like(y_true)))\nGLL_mean = np.sum(scipy.stats.norm.logpdf(y_true, loc=naive_mean * np.ones_like(y_true), scale=naive_sigma * np.ones_like(y_true)))\nprint((GLL_pred-GLL_mean)/(GLL_true-GLL_mean))\n```",
              "votes": 2
            }
          ]
        },
        {
          "id": 3020307,
          "postDate": "2024-10-17T11:41:40.913Z",
          "content": "<p>I use the same scoring function and get 0.495 on the training data, whereas 0.567 is on the LB. </p>",
          "rawMarkdown": "I use the same scoring function and get 0.495 on the training data, whereas 0.567 is on the LB. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 3018656,
      "postDate": "2024-10-16T01:07:22.780Z",
      "content": "<p>Are you using 1 star for training/1 star for validation?</p>",
      "rawMarkdown": "Are you using 1 star for training/1 star for validation?",
      "replies": [
        {
          "id": 3018987,
          "postDate": "2024-10-16T07:30:26.030Z",
          "content": "<p>Does it correlate better with the LB?  No, just out-of-fold prediction in some sense. Would you mind to share your score?</p>",
          "rawMarkdown": "Does it correlate better with the LB?  No, just out-of-fold prediction in some sense. Would you mind to share your score?",
          "replies": [
            {
              "id": 3019046,
              "postDate": "2024-10-16T08:00:02.230Z",
              "content": "<p>I used random split 5fold as cv, ~0.69/0.56 😅. Idk what's wrong with my signal preprocessing …</p>",
              "rawMarkdown": "I used random split 5fold as cv, ~0.69/0.56 😅. Idk what's wrong with my signal preprocessing ...",
              "votes": 1
            },
            {
              "id": 3019085,
              "postDate": "2024-10-16T08:46:12.720Z",
              "content": "<p>Yes, it looks like my second submission. I have no idea what in the test leads to this.</p>",
              "rawMarkdown": "Yes, it looks like my second submission. I have no idea what in the test leads to this.",
              "votes": 1
            },
            {
              "id": 3019162,
              "postDate": "2024-10-16T10:17:55.737Z",
              "content": "<p>So much frustration … my model only has few parameters 🥲</p>",
              "rawMarkdown": "So much frustration ... my model only has few parameters 🥲"
            },
            {
              "id": 3019172,
              "postDate": "2024-10-16T10:31:19.327Z",
              "content": "<p>My score is similar to yours ,  -- 0.68/0.56, I guess it's due to other stars</p>",
              "rawMarkdown": "My score is similar to yours ,  -- 0.68/0.56, I guess it's due to other stars"
            },
            {
              "id": 3019196,
              "postDate": "2024-10-16T10:50:57.050Z",
              "content": "<p>I guess it is some tricks like what had happened in the comeptition below: <br>\n<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/leaderboard\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/leaderboard</a></p>",
              "rawMarkdown": "I guess it is some tricks like what had happened in the comeptition below: \nhttps://www.kaggle.com/competitions/uspto-explainable-ai/leaderboard"
            },
            {
              "id": 3019217,
              "postDate": "2024-10-16T11:14:02.753Z",
              "content": "<p>I don't think so. I tried to mix the submits in this way: for the stars 0 and 1 are the best in CV, for the rest the best in LB. And score gets worse than the best in public. It's about something else. </p>",
              "rawMarkdown": "I don't think so. I tried to mix the submits in this way: for the stars 0 and 1 are the best in CV, for the rest the best in LB. And score gets worse than the best in public. It's about something else. ",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3018504,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-10-15T19:50:44.697000",
      "content": "<p>0.641/0.675 (my 0.678 is LB overfitting; let's not consider it). Funny story: I accidentally used sigma_true=1e-6 instead of 1e-5 from the beginning, so all my scores are ~0.4, and my intuition is built around this range. I used the correct sigma only for this post.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3018507,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-10-15T19:56:03.903000",
          "content": "<p>I noticed that when you start chatting on the forum, you reveal your own mistakes. Just don't say too much. Thanks for sharing! It looks similar. Edit: I'm not sure if this works the same way in English as in Russian. I don't mean you specifically, but I mean in general. Someday I'll learn English too.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3018515,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-10-15T20:10:07.647000",
              "content": "<p>Hehe, it can be interpreted in both ways, but yeah, it sounds more like the first option. Well, English is not my first language, either, but I would probably go with, 'I noticed that when one starts chatting on the forum, one reveals his own mistakes.'</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3018519,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-15T20:14:21.863000",
              "content": "<p>Yes, it sounds good. Thanks)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3020126,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-10-17T07:19:03.830000",
      "content": "<p>Mine is 0.679/0.624. I've recently made several \"improvements\" which turned out to be pure overfitting, decreasing public scores 😪<br>\nMy assumption: train and test sets have very different spectra distributions (and by that I mean the \"shape\" of y for each planet). Whatever you do that improves specifically fitting the train distributions, is an overfit, and decreases test score.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3020132,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2024-10-17T07:28:06.323000",
          "content": "<p>It seems that fitting raw signal will not generalize to LB?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3020308,
              "author_name": "Natan Labarrère",
              "author_url": "",
              "post_date": "2024-10-17T11:42:32.263000",
              "content": "<p>What do you mean?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3019773,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2024-10-16T20:47:39.037000",
      "content": "<p><a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> what is your CV mse? Mine sits at around 4e-9 - 5e-9</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3019784,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-10-16T21:07:54.980000",
          "content": "<p>both about 2e-9. It's overfitting for sure.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3019833,
              "author_name": "Fritz Cremer",
              "author_url": "",
              "post_date": "2024-10-16T22:50:47.417000",
              "content": "<p>Wow okay, so it's not just my sigma estimation that is bad 😄</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3024340,
              "author_name": "ZHANG XINGHAN",
              "author_url": "",
              "post_date": "2024-10-21T14:34:16.780000",
              "content": "<p>That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3024372,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-10-21T15:10:14.563000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3024375,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-21T15:15:23.327000",
              "content": "<blockquote>\n  <p>That's amazing, I can only go around 4e-5~5e-5. btw my score is 0.63/0.565 (cv/lb)</p>\n</blockquote>\n<p>It seems to me that you are talking about Rmse, not MSE. My RMSE is approximately the same as yours.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3024942,
              "author_name": "ZHANG XINGHAN",
              "author_url": "",
              "post_date": "2024-10-22T07:15:39.140000",
              "content": "<p>Oh yeah, right, my mistake</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3019336,
      "author_name": "Dmitry Leontyev",
      "author_url": "",
      "post_date": "2024-10-16T13:27:08.687000",
      "content": "<p>How do you estimate training score? I use the following function and it returns my training score near 0.9 which has no correlation with lb.</p>\n<p><em>gll_score = competition_score(        \n            train_labels.copy(),\n            submission.copy().reset_index(), \n            train_labels.values.mean(),\n            train_labels.values.std(),\n            sigma_true=1e-5).</em><br>\nDo you use lower sigma_true?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3019376,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-10-16T14:23:15.680000",
          "content": "<p>Your code is correct. 0.9 is an unrealistically high score. There is probably a leakage in the training process or model is overfitted </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3019441,
              "author_name": "Dmitry Leontyev",
              "author_url": "",
              "post_date": "2024-10-16T15:19:40.490000",
              "content": "<p>Thanks, that seems a bit weird. I'm currently trying to develop your unsupervised search by polynomial approximation with more careful and stable transit zone handling. As a baseline I forked your great ariel_only_correlation notebook and got local score 0.897. Did you observe the same score when testing this idea? Looking at this discussion, still can't understand if this method may overfit so much somehow or I can't find a bug in my scoring :(</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3019464,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-16T15:49:09.913000",
              "content": "<p>As I can remember, score was 0.48+ with 1.6e-4 sigma. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3019667,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-10-16T18:33:30.603000",
              "content": "<p>Chiming in as well, I believe there's a mistake in the way you're applying your scoring formula. </p>\n<p>Assuming you make a perfect prediction of the mean and copy that across the 283 wavelengths (and Sergei's baseline does that, however without a perfect prediction of the mean, so you'd implicitly expect a lower score), your maximum score could be around 0.53 with a sigma of ~9e-5. Getting to around 0.9 would assume an almost perfect prediction of not only the mean Rp/Rs squared, but also of the distribution across the 283 wavelengths. </p>\n<p>A model that simply predicts the mean value for a planet and copies that across the other wavelengths essentially can't reach a score of 0.9.</p>\n<pre><code>train_target = pd.read_csv('/kaggle/input/ariel-data-challenge-/train_labels.csv')\ny_true = train_target.loc[:, train_target. != 'planet_id'].\nnaive_mean=y_true.(),\nnaive_sigma=y_true.(),\nsigma_true=\nsigma_pred=\n = []\n i  (len(train_target)):\n    .(.repeat(y_true[i].(),))\ny_pred = .()\nGLL_pred = .(scipy.stats.norm.logpdf(y_true, loc=y_pred, =sigma_pred))\nGLL_true = .(scipy.stats.norm.logpdf(y_true, loc=y_true, =sigma_true * .ones_like(y_true)))\nGLL_mean = .(scipy.stats.norm.logpdf(y_true, loc=naive_mean * .ones_like(y_true), =naive_sigma * .ones_like(y_true)))\n((GLL_pred-GLL_mean)/(GLL_true-GLL_mean))\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3020307,
          "author_name": "Oleh Kivernyk",
          "author_url": "",
          "post_date": "2024-10-17T11:41:40.913000",
          "content": "<p>I use the same scoring function and get 0.495 on the training data, whereas 0.567 is on the LB. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3018656,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-10-16T01:07:22.780000",
      "content": "<p>Are you using 1 star for training/1 star for validation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3018987,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-10-16T07:30:26.030000",
          "content": "<p>Does it correlate better with the LB?  No, just out-of-fold prediction in some sense. Would you mind to share your score?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3019046,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-10-16T08:00:02.230000",
              "content": "<p>I used random split 5fold as cv, ~0.69/0.56 😅. Idk what's wrong with my signal preprocessing …</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3019085,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-16T08:46:12.720000",
              "content": "<p>Yes, it looks like my second submission. I have no idea what in the test leads to this.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3019162,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-10-16T10:17:55.737000",
              "content": "<p>So much frustration … my model only has few parameters 🥲</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3019172,
              "author_name": "zihan",
              "author_url": "",
              "post_date": "2024-10-16T10:31:19.327000",
              "content": "<p>My score is similar to yours ,  -- 0.68/0.56, I guess it's due to other stars</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3019196,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-10-16T10:50:57.050000",
              "content": "<p>I guess it is some tricks like what had happened in the comeptition below: <br>\n<a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/leaderboard\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/leaderboard</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3019217,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-10-16T11:14:02.753000",
              "content": "<p>I don't think so. I tried to mix the submits in this way: for the stars 0 and 1 are the best in CV, for the rest the best in LB. And score gets worse than the best in public. It's about something else. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3018191": "Mine 0.660/0.685 and 0.781/0.679. Still have no idea what's wrong with the last one.",
    "3018504": "0.641/0.675 (my 0.678 is LB overfitting; let's not consider it). Funny story: I accidentally used sigma_true=1e-6 instead of 1e-5 from the beginning, so all my scores are ~0.4, and my intuition is built around this range. I used the correct sigma only for this post.",
    "3020126": "Mine is 0.679/0.624. I've recently made several \"improvements\" which turned out to be pure overfitting, decreasing public scores 😪\nMy assumption: train and test sets have very different spectra distributions (and by that I mean the \"shape\" of y for each planet). Whatever you do that improves specifically fitting the train distributions, is an overfit, and decreases test score.",
    "3019773": "@sergeifironov what is your CV mse? Mine sits at around 4e-9 - 5e-9",
    "3019336": "How do you estimate training score? I use the following function and it returns my training score near 0.9 which has no correlation with lb.\n\n*gll_score = competition_score(        \n            train_labels.copy(),\n            submission.copy().reset_index(), \n            train_labels.values.mean(),\n            train_labels.values.std(),\n            sigma_true=1e-5).*\nDo you use lower sigma_true?",
    "3018656": "Are you using 1 star for training/1 star for validation?"
  }
}