{
  "id": 540248,
  "title": "recipe sharing ",
  "url": "/competitions/ariel-data-challenge-2024/discussion/540248",
  "author_name": "Lien",
  "post_date": "2024-10-13T17:51:27.388000",
  "votes": 34,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Some checkpoints for people who are interested. Also would love to get inspirations from the comments.<br>\n(i) Sergei's starter kit can give you score ~536, outpacing half of the teams on the current leaderboard.<br>\n(ii) I see hundreds of teams at 545 Bronze Medal, which suggests there may be already some shared notebook you can directly use to get here.<br>\n(iii) To get 569, I figured out a way to extend the mean estimation. This score can get you #38 on the current leaderboard.<br>\n(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.<br>\n(v) A second method to improve sigma leads me to 619, ranked top20 on current leaderboard.</p>",
  "messages": [
    {
      "id": 3016421,
      "postDate": "2024-10-13T17:51:27.390Z",
      "content": "<p>Some checkpoints for people who are interested. Also would love to get inspirations from the comments.<br>\n(i) Sergei's starter kit can give you score ~536, outpacing half of the teams on the current leaderboard.<br>\n(ii) I see hundreds of teams at 545 Bronze Medal, which suggests there may be already some shared notebook you can directly use to get here.<br>\n(iii) To get 569, I figured out a way to extend the mean estimation. This score can get you #38 on the current leaderboard.<br>\n(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.<br>\n(v) A second method to improve sigma leads me to 619, ranked top20 on current leaderboard.</p>",
      "rawMarkdown": "Some checkpoints for people who are interested. Also would love to get inspirations from the comments.\n(i) Sergei's starter kit can give you score ~536, outpacing half of the teams on the current leaderboard.\n(ii) I see hundreds of teams at 545 Bronze Medal, which suggests there may be already some shared notebook you can directly use to get here.\n(iii) To get 569, I figured out a way to extend the mean estimation. This score can get you #38 on the current leaderboard.\n(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.\n(v) A second method to improve sigma leads me to 619, ranked top20 on current leaderboard.\n\n",
      "votes": 34
    },
    {
      "id": 3016768,
      "postDate": "2024-10-14T06:19:35.310Z",
      "content": "<blockquote>\n  <p>(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.</p>\n</blockquote>\n<p>My local CV by cnn can approach 0.65,  and sigma is my bottleneck, sad. However,  Whether sigma estimation to the private test set is also questionable.</p>",
      "rawMarkdown": ">(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.\n\nMy local CV by cnn can approach 0.65,  and sigma is my bottleneck, sad. However,  Whether sigma estimation to the private test set is also questionable.\n",
      "votes": 3,
      "replies": [
        {
          "id": 3016797,
          "postDate": "2024-10-14T07:09:23.487Z",
          "content": "<p>My local CV also have a great score, with sigma ~ 4e-5, but somehow the sigma in test dataset is ~ 1e-4</p>",
          "rawMarkdown": "My local CV also have a great score, with sigma ~ 4e-5, but somehow the sigma in test dataset is ~ 1e-4",
          "votes": 3,
          "replies": [
            {
              "id": 3030704,
              "postDate": "2024-10-28T19:32:23.870Z",
              "content": "<p>so how do you adjust for test data sigma</p>",
              "rawMarkdown": "so how do you adjust for test data sigma"
            }
          ]
        }
      ]
    },
    {
      "id": 3020935,
      "postDate": "2024-10-18T04:22:49.477Z",
      "content": "<p>Thanks for sharing. Do you mind if I ask what is your mse? Because the max score is highly dependent on what your mse is. Mine have a mse of 91 [ppm] and score=566, cheers.</p>",
      "rawMarkdown": "Thanks for sharing. Do you mind if I ask what is your mse? Because the max score is highly dependent on what your mse is. Mine have a mse of 91 [ppm] and score=566, cheers.",
      "votes": 4,
      "replies": [
        {
          "id": 3029146,
          "postDate": "2024-10-27T00:22:22.953Z",
          "content": "<p>mine is 70 ppm</p>",
          "rawMarkdown": "mine is 70 ppm",
          "votes": 1
        }
      ]
    },
    {
      "id": 3027507,
      "postDate": "2024-10-24T23:54:34.207Z",
      "content": "<p>How did you extend the mean, using a Ridge Regression or down the NN path? I am struggling with overfitting, unfortunately. My RMSE is just a bit over 90, yet any method that gets me below that overfits.</p>",
      "rawMarkdown": "How did you extend the mean, using a Ridge Regression or down the NN path? I am struggling with overfitting, unfortunately. My RMSE is just a bit over 90, yet any method that gets me below that overfits.",
      "votes": 1
    },
    {
      "id": 3016483,
      "postDate": "2024-10-13T19:58:53.523Z",
      "content": "<p>Nice! I have one question - did you use Sergei's solution as a base, or did you use some completely different approach? </p>",
      "rawMarkdown": "Nice! I have one question - did you use Sergei's solution as a base, or did you use some completely different approach? ",
      "votes": 2,
      "replies": [
        {
          "id": 3016717,
          "postDate": "2024-10-14T05:02:00.537Z",
          "content": "<p>Yes, Sergei's solution.</p>",
          "rawMarkdown": "Yes, Sergei's solution.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3017973,
      "postDate": "2024-10-15T12:25:21.973Z",
      "content": "<p>Hello :) By \"extend the mean\", you mean predict more wavelength than just a single mean value ?</p>",
      "rawMarkdown": "Hello :) By \"extend the mean\", you mean predict more wavelength than just a single mean value ?",
      "replies": [
        {
          "id": 3018573,
          "postDate": "2024-10-15T21:18:23.180Z",
          "content": "<p>Yes, that's what I meant. </p>",
          "rawMarkdown": "Yes, that's what I meant. ",
          "votes": 1,
          "replies": [
            {
              "id": 3022626,
              "postDate": "2024-10-19T19:06:47.697Z",
              "content": "<p>For predicting the mean value of transit depth (same for all wavelengths), was there any need to use model with parameters? <br>\nthank you Lien</p>",
              "rawMarkdown": "For predicting the mean value of transit depth (same for all wavelengths), was there any need to use model with parameters? \nthank you Lien",
              "votes": 1
            },
            {
              "id": 3029148,
              "postDate": "2024-10-27T00:32:18.183Z",
              "content": "<p>Hi, I think your prediction looks ok on training set. It may be just a problem of generalization. And not surprisingly the more parameters you have in the model, the higher change it gonna suffer overfitting. </p>",
              "rawMarkdown": "Hi, I think your prediction looks ok on training set. It may be just a problem of generalization. And not surprisingly the more parameters you have in the model, the higher change it gonna suffer overfitting. "
            }
          ]
        }
      ]
    },
    {
      "id": 3025956,
      "postDate": "2024-10-23T09:53:57.107Z",
      "content": "<p>I really look forward to see how you handled the sigma estimate ! </p>\n<p>On my side, the estimates I obtain from the training set do not match at all with what I could infer from the LB … </p>",
      "rawMarkdown": "I really look forward to see how you handled the sigma estimate ! \n\nOn my side, the estimates I obtain from the training set do not match at all with what I could infer from the LB ... ",
      "replies": [
        {
          "id": 3025994,
          "postDate": "2024-10-23T10:51:42.520Z",
          "content": "<p>Due to the difference between train and test set distributions, test set sigma will be higher than train estimate in any case.</p>",
          "rawMarkdown": "Due to the difference between train and test set distributions, test set sigma will be higher than train estimate in any case.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3018911,
      "postDate": "2024-10-16T05:56:18.953Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3016768,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-10-14T06:19:35.310000",
      "content": "<blockquote>\n  <p>(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.</p>\n</blockquote>\n<p>My local CV by cnn can approach 0.65,  and sigma is my bottleneck, sad. However,  Whether sigma estimation to the private test set is also questionable.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3016797,
          "author_name": "ChingYinNg",
          "author_url": "",
          "post_date": "2024-10-14T07:09:23.487000",
          "content": "<p>My local CV also have a great score, with sigma ~ 4e-5, but somehow the sigma in test dataset is ~ 1e-4</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3030704,
              "author_name": "Ajay Varghese",
              "author_url": "",
              "post_date": "2024-10-28T19:32:23.870000",
              "content": "<p>so how do you adjust for test data sigma</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3020935,
      "author_name": "Wei-Kuo Li",
      "author_url": "",
      "post_date": "2024-10-18T04:22:49.477000",
      "content": "<p>Thanks for sharing. Do you mind if I ask what is your mse? Because the max score is highly dependent on what your mse is. Mine have a mse of 91 [ppm] and score=566, cheers.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3029146,
          "author_name": "Lien",
          "author_url": "",
          "post_date": "2024-10-27T00:22:22.953000",
          "content": "<p>mine is 70 ppm</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3027507,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-10-24T23:54:34.207000",
      "content": "<p>How did you extend the mean, using a Ridge Regression or down the NN path? I am struggling with overfitting, unfortunately. My RMSE is just a bit over 90, yet any method that gets me below that overfits.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3016483,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2024-10-13T19:58:53.523000",
      "content": "<p>Nice! I have one question - did you use Sergei's solution as a base, or did you use some completely different approach? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3016717,
          "author_name": "Lien",
          "author_url": "",
          "post_date": "2024-10-14T05:02:00.537000",
          "content": "<p>Yes, Sergei's solution.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3017973,
      "author_name": "dafram2r",
      "author_url": "",
      "post_date": "2024-10-15T12:25:21.973000",
      "content": "<p>Hello :) By \"extend the mean\", you mean predict more wavelength than just a single mean value ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3018573,
          "author_name": "Lien",
          "author_url": "",
          "post_date": "2024-10-15T21:18:23.180000",
          "content": "<p>Yes, that's what I meant. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3022626,
              "author_name": "highDopamine",
              "author_url": "",
              "post_date": "2024-10-19T19:06:47.697000",
              "content": "<p>For predicting the mean value of transit depth (same for all wavelengths), was there any need to use model with parameters? <br>\nthank you Lien</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3029148,
              "author_name": "Lien",
              "author_url": "",
              "post_date": "2024-10-27T00:32:18.183000",
              "content": "<p>Hi, I think your prediction looks ok on training set. It may be just a problem of generalization. And not surprisingly the more parameters you have in the model, the higher change it gonna suffer overfitting. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3025956,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2024-10-23T09:53:57.107000",
      "content": "<p>I really look forward to see how you handled the sigma estimate ! </p>\n<p>On my side, the estimates I obtain from the training set do not match at all with what I could infer from the LB … </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3025994,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2024-10-23T10:51:42.520000",
          "content": "<p>Due to the difference between train and test set distributions, test set sigma will be higher than train estimate in any case.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3018911,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-16T05:56:18.953000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3016421": "Some checkpoints for people who are interested. Also would love to get inspirations from the comments.\n(i) Sergei's starter kit can give you score ~536, outpacing half of the teams on the current leaderboard.\n(ii) I see hundreds of teams at 545 Bronze Medal, which suggests there may be already some shared notebook you can directly use to get here.\n(iii) To get 569, I figured out a way to extend the mean estimation. This score can get you #38 on the current leaderboard.\n(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.\n(v) A second method to improve sigma leads me to 619, ranked top20 on current leaderboard.\n\n",
    "3016768": ">(iv) To get 602, I figured out a way to estimate sigma. This score can get you at #25 on the current leaderboard.\n\nMy local CV by cnn can approach 0.65,  and sigma is my bottleneck, sad. However,  Whether sigma estimation to the private test set is also questionable.\n",
    "3020935": "Thanks for sharing. Do you mind if I ask what is your mse? Because the max score is highly dependent on what your mse is. Mine have a mse of 91 [ppm] and score=566, cheers.",
    "3027507": "How did you extend the mean, using a Ridge Regression or down the NN path? I am struggling with overfitting, unfortunately. My RMSE is just a bit over 90, yet any method that gets me below that overfits.",
    "3016483": "Nice! I have one question - did you use Sergei's solution as a base, or did you use some completely different approach? ",
    "3017973": "Hello :) By \"extend the mean\", you mean predict more wavelength than just a single mean value ?",
    "3025956": "I really look forward to see how you handled the sigma estimate ! \n\nOn my side, the estimates I obtain from the training set do not match at all with what I could infer from the LB ... ",
    "3018911": ""
  }
}