{
  "id": 534376,
  "title": "What is the secret of top 10 LB places?",
  "url": "/competitions/ariel-data-challenge-2024/discussion/534376",
  "author_name": "",
  "post_date": "2024-09-16T10:52:39.280000",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>What is the secret to achieving the top 10 LB places?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20811297%2Fb25db0d72721d33ff2822c44bd4113a1%2Fr.png?generation=1726483662607168&amp;alt=media\" alt=\"\"></p>\n<p>Perhaps there is some information already in this discussion.<br>\nWho can review everything we know about competition?<br>\nIs there any important post list?</p>",
  "messages": [
    {
      "id": 2990550,
      "postDate": "2024-09-16T14:06:01.330Z",
      "content": "<p>Without giving anything away, my approach is along the lines of:</p>\n<ul>\n<li>Get a \"cleaner\" signal.</li>\n<li>Use optimization to get a best fit for depth.</li>\n<li>Get a theoretical estimate of sigma from your best fit.</li>\n<li>Model the noise given some basic domain knowledge. Get a different estimate of sigma.</li>\n<li>Find a consensus estimate of sigma from your two approaches.</li>\n</ul>\n<p>Then, the killer, reimplement everything in that pipeline to take advantage of this problem's structure to avoid bottlenecks in scipy/numpy. Otherwise, the test set would take ~200 hours to solve for.</p>\n<p>I think my method will top out at about 0.63 unless I have some sort of additional insight.</p>",
      "rawMarkdown": "Without giving anything away, my approach is along the lines of:\n- Get a \"cleaner\" signal.\n- Use optimization to get a best fit for depth.\n- Get a theoretical estimate of sigma from your best fit.\n- Model the noise given some basic domain knowledge. Get a different estimate of sigma.\n- Find a consensus estimate of sigma from your two approaches.\n\nThen, the killer, reimplement everything in that pipeline to take advantage of this problem's structure to avoid bottlenecks in scipy/numpy. Otherwise, the test set would take ~200 hours to solve for.\n\nI think my method will top out at about 0.63 unless I have some sort of additional insight.",
      "votes": 18,
      "replies": [
        {
          "id": 2990694,
          "postDate": "2024-09-16T16:09:58.353Z",
          "content": "<p>Thanks, how much \"cleaner\" signal improves your score? Or what do you mean by \"cleaner\" signal? As for me, no denoising algorithm has done well.</p>",
          "rawMarkdown": "Thanks, how much \"cleaner\" signal improves your score? Or what do you mean by \"cleaner\" signal? As for me, no denoising algorithm has done well.",
          "votes": 6,
          "replies": [
            {
              "id": 2991013,
              "postDate": "2024-09-16T23:28:26.477Z",
              "content": "<p>I've tried close to a dozen approaches so far. There are small gains from some of them. The reason I put \"cleaner\" in quotes is, in some sense, simply averaging all the wavelengths provides you with a cleaner signal.</p>",
              "rawMarkdown": "I've tried close to a dozen approaches so far. There are small gains from some of them. The reason I put \"cleaner\" in quotes is, in some sense, simply averaging all the wavelengths provides you with a cleaner signal.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2990378,
      "postDate": "2024-09-16T10:52:39.280Z",
      "content": "<p>What is the secret to achieving the top 10 LB places?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20811297%2Fb25db0d72721d33ff2822c44bd4113a1%2Fr.png?generation=1726483662607168&amp;alt=media\" alt=\"\"></p>\n<p>Perhaps there is some information already in this discussion.<br>\nWho can review everything we know about competition?<br>\nIs there any important post list?</p>",
      "rawMarkdown": "What is the secret to achieving the top 10 LB places?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20811297%2Fb25db0d72721d33ff2822c44bd4113a1%2Fr.png?generation=1726483662607168&alt=media)\n\nPerhaps there is some information already in this discussion.\nWho can review everything we know about competition?\nIs there any important post list?",
      "votes": 12
    },
    {
      "id": 2990498,
      "postDate": "2024-09-16T12:50:02.527Z",
      "content": "<p>I'm half hoping no one will reveal any 'magic' so I can stay in top 5, half hoping that someone will reveal exactly the 'magic' that I'm missing to get to 0.65+ range 🤣</p>",
      "rawMarkdown": "I'm half hoping no one will reveal any 'magic' so I can stay in top 5, half hoping that someone will reveal exactly the 'magic' that I'm missing to get to 0.65+ range 🤣",
      "votes": 7
    },
    {
      "id": 2990478,
      "postDate": "2024-09-16T12:22:28.070Z",
      "content": "<p>I'm pretty sure it has to do with a good estimation of the uncertainty part.</p>",
      "rawMarkdown": "I'm pretty sure it has to do with a good estimation of the uncertainty part.",
      "votes": 3,
      "replies": [
        {
          "id": 2991112,
          "postDate": "2024-09-17T04:15:50.533Z",
          "content": "<p>Yep. A very rough estimate of sigma boosted my score by 0.03</p>",
          "rawMarkdown": "Yep. A very rough estimate of sigma boosted my score by 0.03",
          "votes": 3
        }
      ]
    },
    {
      "id": 2990398,
      "postDate": "2024-09-16T11:01:46.637Z",
      "content": "<p><a href=\"https://www.kaggle.com/arabidopsisthalian\" target=\"_blank\">@arabidopsisthalian</a> we are also struggling to find the magic. at present -&gt; training  is not stability -&gt; sigma in inference is our team struggle.</p>",
      "rawMarkdown": "@arabidopsisthalian we are also struggling to find the magic. at present -> training  is not stability -> sigma in inference is our team struggle.",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 2990550,
      "author_name": "Andrew Matteson",
      "author_url": "",
      "post_date": "2024-09-16T14:06:01.330000",
      "content": "<p>Without giving anything away, my approach is along the lines of:</p>\n<ul>\n<li>Get a \"cleaner\" signal.</li>\n<li>Use optimization to get a best fit for depth.</li>\n<li>Get a theoretical estimate of sigma from your best fit.</li>\n<li>Model the noise given some basic domain knowledge. Get a different estimate of sigma.</li>\n<li>Find a consensus estimate of sigma from your two approaches.</li>\n</ul>\n<p>Then, the killer, reimplement everything in that pipeline to take advantage of this problem's structure to avoid bottlenecks in scipy/numpy. Otherwise, the test set would take ~200 hours to solve for.</p>\n<p>I think my method will top out at about 0.63 unless I have some sort of additional insight.</p>",
      "votes": 18,
      "replies": [
        {
          "id": 2990694,
          "author_name": "Georgii Aparin",
          "author_url": "",
          "post_date": "2024-09-16T16:09:58.353000",
          "content": "<p>Thanks, how much \"cleaner\" signal improves your score? Or what do you mean by \"cleaner\" signal? As for me, no denoising algorithm has done well.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2991013,
              "author_name": "Andrew Matteson",
              "author_url": "",
              "post_date": "2024-09-16T23:28:26.477000",
              "content": "<p>I've tried close to a dozen approaches so far. There are small gains from some of them. The reason I put \"cleaner\" in quotes is, in some sense, simply averaging all the wavelengths provides you with a cleaner signal.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2990498,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-09-16T12:50:02.527000",
      "content": "<p>I'm half hoping no one will reveal any 'magic' so I can stay in top 5, half hoping that someone will reveal exactly the 'magic' that I'm missing to get to 0.65+ range 🤣</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2990478,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-09-16T12:22:28.070000",
      "content": "<p>I'm pretty sure it has to do with a good estimation of the uncertainty part.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2991112,
          "author_name": "ChingYinNg",
          "author_url": "",
          "post_date": "2024-09-17T04:15:50.533000",
          "content": "<p>Yep. A very rough estimate of sigma boosted my score by 0.03</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2990398,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-09-16T11:01:46.637000",
      "content": "<p><a href=\"https://www.kaggle.com/arabidopsisthalian\" target=\"_blank\">@arabidopsisthalian</a> we are also struggling to find the magic. at present -&gt; training  is not stability -&gt; sigma in inference is our team struggle.</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2990550": "Without giving anything away, my approach is along the lines of:\n- Get a \"cleaner\" signal.\n- Use optimization to get a best fit for depth.\n- Get a theoretical estimate of sigma from your best fit.\n- Model the noise given some basic domain knowledge. Get a different estimate of sigma.\n- Find a consensus estimate of sigma from your two approaches.\n\nThen, the killer, reimplement everything in that pipeline to take advantage of this problem's structure to avoid bottlenecks in scipy/numpy. Otherwise, the test set would take ~200 hours to solve for.\n\nI think my method will top out at about 0.63 unless I have some sort of additional insight.",
    "2990378": "What is the secret to achieving the top 10 LB places?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20811297%2Fb25db0d72721d33ff2822c44bd4113a1%2Fr.png?generation=1726483662607168&alt=media)\n\nPerhaps there is some information already in this discussion.\nWho can review everything we know about competition?\nIs there any important post list?",
    "2990498": "I'm half hoping no one will reveal any 'magic' so I can stay in top 5, half hoping that someone will reveal exactly the 'magic' that I'm missing to get to 0.65+ range 🤣",
    "2990478": "I'm pretty sure it has to do with a good estimation of the uncertainty part.",
    "2990398": "@arabidopsisthalian we are also struggling to find the magic. at present -> training  is not stability -> sigma in inference is our team struggle."
  }
}