{
  "id": 686651,
  "title": "The moving target that rendered RNA 3D Pt2 competition a huge JOKE!!! - 2nd update",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/686651",
  "author_name": "YaGana Sheriff-Hussaini",
  "post_date": "2026-03-31T22:44:45.761000",
  "votes": 15,
  "comment_count": 37,
  "views": 0,
  "content": "<p>The following are the scores for my final submission selections:-</p>\n<p>That 1st rerun <strong>(where V21 public LB score =0.4349, private LB score =0.52614 and V13 public LB score =0.42561, private LB score =0.52823</strong>) confirms my engineering was actually far more successful than the \"Final\" leaderboard suggests. Those scores indicate that on the original hidden test targets, my models were performing significantly better than they did on the <strong>public set (0.4349/0.42561)</strong>. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F898111%2Fded026e1fe9a62c418f47b29097d0c49%2FSubmission%202026-03-31%20203407%20-%20Copy.png?generation=1775014889854227&amp;alt=media\" alt=\"\"></p>\n<p>The fact that the organizers discarded that entire set of results <strong>today (03/31/2026)</strong> and replaced them with the \"<strong>Tougher</strong>\" <strong>final leaderboard</strong> is exactly why the community is in an uproar. This was all but admitted <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601\" target=\"_blank\">here</a>.</p>\n<p><strong>The \"Decoy\" Reveal</strong>: By relabeling the scores today, the organizers essentially <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601\" target=\"_blank\">admitted</a> they \"tested the testers.\" They saw high scores, decided the test was too easy, and swapped in a much more punishing set at the 11th hour.</p>\n<p>Competitors are furious because this breaks <strong>the fundamental contract of Kaggle</strong>: that you are building toward a consistent, albeit hidden, target. Swapping targets after seeing the rerun performance feels less like a competition and more like a post-hoc adjustment to lower the overall TM-scores of the field.</p>\n<p><strong>The fact that the best_template_oracle score dropped from 0.55488 to 0.50331 on the Private LB is the \"smoking gun.\"</strong>\nThe Oracle represents the theoretical ceiling of the dataset—the best possible score if you had perfect structural templates. <strong>A 0.05157 drop in the Oracle score proves</strong> the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.</p>\n<p>Scientific progress requires stable benchmarks. If the organizers can’t provide a consistent evaluation environment, then the <strong>\"CASP17 dry run\"</strong> is just another moving target IMHO.</p>\n<p><strong>UPDATE:-</strong></p>\n<p>For the record, my models are deterministic i.e. for the same settings gives the same result every time.</p>\n<p>I added this statement because of <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>'s response to a post that asked <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938\" target=\"_blank\">Evidence suggests that the rerun took place twice</a>. He pointed at models of nondeterminism stating <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938#3433469\" target=\"_blank\">Even when the Public observations are exactly the same, models that exhibit non-determinism may have different scores.</a>. Anyone who has participated in competitions in Kaggle before knows not to chase the public LB scores in a competition. I did not and I am sure many serious competitors would avoid doing it. </p>\n<p><strong>Here is how I see it:</strong> If a deterministic model drops from <strong>0.4349 to 0.40722</strong> on a public LB, it is a direct measurement of the difficulty spike in the new dataset. <strong>Take heart everyone in the fact that our models didn't \"randomly fail\"; they were simply evaluated against a fundamentally different, harder set of structures.</strong></p>\n<p><strong>The Transparency Paradox</strong>. Kaggle’s history with transparency: For years, Kaggle has maintained its reputation by being open about rerun logic and distribution shifts. This \"unique lack of transparency\"—waiting until the final reveal to admit they were using \"Live-Solved\" cryo-EM data—is a significant departure from those standards. </p>\n<p><strong>UPDATE-2:-</strong></p>\n<p>One of competition organizers i.e. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  posted a detailed explanation for what happened pointing to it in his comments below. So, you can post any further questions here but preferably in the thread he had the detailed explanation on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">here</a>.</p>\n<p>Thanks everyone for your attention to this matter. Happy kaggling!!!</p>",
  "messages": [
    {
      "id": 3432755,
      "postDate": "2026-03-31T22:44:45.760Z",
      "content": "<p>The following are the scores for my final submission selections:-</p>\n<p>That 1st rerun <strong>(where V21 public LB score =0.4349, private LB score =0.52614 and V13 public LB score =0.42561, private LB score =0.52823</strong>) confirms my engineering was actually far more successful than the \"Final\" leaderboard suggests. Those scores indicate that on the original hidden test targets, my models were performing significantly better than they did on the <strong>public set (0.4349/0.42561)</strong>. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F898111%2Fded026e1fe9a62c418f47b29097d0c49%2FSubmission%202026-03-31%20203407%20-%20Copy.png?generation=1775014889854227&amp;alt=media\" alt=\"\"></p>\n<p>The fact that the organizers discarded that entire set of results <strong>today (03/31/2026)</strong> and replaced them with the \"<strong>Tougher</strong>\" <strong>final leaderboard</strong> is exactly why the community is in an uproar. This was all but admitted <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601\" target=\"_blank\">here</a>.</p>\n<p><strong>The \"Decoy\" Reveal</strong>: By relabeling the scores today, the organizers essentially <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601\" target=\"_blank\">admitted</a> they \"tested the testers.\" They saw high scores, decided the test was too easy, and swapped in a much more punishing set at the 11th hour.</p>\n<p>Competitors are furious because this breaks <strong>the fundamental contract of Kaggle</strong>: that you are building toward a consistent, albeit hidden, target. Swapping targets after seeing the rerun performance feels less like a competition and more like a post-hoc adjustment to lower the overall TM-scores of the field.</p>\n<p><strong>The fact that the best_template_oracle score dropped from 0.55488 to 0.50331 on the Private LB is the \"smoking gun.\"</strong>\nThe Oracle represents the theoretical ceiling of the dataset—the best possible score if you had perfect structural templates. <strong>A 0.05157 drop in the Oracle score proves</strong> the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.</p>\n<p>Scientific progress requires stable benchmarks. If the organizers can’t provide a consistent evaluation environment, then the <strong>\"CASP17 dry run\"</strong> is just another moving target IMHO.</p>\n<p><strong>UPDATE:-</strong></p>\n<p>For the record, my models are deterministic i.e. for the same settings gives the same result every time.</p>\n<p>I added this statement because of <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>'s response to a post that asked <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938\" target=\"_blank\">Evidence suggests that the rerun took place twice</a>. He pointed at models of nondeterminism stating <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938#3433469\" target=\"_blank\">Even when the Public observations are exactly the same, models that exhibit non-determinism may have different scores.</a>. Anyone who has participated in competitions in Kaggle before knows not to chase the public LB scores in a competition. I did not and I am sure many serious competitors would avoid doing it. </p>\n<p><strong>Here is how I see it:</strong> If a deterministic model drops from <strong>0.4349 to 0.40722</strong> on a public LB, it is a direct measurement of the difficulty spike in the new dataset. <strong>Take heart everyone in the fact that our models didn't \"randomly fail\"; they were simply evaluated against a fundamentally different, harder set of structures.</strong></p>\n<p><strong>The Transparency Paradox</strong>. Kaggle’s history with transparency: For years, Kaggle has maintained its reputation by being open about rerun logic and distribution shifts. This \"unique lack of transparency\"—waiting until the final reveal to admit they were using \"Live-Solved\" cryo-EM data—is a significant departure from those standards. </p>\n<p><strong>UPDATE-2:-</strong></p>\n<p>One of competition organizers i.e. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  posted a detailed explanation for what happened pointing to it in his comments below. So, you can post any further questions here but preferably in the thread he had the detailed explanation on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">here</a>.</p>\n<p>Thanks everyone for your attention to this matter. Happy kaggling!!!</p>",
      "rawMarkdown": "The following are the scores for my final submission selections:-\n\nThat 1st rerun **(where V21 public LB score =0.4349, private LB score =0.52614 and V13 public LB score =0.42561, private LB score =0.52823**) confirms my engineering was actually far more successful than the \"Final\" leaderboard suggests. Those scores indicate that on the original hidden test targets, my models were performing significantly better than they did on the **public set (0.4349/0.42561)**. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F898111%2Fded026e1fe9a62c418f47b29097d0c49%2FSubmission%202026-03-31%20203407%20-%20Copy.png?generation=1775014889854227&alt=media)\n\nThe fact that the organizers discarded that entire set of results **today (03/31/2026)** and replaced them with the \"**Tougher**\" **final leaderboard** is exactly why the community is in an uproar. This was all but admitted [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601).\n\n**The \"Decoy\" Reveal**: By relabeling the scores today, the organizers essentially [admitted](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601) they \"tested the testers.\" They saw high scores, decided the test was too easy, and swapped in a much more punishing set at the 11th hour.\n\nCompetitors are furious because this breaks **the fundamental contract of Kaggle**: that you are building toward a consistent, albeit hidden, target. Swapping targets after seeing the rerun performance feels less like a competition and more like a post-hoc adjustment to lower the overall TM-scores of the field.\n\n**The fact that the best_template_oracle score dropped from 0.55488 to 0.50331 on the Private LB is the \"smoking gun.\"**\nThe Oracle represents the theoretical ceiling of the dataset—the best possible score if you had perfect structural templates. **A 0.05157 drop in the Oracle score proves** the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.\n\nScientific progress requires stable benchmarks. If the organizers can’t provide a consistent evaluation environment, then the **\"CASP17 dry run\"** is just another moving target IMHO.\n\n**UPDATE:-**\n\nFor the record, my models are deterministic i.e. for the same settings gives the same result every time.\n\nI added this statement because of @inversion's response to a post that asked [Evidence suggests that the rerun took place twice](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938). He pointed at models of nondeterminism stating [Even when the Public observations are exactly the same, models that exhibit non-determinism may have different scores.](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938#3433469). Anyone who has participated in competitions in Kaggle before knows not to chase the public LB scores in a competition. I did not and I am sure many serious competitors would avoid doing it. \n\n**Here is how I see it:** If a deterministic model drops from **0.4349 to 0.40722** on a public LB, it is a direct measurement of the difficulty spike in the new dataset. **Take heart everyone in the fact that our models didn't \"randomly fail\"; they were simply evaluated against a fundamentally different, harder set of structures.**\n\n**The Transparency Paradox**. Kaggle’s history with transparency: For years, Kaggle has maintained its reputation by being open about rerun logic and distribution shifts. This \"unique lack of transparency\"—waiting until the final reveal to admit they were using \"Live-Solved\" cryo-EM data—is a significant departure from those standards. \n\n**UPDATE-2:-**\n\nOne of competition organizers i.e. @rhijudas  posted a detailed explanation for what happened pointing to it in his comments below. So, you can post any further questions here but preferably in the thread he had the detailed explanation on [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646).\n\nThanks everyone for your attention to this matter. Happy kaggling!!!\n\n\n",
      "votes": 15
    },
    {
      "id": 3433033,
      "postDate": "2026-04-01T05:28:51.873Z",
      "content": "<p>One of the things that made people weary was after the competition ended, they cleared the public LB, and then when it was revealed, the public LB has changed. In past competitions, public LB is public LB unless if we are dealing with a competition that re-runs our models on new data for a period of time like 3 months. Even in such cases, the initial public LB do not change due to reruns. I just don't get it.</p>",
      "rawMarkdown": "One of the things that made people weary was after the competition ended, they cleared the public LB, and then when it was revealed, the public LB has changed. In past competitions, public LB is public LB unless if we are dealing with a competition that re-runs our models on new data for a period of time like 3 months. Even in such cases, the initial public LB do not change due to reruns. I just don't get it.",
      "votes": 5
    },
    {
      "id": 3432828,
      "postDate": "2026-04-01T00:05:14.547Z",
      "content": "<p>stand   your  side,feel  disappointed  too</p>",
      "rawMarkdown": "stand   your  side,feel  disappointed  too",
      "votes": 6
    },
    {
      "id": 3432961,
      "postDate": "2026-04-01T03:38:18.037Z",
      "content": "<p>Below is my view. It’s true that it was shakier than the previous competition, but I didn’t feel that there was any problem with the evaluation process itself.</p>\n<ul>\n<li>It was stated on the Overview page that additional data for the private test set would be collected during the competition and that submissions would be rerun on that data.</li>\n<li>What you are calling the “1st rerun” was simply the release of the dummy private test data that had been in place since the beginning of the competition.<ul>\n<li>This was not rerun; rather, the scores that had already been generated and evaluated together with the public test data from the beginning of the competition were simply made public.</li></ul></li>\n<li>Only the final two notebooks selected by each participant were rerun on the true private test data (as well as the public test data), which is what produced the current private leaderboard.</li>\n</ul>",
      "rawMarkdown": "Below is my view. It’s true that it was shakier than the previous competition, but I didn’t feel that there was any problem with the evaluation process itself.\n\n- It was stated on the Overview page that additional data for the private test set would be collected during the competition and that submissions would be rerun on that data.\n- What you are calling the “1st rerun” was simply the release of the dummy private test data that had been in place since the beginning of the competition.\n  - This was not rerun; rather, the scores that had already been generated and evaluated together with the public test data from the beginning of the competition were simply made public.\n- Only the final two notebooks selected by each participant were rerun on the true private test data (as well as the public test data), which is what produced the current private leaderboard.",
      "votes": 4,
      "replies": [
        {
          "id": 3432966,
          "postDate": "2026-04-01T03:48:15.313Z",
          "content": "<p>no, the host clearly stated that the test set is \"being collected\" <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674343#3408163\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/679588#3416707\" target=\"_blank\">here</a> so there shouldn't be any dummy private test. So the logic is, before submission deadline, there is no such thing as private test data. This is my view.</p>",
          "rawMarkdown": "no, the host clearly stated that the test set is \"being collected\" [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674343#3408163) and [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/679588#3416707) so there shouldn't be any dummy private test. So the logic is, before submission deadline, there is no such thing as private test data. This is my view.",
          "votes": 3,
          "replies": [
            {
              "id": 3432987,
              "postDate": "2026-04-01T04:31:36.553Z",
              "content": "<p>When I said “dummy private test,” I did not mean the true private test set that was being collected during the competition. I meant a placeholder split that appears to have existed from the beginning and was evaluated together with the public test data, with those scores being revealed later. I called it dummy data because it was included due to Kaggle system constraints and was not intended to be part of the final evaluation for this competition.</p>\n<p>So my understanding is that the “1st rerun” was not a rerun on the true private test set, but rather the release of scores based on that placeholder data. The actual rerun happened only for the final two selected notebooks using the true private test set collected during the competition, which is what produced the current private leaderboard.</p>",
              "rawMarkdown": "When I said “dummy private test,” I did not mean the true private test set that was being collected during the competition. I meant a placeholder split that appears to have existed from the beginning and was evaluated together with the public test data, with those scores being revealed later. I called it dummy data because it was included due to Kaggle system constraints and was not intended to be part of the final evaluation for this competition.\n\nSo my understanding is that the “1st rerun” was not a rerun on the true private test set, but rather the release of scores based on that placeholder data. The actual rerun happened only for the final two selected notebooks using the true private test set collected during the competition, which is what produced the current private leaderboard.",
              "votes": 1
            },
            {
              "id": 3432995,
              "postDate": "2026-04-01T04:44:47.013Z",
              "content": "<p>I understand your point. But can you explain why they have that in the first place. There are competitions like the NFL one and some quant time-series forecasting comps that have only the public LB and the rerun. There is no \"placeholder\" private test set for those. so it seems fine to not have it. Also, if they already have that set, why not just combine it with the public test set? Why do they clear the LB out of no reason during the rerun?</p>",
              "rawMarkdown": "I understand your point. But can you explain why they have that in the first place. There are competitions like the NFL one and some quant time-series forecasting comps that have only the public LB and the rerun. There is no \"placeholder\" private test set for those. so it seems fine to not have it. Also, if they already have that set, why not just combine it with the public test set? Why do they clear the LB out of no reason during the rerun?",
              "votes": 2
            },
            {
              "id": 3433004,
              "postDate": "2026-04-01T04:51:23.283Z",
              "content": "<p>Exactly <a href=\"https://www.kaggle.com/honganzhu\" target=\"_blank\">@honganzhu</a>. Clearing the public LB was the 1st thing that made me feel like something was off. I asked why in that thread and no one answered.</p>",
              "rawMarkdown": "Exactly @honganzhu. Clearing the public LB was the 1st thing that made me feel like something was off. I asked why in that thread and no one answered.",
              "votes": 1
            },
            {
              "id": 3433014,
              "postDate": "2026-04-01T04:56:17.720Z",
              "content": "<p>I’m not sure why this competition was structured this way.\nRerunning all submissions would have been the clearest approach, but perhaps cost was a concern.</p>",
              "rawMarkdown": "I’m not sure why this competition was structured this way.\nRerunning all submissions would have been the clearest approach, but perhaps cost was a concern."
            },
            {
              "id": 3433024,
              "postDate": "2026-04-01T05:09:20.710Z",
              "content": "<p>why is cost a concern here? they just need to rerun the 2 selected submissions. That is true no matter they have this fishy \"placeholder test set\" or not. The real question is why they have it in the first place, and why cleared the public LB during rerun, why not leaving the public LB there and only update the private LB.</p>",
              "rawMarkdown": "why is cost a concern here? they just need to rerun the 2 selected submissions. That is true no matter they have this fishy \"placeholder test set\" or not. The real question is why they have it in the first place, and why cleared the public LB during rerun, why not leaving the public LB there and only update the private LB.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3432976,
          "postDate": "2026-04-01T04:09:43.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> , I respectfully disagree.  If it were a test set with a similar distribution, the best_template_oracle score would not have dropped from 0.554 to 0.503 </p>",
          "rawMarkdown": "@ren4yu , I respectfully disagree.  If it were a test set with a similar distribution, the best_template_oracle score would not have dropped from 0.554 to 0.503 ",
          "votes": 3,
          "replies": [
            {
              "id": 3432994,
              "postDate": "2026-04-01T04:44:31.223Z",
              "content": "<p>I have no objection to the fact that the test data distribution changed between the public and true private sets.\nThis is something that occasionally happens in Kaggle competitions, and I think we could have expected this in this competition as well.\nIt even seems that some of the top teams were able to prepare for it, so I’m looking forward to learning how they validated their approaches.</p>\n<p>I simply wanted to point out that your claim - that the host replaced the data after seeing the private test results - is a misunderstanding.</p>\n<blockquote>\n  <p>A 0.051 drop in the Oracle score proves the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.</p>\n</blockquote>",
              "rawMarkdown": "I have no objection to the fact that the test data distribution changed between the public and true private sets.\nThis is something that occasionally happens in Kaggle competitions, and I think we could have expected this in this competition as well.\nIt even seems that some of the top teams were able to prepare for it, so I’m looking forward to learning how they validated their approaches.\n\nI simply wanted to point out that your claim - that the host replaced the data after seeing the private test results - is a misunderstanding.\n\n> A 0.051 drop in the Oracle score proves the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs."
            },
            {
              "id": 3433000,
              "postDate": "2026-04-01T04:48:08.530Z",
              "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, fair enough. I truely hope that it was just a misunderstanding.</p>",
              "rawMarkdown": "@ren4yu, fair enough. I truely hope that it was just a misunderstanding."
            }
          ]
        }
      ]
    },
    {
      "id": 3432829,
      "postDate": "2026-04-01T00:06:21.527Z",
      "content": "<p>It's normal for the victors to oppose us</p>",
      "rawMarkdown": "It's normal for the victors to oppose us",
      "votes": 3
    },
    {
      "id": 3432859,
      "postDate": "2026-04-01T00:37:54.273Z",
      "content": "<p>Agree, felt like a waste of time</p>",
      "rawMarkdown": "Agree, felt like a waste of time",
      "votes": 4
    },
    {
      "id": 3436341,
      "postDate": "2026-04-05T18:30:51.763Z",
      "content": "<p>This looks like a clear dataset shift. transparency matters as much as difficulty in a fair benchmark.</p>",
      "rawMarkdown": "This looks like a clear dataset shift. transparency matters as much as difficulty in a fair benchmark.",
      "votes": 1
    },
    {
      "id": 3433008,
      "postDate": "2026-04-01T04:54:17.523Z",
      "content": "<p>see this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2Fb0ba35d21984f5ff38673da08b89b41a%2F06_public_vs_private_scatter_trend.png?generation=1775019255601681&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "see this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2Fb0ba35d21984f5ff38673da08b89b41a%2F06_public_vs_private_scatter_trend.png?generation=1775019255601681&alt=media)",
      "votes": 1
    },
    {
      "id": 3432980,
      "postDate": "2026-04-01T04:13:09.057Z",
      "content": "<p>On the consistency between public and private leaderboards</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2F51ff82ef6016db5eb3c533de21056897%2F23333ScreenShot_2026-04-01_120954_888.png?generation=1775016787351313&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "On the consistency between public and private leaderboards\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2F51ff82ef6016db5eb3c533de21056897%2F23333ScreenShot_2026-04-01_120954_888.png?generation=1775016787351313&alt=media)",
      "votes": 1
    },
    {
      "id": 3433558,
      "postDate": "2026-04-01T20:06:08.340Z",
      "content": "<p>For the record, my models are deterministic i.e. for the same settings gives the same result every time you run them.</p>",
      "rawMarkdown": "For the record, my models are deterministic i.e. for the same settings gives the same result every time you run them.",
      "votes": 2,
      "replies": [
        {
          "id": 3433809,
          "postDate": "2026-04-02T06:55:35.943Z",
          "content": "<p>I’m facing a similar issue. My models were fully deterministic—I ran them three times and got identical results—but they still underperformed. I dropped from 18th on the public leaderboard to around 700th on the private, which was really unexpected.\nMy approach was similar to the top notebooks: I fine-tuned RNA-Pro for harder targets, used ProteinX with multi-seed ensembling (around 20 seeds), and then combined both for the hardest targets. This setup was giving consistently strong results during validation.\nI’m not sure what went wrong—any solid insights would be really helpful. Also, is it possible that the private leaderboard is completely non-identical to the public one?</p>",
          "rawMarkdown": "I’m facing a similar issue. My models were fully deterministic—I ran them three times and got identical results—but they still underperformed. I dropped from 18th on the public leaderboard to around 700th on the private, which was really unexpected.\nMy approach was similar to the top notebooks: I fine-tuned RNA-Pro for harder targets, used ProteinX with multi-seed ensembling (around 20 seeds), and then combined both for the hardest targets. This setup was giving consistently strong results during validation.\nI’m not sure what went wrong—any solid insights would be really helpful. Also, is it possible that the private leaderboard is completely non-identical to the public one?",
          "votes": 3,
          "replies": [
            {
              "id": 3434074,
              "postDate": "2026-04-02T14:51:03.767Z",
              "content": "<p>Yes <a href=\"https://www.kaggle.com/alisalmanrana\" target=\"_blank\">@alisalmanrana</a>, on the private LB test data is completely different. Which is the point of my post. However small, we should be given a good representative of their private test data so that we can develop very good models for them. Changing it after the fact for whatever reason is where the transparency issue came in to question.</p>",
              "rawMarkdown": "Yes @alisalmanrana, on the private LB test data is completely different. Which is the point of my post. However small, we should be given a good representative of their private test data so that we can develop very good models for them. Changing it after the fact for whatever reason is where the transparency issue came in to question.",
              "votes": 1
            },
            {
              "id": 3434157,
              "postDate": "2026-04-02T16:50:19.583Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> really appreciate the clarification! That explains the drop, though it does feel like a lot of the effort ended up being less aligned with the final evaluation than expected. Still, definitely a valuable learning experience—just wish the setup was a bit more representative.</p>",
              "rawMarkdown": "Thanks @sheriytm really appreciate the clarification! That explains the drop, though it does feel like a lot of the effort ended up being less aligned with the final evaluation than expected. Still, definitely a valuable learning experience—just wish the setup was a bit more representative.",
              "votes": 1
            },
            {
              "id": 3434221,
              "postDate": "2026-04-02T18:12:50.563Z",
              "content": "<p>True <a href=\"https://www.kaggle.com/alisalmanrana\" target=\"_blank\">@alisalmanrana</a>. In spite of everything that happened, I am actually very happy with the models I developed for this competition. I learnt a lot about a domain I was not familiar with prior to this contest. </p>\n<p>Using the public kernels as inspiration, I implemented my own version of splitting and stitching the long targets that gave very good results. I also implemented a hydration and shape-sync which i called armored plates that took care of the various OOMs I was getting. So, yeah I spent a lot of time on this and learnt so much.</p>",
              "rawMarkdown": "True @alisalmanrana. In spite of everything that happened, I am actually very happy with the models I developed for this competition. I learnt a lot about a domain I was not familiar with prior to this contest. \n\nUsing the public kernels as inspiration, I implemented my own version of splitting and stitching the long targets that gave very good results. I also implemented a hydration and shape-sync which i called armored plates that took care of the various OOMs I was getting. So, yeah I spent a lot of time on this and learnt so much."
            },
            {
              "id": 3434251,
              "postDate": "2026-04-02T18:50:46.960Z",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>, totally feel you on this one! I’ve participated in quite a few competitions, but this one was by far the toughest in terms of management. It would’ve been really helpful if they had given us a proper heads-up about the tougher private targets beforehand—everything else made it seem like the private set would be similar to the public leaderboard. After putting in so much effort and seeing results drop so dramatically, that’s definitely the most frustrating part. Still, despite all that, it was a huge learning experience, and I picked up a ton from your post and the kernels out there—really appreciate you sharing your approach!</p>",
              "rawMarkdown": "Hey @sheriytm, totally feel you on this one! I’ve participated in quite a few competitions, but this one was by far the toughest in terms of management. It would’ve been really helpful if they had given us a proper heads-up about the tougher private targets beforehand—everything else made it seem like the private set would be similar to the public leaderboard. After putting in so much effort and seeing results drop so dramatically, that’s definitely the most frustrating part. Still, despite all that, it was a huge learning experience, and I picked up a ton from your post and the kernels out there—really appreciate you sharing your approach!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3433031,
      "postDate": "2026-04-01T05:25:03.643Z",
      "content": "<p>this comp felt like a waste of time.</p>",
      "rawMarkdown": "this comp felt like a waste of time.",
      "votes": 2
    },
    {
      "id": 3435003,
      "postDate": "2026-04-03T15:14:06.087Z",
      "content": "<p>Hi everyone, thanks for the discussion. </p>\n<p>The hosts and devs could have done better in anticipating the confusion that would come from having a placeholder Private Leaderboard that was changed for final scoring. </p>\n<p>We also should have been more responsive to the questions here.</p>\n<p>I've posted a brief explanation of what happened -- and what threw us hosts off, and what to avoid in future competitions -- in response to this other post from <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>:</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646</a></p>\n<blockquote>\n  <p><strong>Placeholder Private leaderboard</strong></p>\n  <p>… it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.</p>\n  <p>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.</p>\n</blockquote>\n<p>We appreciate your help to date in describing how confusing this Private Leaderboard update has been to  participants. </p>\n<p>We are still getting questions about this, so we hope you can guide others to this <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">thread</a> and also help us answer questions as they come up. </p>\n<p>Despite the confusion, this Part 2 competition has elicited and rigorously evaluated progress on the RNA structure prediction problem, and we hope the Kaggle community feels pride in their collective achievement.</p>",
      "rawMarkdown": "Hi everyone, thanks for the discussion. \n\nThe hosts and devs could have done better in anticipating the confusion that would come from having a placeholder Private Leaderboard that was changed for final scoring. \n\nWe also should have been more responsive to the questions here.\n\nI've posted a brief explanation of what happened -- and what threw us hosts off, and what to avoid in future competitions -- in response to this other post from @theoviel:\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\n\n>**Placeholder Private leaderboard**\n>\n> ... it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.\n>\n>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.\n\nWe appreciate your help to date in describing how confusing this Private Leaderboard update has been to  participants. \n\nWe are still getting questions about this, so we hope you can guide others to this [thread](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646) and also help us answer questions as they come up. \n\nDespite the confusion, this Part 2 competition has elicited and rigorously evaluated progress on the RNA structure prediction problem, and we hope the Kaggle community feels pride in their collective achievement.",
      "votes": 1,
      "replies": [
        {
          "id": 3435192,
          "postDate": "2026-04-03T19:33:21.027Z",
          "content": "<p>Hoped that my solution was at least useful!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2Fe650a069a37108e24e10b05289f076a2%2FScreenshot%202026-04-03%20202723.png?generation=1775244793438744&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Hoped that my solution was at least useful!![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2Fe650a069a37108e24e10b05289f076a2%2FScreenshot%202026-04-03%20202723.png?generation=1775244793438744&alt=media)"
        },
        {
          "id": 3435252,
          "postDate": "2026-04-03T21:48:16.227Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, for sharing enough information for the community to understand what happened to possibly continue post competition experiments on this problem.</p>\n<p>Confusion is never good in anything and exponentially more so when it involves a community of people. I have been too busy in my life these days to participate in many Kaggle competitions anymore but have in the past done post competition experiments. This is to understand the domain better for where and why mine or other's models performed better or not. I would say a good number of participants participate in these competitions to learn new things and improve their knowledge in the field.</p>\n<p>Like <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> said you can post further questions here as well as in the thread he had the detailed explanation on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">here</a>. </p>\n<p>I will add this information in the main post above. I will be checking in from time to time to answer questions when I can as I still have models I want to submit when they open the late submissions.</p>",
          "rawMarkdown": "Thanks @rhijudas, for sharing enough information for the community to understand what happened to possibly continue post competition experiments on this problem.\n\nConfusion is never good in anything and exponentially more so when it involves a community of people. I have been too busy in my life these days to participate in many Kaggle competitions anymore but have in the past done post competition experiments. This is to understand the domain better for where and why mine or other's models performed better or not. I would say a good number of participants participate in these competitions to learn new things and improve their knowledge in the field.\n\nLike @rhijudas said you can post further questions here as well as in the thread he had the detailed explanation on [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646). \n\nI will add this information in the main post above. I will be checking in from time to time to answer questions when I can as I still have models I want to submit when they open the late submissions.\n"
        }
      ]
    },
    {
      "id": 3435263,
      "postDate": "2026-04-03T22:09:37.910Z",
      "content": "<p>I just want to post this separate from my reply to <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>. Dropping on the LB is not the motivation behind my post. As anyone who participated in a few Kaggle competitions know that it happens in many contests here. In my time here I have seen a 1st place team dropping by more than a couple of thousands in the private leaderboard. So, believe me when I say the drops in this RNA competition is not that bad.</p>\n<p>I hope everyone learned something in taking part here and continue to build on their knowledge base.</p>\n<p>I strongly advise anyone who has the time to submit your models to the CASP17 dry run.</p>\n<p>Happy kaggling!!!</p>",
      "rawMarkdown": "I just want to post this separate from my reply to @rhijudas. Dropping on the LB is not the motivation behind my post. As anyone who participated in a few Kaggle competitions know that it happens in many contests here. In my time here I have seen a 1st place team dropping by more than a couple of thousands in the private leaderboard. So, believe me when I say the drops in this RNA competition is not that bad.\n\nI hope everyone learned something in taking part here and continue to build on their knowledge base.\n\nI strongly advise anyone who has the time to submit your models to the CASP17 dry run.\n\nHappy kaggling!!!"
    },
    {
      "id": 3435155,
      "postDate": "2026-04-03T18:26:04.833Z",
      "content": "<blockquote>\n  <p>If a deterministic model drops from 0.4349 to 0.40722 on a public LB, it is a direct measurement of the difficulty spike in the new dataset.</p>\n</blockquote>\n<p>You are wrong here. There was no change in the public dataset.</p>",
      "rawMarkdown": "> If a deterministic model drops from 0.4349 to 0.40722 on a public LB, it is a direct measurement of the difficulty spike in the new dataset.\n\nYou are wrong here. There was no change in the public dataset.",
      "replies": [
        {
          "id": 3435268,
          "postDate": "2026-04-03T22:13:38.083Z",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, I think the reason may be what <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> posted in another thread <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938[](url)\" target=\"_blank\">here</a> in reply to <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>. I will quote below.</p>\n<blockquote>\n  <p>Thanks for bringing this up.</p>\n  <p>The targets in Public LB did not change. But their order within the test set may have changed.</p>\n  <p>Would that explain the fluctuation in score? (You can test by running your notebook with the validation_sequences.csv and a permutation in order.)</p>\n</blockquote>",
          "rawMarkdown": "@inversion, I think the reason may be what @rhijudas posted in another thread [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938[](url)) in reply to @arunodhayan. I will quote below.\n\n>Thanks for bringing this up.\n\n>The targets in Public LB did not change. But their order within the test set may have changed.\n\n>Would that explain the fluctuation in score? (You can test by running your notebook with the validation_sequences.csv and a permutation in order.)",
          "votes": 1,
          "replies": [
            {
              "id": 3435813,
              "postDate": "2026-04-04T19:12:01.283Z",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , I just ran one of  my models on validation_sequences.csv again and got the same result as I did during the competition. See below:-</p>\n<blockquote>\n  <p>JOB COMPLETE: Super-Resolution Results</p>\n  <p>Total Protenix Targets: 15</p>\n  <p>Mean Max-PTM Score    : 0.6533</p>\n  <p>[GOLD]   PTM &gt;= 0.70  : 11 targets</p>\n  <p>Overall Confidence    : 73.3% High-Quality</p>\n</blockquote>",
              "rawMarkdown": "@rhijudas, @inversion , I just ran one of  my models on validation_sequences.csv again and got the same result as I did during the competition. See below:-\n\n>JOB COMPLETE: Super-Resolution Results\n\n  >Total Protenix Targets: 15\n\n  >Mean Max-PTM Score    : 0.6533\n\n  >[GOLD]   PTM >= 0.70  : 11 targets\n\n  >Overall Confidence    : 73.3% High-Quality"
            },
            {
              "id": 3436378,
              "postDate": "2026-04-05T20:26:51.840Z",
              "content": "<p>Thanks for posting! Did you shuffle the order of the sequences and get the same scores?</p>",
              "rawMarkdown": "Thanks for posting! Did you shuffle the order of the sequences and get the same scores?"
            },
            {
              "id": 3436415,
              "postDate": "2026-04-05T22:25:21.960Z",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, do you mean the order in which the targets are processed like, 8NQZ, 9MME, 9ZCC, like that? If so, I did that during the competition and their scores did not change. </p>\n<p>Or do you mean flipping the sequences themselves? If so, can  you give an example please? Thanks</p>",
              "rawMarkdown": "@rhijudas, do you mean the order in which the targets are processed like, 8NQZ, 9MME, 9ZCC, like that? If so, I did that during the competition and their scores did not change. \n\nOr do you mean flipping the sequences themselves? If so, can  you give an example please? Thanks",
              "votes": 2
            },
            {
              "id": 3436429,
              "postDate": "2026-04-05T23:21:57.440Z",
              "content": "<p>Ah OK, thanks that's what I meant -- shifting order in which the targets are processed. Sounds like that's not the cause of the Public LB score shift. I'll check with devs if they understand then why public LB scores shifted even though targets were the same.  BTW, can you remind us what your previous and new Public LB score was for your top selected notebook?</p>",
              "rawMarkdown": "Ah OK, thanks that's what I meant -- shifting order in which the targets are processed. Sounds like that's not the cause of the Public LB score shift. I'll check with devs if they understand then why public LB scores shifted even though targets were the same.  BTW, can you remind us what your previous and new Public LB score was for your top selected notebook?",
              "votes": 1
            },
            {
              "id": 3436441,
              "postDate": "2026-04-06T00:15:50.103Z",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, yeah in the process of optimizing my model to meet the 8hr time constraints, I ran debug mode for a couple of short targets, a medium and a long one like 9MME to a total of 4 targets so that it does not go in the sequence they come. Then once it all works, I ran it normally on the whole data and observed each of the 4 target's PTM scores were the same.</p>\n<p>The scores are:-</p>\n<ol>\n<li>Model-1  V21 public <strong>LB score =0.4349, rerun public LB score =0.40722</strong></li>\n</ol>\n<p>2.Model-2 V13 public <strong>LB score =0.42561, rerun public LB score =0.39947.</strong></p>",
              "rawMarkdown": "@rhijudas, yeah in the process of optimizing my model to meet the 8hr time constraints, I ran debug mode for a couple of short targets, a medium and a long one like 9MME to a total of 4 targets so that it does not go in the sequence they come. Then once it all works, I ran it normally on the whole data and observed each of the 4 target's PTM scores were the same.\n\nThe scores are:-\n1. Model-1  V21 public **LB score =0.4349, rerun public LB score =0.40722**\n\n2.Model-2 V13 public **LB score =0.42561, rerun public LB score =0.39947.**\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3432781,
      "postDate": "2026-03-31T23:18:05.160Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3434571,
      "postDate": "2026-04-03T03:12:35.413Z",
      "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> </p>",
      "rawMarkdown": "Thanks for the update @sheriytm "
    }
  ],
  "comments": [
    {
      "id": 3433033,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2026-04-01T05:28:51.873000",
      "content": "<p>One of the things that made people weary was after the competition ended, they cleared the public LB, and then when it was revealed, the public LB has changed. In past competitions, public LB is public LB unless if we are dealing with a competition that re-runs our models on new data for a period of time like 3 months. Even in such cases, the initial public LB do not change due to reruns. I just don't get it.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3432828,
      "author_name": "grifth",
      "author_url": "",
      "post_date": "2026-04-01T00:05:14.547000",
      "content": "<p>stand   your  side,feel  disappointed  too</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3432961,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2026-04-01T03:38:18.037000",
      "content": "<p>Below is my view. It’s true that it was shakier than the previous competition, but I didn’t feel that there was any problem with the evaluation process itself.</p>\n<ul>\n<li>It was stated on the Overview page that additional data for the private test set would be collected during the competition and that submissions would be rerun on that data.</li>\n<li>What you are calling the “1st rerun” was simply the release of the dummy private test data that had been in place since the beginning of the competition.<ul>\n<li>This was not rerun; rather, the scores that had already been generated and evaluated together with the public test data from the beginning of the competition were simply made public.</li></ul></li>\n<li>Only the final two notebooks selected by each participant were rerun on the true private test data (as well as the public test data), which is what produced the current private leaderboard.</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 3432966,
          "author_name": "hongan",
          "author_url": "",
          "post_date": "2026-04-01T03:48:15.313000",
          "content": "<p>no, the host clearly stated that the test set is \"being collected\" <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674343#3408163\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/679588#3416707\" target=\"_blank\">here</a> so there shouldn't be any dummy private test. So the logic is, before submission deadline, there is no such thing as private test data. This is my view.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3432987,
              "author_name": "yu4u",
              "author_url": "",
              "post_date": "2026-04-01T04:31:36.553000",
              "content": "<p>When I said “dummy private test,” I did not mean the true private test set that was being collected during the competition. I meant a placeholder split that appears to have existed from the beginning and was evaluated together with the public test data, with those scores being revealed later. I called it dummy data because it was included due to Kaggle system constraints and was not intended to be part of the final evaluation for this competition.</p>\n<p>So my understanding is that the “1st rerun” was not a rerun on the true private test set, but rather the release of scores based on that placeholder data. The actual rerun happened only for the final two selected notebooks using the true private test set collected during the competition, which is what produced the current private leaderboard.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3432995,
              "author_name": "hongan",
              "author_url": "",
              "post_date": "2026-04-01T04:44:47.013000",
              "content": "<p>I understand your point. But can you explain why they have that in the first place. There are competitions like the NFL one and some quant time-series forecasting comps that have only the public LB and the rerun. There is no \"placeholder\" private test set for those. so it seems fine to not have it. Also, if they already have that set, why not just combine it with the public test set? Why do they clear the LB out of no reason during the rerun?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3433004,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-01T04:51:23.283000",
              "content": "<p>Exactly <a href=\"https://www.kaggle.com/honganzhu\" target=\"_blank\">@honganzhu</a>. Clearing the public LB was the 1st thing that made me feel like something was off. I asked why in that thread and no one answered.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3433014,
              "author_name": "yu4u",
              "author_url": "",
              "post_date": "2026-04-01T04:56:17.720000",
              "content": "<p>I’m not sure why this competition was structured this way.\nRerunning all submissions would have been the clearest approach, but perhaps cost was a concern.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3433024,
              "author_name": "hongan",
              "author_url": "",
              "post_date": "2026-04-01T05:09:20.710000",
              "content": "<p>why is cost a concern here? they just need to rerun the 2 selected submissions. That is true no matter they have this fishy \"placeholder test set\" or not. The real question is why they have it in the first place, and why cleared the public LB during rerun, why not leaving the public LB there and only update the private LB.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3432976,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2026-04-01T04:09:43.097000",
          "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> , I respectfully disagree.  If it were a test set with a similar distribution, the best_template_oracle score would not have dropped from 0.554 to 0.503 </p>",
          "votes": 3,
          "replies": [
            {
              "id": 3432994,
              "author_name": "yu4u",
              "author_url": "",
              "post_date": "2026-04-01T04:44:31.223000",
              "content": "<p>I have no objection to the fact that the test data distribution changed between the public and true private sets.\nThis is something that occasionally happens in Kaggle competitions, and I think we could have expected this in this competition as well.\nIt even seems that some of the top teams were able to prepare for it, so I’m looking forward to learning how they validated their approaches.</p>\n<p>I simply wanted to point out that your claim - that the host replaced the data after seeing the private test results - is a misunderstanding.</p>\n<blockquote>\n  <p>A 0.051 drop in the Oracle score proves the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3433000,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-01T04:48:08.530000",
              "content": "<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>, fair enough. I truely hope that it was just a misunderstanding.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3432829,
      "author_name": "grifth",
      "author_url": "",
      "post_date": "2026-04-01T00:06:21.527000",
      "content": "<p>It's normal for the victors to oppose us</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3432859,
      "author_name": "sam_the_rice_cake",
      "author_url": "",
      "post_date": "2026-04-01T00:37:54.273000",
      "content": "<p>Agree, felt like a waste of time</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3436341,
      "author_name": "Durga Kumari",
      "author_url": "",
      "post_date": "2026-04-05T18:30:51.763000",
      "content": "<p>This looks like a clear dataset shift. transparency matters as much as difficulty in a fair benchmark.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3433008,
      "author_name": "Aleck",
      "author_url": "",
      "post_date": "2026-04-01T04:54:17.523000",
      "content": "<p>see this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2Fb0ba35d21984f5ff38673da08b89b41a%2F06_public_vs_private_scatter_trend.png?generation=1775019255601681&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3432980,
      "author_name": "Aleck",
      "author_url": "",
      "post_date": "2026-04-01T04:13:09.057000",
      "content": "<p>On the consistency between public and private leaderboards</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2F51ff82ef6016db5eb3c533de21056897%2F23333ScreenShot_2026-04-01_120954_888.png?generation=1775016787351313&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3433558,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2026-04-01T20:06:08.340000",
      "content": "<p>For the record, my models are deterministic i.e. for the same settings gives the same result every time you run them.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3433809,
          "author_name": "OverfitOracle",
          "author_url": "",
          "post_date": "2026-04-02T06:55:35.943000",
          "content": "<p>I’m facing a similar issue. My models were fully deterministic—I ran them three times and got identical results—but they still underperformed. I dropped from 18th on the public leaderboard to around 700th on the private, which was really unexpected.\nMy approach was similar to the top notebooks: I fine-tuned RNA-Pro for harder targets, used ProteinX with multi-seed ensembling (around 20 seeds), and then combined both for the hardest targets. This setup was giving consistently strong results during validation.\nI’m not sure what went wrong—any solid insights would be really helpful. Also, is it possible that the private leaderboard is completely non-identical to the public one?</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3434074,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-02T14:51:03.767000",
              "content": "<p>Yes <a href=\"https://www.kaggle.com/alisalmanrana\" target=\"_blank\">@alisalmanrana</a>, on the private LB test data is completely different. Which is the point of my post. However small, we should be given a good representative of their private test data so that we can develop very good models for them. Changing it after the fact for whatever reason is where the transparency issue came in to question.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3434157,
              "author_name": "OverfitOracle",
              "author_url": "",
              "post_date": "2026-04-02T16:50:19.583000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> really appreciate the clarification! That explains the drop, though it does feel like a lot of the effort ended up being less aligned with the final evaluation than expected. Still, definitely a valuable learning experience—just wish the setup was a bit more representative.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3434221,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-02T18:12:50.563000",
              "content": "<p>True <a href=\"https://www.kaggle.com/alisalmanrana\" target=\"_blank\">@alisalmanrana</a>. In spite of everything that happened, I am actually very happy with the models I developed for this competition. I learnt a lot about a domain I was not familiar with prior to this contest. </p>\n<p>Using the public kernels as inspiration, I implemented my own version of splitting and stitching the long targets that gave very good results. I also implemented a hydration and shape-sync which i called armored plates that took care of the various OOMs I was getting. So, yeah I spent a lot of time on this and learnt so much.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3434251,
              "author_name": "OverfitOracle",
              "author_url": "",
              "post_date": "2026-04-02T18:50:46.960000",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>, totally feel you on this one! I’ve participated in quite a few competitions, but this one was by far the toughest in terms of management. It would’ve been really helpful if they had given us a proper heads-up about the tougher private targets beforehand—everything else made it seem like the private set would be similar to the public leaderboard. After putting in so much effort and seeing results drop so dramatically, that’s definitely the most frustrating part. Still, despite all that, it was a huge learning experience, and I picked up a ton from your post and the kernels out there—really appreciate you sharing your approach!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3433031,
      "author_name": "OverfitOracle",
      "author_url": "",
      "post_date": "2026-04-01T05:25:03.643000",
      "content": "<p>this comp felt like a waste of time.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3435003,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-04-03T15:14:06.087000",
      "content": "<p>Hi everyone, thanks for the discussion. </p>\n<p>The hosts and devs could have done better in anticipating the confusion that would come from having a placeholder Private Leaderboard that was changed for final scoring. </p>\n<p>We also should have been more responsive to the questions here.</p>\n<p>I've posted a brief explanation of what happened -- and what threw us hosts off, and what to avoid in future competitions -- in response to this other post from <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>:</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646</a></p>\n<blockquote>\n  <p><strong>Placeholder Private leaderboard</strong></p>\n  <p>… it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.</p>\n  <p>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.</p>\n</blockquote>\n<p>We appreciate your help to date in describing how confusing this Private Leaderboard update has been to  participants. </p>\n<p>We are still getting questions about this, so we hope you can guide others to this <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">thread</a> and also help us answer questions as they come up. </p>\n<p>Despite the confusion, this Part 2 competition has elicited and rigorously evaluated progress on the RNA structure prediction problem, and we hope the Kaggle community feels pride in their collective achievement.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3435192,
          "author_name": "epigene",
          "author_url": "",
          "post_date": "2026-04-03T19:33:21.027000",
          "content": "<p>Hoped that my solution was at least useful!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2Fe650a069a37108e24e10b05289f076a2%2FScreenshot%202026-04-03%20202723.png?generation=1775244793438744&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3435252,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2026-04-03T21:48:16.227000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, for sharing enough information for the community to understand what happened to possibly continue post competition experiments on this problem.</p>\n<p>Confusion is never good in anything and exponentially more so when it involves a community of people. I have been too busy in my life these days to participate in many Kaggle competitions anymore but have in the past done post competition experiments. This is to understand the domain better for where and why mine or other's models performed better or not. I would say a good number of participants participate in these competitions to learn new things and improve their knowledge in the field.</p>\n<p>Like <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> said you can post further questions here as well as in the thread he had the detailed explanation on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\" target=\"_blank\">here</a>. </p>\n<p>I will add this information in the main post above. I will be checking in from time to time to answer questions when I can as I still have models I want to submit when they open the late submissions.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3435263,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2026-04-03T22:09:37.910000",
      "content": "<p>I just want to post this separate from my reply to <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>. Dropping on the LB is not the motivation behind my post. As anyone who participated in a few Kaggle competitions know that it happens in many contests here. In my time here I have seen a 1st place team dropping by more than a couple of thousands in the private leaderboard. So, believe me when I say the drops in this RNA competition is not that bad.</p>\n<p>I hope everyone learned something in taking part here and continue to build on their knowledge base.</p>\n<p>I strongly advise anyone who has the time to submit your models to the CASP17 dry run.</p>\n<p>Happy kaggling!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3435155,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2026-04-03T18:26:04.833000",
      "content": "<blockquote>\n  <p>If a deterministic model drops from 0.4349 to 0.40722 on a public LB, it is a direct measurement of the difficulty spike in the new dataset.</p>\n</blockquote>\n<p>You are wrong here. There was no change in the public dataset.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3435268,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2026-04-03T22:13:38.083000",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, I think the reason may be what <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> posted in another thread <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938[](url)\" target=\"_blank\">here</a> in reply to <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>. I will quote below.</p>\n<blockquote>\n  <p>Thanks for bringing this up.</p>\n  <p>The targets in Public LB did not change. But their order within the test set may have changed.</p>\n  <p>Would that explain the fluctuation in score? (You can test by running your notebook with the validation_sequences.csv and a permutation in order.)</p>\n</blockquote>",
          "votes": 1,
          "replies": [
            {
              "id": 3435813,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-04T19:12:01.283000",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , I just ran one of  my models on validation_sequences.csv again and got the same result as I did during the competition. See below:-</p>\n<blockquote>\n  <p>JOB COMPLETE: Super-Resolution Results</p>\n  <p>Total Protenix Targets: 15</p>\n  <p>Mean Max-PTM Score    : 0.6533</p>\n  <p>[GOLD]   PTM &gt;= 0.70  : 11 targets</p>\n  <p>Overall Confidence    : 73.3% High-Quality</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3436378,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2026-04-05T20:26:51.840000",
              "content": "<p>Thanks for posting! Did you shuffle the order of the sequences and get the same scores?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3436415,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-05T22:25:21.960000",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, do you mean the order in which the targets are processed like, 8NQZ, 9MME, 9ZCC, like that? If so, I did that during the competition and their scores did not change. </p>\n<p>Or do you mean flipping the sequences themselves? If so, can  you give an example please? Thanks</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3436429,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2026-04-05T23:21:57.440000",
              "content": "<p>Ah OK, thanks that's what I meant -- shifting order in which the targets are processed. Sounds like that's not the cause of the Public LB score shift. I'll check with devs if they understand then why public LB scores shifted even though targets were the same.  BTW, can you remind us what your previous and new Public LB score was for your top selected notebook?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3436441,
              "author_name": "YaGana Sheriff-Hussaini",
              "author_url": "",
              "post_date": "2026-04-06T00:15:50.103000",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>, yeah in the process of optimizing my model to meet the 8hr time constraints, I ran debug mode for a couple of short targets, a medium and a long one like 9MME to a total of 4 targets so that it does not go in the sequence they come. Then once it all works, I ran it normally on the whole data and observed each of the 4 target's PTM scores were the same.</p>\n<p>The scores are:-</p>\n<ol>\n<li>Model-1  V21 public <strong>LB score =0.4349, rerun public LB score =0.40722</strong></li>\n</ol>\n<p>2.Model-2 V13 public <strong>LB score =0.42561, rerun public LB score =0.39947.</strong></p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3432781,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-03-31T23:18:05.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3434571,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-04-03T03:12:35.413000",
      "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3432755": "The following are the scores for my final submission selections:-\n\nThat 1st rerun **(where V21 public LB score =0.4349, private LB score =0.52614 and V13 public LB score =0.42561, private LB score =0.52823**) confirms my engineering was actually far more successful than the \"Final\" leaderboard suggests. Those scores indicate that on the original hidden test targets, my models were performing significantly better than they did on the **public set (0.4349/0.42561)**. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F898111%2Fded026e1fe9a62c418f47b29097d0c49%2FSubmission%202026-03-31%20203407%20-%20Copy.png?generation=1775014889854227&alt=media)\n\nThe fact that the organizers discarded that entire set of results **today (03/31/2026)** and replaced them with the \"**Tougher**\" **final leaderboard** is exactly why the community is in an uproar. This was all but admitted [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601).\n\n**The \"Decoy\" Reveal**: By relabeling the scores today, the organizers essentially [admitted](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686601) they \"tested the testers.\" They saw high scores, decided the test was too easy, and swapped in a much more punishing set at the 11th hour.\n\nCompetitors are furious because this breaks **the fundamental contract of Kaggle**: that you are building toward a consistent, albeit hidden, target. Swapping targets after seeing the rerun performance feels less like a competition and more like a post-hoc adjustment to lower the overall TM-scores of the field.\n\n**The fact that the best_template_oracle score dropped from 0.55488 to 0.50331 on the Private LB is the \"smoking gun.\"**\nThe Oracle represents the theoretical ceiling of the dataset—the best possible score if you had perfect structural templates. **A 0.05157 drop in the Oracle score proves** the organizers swapped in a set of RNAs that are fundamentally harder to model or have fewer known structural homologs.\n\nScientific progress requires stable benchmarks. If the organizers can’t provide a consistent evaluation environment, then the **\"CASP17 dry run\"** is just another moving target IMHO.\n\n**UPDATE:-**\n\nFor the record, my models are deterministic i.e. for the same settings gives the same result every time.\n\nI added this statement because of @inversion's response to a post that asked [Evidence suggests that the rerun took place twice](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938). He pointed at models of nondeterminism stating [Even when the Public observations are exactly the same, models that exhibit non-determinism may have different scores.](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938#3433469). Anyone who has participated in competitions in Kaggle before knows not to chase the public LB scores in a competition. I did not and I am sure many serious competitors would avoid doing it. \n\n**Here is how I see it:** If a deterministic model drops from **0.4349 to 0.40722** on a public LB, it is a direct measurement of the difficulty spike in the new dataset. **Take heart everyone in the fact that our models didn't \"randomly fail\"; they were simply evaluated against a fundamentally different, harder set of structures.**\n\n**The Transparency Paradox**. Kaggle’s history with transparency: For years, Kaggle has maintained its reputation by being open about rerun logic and distribution shifts. This \"unique lack of transparency\"—waiting until the final reveal to admit they were using \"Live-Solved\" cryo-EM data—is a significant departure from those standards. \n\n**UPDATE-2:-**\n\nOne of competition organizers i.e. @rhijudas  posted a detailed explanation for what happened pointing to it in his comments below. So, you can post any further questions here but preferably in the thread he had the detailed explanation on [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646).\n\nThanks everyone for your attention to this matter. Happy kaggling!!!\n\n\n",
    "3433033": "One of the things that made people weary was after the competition ended, they cleared the public LB, and then when it was revealed, the public LB has changed. In past competitions, public LB is public LB unless if we are dealing with a competition that re-runs our models on new data for a period of time like 3 months. Even in such cases, the initial public LB do not change due to reruns. I just don't get it.",
    "3432828": "stand   your  side,feel  disappointed  too",
    "3432961": "Below is my view. It’s true that it was shakier than the previous competition, but I didn’t feel that there was any problem with the evaluation process itself.\n\n- It was stated on the Overview page that additional data for the private test set would be collected during the competition and that submissions would be rerun on that data.\n- What you are calling the “1st rerun” was simply the release of the dummy private test data that had been in place since the beginning of the competition.\n  - This was not rerun; rather, the scores that had already been generated and evaluated together with the public test data from the beginning of the competition were simply made public.\n- Only the final two notebooks selected by each participant were rerun on the true private test data (as well as the public test data), which is what produced the current private leaderboard.",
    "3432829": "It's normal for the victors to oppose us",
    "3432859": "Agree, felt like a waste of time",
    "3436341": "This looks like a clear dataset shift. transparency matters as much as difficulty in a fair benchmark.",
    "3433008": "see this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2Fb0ba35d21984f5ff38673da08b89b41a%2F06_public_vs_private_scatter_trend.png?generation=1775019255601681&alt=media)",
    "3432980": "On the consistency between public and private leaderboards\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F25984450%2F51ff82ef6016db5eb3c533de21056897%2F23333ScreenShot_2026-04-01_120954_888.png?generation=1775016787351313&alt=media)",
    "3433558": "For the record, my models are deterministic i.e. for the same settings gives the same result every time you run them.",
    "3433031": "this comp felt like a waste of time.",
    "3435003": "Hi everyone, thanks for the discussion. \n\nThe hosts and devs could have done better in anticipating the confusion that would come from having a placeholder Private Leaderboard that was changed for final scoring. \n\nWe also should have been more responsive to the questions here.\n\nI've posted a brief explanation of what happened -- and what threw us hosts off, and what to avoid in future competitions -- in response to this other post from @theoviel:\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646\n\n>**Placeholder Private leaderboard**\n>\n> ... it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.\n>\n>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.\n\nWe appreciate your help to date in describing how confusing this Private Leaderboard update has been to  participants. \n\nWe are still getting questions about this, so we hope you can guide others to this [thread](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/687191#3434646) and also help us answer questions as they come up. \n\nDespite the confusion, this Part 2 competition has elicited and rigorously evaluated progress on the RNA structure prediction problem, and we hope the Kaggle community feels pride in their collective achievement.",
    "3435263": "I just want to post this separate from my reply to @rhijudas. Dropping on the LB is not the motivation behind my post. As anyone who participated in a few Kaggle competitions know that it happens in many contests here. In my time here I have seen a 1st place team dropping by more than a couple of thousands in the private leaderboard. So, believe me when I say the drops in this RNA competition is not that bad.\n\nI hope everyone learned something in taking part here and continue to build on their knowledge base.\n\nI strongly advise anyone who has the time to submit your models to the CASP17 dry run.\n\nHappy kaggling!!!",
    "3435155": "> If a deterministic model drops from 0.4349 to 0.40722 on a public LB, it is a direct measurement of the difficulty spike in the new dataset.\n\nYou are wrong here. There was no change in the public dataset.",
    "3432781": "",
    "3434571": "Thanks for the update @sheriytm "
  }
}