{
  "id": 666382,
  "title": "Welcome to Part 2 of the Stanford RNA 3D Folding Challenge!",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/666382",
  "author_name": "Rhiju Das",
  "post_date": "2026-01-06T18:01:23.320000",
  "votes": 35,
  "comment_count": 35,
  "views": 0,
  "content": "<p>In 2025, in <a href=\"https://www.kaggle.com/c/stanford-rna-3d-folding\" target=\"_blank\">Part 1 of this competition</a>, Kagglers took on the fundamental problem of RNA 3D structure prediction.</p>\n<p>RNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.</p>\n<p>Unfortunately, the world's efforts to advance biology and biotechnology have been slowed down by our inability to computationally predict these RNA 3D structures.</p>\n<p>In 2025, Kaggle teams tackled this problem and achieved automated RNA 3D structure prediction methods that, for the first time, tied with a top human expert group!</p>\n<p>A big surprise was the revival of template-based modeling, a non-deep-learning approach. Check out our summary paper at <a href=\"https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full\" target=\"_blank\">Template-based RNA structure prediction advanced through a blind code competition</a>.</p>\n<p>Part 1 focused on single RNA chains under 1000 nts forming largely single structures. </p>\n<p>Part 2 now adds more  use cases important for biology and biotechnology:</p>\n<ul>\n<li>Single chain RNAs whose structures are different from previously available templates.</li>\n<li>RNA complexes with multiple RNA chains</li>\n<li>RNA 'machines' with multiple states - try to capture each of the states in your five predictions.</li>\n<li>RNAs whose conformations depend on proteins, other nucleic acids, and/or small molecule ligands.</li>\n<li>RNA systems with sizes up to 5500 nucleotides.</li>\n</ul>\n<p>We've also updated the evaluation metric to be more stringent. The TM-score metric now only counts residues in the prediction and ground truth as ‘aligned’ if the residues have matching residue numbers.</p>\n<p>Last, we've updated our baseline. In Part 1, for each target, the best-of-Kaggle prediction was similar or just barely better than the best publicly available prior 3D template, as discovered by an oracle with knowledge of the ground truth. This time, we’re making the <code>best_template_oracle</code> score available on the Public leaderboard (0.554). </p>\n<p>There's a lot we're curious about:</p>\n<ul>\n<li>How hard will it be to generalize Part 1's notebooks to multiple chains or proteins?</li>\n<li>How long will it take to beat the best template oracle on the Public leaderboard? Will top models generalize to the Private leaderboard?</li>\n<li>Will literature-aware LLMs that read the English language <code>description</code> fields be able to find templates missed last time?</li>\n<li>Will Kaggle notebooks be able to leverage chemical mapping profiles for millions of RNAs, available in the 2024 Ribonanza Kaggle competition <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">train</a> and <a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-solutions/\" target=\"_blank\">test</a> data?</li>\n<li>The maximum possible mean TM-score allowed by the inherent flexibility of these RNA targets is about 0.8. Can codes achieve this milestone in the RNA structure prediction problem? </li>\n<li>Can a new generation of agentic AI approaches and problem solvers –&nbsp;here receiving feedback from the Public leaderboard – accelerate research in Part 2?</li>\n</ul>\n<p>We’re looking forward to seeing what Kagglers come up with. We welcome back all who participated in Part 1 – your notebooks should be submittable to Part 2 with few or no updates. And we are excited to meet Kagglers just approaching RNA for the first time -- the fresh insights and high-placing models of teams with a \"beginner's mindset\" have had a huge impact in prior RNA competitions.</p>\n<p>For the hosts, </p>\n<p>Rhiju Das&nbsp;@rhijudas, Przemek Porebski <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a> </p>\n<p>And many thanks to:</p>\n<ul>\n<li>The AI@HHMI initiative at Janelia Research Center as well as several experimental RNA cryo-EM groups providing blind targets.</li>\n<li>Our wonderful collaborators at Kaggle: Walter Reade&nbsp;@inversion and Ashley Oldacre <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a></li>\n<li>The co-authors of our <a href=\"https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full\" target=\"_blank\">Part 1 paper</a>, with special thanks to CASP and RNA-Puzzle organizers who advised on the design of Part 2. </li>\n</ul>\n<p>P.S. In case you were wondering, the top Kaggle teams from Part 1 were <em>not</em> informed in advance about Part 2!</p>",
  "messages": [
    {
      "id": 3387313,
      "postDate": "2026-01-06T18:01:23.320Z",
      "content": "<p>In 2025, in <a href=\"https://www.kaggle.com/c/stanford-rna-3d-folding\" target=\"_blank\">Part 1 of this competition</a>, Kagglers took on the fundamental problem of RNA 3D structure prediction.</p>\n<p>RNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.</p>\n<p>Unfortunately, the world's efforts to advance biology and biotechnology have been slowed down by our inability to computationally predict these RNA 3D structures.</p>\n<p>In 2025, Kaggle teams tackled this problem and achieved automated RNA 3D structure prediction methods that, for the first time, tied with a top human expert group!</p>\n<p>A big surprise was the revival of template-based modeling, a non-deep-learning approach. Check out our summary paper at <a href=\"https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full\" target=\"_blank\">Template-based RNA structure prediction advanced through a blind code competition</a>.</p>\n<p>Part 1 focused on single RNA chains under 1000 nts forming largely single structures. </p>\n<p>Part 2 now adds more  use cases important for biology and biotechnology:</p>\n<ul>\n<li>Single chain RNAs whose structures are different from previously available templates.</li>\n<li>RNA complexes with multiple RNA chains</li>\n<li>RNA 'machines' with multiple states - try to capture each of the states in your five predictions.</li>\n<li>RNAs whose conformations depend on proteins, other nucleic acids, and/or small molecule ligands.</li>\n<li>RNA systems with sizes up to 5500 nucleotides.</li>\n</ul>\n<p>We've also updated the evaluation metric to be more stringent. The TM-score metric now only counts residues in the prediction and ground truth as ‘aligned’ if the residues have matching residue numbers.</p>\n<p>Last, we've updated our baseline. In Part 1, for each target, the best-of-Kaggle prediction was similar or just barely better than the best publicly available prior 3D template, as discovered by an oracle with knowledge of the ground truth. This time, we’re making the <code>best_template_oracle</code> score available on the Public leaderboard (0.554). </p>\n<p>There's a lot we're curious about:</p>\n<ul>\n<li>How hard will it be to generalize Part 1's notebooks to multiple chains or proteins?</li>\n<li>How long will it take to beat the best template oracle on the Public leaderboard? Will top models generalize to the Private leaderboard?</li>\n<li>Will literature-aware LLMs that read the English language <code>description</code> fields be able to find templates missed last time?</li>\n<li>Will Kaggle notebooks be able to leverage chemical mapping profiles for millions of RNAs, available in the 2024 Ribonanza Kaggle competition <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">train</a> and <a href=\"https://www.kaggle.com/datasets/rhijudas/ribonanza-solutions/\" target=\"_blank\">test</a> data?</li>\n<li>The maximum possible mean TM-score allowed by the inherent flexibility of these RNA targets is about 0.8. Can codes achieve this milestone in the RNA structure prediction problem? </li>\n<li>Can a new generation of agentic AI approaches and problem solvers –&nbsp;here receiving feedback from the Public leaderboard – accelerate research in Part 2?</li>\n</ul>\n<p>We’re looking forward to seeing what Kagglers come up with. We welcome back all who participated in Part 1 – your notebooks should be submittable to Part 2 with few or no updates. And we are excited to meet Kagglers just approaching RNA for the first time -- the fresh insights and high-placing models of teams with a \"beginner's mindset\" have had a huge impact in prior RNA competitions.</p>\n<p>For the hosts, </p>\n<p>Rhiju Das&nbsp;@rhijudas, Przemek Porebski <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a> </p>\n<p>And many thanks to:</p>\n<ul>\n<li>The AI@HHMI initiative at Janelia Research Center as well as several experimental RNA cryo-EM groups providing blind targets.</li>\n<li>Our wonderful collaborators at Kaggle: Walter Reade&nbsp;@inversion and Ashley Oldacre <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a></li>\n<li>The co-authors of our <a href=\"https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full\" target=\"_blank\">Part 1 paper</a>, with special thanks to CASP and RNA-Puzzle organizers who advised on the design of Part 2. </li>\n</ul>\n<p>P.S. In case you were wondering, the top Kaggle teams from Part 1 were <em>not</em> informed in advance about Part 2!</p>",
      "rawMarkdown": "In 2025, in [Part 1 of this competition](https://www.kaggle.com/c/stanford-rna-3d-folding), Kagglers took on the fundamental problem of RNA 3D structure prediction.\n\nRNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.\n\nUnfortunately, the world's efforts to advance biology and biotechnology have been slowed down by our inability to computationally predict these RNA 3D structures.\n\nIn 2025, Kaggle teams tackled this problem and achieved automated RNA 3D structure prediction methods that, for the first time, tied with a top human expert group!\n\nA big surprise was the revival of template-based modeling, a non-deep-learning approach. Check out our summary paper at [Template-based RNA structure prediction advanced through a blind code competition](https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full).\n\nPart 1 focused on single RNA chains under 1000 nts forming largely single structures. \n\nPart 2 now adds more  use cases important for biology and biotechnology:\n- Single chain RNAs whose structures are different from previously available templates.\n- RNA complexes with multiple RNA chains\n- RNA 'machines' with multiple states - try to capture each of the states in your five predictions.\n- RNAs whose conformations depend on proteins, other nucleic acids, and/or small molecule ligands.\n- RNA systems with sizes up to 5500 nucleotides.\n\nWe've also updated the evaluation metric to be more stringent. The TM-score metric now only counts residues in the prediction and ground truth as ‘aligned’ if the residues have matching residue numbers.\n\nLast, we've updated our baseline. In Part 1, for each target, the best-of-Kaggle prediction was similar or just barely better than the best publicly available prior 3D template, as discovered by an oracle with knowledge of the ground truth. This time, we’re making the `best_template_oracle` score available on the Public leaderboard (0.554). \n\nThere's a lot we're curious about:\n - How hard will it be to generalize Part 1's notebooks to multiple chains or proteins?\n - How long will it take to beat the best template oracle on the Public leaderboard? Will top models generalize to the Private leaderboard?\n - Will literature-aware LLMs that read the English language `description` fields be able to find templates missed last time?\n - Will Kaggle notebooks be able to leverage chemical mapping profiles for millions of RNAs, available in the 2024 Ribonanza Kaggle competition [train](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) and [test](https://www.kaggle.com/datasets/rhijudas/ribonanza-solutions/) data?\n - The maximum possible mean TM-score allowed by the inherent flexibility of these RNA targets is about 0.8. Can codes achieve this milestone in the RNA structure prediction problem? \n - Can a new generation of agentic AI approaches and problem solvers – here receiving feedback from the Public leaderboard – accelerate research in Part 2?\n\nWe’re looking forward to seeing what Kagglers come up with. We welcome back all who participated in Part 1 – your notebooks should be submittable to Part 2 with few or no updates. And we are excited to meet Kagglers just approaching RNA for the first time -- the fresh insights and high-placing models of teams with a \"beginner's mindset\" have had a huge impact in prior RNA competitions.\n\nFor the hosts, \n\nRhiju Das @rhijudas, Przemek Porebski @przemekporebski \n\nAnd many thanks to:\n* The AI@HHMI initiative at Janelia Research Center as well as several experimental RNA cryo-EM groups providing blind targets.\n* Our wonderful collaborators at Kaggle: Walter Reade @inversion and Ashley Oldacre @ashleyoldacre\n* The co-authors of our [Part 1 paper](https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full), with special thanks to CASP and RNA-Puzzle organizers who advised on the design of Part 2. \n\nP.S. In case you were wondering, the top Kaggle teams from Part 1 were *not* informed in advance about Part 2!\n\n",
      "votes": 35
    },
    {
      "id": 3420677,
      "postDate": "2026-03-13T18:29:20.387Z",
      "content": "<p>i AM HAVING A MENTAL BREAKDOWN.  i BUILT A PHYSICS AND MATH-BASED ENGINE FROM MY ORIGIN OF LIFE EXPERTISE AND i CANNOT EVEN GET MY CODE TO GIVE ME A SCORE.  iT RENDERS BUT i THINK i HAVE IMPROPER FORMATTING.  i AM ONE OF THE WORLD'S LEADING INNOVATORS IN QUANTUM BIOPHYSICS BUT i CANNOT CODE  FOR THE LIFE OF ME AND IT IS SO FRUSTRATING.  i HAVE A WHOLE FRAMEWORK WITH TESTED MECHANISMS THAT i CAN DO BY HAND AND GET A VERY HIGH tm SCORE, A MATHEMATICIAN, BUT i HAVE BEEN WORKING ON THIS FOR 9 DAYS AND HONESTLY, MIGHT NEED TO FIND SOME PROFESSIONAL HELP.  iS IT TOO LATE TO JOIN A TEAM???</p>",
      "rawMarkdown": "i AM HAVING A MENTAL BREAKDOWN.  i BUILT A PHYSICS AND MATH-BASED ENGINE FROM MY ORIGIN OF LIFE EXPERTISE AND i CANNOT EVEN GET MY CODE TO GIVE ME A SCORE.  iT RENDERS BUT i THINK i HAVE IMPROPER FORMATTING.  i AM ONE OF THE WORLD'S LEADING INNOVATORS IN QUANTUM BIOPHYSICS BUT i CANNOT CODE  FOR THE LIFE OF ME AND IT IS SO FRUSTRATING.  i HAVE A WHOLE FRAMEWORK WITH TESTED MECHANISMS THAT i CAN DO BY HAND AND GET A VERY HIGH tm SCORE, A MATHEMATICIAN, BUT i HAVE BEEN WORKING ON THIS FOR 9 DAYS AND HONESTLY, MIGHT NEED TO FIND SOME PROFESSIONAL HELP.  iS IT TOO LATE TO JOIN A TEAM???\n",
      "votes": 1,
      "replies": [
        {
          "id": 3420686,
          "postDate": "2026-03-13T18:42:02.180Z",
          "content": "<p>Subject: Don't let the syntax hide your great physics!\nHello, it’s completely understandable to feel frustrated. Many brilliant researchers in Biophysics face the same \"coding wall\" when they first join Kaggle.\nSince you have a solid mathematical framework and high manual TM-scores, you shouldn't worry about the formatting or the API. Here are two quick tips to get your logic into the leaderboard:\nLeverage AI for Translation: You can use LLMs (like Gemini or ChatGPT) to literally translate your physics formulas into Python code. Just describe the logic or the math, and ask it to \"format this for a Kaggle submission CSV.\"\nUse Public Notebooks: Check the \"Code\" tab of this competition. Look for \"Submission Baselines.\" You can copy one, keep the formatting part, and just plug your physics engine's output into it.\nKaggle is as much about community tools as it is about the science. Don't hesitate to use the existing \"scaffolding\" so your unique insights can shine. Good luck!</p>",
          "rawMarkdown": "Subject: Don't let the syntax hide your great physics!\nHello, it’s completely understandable to feel frustrated. Many brilliant researchers in Biophysics face the same \"coding wall\" when they first join Kaggle.\nSince you have a solid mathematical framework and high manual TM-scores, you shouldn't worry about the formatting or the API. Here are two quick tips to get your logic into the leaderboard:\nLeverage AI for Translation: You can use LLMs (like Gemini or ChatGPT) to literally translate your physics formulas into Python code. Just describe the logic or the math, and ask it to \"format this for a Kaggle submission CSV.\"\nUse Public Notebooks: Check the \"Code\" tab of this competition. Look for \"Submission Baselines.\" You can copy one, keep the formatting part, and just plug your physics engine's output into it.\nKaggle is as much about community tools as it is about the science. Don't hesitate to use the existing \"scaffolding\" so your unique insights can shine. Good luck!",
          "votes": 2,
          "replies": [
            {
              "id": 3420698,
              "postDate": "2026-03-13T18:58:20.370Z",
              "content": "<p>That's really nice of you to say. And I can solve three of the millennium problems but I can't code. Does anyone ever team up with an individual like myself. I would just like to team up with one talented coder I don't need a whole team but I'm so new I just want to show what I can do and help people.</p>",
              "rawMarkdown": "That's really nice of you to say. And I can solve three of the millennium problems but I can't code. Does anyone ever team up with an individual like myself. I would just like to team up with one talented coder I don't need a whole team but I'm so new I just want to show what I can do and help people."
            },
            {
              "id": 3420948,
              "postDate": "2026-03-14T09:18:17.330Z",
              "content": "<p>Hey, sent you a PM</p>",
              "rawMarkdown": "Hey, sent you a PM",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3418654,
      "postDate": "2026-03-08T21:54:21.617Z",
      "content": "<p>Is the residue numbering in the dataset aligned with the original PDB numbering, or are residues renumbered starting from 1 even if the first residue is missing in the PDB structure?</p>\n<p>Are the target coordinates evaluated strictly on the C1' atom positions from the deposited PDB structure, or after renumbering / filtering missing residues?</p>\n<p>Found: /kaggle/input/competitions/stanford-rna-3d-folding-2/PDB_RNA/9iwf.cif\n  chain  resid resname       x          y          z\n0     A      2       G  -0.623  46.888000 -23.393999\n1     A      3       G  -6.260  48.626999 -23.152000\n2     A      4       U  -9.899  50.548000 -19.612000\n3     A      5       G -11.262  52.573002 -14.761000\n4     A      6       U  -9.971  54.890999 -10.007000</p>\n<p>Saved: /kaggle/working/</p>\n<p>Found rows: 69\n          ID resname  resid     x_1     y_1     z_1           x_2  \\\n1307  9IWF_1       G      1  -0.623  46.888 -23.394 -1.000000e+18<br>\n1308  9IWF_2       G      2  -6.260  48.627 -23.152 -1.000000e+18<br>\n1309  9IWF_3       U      3  -9.899  50.548 -19.612 -1.000000e+18<br>\n1310  9IWF_4       G      4 -11.262  52.573 -14.761 -1.000000e+18<br>\n1311  9IWF_5       U      5  -9.971  54.891 -10.007 -1.000000e+18   </p>\n<pre><code>  chain  copy   Usage  target  \n</code></pre>\n<p>1307      A     1  Public    9IWF<br>\n1308      A     1  Public    9IWF<br>\n1309      A     1  Public    9IWF<br>\n1310      A     1  Public    9IWF<br>\n1311      A     1  Public    9IWF  </p>\n<p>[5 rows x 127 columns]\nSaved: /kaggle/working/9IWF_validation_coords.csv</p>\n<p>Subject: Possible residue indexing shift between PDB structures and provided ground truth</p>\n<p>Hi,</p>\n<p>While inspecting the structure 9IWF, I noticed a potential inconsistency between the original PDB residue numbering and the residue indexing used in the provided ground truth files.</p>\n<p>In the original PDB structure, the first resolved residue starts at residue 2, meaning residue 1 is missing in the deposited structure. However, in \"validation_labels.csv\", the coordinates appear to be renumbered starting from residue 1.</p>\n<p>This effectively introduces a global shift of one residue between the PDB coordinates and the dataset indexing. If participants rely on the PDB residue IDs directly, this shift propagates through the entire sequence and can lead to incorrect coordinate alignment for all downstream residues.</p>\n<p>Since the evaluation is performed on C1' atom coordinates, this indexing mismatch may unintentionally penalize otherwise correct predictions if the mapping between sequence position and structural residue index is not handled carefully.</p>\n<p>Could you please clarify:</p>\n<ol>\n<li>Whether all structures in the dataset were renumbered sequentially starting from 1, regardless of the original PDB residue numbering.</li>\n<li>Whether the ground truth coordinates correspond exactly to the deposited PDB C1' atoms, or if any preprocessing (such as residue filtering or renumbering) was applied.</li>\n</ol>\n<p>Clarification on this would help ensure participants interpret the ground truth correctly and avoid systematic alignment shifts.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Is the residue numbering in the dataset aligned with the original PDB numbering, or are residues renumbered starting from 1 even if the first residue is missing in the PDB structure?\n\n\nAre the target coordinates evaluated strictly on the C1' atom positions from the deposited PDB structure, or after renumbering / filtering missing residues?\n\n\nFound: /kaggle/input/competitions/stanford-rna-3d-folding-2/PDB_RNA/9iwf.cif\n  chain  resid resname       x          y          z\n0     A      2       G  -0.623  46.888000 -23.393999\n1     A      3       G  -6.260  48.626999 -23.152000\n2     A      4       U  -9.899  50.548000 -19.612000\n3     A      5       G -11.262  52.573002 -14.761000\n4     A      6       U  -9.971  54.890999 -10.007000\n\nSaved: /kaggle/working/\n\n\n\nFound rows: 69\n          ID resname  resid     x_1     y_1     z_1           x_2  \\\n1307  9IWF_1       G      1  -0.623  46.888 -23.394 -1.000000e+18   \n1308  9IWF_2       G      2  -6.260  48.627 -23.152 -1.000000e+18   \n1309  9IWF_3       U      3  -9.899  50.548 -19.612 -1.000000e+18   \n1310  9IWF_4       G      4 -11.262  52.573 -14.761 -1.000000e+18   \n1311  9IWF_5       U      5  -9.971  54.891 -10.007 -1.000000e+18   \n\n\n      chain  copy   Usage  target  \n1307      A     1  Public    9IWF  \n1308      A     1  Public    9IWF  \n1309      A     1  Public    9IWF  \n1310      A     1  Public    9IWF  \n1311      A     1  Public    9IWF  \n\n[5 rows x 127 columns]\nSaved: /kaggle/working/9IWF_validation_coords.csv\n\n\nSubject: Possible residue indexing shift between PDB structures and provided ground truth\n\nHi,\n\nWhile inspecting the structure 9IWF, I noticed a potential inconsistency between the original PDB residue numbering and the residue indexing used in the provided ground truth files.\n\nIn the original PDB structure, the first resolved residue starts at residue 2, meaning residue 1 is missing in the deposited structure. However, in \"validation_labels.csv\", the coordinates appear to be renumbered starting from residue 1.\n\nThis effectively introduces a global shift of one residue between the PDB coordinates and the dataset indexing. If participants rely on the PDB residue IDs directly, this shift propagates through the entire sequence and can lead to incorrect coordinate alignment for all downstream residues.\n\nSince the evaluation is performed on C1' atom coordinates, this indexing mismatch may unintentionally penalize otherwise correct predictions if the mapping between sequence position and structural residue index is not handled carefully.\n\nCould you please clarify:\n\n1. Whether all structures in the dataset were renumbered sequentially starting from 1, regardless of the original PDB residue numbering.\n2. Whether the ground truth coordinates correspond exactly to the deposited PDB C1' atoms, or if any preprocessing (such as residue filtering or renumbering) was applied.\n\nClarification on this would help ensure participants interpret the ground truth correctly and avoid systematic alignment shifts.\n\nThanks!\n",
      "votes": 1,
      "replies": [
        {
          "id": 3419830,
          "postDate": "2026-03-11T19:11:19.453Z",
          "content": "<p>The sequence numbering in the <code>*_labels.csv</code> and submission files corresponds to the <code>sequence</code> in the <code>*_sequences.csv</code> file. We disregard residue numbering provided by the authors and renumber residues sequentially according to the <code>sequence</code> field. The <code>sequence</code> field is derived from <code>_entity_poly</code> and <code>_pdbx_poly_seq_scheme</code> records in the CIF files.  Per wwPDB <a href=\"https://www.wwpdb.org/documentation/procedure#toc_3\" target=\"_blank\">guidelines</a> these records should correspond to the sequence of the complete sample used for experiment. If the sequence covers residues that are missing in the model, then these will be reported in the <code>*_labels.csv</code> but will not have assigned coordinates. </p>\n<p>For this particular case authors decided to start numbering from 2, but they didn't report any unobserved residues prior to that. Sometimes that happens when authors decide to keep their numbering consistent with other entries, database or other reference source. </p>\n<p>Submissions are scored using 1:1 match of the residues between submission and solution. If the coordinates were not observed experimentally these will not affect the scoring as explained <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/666382#3414909\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "The sequence numbering in the `*_labels.csv` and submission files corresponds to the `sequence` in the `*_sequences.csv` file. We disregard residue numbering provided by the authors and renumber residues sequentially according to the `sequence` field. The `sequence` field is derived from ` _entity_poly` and ` _pdbx_poly_seq_scheme` records in the CIF files.  Per wwPDB [guidelines](https://www.wwpdb.org/documentation/procedure#toc_3) these records should correspond to the sequence of the complete sample used for experiment. If the sequence covers residues that are missing in the model, then these will be reported in the `*_labels.csv` but will not have assigned coordinates. \n\nFor this particular case authors decided to start numbering from 2, but they didn't report any unobserved residues prior to that. Sometimes that happens when authors decide to keep their numbering consistent with other entries, database or other reference source. \n\nSubmissions are scored using 1:1 match of the residues between submission and solution. If the coordinates were not observed experimentally these will not affect the scoring as explained [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/666382#3414909)",
          "votes": 2
        }
      ]
    },
    {
      "id": 3414048,
      "postDate": "2026-02-25T20:04:47.153Z",
      "content": "<p>Question 1 — Missing Residues / Gaps\nHello,\nI have a question regarding how the evaluation handles missing residues or discontinuities in the chain.\nIf a model predicts coordinates for positions that correspond to gaps or unresolved regions in the ground truth structure, how is the TM-score computed in that case?\nSpecifically:\nAre only residues present in the ground truth considered in the alignment?\nOr are predicted coordinates for non-existing residues also included in the TM-score calculation?\nI’m trying to understand whether predicting coordinates across gaps (e.g., filling missing segments) would negatively affect the evaluation, or if those positions are simply ignored during scoring.\nQuestion 2 — Long Artificial Straight Connections\nI also have a question about structural continuity.\nIf the model predicts a long artificial straight connection between two distant segments (for example, creating an unrealistic straight bridge in 3D space), how would this affect the TM-score?\nSince evaluation is based on the spatial placement of C1' atoms, would such long geometric distortions significantly penalize the score even if the overall global fold is approximately correct?\nIn other words, does TM-score heavily penalize unrealistic long-range geometric artifacts even when local regions are reasonably aligned?</p>",
      "rawMarkdown": "Question 1 — Missing Residues / Gaps\nHello,\nI have a question regarding how the evaluation handles missing residues or discontinuities in the chain.\nIf a model predicts coordinates for positions that correspond to gaps or unresolved regions in the ground truth structure, how is the TM-score computed in that case?\nSpecifically:\nAre only residues present in the ground truth considered in the alignment?\nOr are predicted coordinates for non-existing residues also included in the TM-score calculation?\nI’m trying to understand whether predicting coordinates across gaps (e.g., filling missing segments) would negatively affect the evaluation, or if those positions are simply ignored during scoring.\nQuestion 2 — Long Artificial Straight Connections\nI also have a question about structural continuity.\nIf the model predicts a long artificial straight connection between two distant segments (for example, creating an unrealistic straight bridge in 3D space), how would this affect the TM-score?\nSince evaluation is based on the spatial placement of C1' atoms, would such long geometric distortions significantly penalize the score even if the overall global fold is approximately correct?\nIn other words, does TM-score heavily penalize unrealistic long-range geometric artifacts even when local regions are reasonably aligned?",
      "votes": 1,
      "replies": [
        {
          "id": 3414909,
          "postDate": "2026-02-27T22:51:12.217Z",
          "content": "<p>If the residues were not included in the ground truth they will not be considered in the alignment and modeling those will not affect a TM-score. However, if the residues that are not included in the prediction are present in the ground truth that will negatively  impact TM-scores as it is calculated relative to the number of residues in the ground truth </p>\n<p>I am not sure if I correctly understand the scenario for question 2, so I will provide several examples:</p>\n<ol>\n<li>Current format for the submission requires numbering of the residues according to concatenated sequence for multimeric structures. So it may happen that the residue <code>ID=target_id_200</code> (from one monomer) is distant to <code>ID=target_id_201</code> (from the other monomer). In such case this stretch is not penalized, as there are no atoms in-between </li>\n<li>For single monomer, If there is a gap in the prediction, and the <code>C1'</code> atoms are missing between residue 100 and 120, but those atoms are present in the ground truth then that will be penalized in TM-score</li>\n<li>If those missing residues  100-120 are filled with <code>C1'</code> atoms with unrealistic positions (for example placed uniformly along straight line) and those residues are present in the ground truth, that will be penalized in TM-score. The penalty will depend on the size of RNA and number of mismatched <code>C1'</code> positions.</li>\n<li>Similarly if there is unrealistic distance between <code>C1'</code> 100 and 101 from the same monomer such distance will likely put this and other atoms in wrong positions and that will be penalized in the TM-score. </li>\n</ol>",
          "rawMarkdown": "If the residues were not included in the ground truth they will not be considered in the alignment and modeling those will not affect a TM-score. However, if the residues that are not included in the prediction are present in the ground truth that will negatively  impact TM-scores as it is calculated relative to the number of residues in the ground truth \n\nI am not sure if I correctly understand the scenario for question 2, so I will provide several examples:\n1. Current format for the submission requires numbering of the residues according to concatenated sequence for multimeric structures. So it may happen that the residue `ID=target_id_200` (from one monomer) is distant to `ID=target_id_201` (from the other monomer). In such case this stretch is not penalized, as there are no atoms in-between \n2. For single monomer, If there is a gap in the prediction, and the `C1'` atoms are missing between residue 100 and 120, but those atoms are present in the ground truth then that will be penalized in TM-score\n3. If those missing residues  100-120 are filled with `C1'` atoms with unrealistic positions (for example placed uniformly along straight line) and those residues are present in the ground truth, that will be penalized in TM-score. The penalty will depend on the size of RNA and number of mismatched `C1'` positions.\n4. Similarly if there is unrealistic distance between `C1'` 100 and 101 from the same monomer such distance will likely put this and other atoms in wrong positions and that will be penalized in the TM-score. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3388677,
      "postDate": "2026-01-09T11:08:54.437Z",
      "content": "<p>Dear hosts Rhiju Das <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and Przemek Porebski <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a>,</p>\n<p>Following recent discussions regarding the disqualification of participants for reusing private notebooks from previous years (as seen in <a href=\"https://www.kaggle.com/competitions/cmi-detect-behavior-with-sensor-data/discussion/602648#3278013\" target=\"_blank\">this thread</a>), our team ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>, <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>, and myself) would like to seek clarification on the rules for this transition.</p>\n<p>During Part 1, we worked as a team and shared private notebooks and datasets. As we move into Part 2, we would like to ensure we remain in full compliance under the following three scenarios:</p>\n<p><strong>1. Full Team Continuity:</strong> If all members of our Part 1 team compete together as the same team in Part 2, can we continue to use our private resources from Part 1?</p>\n<p><strong>2. Partial Team Continuity:</strong> If only a subset of our Part 1 team competes in Part 2 (and the remaining members do not participate in Part 2 at all), can the active members still use the private notebooks and datasets created during Part 1?</p>\n<p><strong>3. Individual Participation:</strong> If members of the Part 1 team decide to compete individually or on different teams in Part 2, what is the protocol for using our previous shared resources? (e.g., Must they be made public first?)</p>\n<p>Best,\nHoa</p>",
      "rawMarkdown": "Dear hosts Rhiju Das @rhijudas and Przemek Porebski @przemekporebski,\n\nFollowing recent discussions regarding the disqualification of participants for reusing private notebooks from previous years (as seen in [this thread](https://www.kaggle.com/competitions/cmi-detect-behavior-with-sensor-data/discussion/602648#3278013)), our team ( @hengck23 , @lihaoweicvch, @arunodhayan, and myself) would like to seek clarification on the rules for this transition.\n\nDuring Part 1, we worked as a team and shared private notebooks and datasets. As we move into Part 2, we would like to ensure we remain in full compliance under the following three scenarios:\n\n**1. Full Team Continuity:** If all members of our Part 1 team compete together as the same team in Part 2, can we continue to use our private resources from Part 1?\n\n**2. Partial Team Continuity:** If only a subset of our Part 1 team competes in Part 2 (and the remaining members do not participate in Part 2 at all), can the active members still use the private notebooks and datasets created during Part 1?\n\n**3. Individual Participation:** If members of the Part 1 team decide to compete individually or on different teams in Part 2, what is the protocol for using our previous shared resources? (e.g., Must they be made public first?)\n\nBest,\nHoa",
      "votes": 2,
      "replies": [
        {
          "id": 3388968,
          "postDate": "2026-01-09T22:33:11.097Z",
          "content": "<p>Hi Hoa (and d4t4 team)! Great to have you participating again. I have relayed your question to Kaggle devs. They will need a little more time to complete internal discussions.</p>",
          "rawMarkdown": "Hi Hoa (and d4t4 team)! Great to have you participating again. I have relayed your question to Kaggle devs. They will need a little more time to complete internal discussions.",
          "votes": 1
        },
        {
          "id": 3388978,
          "postDate": "2026-01-09T23:32:03.550Z",
          "content": "<p>Hello Hao, Thank you for your questions. \nYou are able to continue using your notebooks from part 1 and carry them over to part 2 whether your team is continuing as the original members or including new members. Best of luck to you! </p>",
          "rawMarkdown": "Hello Hao, Thank you for your questions. \nYou are able to continue using your notebooks from part 1 and carry them over to part 2 whether your team is continuing as the original members or including new members. Best of luck to you! ",
          "votes": 2,
          "replies": [
            {
              "id": 3391582,
              "postDate": "2026-01-15T08:26:21.023Z",
              "content": "<p>Thank <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and  <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a> for your instant response!</p>\n<p>How about we join this season but in different teams?</p>",
              "rawMarkdown": "Thank @rhijudas and  @ashleyoldacre for your instant response!\n\nHow about we join this season but in different teams?"
            },
            {
              "id": 3392342,
              "postDate": "2026-01-16T18:30:07.667Z",
              "content": "<p>Yes, no problem. You can join in different teams. Thank you for asking and good luck! </p>",
              "rawMarkdown": "Yes, no problem. You can join in different teams. Thank you for asking and good luck! "
            }
          ]
        }
      ]
    },
    {
      "id": 3388019,
      "postDate": "2026-01-08T00:41:41.407Z",
      "content": "<p>Thank you very much, I'm excited to see what comes from this competition. I'm curious, you noted that this competition will have significantly harder structures to predict (multiple chains, ligands, long chains, etc). Can you clarify, is the test set exclusively these difficult cases, or are there still examples of single, short chains in the test set?</p>\n<p>Edit: Also, how big roughly can we expect the test set to be? ~30 like the example, or will it be more like 100+?</p>",
      "rawMarkdown": "Thank you very much, I'm excited to see what comes from this competition. I'm curious, you noted that this competition will have significantly harder structures to predict (multiple chains, ligands, long chains, etc). Can you clarify, is the test set exclusively these difficult cases, or are there still examples of single, short chains in the test set?\n\nEdit: Also, how big roughly can we expect the test set to be? ~30 like the example, or will it be more like 100+?",
      "votes": 2,
      "replies": [
        {
          "id": 3388031,
          "postDate": "2026-01-08T01:56:45.683Z",
          "content": "<p>The hidden test_set.csv will be a lot like the example test_set.csv, including in terms of number of targets. There will indeed be single short chains, but selected to be template-free. I've edited the welcome post above to clarify. Thanks for the question!</p>",
          "rawMarkdown": "The hidden test_set.csv will be a lot like the example test_set.csv, including in terms of number of targets. There will indeed be single short chains, but selected to be template-free. I've edited the welcome post above to clarify. Thanks for the question!",
          "votes": 2,
          "replies": [
            {
              "id": 3388040,
              "postDate": "2026-01-08T03:03:03.163Z",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> template free means no close structural homologs available and template based matching will not help?</p>",
              "rawMarkdown": "@rhijudas template free means no close structural homologs available and template based matching will not help?"
            },
            {
              "id": 3388387,
              "postDate": "2026-01-08T18:13:59.280Z",
              "content": "<p>Yes I think template search may not help much with single chain.</p>",
              "rawMarkdown": "Yes I think template search may not help much with single chain.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3428229,
      "postDate": "2026-03-25T05:33:02.033Z",
      "content": "<p>heyy! just curious.. has anyone here used graph based approach?</p>",
      "rawMarkdown": "heyy! just curious.. has anyone here used graph based approach?",
      "replies": [
        {
          "id": 3429142,
          "postDate": "2026-03-26T12:31:57.307Z",
          "content": "<p>I did. GNN.</p>",
          "rawMarkdown": "I did. GNN."
        }
      ]
    },
    {
      "id": 3425994,
      "postDate": "2026-03-21T17:03:41.167Z",
      "content": "<p>Amazing data </p>",
      "rawMarkdown": "Amazing data "
    },
    {
      "id": 3408052,
      "postDate": "2026-02-19T17:59:22.373Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2\" target=\"_blank\">https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2</a></p>\n<p>please anyone check my notebook why it always give submission scoring error.</p>",
      "rawMarkdown": "https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2\n\nplease anyone check my notebook why it always give submission scoring error.",
      "replies": [
        {
          "id": 3422853,
          "postDate": "2026-03-18T01:49:06.413Z",
          "content": "<p>I'm having issues too, but at the very least \"x,y,z are clipped between -999.999 and 9999.999 before scoring\". You may want to try to do this before submitting and see if it helps</p>",
          "rawMarkdown": "I'm having issues too, but at the very least \"x,y,z are clipped between -999.999 and 9999.999 before scoring\". You may want to try to do this before submitting and see if it helps",
          "replies": [
            {
              "id": 3426420,
              "postDate": "2026-03-22T15:06:39.980Z",
              "content": "<p>Have you fixed it? I need help.</p>",
              "rawMarkdown": "Have you fixed it? I need help."
            }
          ]
        },
        {
          "id": 3426421,
          "postDate": "2026-03-22T15:07:01.687Z",
          "content": "<p>Have you fixed it? help me.</p>",
          "rawMarkdown": "Have you fixed it? help me."
        }
      ]
    },
    {
      "id": 3407880,
      "postDate": "2026-02-19T11:14:04.193Z",
      "content": "<p>I submitted my notebook 16 times code cells runs perfectly but when I submitted to competition it gives scoring error can anyone helps to solve my this problem , Everything is perfect code and files submitted in csv. But always give scoring error.</p>",
      "rawMarkdown": "I submitted my notebook 16 times code cells runs perfectly but when I submitted to competition it gives scoring error can anyone helps to solve my this problem , Everything is perfect code and files submitted in csv. But always give scoring error.\n",
      "replies": [
        {
          "id": 3426422,
          "postDate": "2026-03-22T15:07:34.197Z",
          "content": "<p>Have you fixed it? </p>",
          "rawMarkdown": "Have you fixed it? ",
          "replies": [
            {
              "id": 3426950,
              "postDate": "2026-03-23T14:35:37.860Z",
              "content": "<p>Not yet.. still working on it.</p>",
              "rawMarkdown": "Not yet.. still working on it.\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3403531,
      "postDate": "2026-02-08T17:42:49.213Z",
      "content": "<p>3D prediction that brings real biological complexity and pushes the field beyond expert-level performance.</p>",
      "rawMarkdown": "3D prediction that brings real biological complexity and pushes the field beyond expert-level performance."
    },
    {
      "id": 3403477,
      "postDate": "2026-02-08T15:25:34.347Z",
      "content": "<p>This challenge is incredibly exciting and inspiring. The progress made in Part 1, especially reaching performance comparable to expert human groups, shows how powerful collaborative science and machine learning can be. Expanding the scope in Part 2 to include multi-chain RNAs, RNA complexes, ligand-dependent conformations, and larger systems makes the problem much closer to real biological scenarios, which is very motivating.</p>\n<p>The updated evaluation metric and availability of the template oracle baseline will likely push participants to develop more robust and generalizable methods. I am especially interested in how hybrid approaches combining template-based modeling, deep learning, and chemical mapping data will perform. The possibility of agentic AI systems accelerating discovery is also fascinating.</p>\n<p>Thank you to the organizers for creating such a meaningful competition that directly contributes to advances in biology, medicine, and biotechnology. I am excited to follow the solutions developed by the community and to learn from this challenge.</p>",
      "rawMarkdown": "This challenge is incredibly exciting and inspiring. The progress made in Part 1, especially reaching performance comparable to expert human groups, shows how powerful collaborative science and machine learning can be. Expanding the scope in Part 2 to include multi-chain RNAs, RNA complexes, ligand-dependent conformations, and larger systems makes the problem much closer to real biological scenarios, which is very motivating.\n\nThe updated evaluation metric and availability of the template oracle baseline will likely push participants to develop more robust and generalizable methods. I am especially interested in how hybrid approaches combining template-based modeling, deep learning, and chemical mapping data will perform. The possibility of agentic AI systems accelerating discovery is also fascinating.\n\nThank you to the organizers for creating such a meaningful competition that directly contributes to advances in biology, medicine, and biotechnology. I am excited to follow the solutions developed by the community and to learn from this challenge."
    },
    {
      "id": 3397222,
      "postDate": "2026-01-26T18:25:14.320Z",
      "content": "<p>مرحبا معكم  رضا من الجزائر مشارك لأول مرة ، واجهت صعوبات كبيرة قبل ارسال اول دفنر ملاحظات وبعد محاولات عديدة بقيت علامتي عندد القيمة .0.098 هل تعتبر جيدة كمشارك هاو لأول  مرة وعندي طلب لماذا لا يتم توسيع عدد مرات ارسال دفتر البيانات او عدم حساب حالة الرفض حتى نتمكن من فهم نقاط الضعف فاهدف هنا إنساني اكبر منه حافز مالي. وشكرا</p>",
      "rawMarkdown": "مرحبا معكم  رضا من الجزائر مشارك لأول مرة ، واجهت صعوبات كبيرة قبل ارسال اول دفنر ملاحظات وبعد محاولات عديدة بقيت علامتي عندد القيمة .0.098 هل تعتبر جيدة كمشارك هاو لأول  مرة وعندي طلب لماذا لا يتم توسيع عدد مرات ارسال دفتر البيانات او عدم حساب حالة الرفض حتى نتمكن من فهم نقاط الضعف فاهدف هنا إنساني اكبر منه حافز مالي. وشكرا"
    },
    {
      "id": 3391366,
      "postDate": "2026-01-14T19:41:47.320Z",
      "content": "<p>should I use just \"train_sequences.csv\" for the training? what are the purpose of MSA and PDB_RNA folders? that's not really clear</p>",
      "rawMarkdown": "should I use just \"train_sequences.csv\" for the training? what are the purpose of MSA and PDB_RNA folders? that's not really clear\n"
    },
    {
      "id": 3390775,
      "postDate": "2026-01-13T23:32:41.357Z",
      "content": "<p>Hi,\nI don't understand why I'm getting such widely divergent results using the theoretically identical measurement method provided by Kaggle at <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score</a>.\nLocal measurements show that the model is learning and the TM-Score is increasing, reaching 0.201 by the 10th epoch, while the same model submitted to Kaggle only achieved a TM-Score of 0.067.\nThere shouldn't be such drastic differences. Can someone explain this to me?</p>",
      "rawMarkdown": "Hi,\nI don't understand why I'm getting such widely divergent results using the theoretically identical measurement method provided by Kaggle at https://www.kaggle.com/code/metric/ribonanza-tm-score.\nLocal measurements show that the model is learning and the TM-Score is increasing, reaching 0.201 by the 10th epoch, while the same model submitted to Kaggle only achieved a TM-Score of 0.067.\nThere shouldn't be such drastic differences. Can someone explain this to me?",
      "replies": [
        {
          "id": 3407562,
          "postDate": "2026-02-18T16:13:16.747Z",
          "content": "<p>You likely computed TM score for the training dataset, while 0.067 is TM-score for a unseen data?</p>",
          "rawMarkdown": "You likely computed TM score for the training dataset, while 0.067 is TM-score for a unseen data?"
        }
      ]
    },
    {
      "id": 3416497,
      "postDate": "2026-03-03T02:48:03.463Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3388056,
      "postDate": "2026-01-08T04:27:22.207Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> </p>",
      "rawMarkdown": "Thanks @rhijudas "
    }
  ],
  "comments": [
    {
      "id": 3420677,
      "author_name": "The Acquired Savant",
      "author_url": "",
      "post_date": "2026-03-13T18:29:20.387000",
      "content": "<p>i AM HAVING A MENTAL BREAKDOWN.  i BUILT A PHYSICS AND MATH-BASED ENGINE FROM MY ORIGIN OF LIFE EXPERTISE AND i CANNOT EVEN GET MY CODE TO GIVE ME A SCORE.  iT RENDERS BUT i THINK i HAVE IMPROPER FORMATTING.  i AM ONE OF THE WORLD'S LEADING INNOVATORS IN QUANTUM BIOPHYSICS BUT i CANNOT CODE  FOR THE LIFE OF ME AND IT IS SO FRUSTRATING.  i HAVE A WHOLE FRAMEWORK WITH TESTED MECHANISMS THAT i CAN DO BY HAND AND GET A VERY HIGH tm SCORE, A MATHEMATICIAN, BUT i HAVE BEEN WORKING ON THIS FOR 9 DAYS AND HONESTLY, MIGHT NEED TO FIND SOME PROFESSIONAL HELP.  iS IT TOO LATE TO JOIN A TEAM???</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3420686,
          "author_name": "Mohamed Hashem (GoxD)",
          "author_url": "",
          "post_date": "2026-03-13T18:42:02.180000",
          "content": "<p>Subject: Don't let the syntax hide your great physics!\nHello, it’s completely understandable to feel frustrated. Many brilliant researchers in Biophysics face the same \"coding wall\" when they first join Kaggle.\nSince you have a solid mathematical framework and high manual TM-scores, you shouldn't worry about the formatting or the API. Here are two quick tips to get your logic into the leaderboard:\nLeverage AI for Translation: You can use LLMs (like Gemini or ChatGPT) to literally translate your physics formulas into Python code. Just describe the logic or the math, and ask it to \"format this for a Kaggle submission CSV.\"\nUse Public Notebooks: Check the \"Code\" tab of this competition. Look for \"Submission Baselines.\" You can copy one, keep the formatting part, and just plug your physics engine's output into it.\nKaggle is as much about community tools as it is about the science. Don't hesitate to use the existing \"scaffolding\" so your unique insights can shine. Good luck!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3420698,
              "author_name": "The Acquired Savant",
              "author_url": "",
              "post_date": "2026-03-13T18:58:20.370000",
              "content": "<p>That's really nice of you to say. And I can solve three of the millennium problems but I can't code. Does anyone ever team up with an individual like myself. I would just like to team up with one talented coder I don't need a whole team but I'm so new I just want to show what I can do and help people.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3420948,
              "author_name": "Andres H. Zapke",
              "author_url": "",
              "post_date": "2026-03-14T09:18:17.330000",
              "content": "<p>Hey, sent you a PM</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3418654,
      "author_name": "Mohamed Hashem (GoxD)",
      "author_url": "",
      "post_date": "2026-03-08T21:54:21.617000",
      "content": "<p>Is the residue numbering in the dataset aligned with the original PDB numbering, or are residues renumbered starting from 1 even if the first residue is missing in the PDB structure?</p>\n<p>Are the target coordinates evaluated strictly on the C1' atom positions from the deposited PDB structure, or after renumbering / filtering missing residues?</p>\n<p>Found: /kaggle/input/competitions/stanford-rna-3d-folding-2/PDB_RNA/9iwf.cif\n  chain  resid resname       x          y          z\n0     A      2       G  -0.623  46.888000 -23.393999\n1     A      3       G  -6.260  48.626999 -23.152000\n2     A      4       U  -9.899  50.548000 -19.612000\n3     A      5       G -11.262  52.573002 -14.761000\n4     A      6       U  -9.971  54.890999 -10.007000</p>\n<p>Saved: /kaggle/working/</p>\n<p>Found rows: 69\n          ID resname  resid     x_1     y_1     z_1           x_2  \\\n1307  9IWF_1       G      1  -0.623  46.888 -23.394 -1.000000e+18<br>\n1308  9IWF_2       G      2  -6.260  48.627 -23.152 -1.000000e+18<br>\n1309  9IWF_3       U      3  -9.899  50.548 -19.612 -1.000000e+18<br>\n1310  9IWF_4       G      4 -11.262  52.573 -14.761 -1.000000e+18<br>\n1311  9IWF_5       U      5  -9.971  54.891 -10.007 -1.000000e+18   </p>\n<pre><code>  chain  copy   Usage  target  \n</code></pre>\n<p>1307      A     1  Public    9IWF<br>\n1308      A     1  Public    9IWF<br>\n1309      A     1  Public    9IWF<br>\n1310      A     1  Public    9IWF<br>\n1311      A     1  Public    9IWF  </p>\n<p>[5 rows x 127 columns]\nSaved: /kaggle/working/9IWF_validation_coords.csv</p>\n<p>Subject: Possible residue indexing shift between PDB structures and provided ground truth</p>\n<p>Hi,</p>\n<p>While inspecting the structure 9IWF, I noticed a potential inconsistency between the original PDB residue numbering and the residue indexing used in the provided ground truth files.</p>\n<p>In the original PDB structure, the first resolved residue starts at residue 2, meaning residue 1 is missing in the deposited structure. However, in \"validation_labels.csv\", the coordinates appear to be renumbered starting from residue 1.</p>\n<p>This effectively introduces a global shift of one residue between the PDB coordinates and the dataset indexing. If participants rely on the PDB residue IDs directly, this shift propagates through the entire sequence and can lead to incorrect coordinate alignment for all downstream residues.</p>\n<p>Since the evaluation is performed on C1' atom coordinates, this indexing mismatch may unintentionally penalize otherwise correct predictions if the mapping between sequence position and structural residue index is not handled carefully.</p>\n<p>Could you please clarify:</p>\n<ol>\n<li>Whether all structures in the dataset were renumbered sequentially starting from 1, regardless of the original PDB residue numbering.</li>\n<li>Whether the ground truth coordinates correspond exactly to the deposited PDB C1' atoms, or if any preprocessing (such as residue filtering or renumbering) was applied.</li>\n</ol>\n<p>Clarification on this would help ensure participants interpret the ground truth correctly and avoid systematic alignment shifts.</p>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3419830,
          "author_name": "Przemek Porebski",
          "author_url": "",
          "post_date": "2026-03-11T19:11:19.453000",
          "content": "<p>The sequence numbering in the <code>*_labels.csv</code> and submission files corresponds to the <code>sequence</code> in the <code>*_sequences.csv</code> file. We disregard residue numbering provided by the authors and renumber residues sequentially according to the <code>sequence</code> field. The <code>sequence</code> field is derived from <code>_entity_poly</code> and <code>_pdbx_poly_seq_scheme</code> records in the CIF files.  Per wwPDB <a href=\"https://www.wwpdb.org/documentation/procedure#toc_3\" target=\"_blank\">guidelines</a> these records should correspond to the sequence of the complete sample used for experiment. If the sequence covers residues that are missing in the model, then these will be reported in the <code>*_labels.csv</code> but will not have assigned coordinates. </p>\n<p>For this particular case authors decided to start numbering from 2, but they didn't report any unobserved residues prior to that. Sometimes that happens when authors decide to keep their numbering consistent with other entries, database or other reference source. </p>\n<p>Submissions are scored using 1:1 match of the residues between submission and solution. If the coordinates were not observed experimentally these will not affect the scoring as explained <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/666382#3414909\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3414048,
      "author_name": "Mohamed Hashem (GoxD)",
      "author_url": "",
      "post_date": "2026-02-25T20:04:47.153000",
      "content": "<p>Question 1 — Missing Residues / Gaps\nHello,\nI have a question regarding how the evaluation handles missing residues or discontinuities in the chain.\nIf a model predicts coordinates for positions that correspond to gaps or unresolved regions in the ground truth structure, how is the TM-score computed in that case?\nSpecifically:\nAre only residues present in the ground truth considered in the alignment?\nOr are predicted coordinates for non-existing residues also included in the TM-score calculation?\nI’m trying to understand whether predicting coordinates across gaps (e.g., filling missing segments) would negatively affect the evaluation, or if those positions are simply ignored during scoring.\nQuestion 2 — Long Artificial Straight Connections\nI also have a question about structural continuity.\nIf the model predicts a long artificial straight connection between two distant segments (for example, creating an unrealistic straight bridge in 3D space), how would this affect the TM-score?\nSince evaluation is based on the spatial placement of C1' atoms, would such long geometric distortions significantly penalize the score even if the overall global fold is approximately correct?\nIn other words, does TM-score heavily penalize unrealistic long-range geometric artifacts even when local regions are reasonably aligned?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3414909,
          "author_name": "Przemek Porebski",
          "author_url": "",
          "post_date": "2026-02-27T22:51:12.217000",
          "content": "<p>If the residues were not included in the ground truth they will not be considered in the alignment and modeling those will not affect a TM-score. However, if the residues that are not included in the prediction are present in the ground truth that will negatively  impact TM-scores as it is calculated relative to the number of residues in the ground truth </p>\n<p>I am not sure if I correctly understand the scenario for question 2, so I will provide several examples:</p>\n<ol>\n<li>Current format for the submission requires numbering of the residues according to concatenated sequence for multimeric structures. So it may happen that the residue <code>ID=target_id_200</code> (from one monomer) is distant to <code>ID=target_id_201</code> (from the other monomer). In such case this stretch is not penalized, as there are no atoms in-between </li>\n<li>For single monomer, If there is a gap in the prediction, and the <code>C1'</code> atoms are missing between residue 100 and 120, but those atoms are present in the ground truth then that will be penalized in TM-score</li>\n<li>If those missing residues  100-120 are filled with <code>C1'</code> atoms with unrealistic positions (for example placed uniformly along straight line) and those residues are present in the ground truth, that will be penalized in TM-score. The penalty will depend on the size of RNA and number of mismatched <code>C1'</code> positions.</li>\n<li>Similarly if there is unrealistic distance between <code>C1'</code> 100 and 101 from the same monomer such distance will likely put this and other atoms in wrong positions and that will be penalized in the TM-score. </li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3388677,
      "author_name": "Pi",
      "author_url": "",
      "post_date": "2026-01-09T11:08:54.437000",
      "content": "<p>Dear hosts Rhiju Das <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and Przemek Porebski <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a>,</p>\n<p>Following recent discussions regarding the disqualification of participants for reusing private notebooks from previous years (as seen in <a href=\"https://www.kaggle.com/competitions/cmi-detect-behavior-with-sensor-data/discussion/602648#3278013\" target=\"_blank\">this thread</a>), our team ( <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a>, <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>, and myself) would like to seek clarification on the rules for this transition.</p>\n<p>During Part 1, we worked as a team and shared private notebooks and datasets. As we move into Part 2, we would like to ensure we remain in full compliance under the following three scenarios:</p>\n<p><strong>1. Full Team Continuity:</strong> If all members of our Part 1 team compete together as the same team in Part 2, can we continue to use our private resources from Part 1?</p>\n<p><strong>2. Partial Team Continuity:</strong> If only a subset of our Part 1 team competes in Part 2 (and the remaining members do not participate in Part 2 at all), can the active members still use the private notebooks and datasets created during Part 1?</p>\n<p><strong>3. Individual Participation:</strong> If members of the Part 1 team decide to compete individually or on different teams in Part 2, what is the protocol for using our previous shared resources? (e.g., Must they be made public first?)</p>\n<p>Best,\nHoa</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3388968,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2026-01-09T22:33:11.097000",
          "content": "<p>Hi Hoa (and d4t4 team)! Great to have you participating again. I have relayed your question to Kaggle devs. They will need a little more time to complete internal discussions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3388978,
          "author_name": "Ashley Oldacre",
          "author_url": "",
          "post_date": "2026-01-09T23:32:03.550000",
          "content": "<p>Hello Hao, Thank you for your questions. \nYou are able to continue using your notebooks from part 1 and carry them over to part 2 whether your team is continuing as the original members or including new members. Best of luck to you! </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3391582,
              "author_name": "Pi",
              "author_url": "",
              "post_date": "2026-01-15T08:26:21.023000",
              "content": "<p>Thank <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> and  <a href=\"https://www.kaggle.com/ashleyoldacre\" target=\"_blank\">@ashleyoldacre</a> for your instant response!</p>\n<p>How about we join this season but in different teams?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3392342,
              "author_name": "Ashley Oldacre",
              "author_url": "",
              "post_date": "2026-01-16T18:30:07.667000",
              "content": "<p>Yes, no problem. You can join in different teams. Thank you for asking and good luck! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3388019,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2026-01-08T00:41:41.407000",
      "content": "<p>Thank you very much, I'm excited to see what comes from this competition. I'm curious, you noted that this competition will have significantly harder structures to predict (multiple chains, ligands, long chains, etc). Can you clarify, is the test set exclusively these difficult cases, or are there still examples of single, short chains in the test set?</p>\n<p>Edit: Also, how big roughly can we expect the test set to be? ~30 like the example, or will it be more like 100+?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3388031,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2026-01-08T01:56:45.683000",
          "content": "<p>The hidden test_set.csv will be a lot like the example test_set.csv, including in terms of number of targets. There will indeed be single short chains, but selected to be template-free. I've edited the welcome post above to clarify. Thanks for the question!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3388040,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2026-01-08T03:03:03.163000",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> template free means no close structural homologs available and template based matching will not help?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3388387,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2026-01-08T18:13:59.280000",
              "content": "<p>Yes I think template search may not help much with single chain.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3428229,
      "author_name": "kurshid basheer",
      "author_url": "",
      "post_date": "2026-03-25T05:33:02.033000",
      "content": "<p>heyy! just curious.. has anyone here used graph based approach?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3429142,
          "author_name": "Sagar Jose",
          "author_url": "",
          "post_date": "2026-03-26T12:31:57.307000",
          "content": "<p>I did. GNN.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3425994,
      "author_name": "Lokendra Singh",
      "author_url": "",
      "post_date": "2026-03-21T17:03:41.167000",
      "content": "<p>Amazing data </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3408052,
      "author_name": "Irfan Ullah Khan",
      "author_url": "",
      "post_date": "2026-02-19T17:59:22.373000",
      "content": "<p><a href=\"https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2\" target=\"_blank\">https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2</a></p>\n<p>please anyone check my notebook why it always give submission scoring error.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3422853,
          "author_name": "Rehan Daya",
          "author_url": "",
          "post_date": "2026-03-18T01:49:06.413000",
          "content": "<p>I'm having issues too, but at the very least \"x,y,z are clipped between -999.999 and 9999.999 before scoring\". You may want to try to do this before submitting and see if it helps</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3426420,
              "author_name": "Jiacheng Ma",
              "author_url": "",
              "post_date": "2026-03-22T15:06:39.980000",
              "content": "<p>Have you fixed it? I need help.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3426421,
          "author_name": "Jiacheng Ma",
          "author_url": "",
          "post_date": "2026-03-22T15:07:01.687000",
          "content": "<p>Have you fixed it? help me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3407880,
      "author_name": "Irfan Ullah Khan",
      "author_url": "",
      "post_date": "2026-02-19T11:14:04.193000",
      "content": "<p>I submitted my notebook 16 times code cells runs perfectly but when I submitted to competition it gives scoring error can anyone helps to solve my this problem , Everything is perfect code and files submitted in csv. But always give scoring error.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3426422,
          "author_name": "Jiacheng Ma",
          "author_url": "",
          "post_date": "2026-03-22T15:07:34.197000",
          "content": "<p>Have you fixed it? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3426950,
              "author_name": "Irfan Ullah Khan",
              "author_url": "",
              "post_date": "2026-03-23T14:35:37.860000",
              "content": "<p>Not yet.. still working on it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3403531,
      "author_name": "AliAsghar",
      "author_url": "",
      "post_date": "2026-02-08T17:42:49.213000",
      "content": "<p>3D prediction that brings real biological complexity and pushes the field beyond expert-level performance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3403477,
      "author_name": "NGABO FREDDY",
      "author_url": "",
      "post_date": "2026-02-08T15:25:34.347000",
      "content": "<p>This challenge is incredibly exciting and inspiring. The progress made in Part 1, especially reaching performance comparable to expert human groups, shows how powerful collaborative science and machine learning can be. Expanding the scope in Part 2 to include multi-chain RNAs, RNA complexes, ligand-dependent conformations, and larger systems makes the problem much closer to real biological scenarios, which is very motivating.</p>\n<p>The updated evaluation metric and availability of the template oracle baseline will likely push participants to develop more robust and generalizable methods. I am especially interested in how hybrid approaches combining template-based modeling, deep learning, and chemical mapping data will perform. The possibility of agentic AI systems accelerating discovery is also fascinating.</p>\n<p>Thank you to the organizers for creating such a meaningful competition that directly contributes to advances in biology, medicine, and biotechnology. I am excited to follow the solutions developed by the community and to learn from this challenge.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3397222,
      "author_name": "Ridha Mostefa from algeria",
      "author_url": "",
      "post_date": "2026-01-26T18:25:14.320000",
      "content": "<p>مرحبا معكم  رضا من الجزائر مشارك لأول مرة ، واجهت صعوبات كبيرة قبل ارسال اول دفنر ملاحظات وبعد محاولات عديدة بقيت علامتي عندد القيمة .0.098 هل تعتبر جيدة كمشارك هاو لأول  مرة وعندي طلب لماذا لا يتم توسيع عدد مرات ارسال دفتر البيانات او عدم حساب حالة الرفض حتى نتمكن من فهم نقاط الضعف فاهدف هنا إنساني اكبر منه حافز مالي. وشكرا</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3391366,
      "author_name": "oscar bell",
      "author_url": "",
      "post_date": "2026-01-14T19:41:47.320000",
      "content": "<p>should I use just \"train_sequences.csv\" for the training? what are the purpose of MSA and PDB_RNA folders? that's not really clear</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3390775,
      "author_name": "Bartosz Bartczak",
      "author_url": "",
      "post_date": "2026-01-13T23:32:41.357000",
      "content": "<p>Hi,\nI don't understand why I'm getting such widely divergent results using the theoretically identical measurement method provided by Kaggle at <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score</a>.\nLocal measurements show that the model is learning and the TM-Score is increasing, reaching 0.201 by the 10th epoch, while the same model submitted to Kaggle only achieved a TM-Score of 0.067.\nThere shouldn't be such drastic differences. Can someone explain this to me?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3407562,
          "author_name": "Andres H. Zapke",
          "author_url": "",
          "post_date": "2026-02-18T16:13:16.747000",
          "content": "<p>You likely computed TM score for the training dataset, while 0.067 is TM-score for a unseen data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3416497,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-03-03T02:48:03.463000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3388056,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-01-08T04:27:22.207000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3387313": "In 2025, in [Part 1 of this competition](https://www.kaggle.com/c/stanford-rna-3d-folding), Kagglers took on the fundamental problem of RNA 3D structure prediction.\n\nRNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.\n\nUnfortunately, the world's efforts to advance biology and biotechnology have been slowed down by our inability to computationally predict these RNA 3D structures.\n\nIn 2025, Kaggle teams tackled this problem and achieved automated RNA 3D structure prediction methods that, for the first time, tied with a top human expert group!\n\nA big surprise was the revival of template-based modeling, a non-deep-learning approach. Check out our summary paper at [Template-based RNA structure prediction advanced through a blind code competition](https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full).\n\nPart 1 focused on single RNA chains under 1000 nts forming largely single structures. \n\nPart 2 now adds more  use cases important for biology and biotechnology:\n- Single chain RNAs whose structures are different from previously available templates.\n- RNA complexes with multiple RNA chains\n- RNA 'machines' with multiple states - try to capture each of the states in your five predictions.\n- RNAs whose conformations depend on proteins, other nucleic acids, and/or small molecule ligands.\n- RNA systems with sizes up to 5500 nucleotides.\n\nWe've also updated the evaluation metric to be more stringent. The TM-score metric now only counts residues in the prediction and ground truth as ‘aligned’ if the residues have matching residue numbers.\n\nLast, we've updated our baseline. In Part 1, for each target, the best-of-Kaggle prediction was similar or just barely better than the best publicly available prior 3D template, as discovered by an oracle with knowledge of the ground truth. This time, we’re making the `best_template_oracle` score available on the Public leaderboard (0.554). \n\nThere's a lot we're curious about:\n - How hard will it be to generalize Part 1's notebooks to multiple chains or proteins?\n - How long will it take to beat the best template oracle on the Public leaderboard? Will top models generalize to the Private leaderboard?\n - Will literature-aware LLMs that read the English language `description` fields be able to find templates missed last time?\n - Will Kaggle notebooks be able to leverage chemical mapping profiles for millions of RNAs, available in the 2024 Ribonanza Kaggle competition [train](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data) and [test](https://www.kaggle.com/datasets/rhijudas/ribonanza-solutions/) data?\n - The maximum possible mean TM-score allowed by the inherent flexibility of these RNA targets is about 0.8. Can codes achieve this milestone in the RNA structure prediction problem? \n - Can a new generation of agentic AI approaches and problem solvers – here receiving feedback from the Public leaderboard – accelerate research in Part 2?\n\nWe’re looking forward to seeing what Kagglers come up with. We welcome back all who participated in Part 1 – your notebooks should be submittable to Part 2 with few or no updates. And we are excited to meet Kagglers just approaching RNA for the first time -- the fresh insights and high-placing models of teams with a \"beginner's mindset\" have had a huge impact in prior RNA competitions.\n\nFor the hosts, \n\nRhiju Das @rhijudas, Przemek Porebski @przemekporebski \n\nAnd many thanks to:\n* The AI@HHMI initiative at Janelia Research Center as well as several experimental RNA cryo-EM groups providing blind targets.\n* Our wonderful collaborators at Kaggle: Walter Reade @inversion and Ashley Oldacre @ashleyoldacre\n* The co-authors of our [Part 1 paper](https://www.biorxiv.org/content/10.64898/2025.12.30.696949v1.full), with special thanks to CASP and RNA-Puzzle organizers who advised on the design of Part 2. \n\nP.S. In case you were wondering, the top Kaggle teams from Part 1 were *not* informed in advance about Part 2!\n\n",
    "3420677": "i AM HAVING A MENTAL BREAKDOWN.  i BUILT A PHYSICS AND MATH-BASED ENGINE FROM MY ORIGIN OF LIFE EXPERTISE AND i CANNOT EVEN GET MY CODE TO GIVE ME A SCORE.  iT RENDERS BUT i THINK i HAVE IMPROPER FORMATTING.  i AM ONE OF THE WORLD'S LEADING INNOVATORS IN QUANTUM BIOPHYSICS BUT i CANNOT CODE  FOR THE LIFE OF ME AND IT IS SO FRUSTRATING.  i HAVE A WHOLE FRAMEWORK WITH TESTED MECHANISMS THAT i CAN DO BY HAND AND GET A VERY HIGH tm SCORE, A MATHEMATICIAN, BUT i HAVE BEEN WORKING ON THIS FOR 9 DAYS AND HONESTLY, MIGHT NEED TO FIND SOME PROFESSIONAL HELP.  iS IT TOO LATE TO JOIN A TEAM???\n",
    "3418654": "Is the residue numbering in the dataset aligned with the original PDB numbering, or are residues renumbered starting from 1 even if the first residue is missing in the PDB structure?\n\n\nAre the target coordinates evaluated strictly on the C1' atom positions from the deposited PDB structure, or after renumbering / filtering missing residues?\n\n\nFound: /kaggle/input/competitions/stanford-rna-3d-folding-2/PDB_RNA/9iwf.cif\n  chain  resid resname       x          y          z\n0     A      2       G  -0.623  46.888000 -23.393999\n1     A      3       G  -6.260  48.626999 -23.152000\n2     A      4       U  -9.899  50.548000 -19.612000\n3     A      5       G -11.262  52.573002 -14.761000\n4     A      6       U  -9.971  54.890999 -10.007000\n\nSaved: /kaggle/working/\n\n\n\nFound rows: 69\n          ID resname  resid     x_1     y_1     z_1           x_2  \\\n1307  9IWF_1       G      1  -0.623  46.888 -23.394 -1.000000e+18   \n1308  9IWF_2       G      2  -6.260  48.627 -23.152 -1.000000e+18   \n1309  9IWF_3       U      3  -9.899  50.548 -19.612 -1.000000e+18   \n1310  9IWF_4       G      4 -11.262  52.573 -14.761 -1.000000e+18   \n1311  9IWF_5       U      5  -9.971  54.891 -10.007 -1.000000e+18   \n\n\n      chain  copy   Usage  target  \n1307      A     1  Public    9IWF  \n1308      A     1  Public    9IWF  \n1309      A     1  Public    9IWF  \n1310      A     1  Public    9IWF  \n1311      A     1  Public    9IWF  \n\n[5 rows x 127 columns]\nSaved: /kaggle/working/9IWF_validation_coords.csv\n\n\nSubject: Possible residue indexing shift between PDB structures and provided ground truth\n\nHi,\n\nWhile inspecting the structure 9IWF, I noticed a potential inconsistency between the original PDB residue numbering and the residue indexing used in the provided ground truth files.\n\nIn the original PDB structure, the first resolved residue starts at residue 2, meaning residue 1 is missing in the deposited structure. However, in \"validation_labels.csv\", the coordinates appear to be renumbered starting from residue 1.\n\nThis effectively introduces a global shift of one residue between the PDB coordinates and the dataset indexing. If participants rely on the PDB residue IDs directly, this shift propagates through the entire sequence and can lead to incorrect coordinate alignment for all downstream residues.\n\nSince the evaluation is performed on C1' atom coordinates, this indexing mismatch may unintentionally penalize otherwise correct predictions if the mapping between sequence position and structural residue index is not handled carefully.\n\nCould you please clarify:\n\n1. Whether all structures in the dataset were renumbered sequentially starting from 1, regardless of the original PDB residue numbering.\n2. Whether the ground truth coordinates correspond exactly to the deposited PDB C1' atoms, or if any preprocessing (such as residue filtering or renumbering) was applied.\n\nClarification on this would help ensure participants interpret the ground truth correctly and avoid systematic alignment shifts.\n\nThanks!\n",
    "3414048": "Question 1 — Missing Residues / Gaps\nHello,\nI have a question regarding how the evaluation handles missing residues or discontinuities in the chain.\nIf a model predicts coordinates for positions that correspond to gaps or unresolved regions in the ground truth structure, how is the TM-score computed in that case?\nSpecifically:\nAre only residues present in the ground truth considered in the alignment?\nOr are predicted coordinates for non-existing residues also included in the TM-score calculation?\nI’m trying to understand whether predicting coordinates across gaps (e.g., filling missing segments) would negatively affect the evaluation, or if those positions are simply ignored during scoring.\nQuestion 2 — Long Artificial Straight Connections\nI also have a question about structural continuity.\nIf the model predicts a long artificial straight connection between two distant segments (for example, creating an unrealistic straight bridge in 3D space), how would this affect the TM-score?\nSince evaluation is based on the spatial placement of C1' atoms, would such long geometric distortions significantly penalize the score even if the overall global fold is approximately correct?\nIn other words, does TM-score heavily penalize unrealistic long-range geometric artifacts even when local regions are reasonably aligned?",
    "3388677": "Dear hosts Rhiju Das @rhijudas and Przemek Porebski @przemekporebski,\n\nFollowing recent discussions regarding the disqualification of participants for reusing private notebooks from previous years (as seen in [this thread](https://www.kaggle.com/competitions/cmi-detect-behavior-with-sensor-data/discussion/602648#3278013)), our team ( @hengck23 , @lihaoweicvch, @arunodhayan, and myself) would like to seek clarification on the rules for this transition.\n\nDuring Part 1, we worked as a team and shared private notebooks and datasets. As we move into Part 2, we would like to ensure we remain in full compliance under the following three scenarios:\n\n**1. Full Team Continuity:** If all members of our Part 1 team compete together as the same team in Part 2, can we continue to use our private resources from Part 1?\n\n**2. Partial Team Continuity:** If only a subset of our Part 1 team competes in Part 2 (and the remaining members do not participate in Part 2 at all), can the active members still use the private notebooks and datasets created during Part 1?\n\n**3. Individual Participation:** If members of the Part 1 team decide to compete individually or on different teams in Part 2, what is the protocol for using our previous shared resources? (e.g., Must they be made public first?)\n\nBest,\nHoa",
    "3388019": "Thank you very much, I'm excited to see what comes from this competition. I'm curious, you noted that this competition will have significantly harder structures to predict (multiple chains, ligands, long chains, etc). Can you clarify, is the test set exclusively these difficult cases, or are there still examples of single, short chains in the test set?\n\nEdit: Also, how big roughly can we expect the test set to be? ~30 like the example, or will it be more like 100+?",
    "3428229": "heyy! just curious.. has anyone here used graph based approach?",
    "3425994": "Amazing data ",
    "3408052": "https://www.kaggle.com/code/alvintera/standford-rna-3d-folding-part-2\n\nplease anyone check my notebook why it always give submission scoring error.",
    "3407880": "I submitted my notebook 16 times code cells runs perfectly but when I submitted to competition it gives scoring error can anyone helps to solve my this problem , Everything is perfect code and files submitted in csv. But always give scoring error.\n",
    "3403531": "3D prediction that brings real biological complexity and pushes the field beyond expert-level performance.",
    "3403477": "This challenge is incredibly exciting and inspiring. The progress made in Part 1, especially reaching performance comparable to expert human groups, shows how powerful collaborative science and machine learning can be. Expanding the scope in Part 2 to include multi-chain RNAs, RNA complexes, ligand-dependent conformations, and larger systems makes the problem much closer to real biological scenarios, which is very motivating.\n\nThe updated evaluation metric and availability of the template oracle baseline will likely push participants to develop more robust and generalizable methods. I am especially interested in how hybrid approaches combining template-based modeling, deep learning, and chemical mapping data will perform. The possibility of agentic AI systems accelerating discovery is also fascinating.\n\nThank you to the organizers for creating such a meaningful competition that directly contributes to advances in biology, medicine, and biotechnology. I am excited to follow the solutions developed by the community and to learn from this challenge.",
    "3397222": "مرحبا معكم  رضا من الجزائر مشارك لأول مرة ، واجهت صعوبات كبيرة قبل ارسال اول دفنر ملاحظات وبعد محاولات عديدة بقيت علامتي عندد القيمة .0.098 هل تعتبر جيدة كمشارك هاو لأول  مرة وعندي طلب لماذا لا يتم توسيع عدد مرات ارسال دفتر البيانات او عدم حساب حالة الرفض حتى نتمكن من فهم نقاط الضعف فاهدف هنا إنساني اكبر منه حافز مالي. وشكرا",
    "3391366": "should I use just \"train_sequences.csv\" for the training? what are the purpose of MSA and PDB_RNA folders? that's not really clear\n",
    "3390775": "Hi,\nI don't understand why I'm getting such widely divergent results using the theoretically identical measurement method provided by Kaggle at https://www.kaggle.com/code/metric/ribonanza-tm-score.\nLocal measurements show that the model is learning and the TM-Score is increasing, reaching 0.201 by the 10th epoch, while the same model submitted to Kaggle only achieved a TM-Score of 0.067.\nThere shouldn't be such drastic differences. Can someone explain this to me?",
    "3416497": "",
    "3388056": "Thanks @rhijudas "
  }
}