{
  "id": 689697,
  "title": "3rd Place Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/689697",
  "author_name": "Stefan Stefanov",
  "post_date": "2026-04-09T11:53:13.733000",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>Acknowledgements</h2>\n<p>I would like to thank the competition hosts and Kaggle for the opportunity to work on such an impactful and interesting problem.<br>\nThank you <a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> for sharing strong solutions at the beginning of the competition.<br>\nCredit to Claude models for the code assistance.</p>\n<h2>Overview</h2>\n<p>My strategy for the competition was to build upon the provided strong solutions TBM, RNAPro and Protenix by incorporating into the modeling new challenges added in Part 2: proteins, DNA, ligands, multiple chains and long sequences.<br>\nAnother aspect I considered important was processing speed in order to be able to include diverse predictions in the submission within the runtime constraints.<br>\nThe solution combines two predictions from TBM, two predictions from Protenix, and one prediction from RNAPro for all targets. There is no additional logic for selecting specific predictions per RNA sequence.<br>\nRNAPro and Protenix inferences run in parallel on the two T4 GPUs. </p>\n<h2>Template-based modeling (TBM)</h2>\n<p>For template-based modeling, a LightGBM model predicting TM-score for a given query-template pair is introduced. From the competition train data, the 1000 most recent query RNA sequences and 200 (per query) template sequences with a preceding temporal cutoff are used to create a training dataset of ~198K query-template pairs. TM-score for each pair is computed and used as a label. The model features are:</p>\n<ul>\n<li>RNA sequences similarity: <code>alignment_score</code>, <code>percent_identity</code>, and <code>len_diff_ratio</code>  </li>\n<li>Text embeddings similarity: cosine similarity between query description and template description embeddings.  The embeddings are created with this model: <a href=\"https://huggingface.co/NeuML/pubmedbert-base-embeddings\" target=\"_blank\">https://huggingface.co/NeuML/pubmedbert-base-embeddings</a></li>\n<li>Proteins similarity: alignment scores with a BLOSUM62 substitution matrix. For each query protein, the max alignment score with template proteins is taken. The min, mean, and max of the resulting vector are added as features. For processing speed considerations, only up to 6 query proteins are compared with up to 20 template proteins, and protein sequences are cropped to a max length of 768. These features are named <code>protein_similarity_matrix_(min|mean|max)</code> in the model.  </li>\n<li>DNA similarity: <code>dna_similarity_matrix_(min|mean|max)</code> features calculated with the logic described above for proteins using alignment scores without a substitution matrix.  </li>\n<li>Ligands similarity: <code>ligand_similarity_matrix_(min|mean|max)</code> features using Tanimoto similarities of Morgan fingerprints calculated with the logic described above for proteins.  </li>\n<li>Composition count features: <code>num_query_(proteins|dna|ligands)</code> and <code>num_template_(proteins|dna|ligands)</code>  </li>\n<li>Chains-related features: <code>(query|template)_num_all_chains</code>, <code>(query|template)_num_unique_chains</code>, and <code>chain_counts_match</code></li>\n</ul>\n<p>The top-2 templates according to the scores predicted by this model are used for the final submission.  </p>\n<p>The script for running this step is <code>srna3d.scripts.create_tbm_submission.</code><br>\nThe script for creating the training dataset is <code>srna3d.scripts.create_tmscore_dataset</code>. The dataset was created locally.<br>\nNotebook how to run it: <a href=\"https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset\" target=\"_blank\">https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset</a><br>\nThe training dataset: <a href=\"https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset</a><br>\nThe model was also trained locally. The script for training is <code>srna3d.scripts.train_tmscore_lightgbm_model</code><br>\nTraining notebook: <a href=\"https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model\" target=\"_blank\">https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model</a>  </p>\n<h2>RNAPro</h2>\n<p>One RNAPro model prediction is generated using the top-5 templates from the TBM step as input. To speed up inference, diffusion steps are decreased to 100, MSA depth is limited to 2048, and <code>fast_layernorm</code> and <code>triattention</code> triangle attention kernels are used, as inference is faster with them on longer sequences. As long sequences are split into chunks, the MSA and templates are also sliced to be consistent with the predicted chunk start and end index.</p>\n<h2>Protenix</h2>\n<p>Two predictions are generated with the latest Protenix model <code>protenix_base_20250630_v1.0.0</code>. In addition to the RNA sequence, proteins, DNA, and ligands are also added as input to the Protenix model. To fit in the T4 GPU memory and time constraints, proteins, DNA, and ligands are limited to a maximum of 6 per group. A <code>max_len_non_rna</code> parameter is added and set to <code>128</code>. This non-RNA sequence length budget is divided equally among protein and DNA sequences, and they are cut accordingly. The idea is to model the effect of short protein and DNA sequences on the RNA 3D structure, at least for targets with a small number of proteins or DNA. Similar to RNAPro, MSA depth is limited to 2048, and <code>fast_layernorm</code> and <code>triattention</code> triangle attention kernels are used.</p>\n<h2>Long Sequences</h2>\n<p>For the RNAPro and Protenix models, RNA with a sequence length above 448 is split in two steps:</p>\n<p>1) If it has multiple chains, it is first split into chains, and for each chain, an overlap of 96 from the next chain is added.<br>\n   For the last chain, the overlap is added circularly from the first chain with the idea to try to capture potential circularity in structure.<br>\n   For faster prediction, repeated chain pairs are skipped. For example, <code>9MME</code> which has U:8 chains with a total length of 4640 (8x580) is split and passed to the next step as one sequence of length 676 (580 from the first U + 96 from the second U chain).<br>\n   If the target has multiple chains, but its total sequence length is below 448, no such splitting is performed.\nThis splitting step is performed by the <code>srna3d.scripts.split_into_chains</code> script.</p>\n<p>2) Sequences from the previous step that are longer than 448 are split into chunks of length 448, having a 96-base overlap with the next chunk. This step is performed by the <code>srna3d.scripts.split_sequences_into_chunks</code> script.</p>\n<p>Chunk predictions are first combined from chunks to chains using the <code>srna3d.scripts.combine_chunked_predictions</code> script. Afterward, they are combined from chains to full sequences with the <code>srna3d.scripts.combine_chain_predictions</code> script, applying Kabsch alignment on the overlapping regions. There is an option to randomly select and combine chain predictions from different samples, but it isn’t applied; for RNAPro, one sample is generated, and for Protenix, two samples are generated.</p>\n<p>The solution includes various limits due to GPU memory and runtime constraints. It would be interesting to see what improvements this and other solutions could achieve with faster GPUs having more memory.</p>\n<h2>References</h2>\n<p>Rao, G. John, RNAPro inference with TBM notebook: <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm</a></p>\n<p>Viel, Theo, et al. RNAPro Inference notebook: <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference</a></p>\n<p>Lee, Youhan, et al. (2025). Template-based RNA structure prediction advanced through a blind code competition. <em>bioRxiv</em>. doi: 10.64898/2025.12.30.696949.<br>\n<a href=\"https://github.com/NVIDIA-Digital-Bio/RNAPro\" target=\"_blank\">https://github.com/NVIDIA-Digital-Bio/RNAPro</a></p>\n<p>Zhang, Yuxuan, et al. (2026). Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction. <em>bioRxiv</em>. doi: 10.64898/2026.02.05.703733.<br>\n<a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a>  </p>",
  "messages": [
    {
      "id": 3438585,
      "postDate": "2026-04-09T11:53:13.733Z",
      "content": "<h2>Acknowledgements</h2>\n<p>I would like to thank the competition hosts and Kaggle for the opportunity to work on such an impactful and interesting problem.<br>\nThank you <a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> for sharing strong solutions at the beginning of the competition.<br>\nCredit to Claude models for the code assistance.</p>\n<h2>Overview</h2>\n<p>My strategy for the competition was to build upon the provided strong solutions TBM, RNAPro and Protenix by incorporating into the modeling new challenges added in Part 2: proteins, DNA, ligands, multiple chains and long sequences.<br>\nAnother aspect I considered important was processing speed in order to be able to include diverse predictions in the submission within the runtime constraints.<br>\nThe solution combines two predictions from TBM, two predictions from Protenix, and one prediction from RNAPro for all targets. There is no additional logic for selecting specific predictions per RNA sequence.<br>\nRNAPro and Protenix inferences run in parallel on the two T4 GPUs. </p>\n<h2>Template-based modeling (TBM)</h2>\n<p>For template-based modeling, a LightGBM model predicting TM-score for a given query-template pair is introduced. From the competition train data, the 1000 most recent query RNA sequences and 200 (per query) template sequences with a preceding temporal cutoff are used to create a training dataset of ~198K query-template pairs. TM-score for each pair is computed and used as a label. The model features are:</p>\n<ul>\n<li>RNA sequences similarity: <code>alignment_score</code>, <code>percent_identity</code>, and <code>len_diff_ratio</code>  </li>\n<li>Text embeddings similarity: cosine similarity between query description and template description embeddings.  The embeddings are created with this model: <a href=\"https://huggingface.co/NeuML/pubmedbert-base-embeddings\" target=\"_blank\">https://huggingface.co/NeuML/pubmedbert-base-embeddings</a></li>\n<li>Proteins similarity: alignment scores with a BLOSUM62 substitution matrix. For each query protein, the max alignment score with template proteins is taken. The min, mean, and max of the resulting vector are added as features. For processing speed considerations, only up to 6 query proteins are compared with up to 20 template proteins, and protein sequences are cropped to a max length of 768. These features are named <code>protein_similarity_matrix_(min|mean|max)</code> in the model.  </li>\n<li>DNA similarity: <code>dna_similarity_matrix_(min|mean|max)</code> features calculated with the logic described above for proteins using alignment scores without a substitution matrix.  </li>\n<li>Ligands similarity: <code>ligand_similarity_matrix_(min|mean|max)</code> features using Tanimoto similarities of Morgan fingerprints calculated with the logic described above for proteins.  </li>\n<li>Composition count features: <code>num_query_(proteins|dna|ligands)</code> and <code>num_template_(proteins|dna|ligands)</code>  </li>\n<li>Chains-related features: <code>(query|template)_num_all_chains</code>, <code>(query|template)_num_unique_chains</code>, and <code>chain_counts_match</code></li>\n</ul>\n<p>The top-2 templates according to the scores predicted by this model are used for the final submission.  </p>\n<p>The script for running this step is <code>srna3d.scripts.create_tbm_submission.</code><br>\nThe script for creating the training dataset is <code>srna3d.scripts.create_tmscore_dataset</code>. The dataset was created locally.<br>\nNotebook how to run it: <a href=\"https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset\" target=\"_blank\">https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset</a><br>\nThe training dataset: <a href=\"https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset</a><br>\nThe model was also trained locally. The script for training is <code>srna3d.scripts.train_tmscore_lightgbm_model</code><br>\nTraining notebook: <a href=\"https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model\" target=\"_blank\">https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model</a>  </p>\n<h2>RNAPro</h2>\n<p>One RNAPro model prediction is generated using the top-5 templates from the TBM step as input. To speed up inference, diffusion steps are decreased to 100, MSA depth is limited to 2048, and <code>fast_layernorm</code> and <code>triattention</code> triangle attention kernels are used, as inference is faster with them on longer sequences. As long sequences are split into chunks, the MSA and templates are also sliced to be consistent with the predicted chunk start and end index.</p>\n<h2>Protenix</h2>\n<p>Two predictions are generated with the latest Protenix model <code>protenix_base_20250630_v1.0.0</code>. In addition to the RNA sequence, proteins, DNA, and ligands are also added as input to the Protenix model. To fit in the T4 GPU memory and time constraints, proteins, DNA, and ligands are limited to a maximum of 6 per group. A <code>max_len_non_rna</code> parameter is added and set to <code>128</code>. This non-RNA sequence length budget is divided equally among protein and DNA sequences, and they are cut accordingly. The idea is to model the effect of short protein and DNA sequences on the RNA 3D structure, at least for targets with a small number of proteins or DNA. Similar to RNAPro, MSA depth is limited to 2048, and <code>fast_layernorm</code> and <code>triattention</code> triangle attention kernels are used.</p>\n<h2>Long Sequences</h2>\n<p>For the RNAPro and Protenix models, RNA with a sequence length above 448 is split in two steps:</p>\n<p>1) If it has multiple chains, it is first split into chains, and for each chain, an overlap of 96 from the next chain is added.<br>\n   For the last chain, the overlap is added circularly from the first chain with the idea to try to capture potential circularity in structure.<br>\n   For faster prediction, repeated chain pairs are skipped. For example, <code>9MME</code> which has U:8 chains with a total length of 4640 (8x580) is split and passed to the next step as one sequence of length 676 (580 from the first U + 96 from the second U chain).<br>\n   If the target has multiple chains, but its total sequence length is below 448, no such splitting is performed.\nThis splitting step is performed by the <code>srna3d.scripts.split_into_chains</code> script.</p>\n<p>2) Sequences from the previous step that are longer than 448 are split into chunks of length 448, having a 96-base overlap with the next chunk. This step is performed by the <code>srna3d.scripts.split_sequences_into_chunks</code> script.</p>\n<p>Chunk predictions are first combined from chunks to chains using the <code>srna3d.scripts.combine_chunked_predictions</code> script. Afterward, they are combined from chains to full sequences with the <code>srna3d.scripts.combine_chain_predictions</code> script, applying Kabsch alignment on the overlapping regions. There is an option to randomly select and combine chain predictions from different samples, but it isn’t applied; for RNAPro, one sample is generated, and for Protenix, two samples are generated.</p>\n<p>The solution includes various limits due to GPU memory and runtime constraints. It would be interesting to see what improvements this and other solutions could achieve with faster GPUs having more memory.</p>\n<h2>References</h2>\n<p>Rao, G. John, RNAPro inference with TBM notebook: <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm</a></p>\n<p>Viel, Theo, et al. RNAPro Inference notebook: <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference</a></p>\n<p>Lee, Youhan, et al. (2025). Template-based RNA structure prediction advanced through a blind code competition. <em>bioRxiv</em>. doi: 10.64898/2025.12.30.696949.<br>\n<a href=\"https://github.com/NVIDIA-Digital-Bio/RNAPro\" target=\"_blank\">https://github.com/NVIDIA-Digital-Bio/RNAPro</a></p>\n<p>Zhang, Yuxuan, et al. (2026). Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction. <em>bioRxiv</em>. doi: 10.64898/2026.02.05.703733.<br>\n<a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a>  </p>",
      "rawMarkdown": "## Acknowledgements\n\nI would like to thank the competition hosts and Kaggle for the opportunity to work on such an impactful and interesting problem.  \nThank you @jaejohn and @theoviel for sharing strong solutions at the beginning of the competition.  \nCredit to Claude models for the code assistance.\n\n## Overview \n\nMy strategy for the competition was to build upon the provided strong solutions TBM, RNAPro and Protenix by incorporating into the modeling new challenges added in Part 2: proteins, DNA, ligands, multiple chains and long sequences.  \nAnother aspect I considered important was processing speed in order to be able to include diverse predictions in the submission within the runtime constraints.  \nThe solution combines two predictions from TBM, two predictions from Protenix, and one prediction from RNAPro for all targets. There is no additional logic for selecting specific predictions per RNA sequence.  \nRNAPro and Protenix inferences run in parallel on the two T4 GPUs. \n\n## Template-based modeling (TBM)\n\nFor template-based modeling, a LightGBM model predicting TM-score for a given query-template pair is introduced. From the competition train data, the 1000 most recent query RNA sequences and 200 (per query) template sequences with a preceding temporal cutoff are used to create a training dataset of ~198K query-template pairs. TM-score for each pair is computed and used as a label. The model features are:\n\n* RNA sequences similarity: `alignment_score`, `percent_identity`, and `len_diff_ratio`  \n* Text embeddings similarity: cosine similarity between query description and template description embeddings.  The embeddings are created with this model: https://huggingface.co/NeuML/pubmedbert-base-embeddings\n* Proteins similarity: alignment scores with a BLOSUM62 substitution matrix. For each query protein, the max alignment score with template proteins is taken. The min, mean, and max of the resulting vector are added as features. For processing speed considerations, only up to 6 query proteins are compared with up to 20 template proteins, and protein sequences are cropped to a max length of 768\\. These features are named `protein_similarity_matrix_(min|mean|max)` in the model.  \n* DNA similarity: `dna_similarity_matrix_(min|mean|max)` features calculated with the logic described above for proteins using alignment scores without a substitution matrix.  \n* Ligands similarity: `ligand_similarity_matrix_(min|mean|max)` features using Tanimoto similarities of Morgan fingerprints calculated with the logic described above for proteins.  \n* Composition count features: `num_query_(proteins|dna|ligands)` and `num_template_(proteins|dna|ligands)`  \n* Chains-related features: `(query|template)_num_all_chains`, `(query|template)_num_unique_chains`, and `chain_counts_match`\n\nThe top-2 templates according to the scores predicted by this model are used for the final submission.  \n\nThe script for running this step is `srna3d.scripts.create_tbm_submission.`  \nThe script for creating the training dataset is `srna3d.scripts.create_tmscore_dataset`. The dataset was created locally.  \nNotebook how to run it: https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset  \nThe training dataset: https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset  \nThe model was also trained locally. The script for training is `srna3d.scripts.train_tmscore_lightgbm_model`   \nTraining notebook: https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model  \n\n## RNAPro\n\nOne RNAPro model prediction is generated using the top-5 templates from the TBM step as input. To speed up inference, diffusion steps are decreased to 100, MSA depth is limited to 2048, and `fast_layernorm` and `triattention` triangle attention kernels are used, as inference is faster with them on longer sequences. As long sequences are split into chunks, the MSA and templates are also sliced to be consistent with the predicted chunk start and end index.\n\n## Protenix\n\nTwo predictions are generated with the latest Protenix model `protenix_base_20250630_v1.0.0`. In addition to the RNA sequence, proteins, DNA, and ligands are also added as input to the Protenix model. To fit in the T4 GPU memory and time constraints, proteins, DNA, and ligands are limited to a maximum of 6 per group. A `max_len_non_rna` parameter is added and set to `128`. This non-RNA sequence length budget is divided equally among protein and DNA sequences, and they are cut accordingly. The idea is to model the effect of short protein and DNA sequences on the RNA 3D structure, at least for targets with a small number of proteins or DNA. Similar to RNAPro, MSA depth is limited to 2048, and `fast_layernorm` and `triattention` triangle attention kernels are used.\n\n## Long Sequences\n\nFor the RNAPro and Protenix models, RNA with a sequence length above 448 is split in two steps:\n\n1) If it has multiple chains, it is first split into chains, and for each chain, an overlap of 96 from the next chain is added.  \n   For the last chain, the overlap is added circularly from the first chain with the idea to try to capture potential circularity in structure.  \n   For faster prediction, repeated chain pairs are skipped. For example, `9MME` which has U:8 chains with a total length of 4640 (8x580) is split and passed to the next step as one sequence of length 676 (580 from the first U + 96 from the second U chain).  \n   If the target has multiple chains, but its total sequence length is below 448, no such splitting is performed.\nThis splitting step is performed by the `srna3d.scripts.split_into_chains` script.\n\n2) Sequences from the previous step that are longer than 448 are split into chunks of length 448, having a 96-base overlap with the next chunk. This step is performed by the `srna3d.scripts.split_sequences_into_chunks` script.\n\nChunk predictions are first combined from chunks to chains using the `srna3d.scripts.combine_chunked_predictions` script. Afterward, they are combined from chains to full sequences with the `srna3d.scripts.combine_chain_predictions` script, applying Kabsch alignment on the overlapping regions. There is an option to randomly select and combine chain predictions from different samples, but it isn’t applied; for RNAPro, one sample is generated, and for Protenix, two samples are generated.\n\nThe solution includes various limits due to GPU memory and runtime constraints. It would be interesting to see what improvements this and other solutions could achieve with faster GPUs having more memory.\n\n## References\n\nRao, G. John, RNAPro inference with TBM notebook: https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\n\nViel, Theo, et al. RNAPro Inference notebook: https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\n\nLee, Youhan, et al. (2025). Template-based RNA structure prediction advanced through a blind code competition. *bioRxiv*. doi: 10.64898/2025.12.30.696949.   \nhttps://github.com/NVIDIA-Digital-Bio/RNAPro\n\nZhang, Yuxuan, et al. (2026). Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction. *bioRxiv*. doi: 10.64898/2026.02.05.703733.  \nhttps://github.com/bytedance/Protenix  \n",
      "votes": 16
    },
    {
      "id": 3438706,
      "postDate": "2026-04-09T15:48:28.293Z",
      "content": "<p>This is amazing work, you covered a lot of techniques and algorithm settings, thanks for sharing.</p>\n<p>How did you track progress, did you keep an eye on some kind of metric (clearly not on the public LB only:)? </p>",
      "rawMarkdown": "This is amazing work, you covered a lot of techniques and algorithm settings, thanks for sharing.\n\nHow did you track progress, did you keep an eye on some kind of metric (clearly not on the public LB only:)? ",
      "replies": [
        {
          "id": 3438720,
          "postDate": "2026-04-09T16:11:20.333Z",
          "content": "<p>Nothing special, I was only checking TM-scores on validation sequences.</p>",
          "rawMarkdown": "Nothing special, I was only checking TM-scores on validation sequences.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3438706,
      "author_name": "Gabor Balazs",
      "author_url": "",
      "post_date": "2026-04-09T15:48:28.293000",
      "content": "<p>This is amazing work, you covered a lot of techniques and algorithm settings, thanks for sharing.</p>\n<p>How did you track progress, did you keep an eye on some kind of metric (clearly not on the public LB only:)? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3438720,
          "author_name": "Stefan Stefanov",
          "author_url": "",
          "post_date": "2026-04-09T16:11:20.333000",
          "content": "<p>Nothing special, I was only checking TM-scores on validation sequences.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3438585": "## Acknowledgements\n\nI would like to thank the competition hosts and Kaggle for the opportunity to work on such an impactful and interesting problem.  \nThank you @jaejohn and @theoviel for sharing strong solutions at the beginning of the competition.  \nCredit to Claude models for the code assistance.\n\n## Overview \n\nMy strategy for the competition was to build upon the provided strong solutions TBM, RNAPro and Protenix by incorporating into the modeling new challenges added in Part 2: proteins, DNA, ligands, multiple chains and long sequences.  \nAnother aspect I considered important was processing speed in order to be able to include diverse predictions in the submission within the runtime constraints.  \nThe solution combines two predictions from TBM, two predictions from Protenix, and one prediction from RNAPro for all targets. There is no additional logic for selecting specific predictions per RNA sequence.  \nRNAPro and Protenix inferences run in parallel on the two T4 GPUs. \n\n## Template-based modeling (TBM)\n\nFor template-based modeling, a LightGBM model predicting TM-score for a given query-template pair is introduced. From the competition train data, the 1000 most recent query RNA sequences and 200 (per query) template sequences with a preceding temporal cutoff are used to create a training dataset of ~198K query-template pairs. TM-score for each pair is computed and used as a label. The model features are:\n\n* RNA sequences similarity: `alignment_score`, `percent_identity`, and `len_diff_ratio`  \n* Text embeddings similarity: cosine similarity between query description and template description embeddings.  The embeddings are created with this model: https://huggingface.co/NeuML/pubmedbert-base-embeddings\n* Proteins similarity: alignment scores with a BLOSUM62 substitution matrix. For each query protein, the max alignment score with template proteins is taken. The min, mean, and max of the resulting vector are added as features. For processing speed considerations, only up to 6 query proteins are compared with up to 20 template proteins, and protein sequences are cropped to a max length of 768\\. These features are named `protein_similarity_matrix_(min|mean|max)` in the model.  \n* DNA similarity: `dna_similarity_matrix_(min|mean|max)` features calculated with the logic described above for proteins using alignment scores without a substitution matrix.  \n* Ligands similarity: `ligand_similarity_matrix_(min|mean|max)` features using Tanimoto similarities of Morgan fingerprints calculated with the logic described above for proteins.  \n* Composition count features: `num_query_(proteins|dna|ligands)` and `num_template_(proteins|dna|ligands)`  \n* Chains-related features: `(query|template)_num_all_chains`, `(query|template)_num_unique_chains`, and `chain_counts_match`\n\nThe top-2 templates according to the scores predicted by this model are used for the final submission.  \n\nThe script for running this step is `srna3d.scripts.create_tbm_submission.`  \nThe script for creating the training dataset is `srna3d.scripts.create_tmscore_dataset`. The dataset was created locally.  \nNotebook how to run it: https://www.kaggle.com/code/stefanstefanov/srna3d2-create-tmscore-dataset  \nThe training dataset: https://www.kaggle.com/datasets/stefanstefanov/srna3d2-tmscore-dataset  \nThe model was also trained locally. The script for training is `srna3d.scripts.train_tmscore_lightgbm_model`   \nTraining notebook: https://www.kaggle.com/code/stefanstefanov/srna3d2-train-tmscore-model  \n\n## RNAPro\n\nOne RNAPro model prediction is generated using the top-5 templates from the TBM step as input. To speed up inference, diffusion steps are decreased to 100, MSA depth is limited to 2048, and `fast_layernorm` and `triattention` triangle attention kernels are used, as inference is faster with them on longer sequences. As long sequences are split into chunks, the MSA and templates are also sliced to be consistent with the predicted chunk start and end index.\n\n## Protenix\n\nTwo predictions are generated with the latest Protenix model `protenix_base_20250630_v1.0.0`. In addition to the RNA sequence, proteins, DNA, and ligands are also added as input to the Protenix model. To fit in the T4 GPU memory and time constraints, proteins, DNA, and ligands are limited to a maximum of 6 per group. A `max_len_non_rna` parameter is added and set to `128`. This non-RNA sequence length budget is divided equally among protein and DNA sequences, and they are cut accordingly. The idea is to model the effect of short protein and DNA sequences on the RNA 3D structure, at least for targets with a small number of proteins or DNA. Similar to RNAPro, MSA depth is limited to 2048, and `fast_layernorm` and `triattention` triangle attention kernels are used.\n\n## Long Sequences\n\nFor the RNAPro and Protenix models, RNA with a sequence length above 448 is split in two steps:\n\n1) If it has multiple chains, it is first split into chains, and for each chain, an overlap of 96 from the next chain is added.  \n   For the last chain, the overlap is added circularly from the first chain with the idea to try to capture potential circularity in structure.  \n   For faster prediction, repeated chain pairs are skipped. For example, `9MME` which has U:8 chains with a total length of 4640 (8x580) is split and passed to the next step as one sequence of length 676 (580 from the first U + 96 from the second U chain).  \n   If the target has multiple chains, but its total sequence length is below 448, no such splitting is performed.\nThis splitting step is performed by the `srna3d.scripts.split_into_chains` script.\n\n2) Sequences from the previous step that are longer than 448 are split into chunks of length 448, having a 96-base overlap with the next chunk. This step is performed by the `srna3d.scripts.split_sequences_into_chunks` script.\n\nChunk predictions are first combined from chunks to chains using the `srna3d.scripts.combine_chunked_predictions` script. Afterward, they are combined from chains to full sequences with the `srna3d.scripts.combine_chain_predictions` script, applying Kabsch alignment on the overlapping regions. There is an option to randomly select and combine chain predictions from different samples, but it isn’t applied; for RNAPro, one sample is generated, and for Protenix, two samples are generated.\n\nThe solution includes various limits due to GPU memory and runtime constraints. It would be interesting to see what improvements this and other solutions could achieve with faster GPUs having more memory.\n\n## References\n\nRao, G. John, RNAPro inference with TBM notebook: https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\n\nViel, Theo, et al. RNAPro Inference notebook: https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\n\nLee, Youhan, et al. (2025). Template-based RNA structure prediction advanced through a blind code competition. *bioRxiv*. doi: 10.64898/2025.12.30.696949.   \nhttps://github.com/NVIDIA-Digital-Bio/RNAPro\n\nZhang, Yuxuan, et al. (2026). Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction. *bioRxiv*. doi: 10.64898/2026.02.05.703733.  \nhttps://github.com/bytedance/Protenix  \n",
    "3438706": "This is amazing work, you covered a lot of techniques and algorithm settings, thanks for sharing.\n\nHow did you track progress, did you keep an eye on some kind of metric (clearly not on the public LB only:)? "
  }
}