{
  "id": 687113,
  "title": "8th Place Solution - TBM and Protenix",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687113",
  "author_name": "Gabor Balazs",
  "post_date": "2026-04-02T08:42:38.909000",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This was my first bioinformatics competition, and I certainly learned a lot, thank you for organizing! And congratulations for the winners!</p>\n<hr>\n<p>These are the scores of my two selected notebooks. They had only slightly different Protenix computation schedule, and different seeds.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation</th>\n<th>Comp. Public LB</th>\n<th>Comp. Private LB</th>\n<th>Final Public LB</th>\n<th>Final Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>v1</td>\n<td>0.497</td>\n<td>0.46411</td>\n<td>0.57233</td>\n<td>0.46496</td>\n<td><strong>0.47641</strong></td>\n</tr>\n<tr>\n<td>v2</td>\n<td>0.493</td>\n<td>0.46853</td>\n<td>0.56621</td>\n<td>0.46828</td>\n<td><strong>0.47196</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Approach Summary</h2>\n<p>My approach is based on template based modeling (TBM) and <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>. I only used Kaggle resources for this competition (2xT4 GPUs), and I tried to optimize the validation + Public LB score (as my GPU quota allowed). The most important extra details compared to public notebooks are the following:</p>\n<ul>\n<li>TBM:<ul>\n<li>computing secondary structure information using ViennaRNA and incorporating it into sequence alignment;</li>\n<li>matching the chains independently within a template and solving the linear sum assignment problem (Hungarian algorithm).</li></ul></li>\n<li>Protenix:<ul>\n<li>for inference of short sequences (length &lt;= 768), splitting chain sequences in JSON, and also the MSA files;</li>\n<li>increasing the token limit from 512 to 768 using the dynamic chunk size feature of Protenix;</li>\n<li>parallelizing inference for 2xT4 GPUs, managing time budget, and choosing only the highest scoring candidates;</li>\n<li>for inference of long sequences (length &gt; 768), only predicting one of each different chains, and treat the others as missing.</li></ul></li>\n<li>Post-processing and candidate selection:<ul>\n<li>ignoring missing chain coordinates by filling them with downscaled coordinates of other chains (avoiding bad chain alignment penalty);</li>\n<li>averaging candidates for which (Kabsch aligned) average RMSD is within 2Å,</li>\n<li>replacing the two lowest scoring TBM candidates by Protenix ones if these TBM candidates do not score over a threshold.</li></ul></li>\n</ul>\n<hr>\n<h2>Template-Based Modeling (TBM)</h2>\n<p>I used <a href=\"https://github.com/jeffdaily/parasail\" target=\"_blank\">Parasail</a> to align two sequences, and <a href=\"https://github.com/ViennaRNA/ViennaRNA\" target=\"_blank\">ViennaRNA</a> to calculate the secondary structure information.</p>\n<h3>Secondary structure information (SSI)</h3>\n<p>For all train sequences, I precomputed the SSI in <a href=\"https://www.tbi.univie.ac.at/RNA/ViennaRNA/doc/html/io/rna_structures.html\" target=\"_blank\">dot-bracket format</a>. For multi-chain sequences, SSI was calculated for each chain independently. For the sequence alignment, I overlayed the SSI information to the sequences by introducing new characters beyond <code>\"AUCG\"</code> as follows:</p>\n<pre><code> 'A' + '.' -&gt; 'A'   ,   'A' + '(' -&gt; 'P'   ,   'A' + ')' -&gt; 'R'   ,\n 'U' + '.' -&gt; 'U'   ,   'U' + '(' -&gt; 'S'   ,   'U' + ')' -&gt; 'T'   ,\n 'C' + '.' -&gt; 'C'   ,   'C' + '(' -&gt; 'D'   ,   'C' + ')' -&gt; 'E'   ,\n 'G' + '.' -&gt; 'G'   ,   'G' + '(' -&gt; 'I'   ,   'G' + ')' -&gt; 'J'   .\n</code></pre>\n<p>For example (target_id: <code>9I9W</code>): <code>\"GGCACUGGAAGUGCGGCACUGGAAGUGC\" + \".(((((...))))).(((((...)))))\" -&gt; \"GIDPDSGGARJTJEGIDPDSGGARJTJE\"</code>.</p>\n<p>Then I replaced the 4x4 distance matrix by a 12x12 one for sequence alignment to score the SSI alignment simultaneously. The sequence aligner configuration included the scores for the sequence match and mismatch (<code>pmatch</code> and <code>pmismatch</code>), the ssi match and mismatch (<code>smatch</code> and <code>smismatch</code>), and the gap opening and extending penalties (<code>gap_open</code> and <code>gap_extend</code>). I used the following two configuration to enhance diversity and selected the five best (aggregated) candidates in a round robin way (where two of which could be replaced later by Protenix):</p>\n<pre><code>pmatch: 30,  pmismatch: -20,  smatch: 10,  smismatch: -10,  gap_open: 95,  gap_extend:  4,\npmatch: 30,  pmismatch: -20,  smatch:  1,  smismatch:  -1,  gap_open: 60,  gap_extend: 10.\n</code></pre>\n<p>Parasail uses integer scores, hence I used larger numbers compared to biopython. More importantly, Parasail expects capital letters in the alphabet by default, an issue which I was lucky to accidentally figure out from its documentation instead of almost discarding the whole idea. The <code>smismatch</code> penalty was doubled for aligning incorrect bracket SSIs (letters related to <code>'('</code> and <code>')'</code>, and vice versa).</p>\n<p>I picked candidates from two different configurations because sometimes using the SSI information could yield worse quality templates (e.g., <code>9G4R</code> and <code>9KGG</code>), and also the different gap structure seemed to provide extra diversity in the candidates. I only used global alignments (<code>nw_</code> prefixed Parasail functions). In order to sort templates, I recalculated the alignment scores ignoring tail gaps in each chains. Still, I could not find a reliable unified way to pick candidates from multiple configurations at once, hence I used round robin.</p>\n<p>I calculated SSI based on the sequence itself without using the MSA information. This is quite fast to compute with ViennaRNA. I also tried to use the consensus-based SSI using the MSA data, but it can take significant time to ViennaRNA to compute (hours for longer sequences). As all attempted runs to use this instead of the cheap no-MSA SSI resulted in worse score for me, I abandoned this idea. </p>\n<h3>Multi-chain alignment</h3>\n<p>Based on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/680551\" target=\"_blank\">this discussion</a>, I wanted a matching algorithm which does not rely on the somewhat arbitrary chain ordering of multi-chain RNA sequences. Therefore, I aligned all chains between the query and a template sequence (N x M alignments for an N-chain query and an M-chain template). The chain mapping was chosen to be the solution to the linear sum assignment problem (as implemented in <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.linear_sum_assignment.html\" target=\"_blank\">scipy</a>).</p>\n<p>I had maintained a big expectation for this feature, but to my surprise, it did not really move the score much. It was kept though, because I really did not like to depend on the provided chain ordering. Although it feels like a lot of computation, Parasail is very fast, and the whole TBM part of my notebook takes only a few minutes to finish overall.</p>\n<hr>\n<h2>Protenix</h2>\n<p>In order to reduce noise, I performed deterministic Protenix inference using <code>protenix.utils.seed.seed_everything(seed, deterministic=True)</code>. If I did not misunderstand the Protenix code, the <code>use_rna_msa</code> switch has no effect without the <code>use_msa</code> switch, which further needs <code>unpairedMsaPath</code> to be specified in the input JSON file. Eventually, I turned on all of these MSA switches along with the <code>use_template</code> switch as well.</p>\n<p>Protenix inference can produce poor quality results, so I tried to generate as many candidates as time allowed, sort them by the <code>ranking_score</code> of Protenix, and pick the 1-2 best as needed. The cheapest config in the <a href=\"docs/PTX_V1_Technical_Report_202602042356.pdf\" target=\"_blank\">Protenix paper</a> used 5 seeds and 5 diffusion samples resulting in 25 candidates, which was already way beyond the time budget here, at least I could not scale up Protenix that much.\nI used the GPU T4 x 2 accelerator on which I ran two independent Protenix instances in parallel with the same parameters, but with different seeds. At the end, I aggregated the candidates, sorted them, and used the best 1-2 of them to replace the worst 1-2 TBM candidates. Occasionally, I skipped Protenix for a sequence if at least five templates with sufficiently high scores were found during TBM.</p>\n<p>I fed sequences at once into Protenix up to length 768 using the dynamic chunk size setting (related to the attention / pairformer blocks), reducing the chunk size from the default 256 to 128 for sequences of length between 512 and 768, which allowed the inference to fit into the 16GB of VRAM.\nIn this case, I split the sequence and the corresponding MSA file per chain to separate <code>rnaSequence</code> records in the JSON file.</p>\n<p>For sequences longer than the 768 character limit, I predicted the chains independently. Chains being still longer than the 768 limit were split up with overlaps (of size 256), and their predictions were aligned using the Kabsch algorithm, similarly as has been shown by other publicly available notebooks.\nIn this case, I only predicted the different chains once, and if there were chain copies, I left them as missing.</p>\n<p>Instead of leaving the chain copies missing, I tried to copy over the predicted chain and align them using either the best templates from TBM, or by an extra inference run from Protenix which was predicting the sequence at once with all chains being truncated. Unfortunately, all of these attempts ended up with worse scores instead of just leaving these copies as missing. I believe, these methods did not capture the alignment correctly, and USalign was penalizing more such chain misalignments than the missing chains (which I hided \"inside\" another chains, see below more on this).</p>\n<p>I implemented a safeguard which exits the Protenix loop as the 8 hours run-time limit approaches. For sequences of size below 768, I allowed at most 5 inference runs with different seeds up to a 5 or 6 minutes limit per GPU in my v1 or v2 notebooks, respectively. For longer sequences, I only allowed a single run. In all cases, the first run generated as many diffusion samples as was requested based on the quality of the TBM scores, while on subsequent runs I only calculated a single diffusion sample. As I noticed, extra diffusion samples were almost as expensive to compute as completely new samples with a fresh seed.</p>\n<hr>\n<h2>Post-processing and candidate selection</h2>\n<p>The TBM candidate scores usually lay between 0 and 1, but they could sometimes exceed 1 (e.g., <code>9G4J</code>) due to the SSI scores (my normalization did not account for this).</p>\n<p>I implemented various constraint-based post-processing steps based on atom distances, but I did not observe any measurable improvement from using them.</p>\n<h3>Filling in missing coordinates</h3>\n<p>This was the only useful post-processing step I found, and therefore the only one I used. At the end, I applied a standard linear interpolation/extrapolation step to fill in coordinates. However, before doing so, I checked for completely missing chains, which I handled differently. If an identical non-missing chain was available (which was often the case), its coordinates could be simply copied over. I found that it worked slightly better to shrink the copied coordinates toward their average position, which is the approach I ultimately used. I believe it helps because the filled-in chain coordinates do not distort the global alignment in USalign. As a result, the alignment of the original non-missing chains improves, leading to a slightly better score.</p>\n<h3>Averaging candidates</h3>\n<p>To diversify the predictions, I aligned the candidates using the <a href=\"https://en.wikipedia.org/wiki/Kabsch_algorithm\" target=\"_blank\">Kabsch algorithm</a> and averaged those with an average <a href=\"https://en.wikipedia.org/wiki/Root_mean_square_deviation\" target=\"_blank\">RMSD</a> of no more than 2Å.</p>\n<h3>Candidate selection</h3>\n<p>I selected at most five candidates from TBM after averaging. I applied very light filtering (e.g., score &gt;= 0), which usually filled all five candiadate slots. I then replaced up to two of the lowest-scoring TBM candidates with Protenix candidates. The exact number (0-2) depended on the TBM scores (using two thresholds: 0.5 and 0.8) and the number of distinct Protenix candidates after averaging. The v1 notebook runs the entire Protenix inference loop again until it times out (limit: 8 hours - 10 minutes).</p>\n<hr>\n<h2>Acknowledgements</h2>\n<p>The most influential notebooks for me and what I learned from them:</p>\n<ul>\n<li>MMseqs+TBM+Parasail: <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates</a></li>\n<li>Protenix: <a href=\"https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution\" target=\"_blank\">https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution</a></li>\n<li>Evaluation: <a href=\"https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2</a></li>\n</ul>\n<p>I also thank ChatGPT for its endless patience of talking through many ideas and writing me some code.</p>\n<h2>What did not work</h2>\n<p>I implemented, but abandoned the following ideas:</p>\n<ul>\n<li>using MSA consensus-based ViennaRNA secondary structure (consistently worse);</li>\n<li>templates based on MMsegs search (no improvement);</li>\n<li>post-processing the coordinates based on distance constraints and trajectory smoothing (no improvement or worse);</li>\n<li>orienting independent chain predictions based on either templates or truncated multi-chain Protenix predicted candidates (consistently worse)</li>\n</ul>",
  "messages": [
    {
      "id": 3433871,
      "postDate": "2026-04-02T08:42:38.910Z",
      "content": "<p>This was my first bioinformatics competition, and I certainly learned a lot, thank you for organizing! And congratulations for the winners!</p>\n<hr>\n<p>These are the scores of my two selected notebooks. They had only slightly different Protenix computation schedule, and different seeds.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation</th>\n<th>Comp. Public LB</th>\n<th>Comp. Private LB</th>\n<th>Final Public LB</th>\n<th>Final Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>v1</td>\n<td>0.497</td>\n<td>0.46411</td>\n<td>0.57233</td>\n<td>0.46496</td>\n<td><strong>0.47641</strong></td>\n</tr>\n<tr>\n<td>v2</td>\n<td>0.493</td>\n<td>0.46853</td>\n<td>0.56621</td>\n<td>0.46828</td>\n<td><strong>0.47196</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Approach Summary</h2>\n<p>My approach is based on template based modeling (TBM) and <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a>. I only used Kaggle resources for this competition (2xT4 GPUs), and I tried to optimize the validation + Public LB score (as my GPU quota allowed). The most important extra details compared to public notebooks are the following:</p>\n<ul>\n<li>TBM:<ul>\n<li>computing secondary structure information using ViennaRNA and incorporating it into sequence alignment;</li>\n<li>matching the chains independently within a template and solving the linear sum assignment problem (Hungarian algorithm).</li></ul></li>\n<li>Protenix:<ul>\n<li>for inference of short sequences (length &lt;= 768), splitting chain sequences in JSON, and also the MSA files;</li>\n<li>increasing the token limit from 512 to 768 using the dynamic chunk size feature of Protenix;</li>\n<li>parallelizing inference for 2xT4 GPUs, managing time budget, and choosing only the highest scoring candidates;</li>\n<li>for inference of long sequences (length &gt; 768), only predicting one of each different chains, and treat the others as missing.</li></ul></li>\n<li>Post-processing and candidate selection:<ul>\n<li>ignoring missing chain coordinates by filling them with downscaled coordinates of other chains (avoiding bad chain alignment penalty);</li>\n<li>averaging candidates for which (Kabsch aligned) average RMSD is within 2Å,</li>\n<li>replacing the two lowest scoring TBM candidates by Protenix ones if these TBM candidates do not score over a threshold.</li></ul></li>\n</ul>\n<hr>\n<h2>Template-Based Modeling (TBM)</h2>\n<p>I used <a href=\"https://github.com/jeffdaily/parasail\" target=\"_blank\">Parasail</a> to align two sequences, and <a href=\"https://github.com/ViennaRNA/ViennaRNA\" target=\"_blank\">ViennaRNA</a> to calculate the secondary structure information.</p>\n<h3>Secondary structure information (SSI)</h3>\n<p>For all train sequences, I precomputed the SSI in <a href=\"https://www.tbi.univie.ac.at/RNA/ViennaRNA/doc/html/io/rna_structures.html\" target=\"_blank\">dot-bracket format</a>. For multi-chain sequences, SSI was calculated for each chain independently. For the sequence alignment, I overlayed the SSI information to the sequences by introducing new characters beyond <code>\"AUCG\"</code> as follows:</p>\n<pre><code> 'A' + '.' -&gt; 'A'   ,   'A' + '(' -&gt; 'P'   ,   'A' + ')' -&gt; 'R'   ,\n 'U' + '.' -&gt; 'U'   ,   'U' + '(' -&gt; 'S'   ,   'U' + ')' -&gt; 'T'   ,\n 'C' + '.' -&gt; 'C'   ,   'C' + '(' -&gt; 'D'   ,   'C' + ')' -&gt; 'E'   ,\n 'G' + '.' -&gt; 'G'   ,   'G' + '(' -&gt; 'I'   ,   'G' + ')' -&gt; 'J'   .\n</code></pre>\n<p>For example (target_id: <code>9I9W</code>): <code>\"GGCACUGGAAGUGCGGCACUGGAAGUGC\" + \".(((((...))))).(((((...)))))\" -&gt; \"GIDPDSGGARJTJEGIDPDSGGARJTJE\"</code>.</p>\n<p>Then I replaced the 4x4 distance matrix by a 12x12 one for sequence alignment to score the SSI alignment simultaneously. The sequence aligner configuration included the scores for the sequence match and mismatch (<code>pmatch</code> and <code>pmismatch</code>), the ssi match and mismatch (<code>smatch</code> and <code>smismatch</code>), and the gap opening and extending penalties (<code>gap_open</code> and <code>gap_extend</code>). I used the following two configuration to enhance diversity and selected the five best (aggregated) candidates in a round robin way (where two of which could be replaced later by Protenix):</p>\n<pre><code>pmatch: 30,  pmismatch: -20,  smatch: 10,  smismatch: -10,  gap_open: 95,  gap_extend:  4,\npmatch: 30,  pmismatch: -20,  smatch:  1,  smismatch:  -1,  gap_open: 60,  gap_extend: 10.\n</code></pre>\n<p>Parasail uses integer scores, hence I used larger numbers compared to biopython. More importantly, Parasail expects capital letters in the alphabet by default, an issue which I was lucky to accidentally figure out from its documentation instead of almost discarding the whole idea. The <code>smismatch</code> penalty was doubled for aligning incorrect bracket SSIs (letters related to <code>'('</code> and <code>')'</code>, and vice versa).</p>\n<p>I picked candidates from two different configurations because sometimes using the SSI information could yield worse quality templates (e.g., <code>9G4R</code> and <code>9KGG</code>), and also the different gap structure seemed to provide extra diversity in the candidates. I only used global alignments (<code>nw_</code> prefixed Parasail functions). In order to sort templates, I recalculated the alignment scores ignoring tail gaps in each chains. Still, I could not find a reliable unified way to pick candidates from multiple configurations at once, hence I used round robin.</p>\n<p>I calculated SSI based on the sequence itself without using the MSA information. This is quite fast to compute with ViennaRNA. I also tried to use the consensus-based SSI using the MSA data, but it can take significant time to ViennaRNA to compute (hours for longer sequences). As all attempted runs to use this instead of the cheap no-MSA SSI resulted in worse score for me, I abandoned this idea. </p>\n<h3>Multi-chain alignment</h3>\n<p>Based on <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/680551\" target=\"_blank\">this discussion</a>, I wanted a matching algorithm which does not rely on the somewhat arbitrary chain ordering of multi-chain RNA sequences. Therefore, I aligned all chains between the query and a template sequence (N x M alignments for an N-chain query and an M-chain template). The chain mapping was chosen to be the solution to the linear sum assignment problem (as implemented in <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.linear_sum_assignment.html\" target=\"_blank\">scipy</a>).</p>\n<p>I had maintained a big expectation for this feature, but to my surprise, it did not really move the score much. It was kept though, because I really did not like to depend on the provided chain ordering. Although it feels like a lot of computation, Parasail is very fast, and the whole TBM part of my notebook takes only a few minutes to finish overall.</p>\n<hr>\n<h2>Protenix</h2>\n<p>In order to reduce noise, I performed deterministic Protenix inference using <code>protenix.utils.seed.seed_everything(seed, deterministic=True)</code>. If I did not misunderstand the Protenix code, the <code>use_rna_msa</code> switch has no effect without the <code>use_msa</code> switch, which further needs <code>unpairedMsaPath</code> to be specified in the input JSON file. Eventually, I turned on all of these MSA switches along with the <code>use_template</code> switch as well.</p>\n<p>Protenix inference can produce poor quality results, so I tried to generate as many candidates as time allowed, sort them by the <code>ranking_score</code> of Protenix, and pick the 1-2 best as needed. The cheapest config in the <a href=\"docs/PTX_V1_Technical_Report_202602042356.pdf\" target=\"_blank\">Protenix paper</a> used 5 seeds and 5 diffusion samples resulting in 25 candidates, which was already way beyond the time budget here, at least I could not scale up Protenix that much.\nI used the GPU T4 x 2 accelerator on which I ran two independent Protenix instances in parallel with the same parameters, but with different seeds. At the end, I aggregated the candidates, sorted them, and used the best 1-2 of them to replace the worst 1-2 TBM candidates. Occasionally, I skipped Protenix for a sequence if at least five templates with sufficiently high scores were found during TBM.</p>\n<p>I fed sequences at once into Protenix up to length 768 using the dynamic chunk size setting (related to the attention / pairformer blocks), reducing the chunk size from the default 256 to 128 for sequences of length between 512 and 768, which allowed the inference to fit into the 16GB of VRAM.\nIn this case, I split the sequence and the corresponding MSA file per chain to separate <code>rnaSequence</code> records in the JSON file.</p>\n<p>For sequences longer than the 768 character limit, I predicted the chains independently. Chains being still longer than the 768 limit were split up with overlaps (of size 256), and their predictions were aligned using the Kabsch algorithm, similarly as has been shown by other publicly available notebooks.\nIn this case, I only predicted the different chains once, and if there were chain copies, I left them as missing.</p>\n<p>Instead of leaving the chain copies missing, I tried to copy over the predicted chain and align them using either the best templates from TBM, or by an extra inference run from Protenix which was predicting the sequence at once with all chains being truncated. Unfortunately, all of these attempts ended up with worse scores instead of just leaving these copies as missing. I believe, these methods did not capture the alignment correctly, and USalign was penalizing more such chain misalignments than the missing chains (which I hided \"inside\" another chains, see below more on this).</p>\n<p>I implemented a safeguard which exits the Protenix loop as the 8 hours run-time limit approaches. For sequences of size below 768, I allowed at most 5 inference runs with different seeds up to a 5 or 6 minutes limit per GPU in my v1 or v2 notebooks, respectively. For longer sequences, I only allowed a single run. In all cases, the first run generated as many diffusion samples as was requested based on the quality of the TBM scores, while on subsequent runs I only calculated a single diffusion sample. As I noticed, extra diffusion samples were almost as expensive to compute as completely new samples with a fresh seed.</p>\n<hr>\n<h2>Post-processing and candidate selection</h2>\n<p>The TBM candidate scores usually lay between 0 and 1, but they could sometimes exceed 1 (e.g., <code>9G4J</code>) due to the SSI scores (my normalization did not account for this).</p>\n<p>I implemented various constraint-based post-processing steps based on atom distances, but I did not observe any measurable improvement from using them.</p>\n<h3>Filling in missing coordinates</h3>\n<p>This was the only useful post-processing step I found, and therefore the only one I used. At the end, I applied a standard linear interpolation/extrapolation step to fill in coordinates. However, before doing so, I checked for completely missing chains, which I handled differently. If an identical non-missing chain was available (which was often the case), its coordinates could be simply copied over. I found that it worked slightly better to shrink the copied coordinates toward their average position, which is the approach I ultimately used. I believe it helps because the filled-in chain coordinates do not distort the global alignment in USalign. As a result, the alignment of the original non-missing chains improves, leading to a slightly better score.</p>\n<h3>Averaging candidates</h3>\n<p>To diversify the predictions, I aligned the candidates using the <a href=\"https://en.wikipedia.org/wiki/Kabsch_algorithm\" target=\"_blank\">Kabsch algorithm</a> and averaged those with an average <a href=\"https://en.wikipedia.org/wiki/Root_mean_square_deviation\" target=\"_blank\">RMSD</a> of no more than 2Å.</p>\n<h3>Candidate selection</h3>\n<p>I selected at most five candidates from TBM after averaging. I applied very light filtering (e.g., score &gt;= 0), which usually filled all five candiadate slots. I then replaced up to two of the lowest-scoring TBM candidates with Protenix candidates. The exact number (0-2) depended on the TBM scores (using two thresholds: 0.5 and 0.8) and the number of distinct Protenix candidates after averaging. The v1 notebook runs the entire Protenix inference loop again until it times out (limit: 8 hours - 10 minutes).</p>\n<hr>\n<h2>Acknowledgements</h2>\n<p>The most influential notebooks for me and what I learned from them:</p>\n<ul>\n<li>MMseqs+TBM+Parasail: <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates</a></li>\n<li>Protenix: <a href=\"https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution\" target=\"_blank\">https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution</a></li>\n<li>Evaluation: <a href=\"https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2</a></li>\n</ul>\n<p>I also thank ChatGPT for its endless patience of talking through many ideas and writing me some code.</p>\n<h2>What did not work</h2>\n<p>I implemented, but abandoned the following ideas:</p>\n<ul>\n<li>using MSA consensus-based ViennaRNA secondary structure (consistently worse);</li>\n<li>templates based on MMsegs search (no improvement);</li>\n<li>post-processing the coordinates based on distance constraints and trajectory smoothing (no improvement or worse);</li>\n<li>orienting independent chain predictions based on either templates or truncated multi-chain Protenix predicted candidates (consistently worse)</li>\n</ul>",
      "rawMarkdown": "This was my first bioinformatics competition, and I certainly learned a lot, thank you for organizing! And congratulations for the winners!\n\n---\n\nThese are the scores of my two selected notebooks. They had only slightly different Protenix computation schedule, and different seeds.\n| | Validation | Comp. Public LB | Comp. Private LB | Final Public LB | Final Private LB |\n| --- | --- |\n| v1 | 0.497 | 0.46411 | 0.57233 | 0.46496 | **0.47641** |\n| v2 | 0.493 | 0.46853 | 0.56621 | 0.46828 | **0.47196** |\n\n## Approach Summary\n\nMy approach is based on template based modeling (TBM) and [Protenix](https://github.com/bytedance/Protenix). I only used Kaggle resources for this competition (2xT4 GPUs), and I tried to optimize the validation + Public LB score (as my GPU quota allowed). The most important extra details compared to public notebooks are the following:\n- TBM:\n  - computing secondary structure information using ViennaRNA and incorporating it into sequence alignment;\n  - matching the chains independently within a template and solving the linear sum assignment problem (Hungarian algorithm).\n- Protenix:\n  - for inference of short sequences (length <= 768), splitting chain sequences in JSON, and also the MSA files;\n  - increasing the token limit from 512 to 768 using the dynamic chunk size feature of Protenix;\n  - parallelizing inference for 2xT4 GPUs, managing time budget, and choosing only the highest scoring candidates;\n  - for inference of long sequences (length > 768), only predicting one of each different chains, and treat the others as missing.\n- Post-processing and candidate selection:\n  - ignoring missing chain coordinates by filling them with downscaled coordinates of other chains (avoiding bad chain alignment penalty);\n  - averaging candidates for which (Kabsch aligned) average RMSD is within 2Å,\n  - replacing the two lowest scoring TBM candidates by Protenix ones if these TBM candidates do not score over a threshold.\n\n---\n\n## Template-Based Modeling (TBM)\n\nI used [Parasail](https://github.com/jeffdaily/parasail) to align two sequences, and [ViennaRNA](https://github.com/ViennaRNA/ViennaRNA) to calculate the secondary structure information.\n\n### Secondary structure information (SSI)\n\nFor all train sequences, I precomputed the SSI in [dot-bracket format](https://www.tbi.univie.ac.at/RNA/ViennaRNA/doc/html/io/rna_structures.html). For multi-chain sequences, SSI was calculated for each chain independently. For the sequence alignment, I overlayed the SSI information to the sequences by introducing new characters beyond `\"AUCG\"` as follows:\n\n     'A' + '.' -> 'A'   ,   'A' + '(' -> 'P'   ,   'A' + ')' -> 'R'   ,\n     'U' + '.' -> 'U'   ,   'U' + '(' -> 'S'   ,   'U' + ')' -> 'T'   ,\n     'C' + '.' -> 'C'   ,   'C' + '(' -> 'D'   ,   'C' + ')' -> 'E'   ,\n     'G' + '.' -> 'G'   ,   'G' + '(' -> 'I'   ,   'G' + ')' -> 'J'   .\n\nFor example (target_id: `9I9W`): `\"GGCACUGGAAGUGCGGCACUGGAAGUGC\" + \".(((((...))))).(((((...)))))\" -> \"GIDPDSGGARJTJEGIDPDSGGARJTJE\"`.\n\nThen I replaced the 4x4 distance matrix by a 12x12 one for sequence alignment to score the SSI alignment simultaneously. The sequence aligner configuration included the scores for the sequence match and mismatch (`pmatch` and `pmismatch`), the ssi match and mismatch (`smatch` and `smismatch`), and the gap opening and extending penalties (`gap_open` and `gap_extend`). I used the following two configuration to enhance diversity and selected the five best (aggregated) candidates in a round robin way (where two of which could be replaced later by Protenix):\n\n    pmatch: 30,  pmismatch: -20,  smatch: 10,  smismatch: -10,  gap_open: 95,  gap_extend:  4,\n    pmatch: 30,  pmismatch: -20,  smatch:  1,  smismatch:  -1,  gap_open: 60,  gap_extend: 10.\n\nParasail uses integer scores, hence I used larger numbers compared to biopython. More importantly, Parasail expects capital letters in the alphabet by default, an issue which I was lucky to accidentally figure out from its documentation instead of almost discarding the whole idea. The `smismatch` penalty was doubled for aligning incorrect bracket SSIs (letters related to `'('` and `')'`, and vice versa).\n\nI picked candidates from two different configurations because sometimes using the SSI information could yield worse quality templates (e.g., `9G4R` and `9KGG`), and also the different gap structure seemed to provide extra diversity in the candidates. I only used global alignments (`nw_` prefixed Parasail functions). In order to sort templates, I recalculated the alignment scores ignoring tail gaps in each chains. Still, I could not find a reliable unified way to pick candidates from multiple configurations at once, hence I used round robin.\n\nI calculated SSI based on the sequence itself without using the MSA information. This is quite fast to compute with ViennaRNA. I also tried to use the consensus-based SSI using the MSA data, but it can take significant time to ViennaRNA to compute (hours for longer sequences). As all attempted runs to use this instead of the cheap no-MSA SSI resulted in worse score for me, I abandoned this idea. \n\n### Multi-chain alignment\n\nBased on [this discussion](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/680551), I wanted a matching algorithm which does not rely on the somewhat arbitrary chain ordering of multi-chain RNA sequences. Therefore, I aligned all chains between the query and a template sequence (N x M alignments for an N-chain query and an M-chain template). The chain mapping was chosen to be the solution to the linear sum assignment problem (as implemented in [scipy](https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.linear_sum_assignment.html)).\n\nI had maintained a big expectation for this feature, but to my surprise, it did not really move the score much. It was kept though, because I really did not like to depend on the provided chain ordering. Although it feels like a lot of computation, Parasail is very fast, and the whole TBM part of my notebook takes only a few minutes to finish overall.\n\n---\n\n## Protenix\n\nIn order to reduce noise, I performed deterministic Protenix inference using `protenix.utils.seed.seed_everything(seed, deterministic=True)`. If I did not misunderstand the Protenix code, the `use_rna_msa` switch has no effect without the `use_msa` switch, which further needs `unpairedMsaPath` to be specified in the input JSON file. Eventually, I turned on all of these MSA switches along with the `use_template` switch as well.\n\nProtenix inference can produce poor quality results, so I tried to generate as many candidates as time allowed, sort them by the `ranking_score` of Protenix, and pick the 1-2 best as needed. The cheapest config in the [Protenix paper](docs/PTX_V1_Technical_Report_202602042356.pdf) used 5 seeds and 5 diffusion samples resulting in 25 candidates, which was already way beyond the time budget here, at least I could not scale up Protenix that much.\nI used the GPU T4 x 2 accelerator on which I ran two independent Protenix instances in parallel with the same parameters, but with different seeds. At the end, I aggregated the candidates, sorted them, and used the best 1-2 of them to replace the worst 1-2 TBM candidates. Occasionally, I skipped Protenix for a sequence if at least five templates with sufficiently high scores were found during TBM.\n\nI fed sequences at once into Protenix up to length 768 using the dynamic chunk size setting (related to the attention / pairformer blocks), reducing the chunk size from the default 256 to 128 for sequences of length between 512 and 768, which allowed the inference to fit into the 16GB of VRAM.\nIn this case, I split the sequence and the corresponding MSA file per chain to separate `rnaSequence` records in the JSON file.\n\nFor sequences longer than the 768 character limit, I predicted the chains independently. Chains being still longer than the 768 limit were split up with overlaps (of size 256), and their predictions were aligned using the Kabsch algorithm, similarly as has been shown by other publicly available notebooks.\nIn this case, I only predicted the different chains once, and if there were chain copies, I left them as missing.\n\nInstead of leaving the chain copies missing, I tried to copy over the predicted chain and align them using either the best templates from TBM, or by an extra inference run from Protenix which was predicting the sequence at once with all chains being truncated. Unfortunately, all of these attempts ended up with worse scores instead of just leaving these copies as missing. I believe, these methods did not capture the alignment correctly, and USalign was penalizing more such chain misalignments than the missing chains (which I hided \"inside\" another chains, see below more on this).\n\nI implemented a safeguard which exits the Protenix loop as the 8 hours run-time limit approaches. For sequences of size below 768, I allowed at most 5 inference runs with different seeds up to a 5 or 6 minutes limit per GPU in my v1 or v2 notebooks, respectively. For longer sequences, I only allowed a single run. In all cases, the first run generated as many diffusion samples as was requested based on the quality of the TBM scores, while on subsequent runs I only calculated a single diffusion sample. As I noticed, extra diffusion samples were almost as expensive to compute as completely new samples with a fresh seed.\n\n---\n\n## Post-processing and candidate selection\n\nThe TBM candidate scores usually lay between 0 and 1, but they could sometimes exceed 1 (e.g., `9G4J`) due to the SSI scores (my normalization did not account for this).\n\nI implemented various constraint-based post-processing steps based on atom distances, but I did not observe any measurable improvement from using them.\n\n### Filling in missing coordinates\n\nThis was the only useful post-processing step I found, and therefore the only one I used. At the end, I applied a standard linear interpolation/extrapolation step to fill in coordinates. However, before doing so, I checked for completely missing chains, which I handled differently. If an identical non-missing chain was available (which was often the case), its coordinates could be simply copied over. I found that it worked slightly better to shrink the copied coordinates toward their average position, which is the approach I ultimately used. I believe it helps because the filled-in chain coordinates do not distort the global alignment in USalign. As a result, the alignment of the original non-missing chains improves, leading to a slightly better score.\n\n### Averaging candidates\n\nTo diversify the predictions, I aligned the candidates using the [Kabsch algorithm](https://en.wikipedia.org/wiki/Kabsch_algorithm) and averaged those with an average [RMSD](https://en.wikipedia.org/wiki/Root_mean_square_deviation) of no more than 2Å.\n\n### Candidate selection\n\nI selected at most five candidates from TBM after averaging. I applied very light filtering (e.g., score >= 0), which usually filled all five candiadate slots. I then replaced up to two of the lowest-scoring TBM candidates with Protenix candidates. The exact number (0-2) depended on the TBM scores (using two thresholds: 0.5 and 0.8) and the number of distinct Protenix candidates after averaging. The v1 notebook runs the entire Protenix inference loop again until it times out (limit: 8 hours - 10 minutes).\n\n---\n\n## Acknowledgements\n\nThe most influential notebooks for me and what I learned from them:\n- MMseqs+TBM+Parasail: https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates\n- Protenix: https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution\n- Evaluation: https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2\n\nI also thank ChatGPT for its endless patience of talking through many ideas and writing me some code.\n\n## What did not work\n\nI implemented, but abandoned the following ideas:\n- using MSA consensus-based ViennaRNA secondary structure (consistently worse);\n- templates based on MMsegs search (no improvement);\n- post-processing the coordinates based on distance constraints and trajectory smoothing (no improvement or worse);\n- orienting independent chain predictions based on either templates or truncated multi-chain Protenix predicted candidates (consistently worse)\n\n",
      "votes": 16
    },
    {
      "id": 3433924,
      "postDate": "2026-04-02T10:19:27.537Z",
      "content": "<p>Congratulations! That's a fantastic approach.</p>\n<p>How did you come up with your original TBM idea? Also, what kind of score improvements did you see in your validation and publicly available data?</p>",
      "rawMarkdown": "Congratulations! That's a fantastic approach.\n\nHow did you come up with your original TBM idea? Also, what kind of score improvements did you see in your validation and publicly available data?",
      "votes": 1,
      "replies": [
        {
          "id": 3433954,
          "postDate": "2026-04-02T11:34:01.210Z",
          "content": "<p>The TBM idea emerged from a conversation I had with ChatGPT about ranking sequence alignments. It was bullish on the idea of using secondary structure information and I was pushing for simplicity, so we converged to this. I would be very surprised if the idea were truly novel, and it may not provide any statistically significant advantage after all. I think the sample size is too small to tell.</p>\n<p>Therefore, I did not systematically track the scores, I used them to make an accept/reject decision about new features and moved on. The TBM idea happened quite early in my journey, before I implemented my Protenix logic. At that stage, it improved my scores from Valid:0.397 and Public:0.335 to Valid:0.407 and Public:0.360. And my intuition also liked the idea, so I decided to keep it. :)</p>",
          "rawMarkdown": "The TBM idea emerged from a conversation I had with ChatGPT about ranking sequence alignments. It was bullish on the idea of using secondary structure information and I was pushing for simplicity, so we converged to this. I would be very surprised if the idea were truly novel, and it may not provide any statistically significant advantage after all. I think the sample size is too small to tell.\n\nTherefore, I did not systematically track the scores, I used them to make an accept/reject decision about new features and moved on. The TBM idea happened quite early in my journey, before I implemented my Protenix logic. At that stage, it improved my scores from Valid:0.397 and Public:0.335 to Valid:0.407 and Public:0.360. And my intuition also liked the idea, so I decided to keep it. :)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3433924,
      "author_name": "koooeo",
      "author_url": "",
      "post_date": "2026-04-02T10:19:27.537000",
      "content": "<p>Congratulations! That's a fantastic approach.</p>\n<p>How did you come up with your original TBM idea? Also, what kind of score improvements did you see in your validation and publicly available data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3433954,
          "author_name": "Gabor Balazs",
          "author_url": "",
          "post_date": "2026-04-02T11:34:01.210000",
          "content": "<p>The TBM idea emerged from a conversation I had with ChatGPT about ranking sequence alignments. It was bullish on the idea of using secondary structure information and I was pushing for simplicity, so we converged to this. I would be very surprised if the idea were truly novel, and it may not provide any statistically significant advantage after all. I think the sample size is too small to tell.</p>\n<p>Therefore, I did not systematically track the scores, I used them to make an accept/reject decision about new features and moved on. The TBM idea happened quite early in my journey, before I implemented my Protenix logic. At that stage, it improved my scores from Valid:0.397 and Public:0.335 to Valid:0.407 and Public:0.360. And my intuition also liked the idea, so I decided to keep it. :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3433871": "This was my first bioinformatics competition, and I certainly learned a lot, thank you for organizing! And congratulations for the winners!\n\n---\n\nThese are the scores of my two selected notebooks. They had only slightly different Protenix computation schedule, and different seeds.\n| | Validation | Comp. Public LB | Comp. Private LB | Final Public LB | Final Private LB |\n| --- | --- |\n| v1 | 0.497 | 0.46411 | 0.57233 | 0.46496 | **0.47641** |\n| v2 | 0.493 | 0.46853 | 0.56621 | 0.46828 | **0.47196** |\n\n## Approach Summary\n\nMy approach is based on template based modeling (TBM) and [Protenix](https://github.com/bytedance/Protenix). I only used Kaggle resources for this competition (2xT4 GPUs), and I tried to optimize the validation + Public LB score (as my GPU quota allowed). The most important extra details compared to public notebooks are the following:\n- TBM:\n  - computing secondary structure information using ViennaRNA and incorporating it into sequence alignment;\n  - matching the chains independently within a template and solving the linear sum assignment problem (Hungarian algorithm).\n- Protenix:\n  - for inference of short sequences (length <= 768), splitting chain sequences in JSON, and also the MSA files;\n  - increasing the token limit from 512 to 768 using the dynamic chunk size feature of Protenix;\n  - parallelizing inference for 2xT4 GPUs, managing time budget, and choosing only the highest scoring candidates;\n  - for inference of long sequences (length > 768), only predicting one of each different chains, and treat the others as missing.\n- Post-processing and candidate selection:\n  - ignoring missing chain coordinates by filling them with downscaled coordinates of other chains (avoiding bad chain alignment penalty);\n  - averaging candidates for which (Kabsch aligned) average RMSD is within 2Å,\n  - replacing the two lowest scoring TBM candidates by Protenix ones if these TBM candidates do not score over a threshold.\n\n---\n\n## Template-Based Modeling (TBM)\n\nI used [Parasail](https://github.com/jeffdaily/parasail) to align two sequences, and [ViennaRNA](https://github.com/ViennaRNA/ViennaRNA) to calculate the secondary structure information.\n\n### Secondary structure information (SSI)\n\nFor all train sequences, I precomputed the SSI in [dot-bracket format](https://www.tbi.univie.ac.at/RNA/ViennaRNA/doc/html/io/rna_structures.html). For multi-chain sequences, SSI was calculated for each chain independently. For the sequence alignment, I overlayed the SSI information to the sequences by introducing new characters beyond `\"AUCG\"` as follows:\n\n     'A' + '.' -> 'A'   ,   'A' + '(' -> 'P'   ,   'A' + ')' -> 'R'   ,\n     'U' + '.' -> 'U'   ,   'U' + '(' -> 'S'   ,   'U' + ')' -> 'T'   ,\n     'C' + '.' -> 'C'   ,   'C' + '(' -> 'D'   ,   'C' + ')' -> 'E'   ,\n     'G' + '.' -> 'G'   ,   'G' + '(' -> 'I'   ,   'G' + ')' -> 'J'   .\n\nFor example (target_id: `9I9W`): `\"GGCACUGGAAGUGCGGCACUGGAAGUGC\" + \".(((((...))))).(((((...)))))\" -> \"GIDPDSGGARJTJEGIDPDSGGARJTJE\"`.\n\nThen I replaced the 4x4 distance matrix by a 12x12 one for sequence alignment to score the SSI alignment simultaneously. The sequence aligner configuration included the scores for the sequence match and mismatch (`pmatch` and `pmismatch`), the ssi match and mismatch (`smatch` and `smismatch`), and the gap opening and extending penalties (`gap_open` and `gap_extend`). I used the following two configuration to enhance diversity and selected the five best (aggregated) candidates in a round robin way (where two of which could be replaced later by Protenix):\n\n    pmatch: 30,  pmismatch: -20,  smatch: 10,  smismatch: -10,  gap_open: 95,  gap_extend:  4,\n    pmatch: 30,  pmismatch: -20,  smatch:  1,  smismatch:  -1,  gap_open: 60,  gap_extend: 10.\n\nParasail uses integer scores, hence I used larger numbers compared to biopython. More importantly, Parasail expects capital letters in the alphabet by default, an issue which I was lucky to accidentally figure out from its documentation instead of almost discarding the whole idea. The `smismatch` penalty was doubled for aligning incorrect bracket SSIs (letters related to `'('` and `')'`, and vice versa).\n\nI picked candidates from two different configurations because sometimes using the SSI information could yield worse quality templates (e.g., `9G4R` and `9KGG`), and also the different gap structure seemed to provide extra diversity in the candidates. I only used global alignments (`nw_` prefixed Parasail functions). In order to sort templates, I recalculated the alignment scores ignoring tail gaps in each chains. Still, I could not find a reliable unified way to pick candidates from multiple configurations at once, hence I used round robin.\n\nI calculated SSI based on the sequence itself without using the MSA information. This is quite fast to compute with ViennaRNA. I also tried to use the consensus-based SSI using the MSA data, but it can take significant time to ViennaRNA to compute (hours for longer sequences). As all attempted runs to use this instead of the cheap no-MSA SSI resulted in worse score for me, I abandoned this idea. \n\n### Multi-chain alignment\n\nBased on [this discussion](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/680551), I wanted a matching algorithm which does not rely on the somewhat arbitrary chain ordering of multi-chain RNA sequences. Therefore, I aligned all chains between the query and a template sequence (N x M alignments for an N-chain query and an M-chain template). The chain mapping was chosen to be the solution to the linear sum assignment problem (as implemented in [scipy](https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.linear_sum_assignment.html)).\n\nI had maintained a big expectation for this feature, but to my surprise, it did not really move the score much. It was kept though, because I really did not like to depend on the provided chain ordering. Although it feels like a lot of computation, Parasail is very fast, and the whole TBM part of my notebook takes only a few minutes to finish overall.\n\n---\n\n## Protenix\n\nIn order to reduce noise, I performed deterministic Protenix inference using `protenix.utils.seed.seed_everything(seed, deterministic=True)`. If I did not misunderstand the Protenix code, the `use_rna_msa` switch has no effect without the `use_msa` switch, which further needs `unpairedMsaPath` to be specified in the input JSON file. Eventually, I turned on all of these MSA switches along with the `use_template` switch as well.\n\nProtenix inference can produce poor quality results, so I tried to generate as many candidates as time allowed, sort them by the `ranking_score` of Protenix, and pick the 1-2 best as needed. The cheapest config in the [Protenix paper](docs/PTX_V1_Technical_Report_202602042356.pdf) used 5 seeds and 5 diffusion samples resulting in 25 candidates, which was already way beyond the time budget here, at least I could not scale up Protenix that much.\nI used the GPU T4 x 2 accelerator on which I ran two independent Protenix instances in parallel with the same parameters, but with different seeds. At the end, I aggregated the candidates, sorted them, and used the best 1-2 of them to replace the worst 1-2 TBM candidates. Occasionally, I skipped Protenix for a sequence if at least five templates with sufficiently high scores were found during TBM.\n\nI fed sequences at once into Protenix up to length 768 using the dynamic chunk size setting (related to the attention / pairformer blocks), reducing the chunk size from the default 256 to 128 for sequences of length between 512 and 768, which allowed the inference to fit into the 16GB of VRAM.\nIn this case, I split the sequence and the corresponding MSA file per chain to separate `rnaSequence` records in the JSON file.\n\nFor sequences longer than the 768 character limit, I predicted the chains independently. Chains being still longer than the 768 limit were split up with overlaps (of size 256), and their predictions were aligned using the Kabsch algorithm, similarly as has been shown by other publicly available notebooks.\nIn this case, I only predicted the different chains once, and if there were chain copies, I left them as missing.\n\nInstead of leaving the chain copies missing, I tried to copy over the predicted chain and align them using either the best templates from TBM, or by an extra inference run from Protenix which was predicting the sequence at once with all chains being truncated. Unfortunately, all of these attempts ended up with worse scores instead of just leaving these copies as missing. I believe, these methods did not capture the alignment correctly, and USalign was penalizing more such chain misalignments than the missing chains (which I hided \"inside\" another chains, see below more on this).\n\nI implemented a safeguard which exits the Protenix loop as the 8 hours run-time limit approaches. For sequences of size below 768, I allowed at most 5 inference runs with different seeds up to a 5 or 6 minutes limit per GPU in my v1 or v2 notebooks, respectively. For longer sequences, I only allowed a single run. In all cases, the first run generated as many diffusion samples as was requested based on the quality of the TBM scores, while on subsequent runs I only calculated a single diffusion sample. As I noticed, extra diffusion samples were almost as expensive to compute as completely new samples with a fresh seed.\n\n---\n\n## Post-processing and candidate selection\n\nThe TBM candidate scores usually lay between 0 and 1, but they could sometimes exceed 1 (e.g., `9G4J`) due to the SSI scores (my normalization did not account for this).\n\nI implemented various constraint-based post-processing steps based on atom distances, but I did not observe any measurable improvement from using them.\n\n### Filling in missing coordinates\n\nThis was the only useful post-processing step I found, and therefore the only one I used. At the end, I applied a standard linear interpolation/extrapolation step to fill in coordinates. However, before doing so, I checked for completely missing chains, which I handled differently. If an identical non-missing chain was available (which was often the case), its coordinates could be simply copied over. I found that it worked slightly better to shrink the copied coordinates toward their average position, which is the approach I ultimately used. I believe it helps because the filled-in chain coordinates do not distort the global alignment in USalign. As a result, the alignment of the original non-missing chains improves, leading to a slightly better score.\n\n### Averaging candidates\n\nTo diversify the predictions, I aligned the candidates using the [Kabsch algorithm](https://en.wikipedia.org/wiki/Kabsch_algorithm) and averaged those with an average [RMSD](https://en.wikipedia.org/wiki/Root_mean_square_deviation) of no more than 2Å.\n\n### Candidate selection\n\nI selected at most five candidates from TBM after averaging. I applied very light filtering (e.g., score >= 0), which usually filled all five candiadate slots. I then replaced up to two of the lowest-scoring TBM candidates with Protenix candidates. The exact number (0-2) depended on the TBM scores (using two thresholds: 0.5 and 0.8) and the number of distinct Protenix candidates after averaging. The v1 notebook runs the entire Protenix inference loop again until it times out (limit: 8 hours - 10 minutes).\n\n---\n\n## Acknowledgements\n\nThe most influential notebooks for me and what I learned from them:\n- MMseqs+TBM+Parasail: https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-templates\n- Protenix: https://www.kaggle.com/code/sigmaborov/stanford-rna-3d-folding-top-1-solution\n- Evaluation: https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2\n\nI also thank ChatGPT for its endless patience of talking through many ideas and writing me some code.\n\n## What did not work\n\nI implemented, but abandoned the following ideas:\n- using MSA consensus-based ViennaRNA secondary structure (consistently worse);\n- templates based on MMsegs search (no improvement);\n- post-processing the coordinates based on distance constraints and trajectory smoothing (no improvement or worse);\n- orienting independent chain predictions based on either templates or truncated multi-chain Protenix predicted candidates (consistently worse)\n\n",
    "3433924": "Congratulations! That's a fantastic approach.\n\nHow did you come up with your original TBM idea? Also, what kind of score improvements did you see in your validation and publicly available data?"
  }
}