{
  "id": 687506,
  "title": "69th Place Solution - TBM + Protenix",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687506",
  "author_name": "Adi_3375",
  "post_date": "2026-04-03T04:55:41.278000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Silver Medal Solution · 68th / 1867 Teams · Stanford RNA 3D Folding Part 2</h2>\n<blockquote>\n  <p><em>How we went from 0.184 to 0.424 — fixing randomness, diagnosing templates, and redesigning the prediction pipeline</em></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Final LB Score: 0.424 · Deterministic · TBM + multi-seed Protenix · Generate → Rank → Select architecture</strong></p>\n</blockquote>\n<hr>\n<h2>1. Problem &amp; Approach</h2>\n<p>The competition asks for 5 predicted 3D structures per RNA sequence. The metric takes the best TM-score among your 5 predictions — so the goal is not one perfect prediction, but five diverse, high-quality candidates.</p>\n<p>We combined two complementary methods:</p>\n<ul>\n<li><strong>Template-Based Modeling (TBM)</strong> — find structurally similar sequences in a database of ~21,760 known solved structures, transfer their coordinates, apply diversity transforms, and score the results</li>\n<li><strong>Protenix</strong> — a diffusion-based neural network (similar to AlphaFold3) that folds RNA from sequence alone, run with multiple seeds for structural diversity</li>\n</ul>\n<p>The 28 test sequences fall into three natural groups that required different handling:</p>\n<ul>\n<li><strong>9 strong sequences (OUR_SEQUENCES)</strong> — excellent TBM templates exist, TM-scores 0.52–0.80</li>\n<li><strong>11 weak sequences</strong> — no good templates, Protenix is the only option</li>\n<li><strong>2 ultra-long sequences</strong> (9MME=4640nt, 9ZCC=1460nt) — too long for single Protenix pass, require chunked inference</li>\n</ul>\n<hr>\n<h2>2. How TBM Became Strong — The Evolution</h2>\n<p>TBM started at 0.184 and grew through several key iterations before becoming the backbone of the final solution.</p>\n<h3>2.1 From Noise to Strategic Diversity (0.184 → 0.233)</h3>\n<p>The original baseline generated 5 predictions by adding Gaussian noise to a single morphed template. The problem: all 5 predictions were nearly identical, wasting the best-of-5 metric.</p>\n<p>The fix was to use genuinely different template subsets for each prediction slot — top-4, templates 3–6, templates 5–8, even-indexed, odd-indexed. Each slot blended its assigned templates by coordinate averaging. This created real structural diversity and jumped the score to <strong>0.233</strong>.</p>\n<h3>2.2 Energy-Based Reranking (0.233 → 0.237)</h3>\n<p>After generating 5 predictions, we sorted them by a physics-based pseudo-energy function so the most plausible structure became prediction 1. The energy combined bond length deviation from the ideal 5.9Å C1'–C1' distance and a steric clash penalty for atoms closer than 3.8Å.</p>\n<h3>2.3 Per-Sequence Routing (0.237 → 0.242)</h3>\n<p>Rather than applying one strategy uniformly, we diagnosed each sequence individually and routed it to the strategy where it performed best. The 7 sequences with no templates at all worked best with a diverse geometry fallback (helix, extended strand, compact globule, sequence-guided, reverse helix).</p>\n<h3>2.4 Composite Template Scoring</h3>\n<p>We replaced the raw alignment score with a composite that factors in length similarity, exact match ratio, and GC content similarity:</p>\n<pre><code>score = (align²) × len_sim × (1 + 0.3×exact) × (0.5 + 0.5×gc_sim)\n</code></pre>\n<p>Squaring the alignment score amplifies strong matches and penalises weak ones.</p>\n<h3>2.5 More Templates and Better Diversity (TOPK 8 → 15)</h3>\n<p>Increasing TOPK from 8 to 15 and generating up to 12 candidates per sequence gave the scoring function more material to work with. We also replaced Gaussian noise with structured diversity transforms applied in rotation:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Transform</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>Clean</td>\n<td>Best template, no noise</td>\n</tr>\n<tr>\n<td>1</td>\n<td>Tiny noise</td>\n<td>Proportional to alignment uncertainty</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Hinge rotation</td>\n<td>Rotates one segment around a pivot</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Jitter chains</td>\n<td>Independent rigid rotation + translation per chain</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Smooth wiggle</td>\n<td>Smooth interpolated displacement field</td>\n</tr>\n<tr>\n<td>5+</td>\n<td>Gaussian noise</td>\n<td>Fallback</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>3. The Randomness Problem</h2>\n<p>Before any architectural improvements were possible, we had to solve a more fundamental problem: <strong>the same notebook was producing scores of 0.392, 0.404, 0.409, and 0.426 on different runs</strong>. Without reproducibility, every experiment was noise.</p>\n<h3>3.1 Finding All 12 Sources of Non-Determinism</h3>\n<p>A systematic audit found 12 distinct sources of randomness. The most critical:</p>\n<ul>\n<li><strong><code>CUBLAS_WORKSPACE_CONFIG</code> set after PyTorch import</strong> — this environment variable must be set before <code>torch</code> is imported, or it has no effect on GPU determinism</li>\n<li><strong><code>np.random.seed()</code> inside prediction functions</strong> with <code>seed=len(predictions)*7+len(seq)</code> — if a morph failed upstream and <code>len(predictions)</code> differed from expected, the seed changed and all downstream noise changed</li>\n<li><strong><code>shortlist_templates</code> sort with no tiebreaker</strong> — equal alignment scores had non-deterministic ordering depending on dataframe memory layout</li>\n<li><strong><code>lru_cache(maxsize=200,000)</code></strong> — when the cache filled, eviction order varied between runs</li>\n<li><strong>Cross-product threshold at 1e-6</strong> — GPU vs CPU floating point differences could flip this condition</li>\n</ul>\n<h3>3.2 The Fix</h3>\n<p>Seven targeted changes restored full determinism:</p>\n<pre><code># Cell 0 — before ANY import\nimport os, random\nos.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'\nos.environ['PYTHONHASHSEED'] = '42'\nrandom.seed(42)\nimport numpy as np; np.random.seed(42)\nimport torch; torch.manual_seed(42)\ntorch.use_deterministic_algorithms(True, warn_only=True)\n</code></pre>\n<ul>\n<li>All <code>np.random.seed()</code> calls replaced with isolated <code>np.random.default_rng(seed=(ridx × 10B + slot × 10007) % 2³²)</code> — seed is globally unique per sequence and slot</li>\n<li>Added <code>tid</code> as secondary sort key in all template ranking functions</li>\n<li>Raised <code>lru_cache</code> to <code>maxsize=None</code></li>\n<li>Raised cross-product threshold to <code>0.1</code></li>\n</ul>\n<blockquote>\n  <p><strong>Result: Three consecutive submissions all scored exactly 0.392 — perfect determinism confirmed.</strong></p>\n</blockquote>\n<hr>\n<h2>4. The Template Diagnosis — Understanding the Ceiling</h2>\n<p>After fixing determinism, every experiment returned 0.392. To understand why, we built a deep diagnostic that computed the TM-score of raw template coordinates <strong>before any pipeline processing</strong>.</p>\n<table>\n<thead>\n<tr>\n<th>Finding</th>\n<th>Sequences Affected</th>\n<th>Conclusion</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>High alignment (&gt;0.4) but template TM &lt; 0.2</td>\n<td>10 sequences</td>\n<td>Same sequence, different 3D fold — conformational isomers</td>\n</tr>\n<tr>\n<td>Alignment &lt; 0.3 — no template exists</td>\n<td>8 sequences</td>\n<td>Genuinely novel folds — no known analog in any database</td>\n</tr>\n<tr>\n<td>morph_delta ≈ 0 for all sequences</td>\n<td>All 28</td>\n<td>Pipeline is NOT destroying good templates — ceiling is the data</td>\n</tr>\n<tr>\n<td>9EBP — only sequence with tmpl_TM &gt; 0.35</td>\n<td>1 sequence</td>\n<td>Good template exists, already handled well by pipeline</td>\n</tr>\n</tbody>\n</table>\n<p>The most striking finding was <strong>9LEL</strong> — alignment score 1.280 (near-identical sequence) but template TM = 0.020 (completely wrong structure). The template in the database folds the same sequence into a different 3D conformation. This is conformational isomerism — common in RNA, and impossible to fix by improving the morphing pipeline.</p>\n<blockquote>\n  <p><strong>Key insight: The pipeline was not the problem. The training database simply does not contain structurally correct templates for 18 of 28 test sequences. These were specifically chosen for the competition because they represent novel folds. No amount of pipeline tuning overcomes this — only a better generative model can.</strong></p>\n</blockquote>\n<hr>\n<h2>5. The Breakthrough — Generate → Rank → Select</h2>\n<h3>5.1 The Old Architecture Problem</h3>\n<p>The original pipeline pre-assigned slots before generating any predictions. It counted good TBM candidates (alignment ≥ 0.4), allocated N TBM slots, and sent the remaining 5-N slots to Protenix. The problem: if TBM slot 3 produced a better structure than Protenix slot 1, TBM slot 3 was discarded anyway. <strong>The pipeline had no quality gate.</strong></p>\n<h3>5.2 The New Architecture</h3>\n<p>We replaced slot allocation with a generate-then-rank approach:</p>\n<ul>\n<li>Generate up to 12 TBM candidates with all 5 diversity transforms</li>\n<li>Generate Protenix candidates with multiple seeds (N_sample=1 per seed)</li>\n<li>Pool all candidates together — source doesn't matter</li>\n<li>Score every candidate with <code>score_structure()</code></li>\n<li>Select top 5 by score, with a diversity check to prevent near-duplicate selections</li>\n</ul>\n<pre><code>def score_structure(coords):\n    diffs     = np.linalg.norm(np.diff(coords, axis=0), axis=1)\n    bond_pen  = np.sum((diffs - 5.9) ** 2)          # ideal C1'-C1' = 5.9Å\n    dm        = distance_matrix(coords, coords)\n    clash_pen = np.sum(dm &lt; 3.8)                    # atomic clash threshold\n    return -(clash_pen * 6.0 + bond_pen * 0.3)      # higher = better\n</code></pre>\n<p>TBM candidates also receive an alignment bonus <code>(align² × 8.0)</code> and a length-dependent bonus, ensuring that high-confidence template morphs rank above uncertain ones.</p>\n<h3>5.3 Dynamic Protenix Seed Schedule</h3>\n<p>Instead of one fixed seed with N_sample=5, Protenix runs N_sample=1 per seed with a sequence-length-dependent number of seeds:</p>\n<table>\n<thead>\n<tr>\n<th>Sequence Type</th>\n<th>Seed Count</th>\n<th>Rationale</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Short (≤400 nt)</td>\n<td>7 seeds</td>\n<td>More exploration — each run is fast</td>\n</tr>\n<tr>\n<td>Medium (400–1000 nt)</td>\n<td>3 seeds</td>\n<td>Balanced exploration vs compute</td>\n</tr>\n<tr>\n<td>9MME, 9ZCC</td>\n<td>2 seeds</td>\n<td>Chunked inference (512nt chunks, 128nt overlap, Kabsch-aligned stitching)</td>\n</tr>\n<tr>\n<td>OUR_SEQUENCES</td>\n<td>0 Protenix</td>\n<td>TBM is already near-optimal</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>6. Score Progression — Full Timeline</h2>\n<table>\n<thead>\n<tr>\n<th>Phase</th>\n<th>Key Change</th>\n<th>LB Score</th>\n<th>Delta</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline TBM</td>\n<td>Template morphing + Gaussian noise</td>\n<td>0.184</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Experiment B</td>\n<td>Strategic template diversity (different subsets per slot)</td>\n<td>0.233</td>\n<td>+0.049</td>\n</tr>\n<tr>\n<td>Experiment 1</td>\n<td>Energy-based reranking of 5 predictions</td>\n<td>0.237</td>\n<td>+0.004</td>\n</tr>\n<tr>\n<td>FINAL routing</td>\n<td>Per-sequence strategy routing</td>\n<td>0.242</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Seed fix</td>\n<td>All randomness sources fixed — deterministic baseline</td>\n<td>0.392</td>\n<td>—</td>\n</tr>\n<tr>\n<td><strong>Generate→Rank→Select</strong></td>\n<td><strong>TOPK=15, all diversity transforms, score_structure()</strong></td>\n<td><strong>0.421</strong></td>\n<td><strong>+0.029</strong></td>\n</tr>\n<tr>\n<td>Protenix tuning</td>\n<td>Smoothness penalty, small-sequence boost, slot control</td>\n<td>0.422</td>\n<td>+0.001</td>\n</tr>\n<tr>\n<td><strong>Seed optimization</strong></td>\n<td><strong>Dual seed for long + 7 seeds for short sequences</strong></td>\n<td><strong>0.424</strong></td>\n<td><strong>+0.002</strong></td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>Note: The jump from 0.242 to 0.392 reflects the addition of Protenix to the pipeline. The apparent gap is because the deterministic Protenix baseline (0.392) replaced earlier non-deterministic scores which included lucky runs up to 0.426.</p>\n</blockquote>\n<hr>\n<h2>7. Seed Strategy Optimization</h2>\n<p>The interaction between seed count and ranking strength was non-obvious. More seeds alone did not help — the ranking function had to be strong enough to select the best from a larger pool.</p>\n<table>\n<thead>\n<tr>\n<th>Setup</th>\n<th>LB Score</th>\n<th>Verdict</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline (5 samples, 1 seed)</td>\n<td>0.422</td>\n<td>—</td>\n</tr>\n<tr>\n<td>+ Dual seed for long sequences only</td>\n<td>0.416</td>\n<td>❌ Worse — expensive, low benefit alone</td>\n</tr>\n<tr>\n<td>+ 7 seeds for short sequences only</td>\n<td>0.420</td>\n<td>⚠️ Slight drop — ranking not strong enough</td>\n</tr>\n<tr>\n<td>Dual seed (long) + 7 seeds (short)</td>\n<td><strong>0.424</strong></td>\n<td>✅ Best — diversity and ranking work together</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p><strong>Key insight: More candidates ≠ better predictions. Better candidates + proper ranking = improvement. The dual-seed and 7-seed strategies only worked together because <code>score_structure()</code> was strong enough to identify the better structures from the larger pool.</strong></p>\n</blockquote>\n<hr>\n<h2>8. What Did Not Help</h2>\n<h3>Validation sequences in template pool</h3>\n<p>Adding validation labels as templates would have given better structural references. However, the alignment search over ~28,000 sequences (train + val) caused a Kaggle timeout every time. The pipeline never finished scoring.</p>\n<h3>Structure blending</h3>\n<p>Blending the top-2 candidates by coordinate averaging: <code>(c1 + c2) / 2</code>. The hypothesis was that averaging would smooth out errors. In practice, the top candidates were already well-optimised and the average degraded topology rather than improving it.</p>\n<h3>Hardcoded diversity per sequence</h3>\n<p>We diagnosed which diversity transform helped each sequence on the validation set and hardcoded the best transform per sequence. The +0.011 improvement on validation was too small to measure on the leaderboard.</p>\n<h3>Chunked Protenix for 9MME / 9ZCC in isolation</h3>\n<p>Splitting ultra-long sequences into 512nt chunks, running Protenix, and stitching with Kabsch alignment. The stitched output was no better than TBM for these sequences in isolation. Only when combined with dual-seed and the overall generate/rank/select architecture did it contribute.</p>\n<hr>\n<h2>9. What Could Have Been Better</h2>\n<ul>\n<li><strong>Fine-tuning Protenix on RNA-specific data</strong> — the 18 weak sequences represent novel folds that a better-trained model would handle. The base Protenix model was not optimised for these specific RNA topologies.</li>\n<li><strong>Confidence-based selection</strong> — Protenix outputs a pLDDT confidence score per residue. We experimented with using this to select the best samples across seeds, but the scores were often zero for long-sequence chunks due to extraction issues.</li>\n<li><strong>MSA depth for weak sequences</strong> — some weak sequences may have had sparse or missing MSA files, meaning Protenix ran without evolutionary information. Checking and augmenting MSA coverage could have improved Protenix quality on these sequences.</li>\n<li><strong>A fundamentally different model for the 18 hard sequences</strong> — AlphaFold3, RoseTTAFold2NA, or a fine-tuned Protenix. The generate/rank/select architecture is sound; the ceiling is the quality of the base models.</li>\n</ul>\n<hr>\n<h2>10. Key Lessons</h2>\n<h3>1. Fix randomness before anything else</h3>\n<p>Until seeds are deterministic, you cannot measure whether a change actually helps. Every experiment before the seed fix was polluted by Protenix variance. The seed fix was not a performance improvement — it was a prerequisite for all further meaningful work.</p>\n<h3>2. Diagnose per-sequence, not overall</h3>\n<p>Average metrics hide what's actually failing. The template TM-score diagnostic — computing TM of raw template coordinates before any pipeline processing — revealed in one experiment that the pipeline was not the bottleneck. Without this, we could have spent days optimising morphing and refinement code that would never move the leaderboard.</p>\n<h3>3. Understand your metric</h3>\n<p>Best-of-5 scoring means diversity across predictions is strictly more valuable than precision of one. The entire generate/rank/select architecture exists to exploit this. A pipeline that generates 12 diverse candidates and selects the best 5 beats a pipeline that generates 5 careful predictions.</p>\n<h3>4. TBM is signal; Protenix is exploration</h3>\n<p>When a good template exists (alignment &gt; 0.5), TBM is near-perfect and Protenix cannot improve on it. When no good template exists, TBM is useless and Protenix is the only option. Understanding this separation and routing accordingly is what kept the strong sequences strong while giving the weak sequences their best chance.</p>\n<h3>5. The database is the real ceiling</h3>\n<p>18 of 28 test sequences are genuinely novel RNA folds with no structurally similar solved structure in any public database. This is not a pipeline problem. These sequences were chosen for the competition precisely because they are hard. The only path to improving predictions on them is a better generative model.</p>\n<h3>6. Pipeline thinking beats model thinking</h3>\n<p>We did not change the underlying models. We changed how they are used — when to trust TBM, when to use Protenix, how many candidates to generate, and how to select among them. The 0.029 improvement from 0.392 to 0.421 came entirely from architectural decisions, not from improving the models themselves.</p>\n<hr>",
  "messages": [
    {
      "id": 3434651,
      "postDate": "2026-04-03T04:55:41.280Z",
      "content": "<h2>Silver Medal Solution · 68th / 1867 Teams · Stanford RNA 3D Folding Part 2</h2>\n<blockquote>\n  <p><em>How we went from 0.184 to 0.424 — fixing randomness, diagnosing templates, and redesigning the prediction pipeline</em></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Final LB Score: 0.424 · Deterministic · TBM + multi-seed Protenix · Generate → Rank → Select architecture</strong></p>\n</blockquote>\n<hr>\n<h2>1. Problem &amp; Approach</h2>\n<p>The competition asks for 5 predicted 3D structures per RNA sequence. The metric takes the best TM-score among your 5 predictions — so the goal is not one perfect prediction, but five diverse, high-quality candidates.</p>\n<p>We combined two complementary methods:</p>\n<ul>\n<li><strong>Template-Based Modeling (TBM)</strong> — find structurally similar sequences in a database of ~21,760 known solved structures, transfer their coordinates, apply diversity transforms, and score the results</li>\n<li><strong>Protenix</strong> — a diffusion-based neural network (similar to AlphaFold3) that folds RNA from sequence alone, run with multiple seeds for structural diversity</li>\n</ul>\n<p>The 28 test sequences fall into three natural groups that required different handling:</p>\n<ul>\n<li><strong>9 strong sequences (OUR_SEQUENCES)</strong> — excellent TBM templates exist, TM-scores 0.52–0.80</li>\n<li><strong>11 weak sequences</strong> — no good templates, Protenix is the only option</li>\n<li><strong>2 ultra-long sequences</strong> (9MME=4640nt, 9ZCC=1460nt) — too long for single Protenix pass, require chunked inference</li>\n</ul>\n<hr>\n<h2>2. How TBM Became Strong — The Evolution</h2>\n<p>TBM started at 0.184 and grew through several key iterations before becoming the backbone of the final solution.</p>\n<h3>2.1 From Noise to Strategic Diversity (0.184 → 0.233)</h3>\n<p>The original baseline generated 5 predictions by adding Gaussian noise to a single morphed template. The problem: all 5 predictions were nearly identical, wasting the best-of-5 metric.</p>\n<p>The fix was to use genuinely different template subsets for each prediction slot — top-4, templates 3–6, templates 5–8, even-indexed, odd-indexed. Each slot blended its assigned templates by coordinate averaging. This created real structural diversity and jumped the score to <strong>0.233</strong>.</p>\n<h3>2.2 Energy-Based Reranking (0.233 → 0.237)</h3>\n<p>After generating 5 predictions, we sorted them by a physics-based pseudo-energy function so the most plausible structure became prediction 1. The energy combined bond length deviation from the ideal 5.9Å C1'–C1' distance and a steric clash penalty for atoms closer than 3.8Å.</p>\n<h3>2.3 Per-Sequence Routing (0.237 → 0.242)</h3>\n<p>Rather than applying one strategy uniformly, we diagnosed each sequence individually and routed it to the strategy where it performed best. The 7 sequences with no templates at all worked best with a diverse geometry fallback (helix, extended strand, compact globule, sequence-guided, reverse helix).</p>\n<h3>2.4 Composite Template Scoring</h3>\n<p>We replaced the raw alignment score with a composite that factors in length similarity, exact match ratio, and GC content similarity:</p>\n<pre><code>score = (align²) × len_sim × (1 + 0.3×exact) × (0.5 + 0.5×gc_sim)\n</code></pre>\n<p>Squaring the alignment score amplifies strong matches and penalises weak ones.</p>\n<h3>2.5 More Templates and Better Diversity (TOPK 8 → 15)</h3>\n<p>Increasing TOPK from 8 to 15 and generating up to 12 candidates per sequence gave the scoring function more material to work with. We also replaced Gaussian noise with structured diversity transforms applied in rotation:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Transform</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>Clean</td>\n<td>Best template, no noise</td>\n</tr>\n<tr>\n<td>1</td>\n<td>Tiny noise</td>\n<td>Proportional to alignment uncertainty</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Hinge rotation</td>\n<td>Rotates one segment around a pivot</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Jitter chains</td>\n<td>Independent rigid rotation + translation per chain</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Smooth wiggle</td>\n<td>Smooth interpolated displacement field</td>\n</tr>\n<tr>\n<td>5+</td>\n<td>Gaussian noise</td>\n<td>Fallback</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>3. The Randomness Problem</h2>\n<p>Before any architectural improvements were possible, we had to solve a more fundamental problem: <strong>the same notebook was producing scores of 0.392, 0.404, 0.409, and 0.426 on different runs</strong>. Without reproducibility, every experiment was noise.</p>\n<h3>3.1 Finding All 12 Sources of Non-Determinism</h3>\n<p>A systematic audit found 12 distinct sources of randomness. The most critical:</p>\n<ul>\n<li><strong><code>CUBLAS_WORKSPACE_CONFIG</code> set after PyTorch import</strong> — this environment variable must be set before <code>torch</code> is imported, or it has no effect on GPU determinism</li>\n<li><strong><code>np.random.seed()</code> inside prediction functions</strong> with <code>seed=len(predictions)*7+len(seq)</code> — if a morph failed upstream and <code>len(predictions)</code> differed from expected, the seed changed and all downstream noise changed</li>\n<li><strong><code>shortlist_templates</code> sort with no tiebreaker</strong> — equal alignment scores had non-deterministic ordering depending on dataframe memory layout</li>\n<li><strong><code>lru_cache(maxsize=200,000)</code></strong> — when the cache filled, eviction order varied between runs</li>\n<li><strong>Cross-product threshold at 1e-6</strong> — GPU vs CPU floating point differences could flip this condition</li>\n</ul>\n<h3>3.2 The Fix</h3>\n<p>Seven targeted changes restored full determinism:</p>\n<pre><code># Cell 0 — before ANY import\nimport os, random\nos.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'\nos.environ['PYTHONHASHSEED'] = '42'\nrandom.seed(42)\nimport numpy as np; np.random.seed(42)\nimport torch; torch.manual_seed(42)\ntorch.use_deterministic_algorithms(True, warn_only=True)\n</code></pre>\n<ul>\n<li>All <code>np.random.seed()</code> calls replaced with isolated <code>np.random.default_rng(seed=(ridx × 10B + slot × 10007) % 2³²)</code> — seed is globally unique per sequence and slot</li>\n<li>Added <code>tid</code> as secondary sort key in all template ranking functions</li>\n<li>Raised <code>lru_cache</code> to <code>maxsize=None</code></li>\n<li>Raised cross-product threshold to <code>0.1</code></li>\n</ul>\n<blockquote>\n  <p><strong>Result: Three consecutive submissions all scored exactly 0.392 — perfect determinism confirmed.</strong></p>\n</blockquote>\n<hr>\n<h2>4. The Template Diagnosis — Understanding the Ceiling</h2>\n<p>After fixing determinism, every experiment returned 0.392. To understand why, we built a deep diagnostic that computed the TM-score of raw template coordinates <strong>before any pipeline processing</strong>.</p>\n<table>\n<thead>\n<tr>\n<th>Finding</th>\n<th>Sequences Affected</th>\n<th>Conclusion</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>High alignment (&gt;0.4) but template TM &lt; 0.2</td>\n<td>10 sequences</td>\n<td>Same sequence, different 3D fold — conformational isomers</td>\n</tr>\n<tr>\n<td>Alignment &lt; 0.3 — no template exists</td>\n<td>8 sequences</td>\n<td>Genuinely novel folds — no known analog in any database</td>\n</tr>\n<tr>\n<td>morph_delta ≈ 0 for all sequences</td>\n<td>All 28</td>\n<td>Pipeline is NOT destroying good templates — ceiling is the data</td>\n</tr>\n<tr>\n<td>9EBP — only sequence with tmpl_TM &gt; 0.35</td>\n<td>1 sequence</td>\n<td>Good template exists, already handled well by pipeline</td>\n</tr>\n</tbody>\n</table>\n<p>The most striking finding was <strong>9LEL</strong> — alignment score 1.280 (near-identical sequence) but template TM = 0.020 (completely wrong structure). The template in the database folds the same sequence into a different 3D conformation. This is conformational isomerism — common in RNA, and impossible to fix by improving the morphing pipeline.</p>\n<blockquote>\n  <p><strong>Key insight: The pipeline was not the problem. The training database simply does not contain structurally correct templates for 18 of 28 test sequences. These were specifically chosen for the competition because they represent novel folds. No amount of pipeline tuning overcomes this — only a better generative model can.</strong></p>\n</blockquote>\n<hr>\n<h2>5. The Breakthrough — Generate → Rank → Select</h2>\n<h3>5.1 The Old Architecture Problem</h3>\n<p>The original pipeline pre-assigned slots before generating any predictions. It counted good TBM candidates (alignment ≥ 0.4), allocated N TBM slots, and sent the remaining 5-N slots to Protenix. The problem: if TBM slot 3 produced a better structure than Protenix slot 1, TBM slot 3 was discarded anyway. <strong>The pipeline had no quality gate.</strong></p>\n<h3>5.2 The New Architecture</h3>\n<p>We replaced slot allocation with a generate-then-rank approach:</p>\n<ul>\n<li>Generate up to 12 TBM candidates with all 5 diversity transforms</li>\n<li>Generate Protenix candidates with multiple seeds (N_sample=1 per seed)</li>\n<li>Pool all candidates together — source doesn't matter</li>\n<li>Score every candidate with <code>score_structure()</code></li>\n<li>Select top 5 by score, with a diversity check to prevent near-duplicate selections</li>\n</ul>\n<pre><code>def score_structure(coords):\n    diffs     = np.linalg.norm(np.diff(coords, axis=0), axis=1)\n    bond_pen  = np.sum((diffs - 5.9) ** 2)          # ideal C1'-C1' = 5.9Å\n    dm        = distance_matrix(coords, coords)\n    clash_pen = np.sum(dm &lt; 3.8)                    # atomic clash threshold\n    return -(clash_pen * 6.0 + bond_pen * 0.3)      # higher = better\n</code></pre>\n<p>TBM candidates also receive an alignment bonus <code>(align² × 8.0)</code> and a length-dependent bonus, ensuring that high-confidence template morphs rank above uncertain ones.</p>\n<h3>5.3 Dynamic Protenix Seed Schedule</h3>\n<p>Instead of one fixed seed with N_sample=5, Protenix runs N_sample=1 per seed with a sequence-length-dependent number of seeds:</p>\n<table>\n<thead>\n<tr>\n<th>Sequence Type</th>\n<th>Seed Count</th>\n<th>Rationale</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Short (≤400 nt)</td>\n<td>7 seeds</td>\n<td>More exploration — each run is fast</td>\n</tr>\n<tr>\n<td>Medium (400–1000 nt)</td>\n<td>3 seeds</td>\n<td>Balanced exploration vs compute</td>\n</tr>\n<tr>\n<td>9MME, 9ZCC</td>\n<td>2 seeds</td>\n<td>Chunked inference (512nt chunks, 128nt overlap, Kabsch-aligned stitching)</td>\n</tr>\n<tr>\n<td>OUR_SEQUENCES</td>\n<td>0 Protenix</td>\n<td>TBM is already near-optimal</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>6. Score Progression — Full Timeline</h2>\n<table>\n<thead>\n<tr>\n<th>Phase</th>\n<th>Key Change</th>\n<th>LB Score</th>\n<th>Delta</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline TBM</td>\n<td>Template morphing + Gaussian noise</td>\n<td>0.184</td>\n<td>—</td>\n</tr>\n<tr>\n<td>Experiment B</td>\n<td>Strategic template diversity (different subsets per slot)</td>\n<td>0.233</td>\n<td>+0.049</td>\n</tr>\n<tr>\n<td>Experiment 1</td>\n<td>Energy-based reranking of 5 predictions</td>\n<td>0.237</td>\n<td>+0.004</td>\n</tr>\n<tr>\n<td>FINAL routing</td>\n<td>Per-sequence strategy routing</td>\n<td>0.242</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Seed fix</td>\n<td>All randomness sources fixed — deterministic baseline</td>\n<td>0.392</td>\n<td>—</td>\n</tr>\n<tr>\n<td><strong>Generate→Rank→Select</strong></td>\n<td><strong>TOPK=15, all diversity transforms, score_structure()</strong></td>\n<td><strong>0.421</strong></td>\n<td><strong>+0.029</strong></td>\n</tr>\n<tr>\n<td>Protenix tuning</td>\n<td>Smoothness penalty, small-sequence boost, slot control</td>\n<td>0.422</td>\n<td>+0.001</td>\n</tr>\n<tr>\n<td><strong>Seed optimization</strong></td>\n<td><strong>Dual seed for long + 7 seeds for short sequences</strong></td>\n<td><strong>0.424</strong></td>\n<td><strong>+0.002</strong></td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p>Note: The jump from 0.242 to 0.392 reflects the addition of Protenix to the pipeline. The apparent gap is because the deterministic Protenix baseline (0.392) replaced earlier non-deterministic scores which included lucky runs up to 0.426.</p>\n</blockquote>\n<hr>\n<h2>7. Seed Strategy Optimization</h2>\n<p>The interaction between seed count and ranking strength was non-obvious. More seeds alone did not help — the ranking function had to be strong enough to select the best from a larger pool.</p>\n<table>\n<thead>\n<tr>\n<th>Setup</th>\n<th>LB Score</th>\n<th>Verdict</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline (5 samples, 1 seed)</td>\n<td>0.422</td>\n<td>—</td>\n</tr>\n<tr>\n<td>+ Dual seed for long sequences only</td>\n<td>0.416</td>\n<td>❌ Worse — expensive, low benefit alone</td>\n</tr>\n<tr>\n<td>+ 7 seeds for short sequences only</td>\n<td>0.420</td>\n<td>⚠️ Slight drop — ranking not strong enough</td>\n</tr>\n<tr>\n<td>Dual seed (long) + 7 seeds (short)</td>\n<td><strong>0.424</strong></td>\n<td>✅ Best — diversity and ranking work together</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p><strong>Key insight: More candidates ≠ better predictions. Better candidates + proper ranking = improvement. The dual-seed and 7-seed strategies only worked together because <code>score_structure()</code> was strong enough to identify the better structures from the larger pool.</strong></p>\n</blockquote>\n<hr>\n<h2>8. What Did Not Help</h2>\n<h3>Validation sequences in template pool</h3>\n<p>Adding validation labels as templates would have given better structural references. However, the alignment search over ~28,000 sequences (train + val) caused a Kaggle timeout every time. The pipeline never finished scoring.</p>\n<h3>Structure blending</h3>\n<p>Blending the top-2 candidates by coordinate averaging: <code>(c1 + c2) / 2</code>. The hypothesis was that averaging would smooth out errors. In practice, the top candidates were already well-optimised and the average degraded topology rather than improving it.</p>\n<h3>Hardcoded diversity per sequence</h3>\n<p>We diagnosed which diversity transform helped each sequence on the validation set and hardcoded the best transform per sequence. The +0.011 improvement on validation was too small to measure on the leaderboard.</p>\n<h3>Chunked Protenix for 9MME / 9ZCC in isolation</h3>\n<p>Splitting ultra-long sequences into 512nt chunks, running Protenix, and stitching with Kabsch alignment. The stitched output was no better than TBM for these sequences in isolation. Only when combined with dual-seed and the overall generate/rank/select architecture did it contribute.</p>\n<hr>\n<h2>9. What Could Have Been Better</h2>\n<ul>\n<li><strong>Fine-tuning Protenix on RNA-specific data</strong> — the 18 weak sequences represent novel folds that a better-trained model would handle. The base Protenix model was not optimised for these specific RNA topologies.</li>\n<li><strong>Confidence-based selection</strong> — Protenix outputs a pLDDT confidence score per residue. We experimented with using this to select the best samples across seeds, but the scores were often zero for long-sequence chunks due to extraction issues.</li>\n<li><strong>MSA depth for weak sequences</strong> — some weak sequences may have had sparse or missing MSA files, meaning Protenix ran without evolutionary information. Checking and augmenting MSA coverage could have improved Protenix quality on these sequences.</li>\n<li><strong>A fundamentally different model for the 18 hard sequences</strong> — AlphaFold3, RoseTTAFold2NA, or a fine-tuned Protenix. The generate/rank/select architecture is sound; the ceiling is the quality of the base models.</li>\n</ul>\n<hr>\n<h2>10. Key Lessons</h2>\n<h3>1. Fix randomness before anything else</h3>\n<p>Until seeds are deterministic, you cannot measure whether a change actually helps. Every experiment before the seed fix was polluted by Protenix variance. The seed fix was not a performance improvement — it was a prerequisite for all further meaningful work.</p>\n<h3>2. Diagnose per-sequence, not overall</h3>\n<p>Average metrics hide what's actually failing. The template TM-score diagnostic — computing TM of raw template coordinates before any pipeline processing — revealed in one experiment that the pipeline was not the bottleneck. Without this, we could have spent days optimising morphing and refinement code that would never move the leaderboard.</p>\n<h3>3. Understand your metric</h3>\n<p>Best-of-5 scoring means diversity across predictions is strictly more valuable than precision of one. The entire generate/rank/select architecture exists to exploit this. A pipeline that generates 12 diverse candidates and selects the best 5 beats a pipeline that generates 5 careful predictions.</p>\n<h3>4. TBM is signal; Protenix is exploration</h3>\n<p>When a good template exists (alignment &gt; 0.5), TBM is near-perfect and Protenix cannot improve on it. When no good template exists, TBM is useless and Protenix is the only option. Understanding this separation and routing accordingly is what kept the strong sequences strong while giving the weak sequences their best chance.</p>\n<h3>5. The database is the real ceiling</h3>\n<p>18 of 28 test sequences are genuinely novel RNA folds with no structurally similar solved structure in any public database. This is not a pipeline problem. These sequences were chosen for the competition precisely because they are hard. The only path to improving predictions on them is a better generative model.</p>\n<h3>6. Pipeline thinking beats model thinking</h3>\n<p>We did not change the underlying models. We changed how they are used — when to trust TBM, when to use Protenix, how many candidates to generate, and how to select among them. The 0.029 improvement from 0.392 to 0.421 came entirely from architectural decisions, not from improving the models themselves.</p>\n<hr>",
      "rawMarkdown": "## Silver Medal Solution · 68th / 1867 Teams · Stanford RNA 3D Folding Part 2\n\n> *How we went from 0.184 to 0.424 — fixing randomness, diagnosing templates, and redesigning the prediction pipeline*\n\n---\n\n> **Final LB Score: 0.424 · Deterministic · TBM + multi-seed Protenix · Generate → Rank → Select architecture**\n\n---\n\n## 1. Problem & Approach\n\nThe competition asks for 5 predicted 3D structures per RNA sequence. The metric takes the best TM-score among your 5 predictions — so the goal is not one perfect prediction, but five diverse, high-quality candidates.\n\nWe combined two complementary methods:\n\n- **Template-Based Modeling (TBM)** — find structurally similar sequences in a database of ~21,760 known solved structures, transfer their coordinates, apply diversity transforms, and score the results\n- **Protenix** — a diffusion-based neural network (similar to AlphaFold3) that folds RNA from sequence alone, run with multiple seeds for structural diversity\n\nThe 28 test sequences fall into three natural groups that required different handling:\n\n- **9 strong sequences (OUR_SEQUENCES)** — excellent TBM templates exist, TM-scores 0.52–0.80\n- **11 weak sequences** — no good templates, Protenix is the only option\n- **2 ultra-long sequences** (9MME=4640nt, 9ZCC=1460nt) — too long for single Protenix pass, require chunked inference\n\n---\n\n## 2. How TBM Became Strong — The Evolution\n\nTBM started at 0.184 and grew through several key iterations before becoming the backbone of the final solution.\n\n### 2.1 From Noise to Strategic Diversity (0.184 → 0.233)\n\nThe original baseline generated 5 predictions by adding Gaussian noise to a single morphed template. The problem: all 5 predictions were nearly identical, wasting the best-of-5 metric.\n\nThe fix was to use genuinely different template subsets for each prediction slot — top-4, templates 3–6, templates 5–8, even-indexed, odd-indexed. Each slot blended its assigned templates by coordinate averaging. This created real structural diversity and jumped the score to **0.233**.\n\n### 2.2 Energy-Based Reranking (0.233 → 0.237)\n\nAfter generating 5 predictions, we sorted them by a physics-based pseudo-energy function so the most plausible structure became prediction 1. The energy combined bond length deviation from the ideal 5.9Å C1'–C1' distance and a steric clash penalty for atoms closer than 3.8Å.\n\n### 2.3 Per-Sequence Routing (0.237 → 0.242)\n\nRather than applying one strategy uniformly, we diagnosed each sequence individually and routed it to the strategy where it performed best. The 7 sequences with no templates at all worked best with a diverse geometry fallback (helix, extended strand, compact globule, sequence-guided, reverse helix).\n\n### 2.4 Composite Template Scoring\n\nWe replaced the raw alignment score with a composite that factors in length similarity, exact match ratio, and GC content similarity:\n\n```python\nscore = (align²) × len_sim × (1 + 0.3×exact) × (0.5 + 0.5×gc_sim)\n```\n\nSquaring the alignment score amplifies strong matches and penalises weak ones.\n\n### 2.5 More Templates and Better Diversity (TOPK 8 → 15)\n\nIncreasing TOPK from 8 to 15 and generating up to 12 candidates per sequence gave the scoring function more material to work with. We also replaced Gaussian noise with structured diversity transforms applied in rotation:\n\n| Slot | Transform | Description |\n|------|-----------|-------------|\n| 0 | Clean | Best template, no noise |\n| 1 | Tiny noise | Proportional to alignment uncertainty |\n| 2 | Hinge rotation | Rotates one segment around a pivot |\n| 3 | Jitter chains | Independent rigid rotation + translation per chain |\n| 4 | Smooth wiggle | Smooth interpolated displacement field |\n| 5+ | Gaussian noise | Fallback |\n\n---\n\n## 3. The Randomness Problem\n\nBefore any architectural improvements were possible, we had to solve a more fundamental problem: **the same notebook was producing scores of 0.392, 0.404, 0.409, and 0.426 on different runs**. Without reproducibility, every experiment was noise.\n\n### 3.1 Finding All 12 Sources of Non-Determinism\n\nA systematic audit found 12 distinct sources of randomness. The most critical:\n\n- **`CUBLAS_WORKSPACE_CONFIG` set after PyTorch import** — this environment variable must be set before `torch` is imported, or it has no effect on GPU determinism\n- **`np.random.seed()` inside prediction functions** with `seed=len(predictions)*7+len(seq)` — if a morph failed upstream and `len(predictions)` differed from expected, the seed changed and all downstream noise changed\n- **`shortlist_templates` sort with no tiebreaker** — equal alignment scores had non-deterministic ordering depending on dataframe memory layout\n- **`lru_cache(maxsize=200,000)`** — when the cache filled, eviction order varied between runs\n- **Cross-product threshold at 1e-6** — GPU vs CPU floating point differences could flip this condition\n\n### 3.2 The Fix\n\nSeven targeted changes restored full determinism:\n\n```python\n# Cell 0 — before ANY import\nimport os, random\nos.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'\nos.environ['PYTHONHASHSEED'] = '42'\nrandom.seed(42)\nimport numpy as np; np.random.seed(42)\nimport torch; torch.manual_seed(42)\ntorch.use_deterministic_algorithms(True, warn_only=True)\n```\n\n- All `np.random.seed()` calls replaced with isolated `np.random.default_rng(seed=(ridx × 10B + slot × 10007) % 2³²)` — seed is globally unique per sequence and slot\n- Added `tid` as secondary sort key in all template ranking functions\n- Raised `lru_cache` to `maxsize=None`\n- Raised cross-product threshold to `0.1`\n\n> **Result: Three consecutive submissions all scored exactly 0.392 — perfect determinism confirmed.**\n\n---\n\n## 4. The Template Diagnosis — Understanding the Ceiling\n\nAfter fixing determinism, every experiment returned 0.392. To understand why, we built a deep diagnostic that computed the TM-score of raw template coordinates **before any pipeline processing**.\n\n| Finding | Sequences Affected | Conclusion |\n|---------|-------------------|------------|\n| High alignment (>0.4) but template TM < 0.2 | 10 sequences | Same sequence, different 3D fold — conformational isomers |\n| Alignment < 0.3 — no template exists | 8 sequences | Genuinely novel folds — no known analog in any database |\n| morph_delta ≈ 0 for all sequences | All 28 | Pipeline is NOT destroying good templates — ceiling is the data |\n| 9EBP — only sequence with tmpl_TM > 0.35 | 1 sequence | Good template exists, already handled well by pipeline |\n\nThe most striking finding was **9LEL** — alignment score 1.280 (near-identical sequence) but template TM = 0.020 (completely wrong structure). The template in the database folds the same sequence into a different 3D conformation. This is conformational isomerism — common in RNA, and impossible to fix by improving the morphing pipeline.\n\n> **Key insight: The pipeline was not the problem. The training database simply does not contain structurally correct templates for 18 of 28 test sequences. These were specifically chosen for the competition because they represent novel folds. No amount of pipeline tuning overcomes this — only a better generative model can.**\n\n---\n\n## 5. The Breakthrough — Generate → Rank → Select\n\n### 5.1 The Old Architecture Problem\n\nThe original pipeline pre-assigned slots before generating any predictions. It counted good TBM candidates (alignment ≥ 0.4), allocated N TBM slots, and sent the remaining 5-N slots to Protenix. The problem: if TBM slot 3 produced a better structure than Protenix slot 1, TBM slot 3 was discarded anyway. **The pipeline had no quality gate.**\n\n### 5.2 The New Architecture\n\nWe replaced slot allocation with a generate-then-rank approach:\n\n- Generate up to 12 TBM candidates with all 5 diversity transforms\n- Generate Protenix candidates with multiple seeds (N_sample=1 per seed)\n- Pool all candidates together — source doesn't matter\n- Score every candidate with `score_structure()`\n- Select top 5 by score, with a diversity check to prevent near-duplicate selections\n\n```python\ndef score_structure(coords):\n    diffs     = np.linalg.norm(np.diff(coords, axis=0), axis=1)\n    bond_pen  = np.sum((diffs - 5.9) ** 2)          # ideal C1'-C1' = 5.9Å\n    dm        = distance_matrix(coords, coords)\n    clash_pen = np.sum(dm < 3.8)                    # atomic clash threshold\n    return -(clash_pen * 6.0 + bond_pen * 0.3)      # higher = better\n```\n\nTBM candidates also receive an alignment bonus `(align² × 8.0)` and a length-dependent bonus, ensuring that high-confidence template morphs rank above uncertain ones.\n\n### 5.3 Dynamic Protenix Seed Schedule\n\nInstead of one fixed seed with N_sample=5, Protenix runs N_sample=1 per seed with a sequence-length-dependent number of seeds:\n\n| Sequence Type | Seed Count | Rationale |\n|--------------|------------|-----------|\n| Short (≤400 nt) | 7 seeds | More exploration — each run is fast |\n| Medium (400–1000 nt) | 3 seeds | Balanced exploration vs compute |\n| 9MME, 9ZCC | 2 seeds | Chunked inference (512nt chunks, 128nt overlap, Kabsch-aligned stitching) |\n| OUR_SEQUENCES | 0 Protenix | TBM is already near-optimal |\n\n---\n\n## 6. Score Progression — Full Timeline\n\n| Phase | Key Change | LB Score | Delta |\n|-------|-----------|----------|-------|\n| Baseline TBM | Template morphing + Gaussian noise | 0.184 | — |\n| Experiment B | Strategic template diversity (different subsets per slot) | 0.233 | +0.049 |\n| Experiment 1 | Energy-based reranking of 5 predictions | 0.237 | +0.004 |\n| FINAL routing | Per-sequence strategy routing | 0.242 | +0.005 |\n| Seed fix | All randomness sources fixed — deterministic baseline | 0.392 | — |\n| **Generate→Rank→Select** | **TOPK=15, all diversity transforms, score_structure()** | **0.421** | **+0.029** |\n| Protenix tuning | Smoothness penalty, small-sequence boost, slot control | 0.422 | +0.001 |\n| **Seed optimization** | **Dual seed for long + 7 seeds for short sequences** | **0.424** | **+0.002** |\n\n> Note: The jump from 0.242 to 0.392 reflects the addition of Protenix to the pipeline. The apparent gap is because the deterministic Protenix baseline (0.392) replaced earlier non-deterministic scores which included lucky runs up to 0.426.\n\n---\n\n## 7. Seed Strategy Optimization\n\nThe interaction between seed count and ranking strength was non-obvious. More seeds alone did not help — the ranking function had to be strong enough to select the best from a larger pool.\n\n| Setup | LB Score | Verdict |\n|-------|----------|---------|\n| Baseline (5 samples, 1 seed) | 0.422 | — |\n| + Dual seed for long sequences only | 0.416 | ❌ Worse — expensive, low benefit alone |\n| + 7 seeds for short sequences only | 0.420 | ⚠️ Slight drop — ranking not strong enough |\n| Dual seed (long) + 7 seeds (short) | **0.424** | ✅ Best — diversity and ranking work together |\n\n> **Key insight: More candidates ≠ better predictions. Better candidates + proper ranking = improvement. The dual-seed and 7-seed strategies only worked together because `score_structure()` was strong enough to identify the better structures from the larger pool.**\n\n---\n\n## 8. What Did Not Help\n\n### Validation sequences in template pool\nAdding validation labels as templates would have given better structural references. However, the alignment search over ~28,000 sequences (train + val) caused a Kaggle timeout every time. The pipeline never finished scoring.\n\n### Structure blending\nBlending the top-2 candidates by coordinate averaging: `(c1 + c2) / 2`. The hypothesis was that averaging would smooth out errors. In practice, the top candidates were already well-optimised and the average degraded topology rather than improving it.\n\n### Hardcoded diversity per sequence\nWe diagnosed which diversity transform helped each sequence on the validation set and hardcoded the best transform per sequence. The +0.011 improvement on validation was too small to measure on the leaderboard.\n\n### Chunked Protenix for 9MME / 9ZCC in isolation\nSplitting ultra-long sequences into 512nt chunks, running Protenix, and stitching with Kabsch alignment. The stitched output was no better than TBM for these sequences in isolation. Only when combined with dual-seed and the overall generate/rank/select architecture did it contribute.\n\n---\n\n## 9. What Could Have Been Better\n\n- **Fine-tuning Protenix on RNA-specific data** — the 18 weak sequences represent novel folds that a better-trained model would handle. The base Protenix model was not optimised for these specific RNA topologies.\n- **Confidence-based selection** — Protenix outputs a pLDDT confidence score per residue. We experimented with using this to select the best samples across seeds, but the scores were often zero for long-sequence chunks due to extraction issues.\n- **MSA depth for weak sequences** — some weak sequences may have had sparse or missing MSA files, meaning Protenix ran without evolutionary information. Checking and augmenting MSA coverage could have improved Protenix quality on these sequences.\n- **A fundamentally different model for the 18 hard sequences** — AlphaFold3, RoseTTAFold2NA, or a fine-tuned Protenix. The generate/rank/select architecture is sound; the ceiling is the quality of the base models.\n\n---\n\n## 10. Key Lessons\n\n### 1. Fix randomness before anything else\nUntil seeds are deterministic, you cannot measure whether a change actually helps. Every experiment before the seed fix was polluted by Protenix variance. The seed fix was not a performance improvement — it was a prerequisite for all further meaningful work.\n\n### 2. Diagnose per-sequence, not overall\nAverage metrics hide what's actually failing. The template TM-score diagnostic — computing TM of raw template coordinates before any pipeline processing — revealed in one experiment that the pipeline was not the bottleneck. Without this, we could have spent days optimising morphing and refinement code that would never move the leaderboard.\n\n### 3. Understand your metric\nBest-of-5 scoring means diversity across predictions is strictly more valuable than precision of one. The entire generate/rank/select architecture exists to exploit this. A pipeline that generates 12 diverse candidates and selects the best 5 beats a pipeline that generates 5 careful predictions.\n\n### 4. TBM is signal; Protenix is exploration\nWhen a good template exists (alignment > 0.5), TBM is near-perfect and Protenix cannot improve on it. When no good template exists, TBM is useless and Protenix is the only option. Understanding this separation and routing accordingly is what kept the strong sequences strong while giving the weak sequences their best chance.\n\n### 5. The database is the real ceiling\n18 of 28 test sequences are genuinely novel RNA folds with no structurally similar solved structure in any public database. This is not a pipeline problem. These sequences were chosen for the competition precisely because they are hard. The only path to improving predictions on them is a better generative model.\n\n### 6. Pipeline thinking beats model thinking\nWe did not change the underlying models. We changed how they are used — when to trust TBM, when to use Protenix, how many candidates to generate, and how to select among them. The 0.029 improvement from 0.392 to 0.421 came entirely from architectural decisions, not from improving the models themselves.\n\n---",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3434651": "## Silver Medal Solution · 68th / 1867 Teams · Stanford RNA 3D Folding Part 2\n\n> *How we went from 0.184 to 0.424 — fixing randomness, diagnosing templates, and redesigning the prediction pipeline*\n\n---\n\n> **Final LB Score: 0.424 · Deterministic · TBM + multi-seed Protenix · Generate → Rank → Select architecture**\n\n---\n\n## 1. Problem & Approach\n\nThe competition asks for 5 predicted 3D structures per RNA sequence. The metric takes the best TM-score among your 5 predictions — so the goal is not one perfect prediction, but five diverse, high-quality candidates.\n\nWe combined two complementary methods:\n\n- **Template-Based Modeling (TBM)** — find structurally similar sequences in a database of ~21,760 known solved structures, transfer their coordinates, apply diversity transforms, and score the results\n- **Protenix** — a diffusion-based neural network (similar to AlphaFold3) that folds RNA from sequence alone, run with multiple seeds for structural diversity\n\nThe 28 test sequences fall into three natural groups that required different handling:\n\n- **9 strong sequences (OUR_SEQUENCES)** — excellent TBM templates exist, TM-scores 0.52–0.80\n- **11 weak sequences** — no good templates, Protenix is the only option\n- **2 ultra-long sequences** (9MME=4640nt, 9ZCC=1460nt) — too long for single Protenix pass, require chunked inference\n\n---\n\n## 2. How TBM Became Strong — The Evolution\n\nTBM started at 0.184 and grew through several key iterations before becoming the backbone of the final solution.\n\n### 2.1 From Noise to Strategic Diversity (0.184 → 0.233)\n\nThe original baseline generated 5 predictions by adding Gaussian noise to a single morphed template. The problem: all 5 predictions were nearly identical, wasting the best-of-5 metric.\n\nThe fix was to use genuinely different template subsets for each prediction slot — top-4, templates 3–6, templates 5–8, even-indexed, odd-indexed. Each slot blended its assigned templates by coordinate averaging. This created real structural diversity and jumped the score to **0.233**.\n\n### 2.2 Energy-Based Reranking (0.233 → 0.237)\n\nAfter generating 5 predictions, we sorted them by a physics-based pseudo-energy function so the most plausible structure became prediction 1. The energy combined bond length deviation from the ideal 5.9Å C1'–C1' distance and a steric clash penalty for atoms closer than 3.8Å.\n\n### 2.3 Per-Sequence Routing (0.237 → 0.242)\n\nRather than applying one strategy uniformly, we diagnosed each sequence individually and routed it to the strategy where it performed best. The 7 sequences with no templates at all worked best with a diverse geometry fallback (helix, extended strand, compact globule, sequence-guided, reverse helix).\n\n### 2.4 Composite Template Scoring\n\nWe replaced the raw alignment score with a composite that factors in length similarity, exact match ratio, and GC content similarity:\n\n```python\nscore = (align²) × len_sim × (1 + 0.3×exact) × (0.5 + 0.5×gc_sim)\n```\n\nSquaring the alignment score amplifies strong matches and penalises weak ones.\n\n### 2.5 More Templates and Better Diversity (TOPK 8 → 15)\n\nIncreasing TOPK from 8 to 15 and generating up to 12 candidates per sequence gave the scoring function more material to work with. We also replaced Gaussian noise with structured diversity transforms applied in rotation:\n\n| Slot | Transform | Description |\n|------|-----------|-------------|\n| 0 | Clean | Best template, no noise |\n| 1 | Tiny noise | Proportional to alignment uncertainty |\n| 2 | Hinge rotation | Rotates one segment around a pivot |\n| 3 | Jitter chains | Independent rigid rotation + translation per chain |\n| 4 | Smooth wiggle | Smooth interpolated displacement field |\n| 5+ | Gaussian noise | Fallback |\n\n---\n\n## 3. The Randomness Problem\n\nBefore any architectural improvements were possible, we had to solve a more fundamental problem: **the same notebook was producing scores of 0.392, 0.404, 0.409, and 0.426 on different runs**. Without reproducibility, every experiment was noise.\n\n### 3.1 Finding All 12 Sources of Non-Determinism\n\nA systematic audit found 12 distinct sources of randomness. The most critical:\n\n- **`CUBLAS_WORKSPACE_CONFIG` set after PyTorch import** — this environment variable must be set before `torch` is imported, or it has no effect on GPU determinism\n- **`np.random.seed()` inside prediction functions** with `seed=len(predictions)*7+len(seq)` — if a morph failed upstream and `len(predictions)` differed from expected, the seed changed and all downstream noise changed\n- **`shortlist_templates` sort with no tiebreaker** — equal alignment scores had non-deterministic ordering depending on dataframe memory layout\n- **`lru_cache(maxsize=200,000)`** — when the cache filled, eviction order varied between runs\n- **Cross-product threshold at 1e-6** — GPU vs CPU floating point differences could flip this condition\n\n### 3.2 The Fix\n\nSeven targeted changes restored full determinism:\n\n```python\n# Cell 0 — before ANY import\nimport os, random\nos.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'\nos.environ['PYTHONHASHSEED'] = '42'\nrandom.seed(42)\nimport numpy as np; np.random.seed(42)\nimport torch; torch.manual_seed(42)\ntorch.use_deterministic_algorithms(True, warn_only=True)\n```\n\n- All `np.random.seed()` calls replaced with isolated `np.random.default_rng(seed=(ridx × 10B + slot × 10007) % 2³²)` — seed is globally unique per sequence and slot\n- Added `tid` as secondary sort key in all template ranking functions\n- Raised `lru_cache` to `maxsize=None`\n- Raised cross-product threshold to `0.1`\n\n> **Result: Three consecutive submissions all scored exactly 0.392 — perfect determinism confirmed.**\n\n---\n\n## 4. The Template Diagnosis — Understanding the Ceiling\n\nAfter fixing determinism, every experiment returned 0.392. To understand why, we built a deep diagnostic that computed the TM-score of raw template coordinates **before any pipeline processing**.\n\n| Finding | Sequences Affected | Conclusion |\n|---------|-------------------|------------|\n| High alignment (>0.4) but template TM < 0.2 | 10 sequences | Same sequence, different 3D fold — conformational isomers |\n| Alignment < 0.3 — no template exists | 8 sequences | Genuinely novel folds — no known analog in any database |\n| morph_delta ≈ 0 for all sequences | All 28 | Pipeline is NOT destroying good templates — ceiling is the data |\n| 9EBP — only sequence with tmpl_TM > 0.35 | 1 sequence | Good template exists, already handled well by pipeline |\n\nThe most striking finding was **9LEL** — alignment score 1.280 (near-identical sequence) but template TM = 0.020 (completely wrong structure). The template in the database folds the same sequence into a different 3D conformation. This is conformational isomerism — common in RNA, and impossible to fix by improving the morphing pipeline.\n\n> **Key insight: The pipeline was not the problem. The training database simply does not contain structurally correct templates for 18 of 28 test sequences. These were specifically chosen for the competition because they represent novel folds. No amount of pipeline tuning overcomes this — only a better generative model can.**\n\n---\n\n## 5. The Breakthrough — Generate → Rank → Select\n\n### 5.1 The Old Architecture Problem\n\nThe original pipeline pre-assigned slots before generating any predictions. It counted good TBM candidates (alignment ≥ 0.4), allocated N TBM slots, and sent the remaining 5-N slots to Protenix. The problem: if TBM slot 3 produced a better structure than Protenix slot 1, TBM slot 3 was discarded anyway. **The pipeline had no quality gate.**\n\n### 5.2 The New Architecture\n\nWe replaced slot allocation with a generate-then-rank approach:\n\n- Generate up to 12 TBM candidates with all 5 diversity transforms\n- Generate Protenix candidates with multiple seeds (N_sample=1 per seed)\n- Pool all candidates together — source doesn't matter\n- Score every candidate with `score_structure()`\n- Select top 5 by score, with a diversity check to prevent near-duplicate selections\n\n```python\ndef score_structure(coords):\n    diffs     = np.linalg.norm(np.diff(coords, axis=0), axis=1)\n    bond_pen  = np.sum((diffs - 5.9) ** 2)          # ideal C1'-C1' = 5.9Å\n    dm        = distance_matrix(coords, coords)\n    clash_pen = np.sum(dm < 3.8)                    # atomic clash threshold\n    return -(clash_pen * 6.0 + bond_pen * 0.3)      # higher = better\n```\n\nTBM candidates also receive an alignment bonus `(align² × 8.0)` and a length-dependent bonus, ensuring that high-confidence template morphs rank above uncertain ones.\n\n### 5.3 Dynamic Protenix Seed Schedule\n\nInstead of one fixed seed with N_sample=5, Protenix runs N_sample=1 per seed with a sequence-length-dependent number of seeds:\n\n| Sequence Type | Seed Count | Rationale |\n|--------------|------------|-----------|\n| Short (≤400 nt) | 7 seeds | More exploration — each run is fast |\n| Medium (400–1000 nt) | 3 seeds | Balanced exploration vs compute |\n| 9MME, 9ZCC | 2 seeds | Chunked inference (512nt chunks, 128nt overlap, Kabsch-aligned stitching) |\n| OUR_SEQUENCES | 0 Protenix | TBM is already near-optimal |\n\n---\n\n## 6. Score Progression — Full Timeline\n\n| Phase | Key Change | LB Score | Delta |\n|-------|-----------|----------|-------|\n| Baseline TBM | Template morphing + Gaussian noise | 0.184 | — |\n| Experiment B | Strategic template diversity (different subsets per slot) | 0.233 | +0.049 |\n| Experiment 1 | Energy-based reranking of 5 predictions | 0.237 | +0.004 |\n| FINAL routing | Per-sequence strategy routing | 0.242 | +0.005 |\n| Seed fix | All randomness sources fixed — deterministic baseline | 0.392 | — |\n| **Generate→Rank→Select** | **TOPK=15, all diversity transforms, score_structure()** | **0.421** | **+0.029** |\n| Protenix tuning | Smoothness penalty, small-sequence boost, slot control | 0.422 | +0.001 |\n| **Seed optimization** | **Dual seed for long + 7 seeds for short sequences** | **0.424** | **+0.002** |\n\n> Note: The jump from 0.242 to 0.392 reflects the addition of Protenix to the pipeline. The apparent gap is because the deterministic Protenix baseline (0.392) replaced earlier non-deterministic scores which included lucky runs up to 0.426.\n\n---\n\n## 7. Seed Strategy Optimization\n\nThe interaction between seed count and ranking strength was non-obvious. More seeds alone did not help — the ranking function had to be strong enough to select the best from a larger pool.\n\n| Setup | LB Score | Verdict |\n|-------|----------|---------|\n| Baseline (5 samples, 1 seed) | 0.422 | — |\n| + Dual seed for long sequences only | 0.416 | ❌ Worse — expensive, low benefit alone |\n| + 7 seeds for short sequences only | 0.420 | ⚠️ Slight drop — ranking not strong enough |\n| Dual seed (long) + 7 seeds (short) | **0.424** | ✅ Best — diversity and ranking work together |\n\n> **Key insight: More candidates ≠ better predictions. Better candidates + proper ranking = improvement. The dual-seed and 7-seed strategies only worked together because `score_structure()` was strong enough to identify the better structures from the larger pool.**\n\n---\n\n## 8. What Did Not Help\n\n### Validation sequences in template pool\nAdding validation labels as templates would have given better structural references. However, the alignment search over ~28,000 sequences (train + val) caused a Kaggle timeout every time. The pipeline never finished scoring.\n\n### Structure blending\nBlending the top-2 candidates by coordinate averaging: `(c1 + c2) / 2`. The hypothesis was that averaging would smooth out errors. In practice, the top candidates were already well-optimised and the average degraded topology rather than improving it.\n\n### Hardcoded diversity per sequence\nWe diagnosed which diversity transform helped each sequence on the validation set and hardcoded the best transform per sequence. The +0.011 improvement on validation was too small to measure on the leaderboard.\n\n### Chunked Protenix for 9MME / 9ZCC in isolation\nSplitting ultra-long sequences into 512nt chunks, running Protenix, and stitching with Kabsch alignment. The stitched output was no better than TBM for these sequences in isolation. Only when combined with dual-seed and the overall generate/rank/select architecture did it contribute.\n\n---\n\n## 9. What Could Have Been Better\n\n- **Fine-tuning Protenix on RNA-specific data** — the 18 weak sequences represent novel folds that a better-trained model would handle. The base Protenix model was not optimised for these specific RNA topologies.\n- **Confidence-based selection** — Protenix outputs a pLDDT confidence score per residue. We experimented with using this to select the best samples across seeds, but the scores were often zero for long-sequence chunks due to extraction issues.\n- **MSA depth for weak sequences** — some weak sequences may have had sparse or missing MSA files, meaning Protenix ran without evolutionary information. Checking and augmenting MSA coverage could have improved Protenix quality on these sequences.\n- **A fundamentally different model for the 18 hard sequences** — AlphaFold3, RoseTTAFold2NA, or a fine-tuned Protenix. The generate/rank/select architecture is sound; the ceiling is the quality of the base models.\n\n---\n\n## 10. Key Lessons\n\n### 1. Fix randomness before anything else\nUntil seeds are deterministic, you cannot measure whether a change actually helps. Every experiment before the seed fix was polluted by Protenix variance. The seed fix was not a performance improvement — it was a prerequisite for all further meaningful work.\n\n### 2. Diagnose per-sequence, not overall\nAverage metrics hide what's actually failing. The template TM-score diagnostic — computing TM of raw template coordinates before any pipeline processing — revealed in one experiment that the pipeline was not the bottleneck. Without this, we could have spent days optimising morphing and refinement code that would never move the leaderboard.\n\n### 3. Understand your metric\nBest-of-5 scoring means diversity across predictions is strictly more valuable than precision of one. The entire generate/rank/select architecture exists to exploit this. A pipeline that generates 12 diverse candidates and selects the best 5 beats a pipeline that generates 5 careful predictions.\n\n### 4. TBM is signal; Protenix is exploration\nWhen a good template exists (alignment > 0.5), TBM is near-perfect and Protenix cannot improve on it. When no good template exists, TBM is useless and Protenix is the only option. Understanding this separation and routing accordingly is what kept the strong sequences strong while giving the weak sequences their best chance.\n\n### 5. The database is the real ceiling\n18 of 28 test sequences are genuinely novel RNA folds with no structurally similar solved structure in any public database. This is not a pipeline problem. These sequences were chosen for the competition precisely because they are hard. The only path to improving predictions on them is a better generative model.\n\n### 6. Pipeline thinking beats model thinking\nWe did not change the underlying models. We changed how they are used — when to trust TBM, when to use Protenix, how many candidates to generate, and how to select among them. The 0.029 improvement from 0.392 to 0.421 came entirely from architectural decisions, not from improving the models themselves.\n\n---"
  }
}