{
  "id": 611908,
  "title": "9th place solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611908",
  "author_name": "Tom",
  "post_date": "2025-10-15T13:34:16.853000",
  "votes": 47,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>Acknowledgement</h2>\n<p>First, we thank the competition hosts <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, <a href=\"https://www.kaggle.com/shosys\" target=\"_blank\">@shosys</a> and every person involved from RSNA in data preparation and competition related process, and Kaggle staff for organizing this competition. </p>\n<p>Below, we introduce the solution of team <strong>Vibes and Genius Trade-Off</strong> -- <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>, <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>, <a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a>, <a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a>!</p>\n<p><strong>TL;DR</strong></p>\n<p>Our solution combines three complementary approaches:</p>\n<ul>\n<li>YOLO 2.5D with different backbones. Derived from <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU competition</a></li>\n<li>3D CenterNet with 2D Effv2s extractor.</li>\n<li>Meta-classifiers (LightGBM, XGBoost, CatBoost)</li>\n</ul>\n<p>The final probabilities are obtained by averaging the outputs of the YOLO 2.5D models, 3D CenterNet with 2D Effv2s extractor and the three meta-classifiers.</p>\n<p>Here is the diagram of our approach.</p>\n<p>We used two different variations of YOLO:</p>\n<ol>\n<li><strong>YOLOv11m</strong> - The standard YOLO11 medium model from Ultralytcs</li>\n<li><strong>Custom YOLO with timm backbone</strong> - YOLO architecture with <code>timm/tf_efficientnetv2_s.in21k_ft_in1k</code> as the backbone</li>\n</ol>\n<p>For details about the YOLO+timm customization, please refer to the <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU writeup</a>.</p>\n<p>We treated each of the <strong>13 vessel locations as separate bounding box classes</strong>:</p>\n<ul>\n<li>Left/Right Infraclinoid Internal Carotid Artery</li>\n<li>Left/Right Supraclinoid Internal Carotid Artery</li>\n<li>Left/Right Middle Cerebral Artery</li>\n<li>Anterior Communicating Artery</li>\n<li>Left/Right Anterior Cerebral Artery</li>\n<li>Left/Right Posterior Communicating Artery</li>\n<li>Basilar Tip</li>\n<li>Other Posterior Circulation</li>\n</ul>\n<h3><strong>2.5D Strategy</strong></h3>\n<p>Images were resized to 512×512, normalized with min-max and converted to 2.5D variation:</p>\n<p><code>R (Red channel)   = slice i-1\nG (Green channel) = slice i\nB (Blue channel)  = slice i+1</code></p>\n<p>The figure below shows examples of slices in 2D vs 2.5D variation:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fbc5e5f0a3d330bfc78d7961160b805f6%2F2D.png?generation=1760534875382926&amp;alt=media\" alt=\"\"></p>\n<p>As can be seen in the examples, the images vary significantly because <strong>we didn't standardize Z spacing</strong>.</p>\n<p>I (@sersasj) spent 1-2 weeks experimenting with Z-axis resize, but results were consistently worse. Perhaps I was doing something wrong, but didn't have time to investigate further.</p>\n<p>We were very reluctant to try 2.5D without proper Z-spacing resize (in my mind didn’t seem a good idea - <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>). Nevertheless, the 2.5D approach even without Z-resampling improved results by 0.02+ in CV.</p>\n<h3><strong>Training Configuration</strong></h3>\n<p>Data Sampling Strategy:</p>\n<p>For negative samples we took 10 evenly sampled slices per series.</p>\n<p>Example: For a 100-slice series → slices [1, 11, 22, 33, 44, 56, 67, 78, 89, 99]</p>\n<p>For positive samples we used all slices containing annotations</p>\n<h3><strong>YOLO with timm/tf_efficientnetv2_s Backbone</strong></h3>\n<pre><code> \n \n \n \n  \n \n \n \n   \n</code></pre>\n<h3><strong>YOLOv11m</strong></h3>\n<pre><code> \n \n \n \n \n \n \n \n   \n</code></pre>\n<p>We also modified the fitness function to include auc metric and prioritize mAP@50.</p>\n<pre><code>fitness =  × mAP@ +  × mAP@- +  × mAUC\n</code></pre>\n<p>Did this help? Maybe a little. In the <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU competition</a>, we had already found that prioritizing mAP@50 over the Ultralytics standard 1.0×mAP@50-95 gave better results.</p>\n<p>Adding AUC seemed like a good idea at the time since it's the competition metric, but we didn't investigate thoroughly whether it significantly improved performance. It's possible the benefit was marginal, but we kept it for consistency.</p>\n<p><strong>Training Hardware and Time</strong></p>\n<table>\n<thead>\n<tr>\n<th>Hardware Setup</th>\n<th>GPU</th>\n<th>CPU</th>\n<th>RAM</th>\n<th>Time per Fold</th>\n<th>Total (5 folds)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a></td>\n<td>RTX 3090</td>\n<td>Intel Core i5-12400F (12) @ 4.4GHz</td>\n<td>32GB</td>\n<td>4-5 hours</td>\n<td>~24 hours</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a></td>\n<td>RTX 4090</td>\n<td>AMD Ryzen 9 7950X (32) @ 5.883GHz</td>\n<td>64GB</td>\n<td>~2 hours</td>\n<td>~10 hours</td>\n</tr>\n</tbody>\n</table>\n<h3>Inference</h3>\n<p>For inference, we sort DICOM slices by spatial position (SliceLocation → ImagePositionPatient → InstanceNumber), create 2.5D triplets for slices 1 to N-1<br>\nRun batch inference across both models, extract max confidence per class and use max(localization_confidences) as overall aneurysm presence probability<br>\nThe process can be observed in the gif below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F7a4a4d9d460b5a6b0cb6eb82a3c4b5ad%2Fyolo_predictions_1.2.826.0.1.3680043.8.498.10022688097731894079510930966432818105%20(2).gif?generation=1761065820372307&amp;alt=media\" alt=\"\"></p>\n<h3>Cross-Validation Scores</h3>\n<p>folds were stratified using <code>MultilabelStratifiedKFold</code> to ensure balanced distribution across:</p>\n<ul>\n<li>Aneurysm presence</li>\n<li>All 13 vessel locations</li>\n<li>Modality</li>\n</ul>\n<h3><strong>YOLO11m 2.5D Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.8184</td>\n<td>0.7549</td>\n<td>0.7867</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8114</td>\n<td>0.7897</td>\n<td>0.8005</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8134</td>\n<td>0.7827</td>\n<td>0.7981</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8162</td>\n<td>0.7872</td>\n<td>0.8017</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8540</td>\n<td>0.8281</td>\n<td>0.8410</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8227</strong></td>\n<td><strong>0.7885</strong></td>\n<td><strong>0.8056</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>EfficientNetV2-S 2.5D Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7978</td>\n<td>0.7790</td>\n<td>0.7884</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8159</td>\n<td>0.8269</td>\n<td>0.8214</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8155</td>\n<td>0.7921</td>\n<td>0.8038</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8153</td>\n<td>0.7907</td>\n<td>0.8030</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8499</td>\n<td>0.8560</td>\n<td>0.8529</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8189</strong></td>\n<td><strong>0.8090</strong></td>\n<td><strong>0.8139</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>Ensemble Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.8393</td>\n<td>0.7849</td>\n<td>0.8121</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8330</td>\n<td>0.8291</td>\n<td>0.8310</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8399</td>\n<td>0.8135</td>\n<td>0.8267</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8370</td>\n<td>0.8093</td>\n<td>0.8232</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8738</td>\n<td>0.8620</td>\n<td>0.8679</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8446</strong></td>\n<td><strong>0.8198</strong></td>\n<td><strong>0.8322</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work:</h2>\n<p>In the initial stages of YOLO development <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> created an ensemble with 3xYolo11m and 0.69LB <a href=\"https://www.kaggle.com/code/yosukeyama/rsna2025-32ch-img-infer-lb-0-69-share\" target=\"_blank\">EfficientNetB2 public notebook</a>. That gave us a score of 0.78 LB and put us in the top 3 in early stages of competition</p>\n<p>After a bit of probing the leaderboard, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> found out that YOLO's aneurysm classification AUROC was very low (around ~0.58). Then we started to develop models that would complement YOLO and boost the aneurysm classification AUROC. (We later discovered that the poor results of early YOLO development were due to an insufficient amount of negative samples in training.)</p>\n<p>One of the class of models that we used was 2.5D models that just detects the presence of aneurysm in the given series with auxiliary segmentation head. These models were inspired by <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a>. These models in combination with yolo11m gave us a LB of 0.80. </p>\n<p>These models then became obsolete when we merged with  <a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a>, <a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a>.</p>\n<p>In the BYU competition, our best-performing backbone was <code>timm/convnextv2_base.fcmae_ft_in22k_in1k</code>. However, in the current competition, training with half precision caused the loss to quickly become NaN. Based on our investigation, this appears to be a common issue with ConvNeXt architectures. We hypothesize that if trained in full precision, this backbone could achieve performance comparable to YOLO11m and <code>timm/tf_efficientnetv2_s.in21k_ft_in1k</code>. Unfortunately, we were unable to verify this due to time constraints.</p>\n<p>Modifying the neck and head was also considered. From what we investigate with <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">awesome reverse engineer notebook</a> P3 features appeared to be the most utilized. We conducted a quick test on a single fold, but results on validation were more or less the same so we decided to move on.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7a4ca157151f82f0436dfdee766ec665%2Fimage%20(2).png?generation=1760534911466456&amp;alt=media\" alt=\"\"></p>\n<p>We also try to use our 3d segmentation result to <a href=\"https://www.kaggle.com/code/tom99763/rsna-iad-vessel-segmentation\" target=\"_blank\">filter out invalid location prediction from yolo</a>, but recall drops a lot.<br>\nMoreover, we tried extracting patches using YOLO and applying wavelet and log-polar transforms to the patches before predicting their classes; however, <a href=\"https://www.kaggle.com/code/tom99763/yolos-phw-mip-lp-rsna-i\" target=\"_blank\">this approach </a> did not yield good results.</p>\n<h2>EfficientV2s + 3D-CenterNet (Flayer)</h2>\n<p>This is a highly <strong>experimental project</strong>, made with a lot of help from AI. We started with the popular “LB 0.69 Public Notebook” , reconstructed the entire <strong>training pipeline and data preprocessing</strong> from it, and then incorporated a <strong>CenterNet-like mechanism to detect the center of the aneurysms</strong>.</p>\n<h3>Model architecture</h3>\n<p>We adopt a <strong>slice-wise 2D encoder + shallow 3D head</strong> design:</p>\n<ul>\n<li><strong>Encoder.</strong> A 2D ImageNet-pretrained <strong>EfficientNetV2-S</strong> (<code>tf_efficientnetv2_s.in21k_ft_in1k</code>) from <code>timm</code> in <em>features_only</em> mode with the <strong>penultimate</strong> feature map (<code>out_indices = (1,-2)</code>).</li>\n<li><strong>Auxiliary 2D head.</strong>The intermediate feature map (early stage) is fed into a lightweight 2D convolutional head to produce per-slice auxiliary logits, which encourage the encoder to capture vessel-relevant spatial cues before 3D aggregation. The auxiliary head is supervised with slice-level targets to stabilize early-layer learning.</li>\n<li><strong>2D→3D fusion.</strong> Each volume <code>(B, C=1, D, H, W)</code> is <strong>rearranged into <code>B·D</code> 2D slices</strong>, encoded independently, then <strong>reassembled</strong> to <code>(B, C_feat, D, H′, W′)</code> before the 3D head. If the backbone expects 3 channels, single-channel inputs are channel-repeated.</li>\n<li><strong>3D temporal head.</strong> A lightweight head (Conv3d-BN-ReLU ×2, <code>base_channels</code>) aggregates signals across depth.</li>\n<li><strong>Outputs.</strong> Two 1×1×1 conv heads:<ul>\n<li><strong>Heatmap head:</strong> <code>num_classes = 13</code> vessel locations (3D centerness maps)</li>\n<li><strong>Offset head:</strong> 3-channel sub-voxel offsets <code>(dz, dy, dx)</code></li></ul></li>\n</ul>\n<h3>Supervision targets</h3>\n<ul>\n<li>Place a <strong>3D Gaussian</strong> peak (stride-aware) at each annotated center on the class-specific heatmap.</li>\n<li>Store <strong>sub-voxel residual offsets</strong> at the nearest grid cell and mark them with an <strong>offset mask</strong>.</li>\n<li>Scale coordinates from the original series size to the current <code>(D, H, W)</code>; map to the output grid using <code>(stride_d=1, stride_h=stride_w=16)</code> so the output is <code>D × H/16 × W/16</code>.</li>\n<li>Default <strong>Gaussian σ = 0.2</strong>.</li>\n</ul>\n<h3>Segmentation masks</h3>\n<p>Vessel masks are produced in physical space by training a 3D DynUNet on volumes resampled to 0.7 mm isotropic spacing. The provided voxel-level annotations of 13 vessel classes are merged into a single binary vessel mask for training. The resulting 3D masks are then converted to slice-wise 2D masks to supervise the auxiliary head.</p>\n<h3>Losses</h3>\n<p>We train with four terms and sum them with weights:</p>\n<ul>\n<li><strong>CenterNet-style focal loss</strong> (weight = 1.0) on heatmaps (α=2, β=4).</li>\n<li><strong>L1 offset loss</strong> (weight = 1.0) (only at valid center cells).</li>\n<li><strong>BCE-with-logits classification loss</strong> (weight = 1.0) from series-level logits .</li>\n<li><strong>Auxiliary Dice &amp; BCE loss</strong> (weight = 0.5) on 2D auxiliary outputs when vessel segmentation masks are available; otherwise, it is set to zero, encouraging slice-wise feature consistency.</li>\n</ul>\n<h3>From voxelwise maps to series-level logits (for AUC &amp; ensembling)</h3>\n<p>Given heatmap logits <code>H ∈ ℝ^{B×13×D×H′×W′}</code>:</p>\n<ul>\n<li><p><strong>Per-class series logits:</strong> spatial <strong>max</strong> over <code>(D, H′, W′)</code> for each of the 13 classes.</p></li>\n<li><p><strong>Aneurysm Present (global) logit:</strong> the <strong>max over the 13</strong> class logits.</p>\n<p>These <strong>14 logits</strong> drive the BCE loss and are exported for meta-ensembling.</p></li>\n</ul>\n<h3>Data pipeline &amp; augmentations</h3>\n<ul>\n<li><strong>Input.</strong> Precomputed <code>.npy</code> volumes; <code>(D,H,W)</code> expanded to <code>(1,D,H,W)</code>— here D = 64.</li>\n<li><strong>Label scaling.</strong> Rescale coordinates to current volume size; apply stride mapping for target grids.</li>\n<li><strong>Augmentations.</strong> A <strong>shared 2D affine warp</strong> per volume (same transform for all slices): random rotation, scale, translation; vertical flip available (off by default).</li>\n<li><strong>Normalization.</strong> Albumentations normalization → tensor conversion; channel repeat to 3-ch if needed.</li>\n</ul>\n<h3>Optimization &amp; schedules</h3>\n<ul>\n<li><strong>Optimizer.</strong> AdamW (LR=2e-4, weight decay=1e-5), <strong>AMP</strong> mixed precision, <strong>grad clipping</strong> (=1.0).</li>\n<li><strong>LR schedule.</strong> Default <strong>CosineAnnealingLR</strong> (<code>T_max = epochs</code>, <code>eta_min = 1e-6</code>); <code>StepLR</code> / <code>ReduceLROnPlateau</code> also supported.</li>\n<li><strong>Checkpointing.</strong> Best model selected by <strong>mean AUC</strong> on validation.</li>\n</ul>\n<h3>Training configuration</h3>\n<ul>\n<li><strong>Epochs:</strong> 16</li>\n<li><strong>Batch size:</strong> 4 (train) / 2 (val)</li>\n<li><strong>accumulate:</strong>2</li>\n<li><strong>Workers:</strong> 16</li>\n<li><strong>Strides:</strong> depth <code>1</code>, lateral <code>16</code></li>\n<li><strong>Gaussian σ:</strong> 0.2</li>\n<li><strong>Backbone:</strong> EfficientNetV2-S, <code>pretrained=True</code>, not frozen by default</li>\n</ul>\n<h3>Cross-validation &amp; metrics</h3>\n<ul>\n<li>K-fold protocol driven by the training CSV split.</li>\n<li>Compute <strong>per-label ROC-AUC</strong>, plus <strong>mean</strong> and <strong>presence-weighted</strong> AUC each epoch.</li>\n<li>Scheduler stepping and best-by-mean-AUC saving per fold; final summary reports per-fold and average AUC.</li>\n</ul>\n<h3><strong>Flayer Results(without aux head)</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7762</td>\n<td>0.7747</td>\n<td>0.7754</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.7676</td>\n<td>0.8069</td>\n<td>0.7873</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.7508</td>\n<td>0.7892</td>\n<td>0.7700</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.7553</td>\n<td>0.7778</td>\n<td>0.7666</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.7734</td>\n<td>0.7758</td>\n<td>0.7746</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.7647</strong></td>\n<td><strong>0.7848</strong></td>\n<td><strong>0.7748</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>Flayer Results(with aux head)</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7779</td>\n<td>0.7991</td>\n<td>0.7885</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.7828</td>\n<td>0.7987</td>\n<td>0.7908</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.7601</td>\n<td>0.8051</td>\n<td>0.7826</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.7683</td>\n<td>0.7826</td>\n<td>0.7755</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.7586</td>\n<td>0.7954</td>\n<td>0.7770</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.7695</strong></td>\n<td><strong>0.7962</strong></td>\n<td><strong>0.7829</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Summary</h3>\n<p><strong>Flayer</strong> converts 2D EfficientNetV2-S slice features into 3D evidence maps with a light 3D head, supervises them using CenterNet-style heatmaps + offsets, and produces robust <strong>14-logit</strong> series-level predictions (including an aggregated <em>aneurysm present</em> logit) that plug cleanly into our stacked meta-classifier.</p>\n<h2>Meta Classifier</h2>\n<p>We designed a stacked ensemble architecture to integrate predictions from <strong>YOLO11m</strong>, <strong>YOLO11-EffV2s</strong>, and <strong>EffV2s-3D-CenterNet</strong> models. Specifically, each meta-classifier receives the concatenated predictions from all base models as input and produces final predictions for both <strong>aneurysm presence</strong> and <strong>location-specific detection</strong>. This setup enables the model to assess the presence of an aneurysm at a given location by considering predictions from all locations across all models, effectively capturing the ensemble’s collective reasoning about aneurysm presence.</p>\n<pre><code>\nmetadata = np.array([age, sex])\nX = np.concatenate([np.array([yolo11m_cls_preds[fold_id]]), yolo11m_loc_preds[fold_id],\n                                np.array([effv2s_cls_preds[fold_id]]), effv2s_loc_preds[fold_id],\n                                flayer_fold_preds[fold_id], metadata], axis=)[, :]\nlgb_pred = predict_prob_lgb(X, fold_id) \nxgb_pred = predict_prob_xgb(X, fold_id)\ncat_pred = predict_prob_cat(X, fold_id)\n</code></pre>\n<p>We experimented with two ensemble structures:</p>\n<ul>\n<li><strong>Series:</strong> A single YOLO prediction combined with one Flayer prediction is fed into the ensemble model to produce the final output. </li>\n</ul>\n<pre><code>( features, [() + flayer() + meta()])\n</code></pre>\n<ul>\n<li><strong>Parallel:</strong> Two predictions from two YOLOs and one predictions from Flayer are processed in parallel and jointly used by the ensemble model to generate the final prediction. </li>\n</ul>\n<pre><code>( features, [yolo11m() + ()+ () + ()])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2Fef14572c1840aec4e9d0ce115fe7adb3%2Fmeta.png?generation=1760538976614470&amp;alt=media\" alt=\"diagram-meta-models\"><br>\nThis comparison allows us to evaluate whether combining multiple feature streams in parallel provides additional complementary information beyond the sequential (series) configuration.</p>\n<p>We found that combining predictions from different YOLO backbones and flayer in parallel provides better improvement than series in the ensemble CV score. This suggests that the GBDT effectively leverages each model’s perception of aneurysm presence across locations to make accurate judgments. We were unable to complete the 5-fold submission due to issues encountered on the final day. Therefore, we present here our final selected result, which is based on the average of 2 folds.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Series</th>\n<th>Parallel</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5 folds CV (average)</td>\n<td>0.84</td>\n<td><strong>0.852</strong></td>\n</tr>\n<tr>\n<td>5 folds CV (nelder-mead optimized)</td>\n<td>0.841</td>\n<td><strong>0.858</strong></td>\n</tr>\n<tr>\n<td>public score (average; 2 folds)</td>\n<td>0.83</td>\n<td>0.83082</td>\n</tr>\n<tr>\n<td>private score (average; 2 folds)</td>\n<td>0.81</td>\n<td><strong>0.8230</strong></td>\n</tr>\n</tbody>\n</table>\n<p>We’ve also tried Keypoint-Based Dual-Graph Predictor in <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/bias-vibes-trade-off-20th-place-solution-keypoint-\" target=\"_blank\">Tom’s BYU approach</a> earlier and found it performed well in terms of ensemble CV performance. However, the YOLO-based feature extraction and preprocessing were too time-consuming and often led to timeouts. Therefore, we did not include this method in our final solution.</p>\n<h2>Codebase &amp; Submission</h2>\n<p>Submission: </p>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-classifier-final-ver?scriptVersionId=267890122\" target=\"_blank\"><strong>2x yolo + flayer + meta classifier (final ver)</strong></a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/tom99763/9th-place-solution-rsna-iad\" target=\"_blank\">kaggle models submission</a></p></li>\n</ul>\n<p>Meta classifier training: <a href=\"https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-training-final-ver\" target=\"_blank\"><strong>2x yolo + flayer meta training (final ver)</strong></a></p>\n<p>Github: <a href=\"https://github.com/tom99763/9th-place-solution-yolo-RSNA-IAD\" target=\"_blank\">9th-place-solution-yolo-RSNA-IAD</a></p>",
  "messages": [
    {
      "id": 3302294,
      "postDate": "2025-10-15T13:34:16.853Z",
      "content": "<h2>Acknowledgement</h2>\n<p>First, we thank the competition hosts <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, <a href=\"https://www.kaggle.com/shosys\" target=\"_blank\">@shosys</a> and every person involved from RSNA in data preparation and competition related process, and Kaggle staff for organizing this competition. </p>\n<p>Below, we introduce the solution of team <strong>Vibes and Genius Trade-Off</strong> -- <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>, <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>, <a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a>, <a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a>!</p>\n<p><strong>TL;DR</strong></p>\n<p>Our solution combines three complementary approaches:</p>\n<ul>\n<li>YOLO 2.5D with different backbones. Derived from <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU competition</a></li>\n<li>3D CenterNet with 2D Effv2s extractor.</li>\n<li>Meta-classifiers (LightGBM, XGBoost, CatBoost)</li>\n</ul>\n<p>The final probabilities are obtained by averaging the outputs of the YOLO 2.5D models, 3D CenterNet with 2D Effv2s extractor and the three meta-classifiers.</p>\n<p>Here is the diagram of our approach.</p>\n<p>We used two different variations of YOLO:</p>\n<ol>\n<li><strong>YOLOv11m</strong> - The standard YOLO11 medium model from Ultralytcs</li>\n<li><strong>Custom YOLO with timm backbone</strong> - YOLO architecture with <code>timm/tf_efficientnetv2_s.in21k_ft_in1k</code> as the backbone</li>\n</ol>\n<p>For details about the YOLO+timm customization, please refer to the <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU writeup</a>.</p>\n<p>We treated each of the <strong>13 vessel locations as separate bounding box classes</strong>:</p>\n<ul>\n<li>Left/Right Infraclinoid Internal Carotid Artery</li>\n<li>Left/Right Supraclinoid Internal Carotid Artery</li>\n<li>Left/Right Middle Cerebral Artery</li>\n<li>Anterior Communicating Artery</li>\n<li>Left/Right Anterior Cerebral Artery</li>\n<li>Left/Right Posterior Communicating Artery</li>\n<li>Basilar Tip</li>\n<li>Other Posterior Circulation</li>\n</ul>\n<h3><strong>2.5D Strategy</strong></h3>\n<p>Images were resized to 512×512, normalized with min-max and converted to 2.5D variation:</p>\n<p><code>R (Red channel)   = slice i-1\nG (Green channel) = slice i\nB (Blue channel)  = slice i+1</code></p>\n<p>The figure below shows examples of slices in 2D vs 2.5D variation:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fbc5e5f0a3d330bfc78d7961160b805f6%2F2D.png?generation=1760534875382926&amp;alt=media\" alt=\"\"></p>\n<p>As can be seen in the examples, the images vary significantly because <strong>we didn't standardize Z spacing</strong>.</p>\n<p>I (@sersasj) spent 1-2 weeks experimenting with Z-axis resize, but results were consistently worse. Perhaps I was doing something wrong, but didn't have time to investigate further.</p>\n<p>We were very reluctant to try 2.5D without proper Z-spacing resize (in my mind didn’t seem a good idea - <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>). Nevertheless, the 2.5D approach even without Z-resampling improved results by 0.02+ in CV.</p>\n<h3><strong>Training Configuration</strong></h3>\n<p>Data Sampling Strategy:</p>\n<p>For negative samples we took 10 evenly sampled slices per series.</p>\n<p>Example: For a 100-slice series → slices [1, 11, 22, 33, 44, 56, 67, 78, 89, 99]</p>\n<p>For positive samples we used all slices containing annotations</p>\n<h3><strong>YOLO with timm/tf_efficientnetv2_s Backbone</strong></h3>\n<pre><code> \n \n \n \n  \n \n \n \n   \n</code></pre>\n<h3><strong>YOLOv11m</strong></h3>\n<pre><code> \n \n \n \n \n \n \n \n   \n</code></pre>\n<p>We also modified the fitness function to include auc metric and prioritize mAP@50.</p>\n<pre><code>fitness =  × mAP@ +  × mAP@- +  × mAUC\n</code></pre>\n<p>Did this help? Maybe a little. In the <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly\" target=\"_blank\">BYU competition</a>, we had already found that prioritizing mAP@50 over the Ultralytics standard 1.0×mAP@50-95 gave better results.</p>\n<p>Adding AUC seemed like a good idea at the time since it's the competition metric, but we didn't investigate thoroughly whether it significantly improved performance. It's possible the benefit was marginal, but we kept it for consistency.</p>\n<p><strong>Training Hardware and Time</strong></p>\n<table>\n<thead>\n<tr>\n<th>Hardware Setup</th>\n<th>GPU</th>\n<th>CPU</th>\n<th>RAM</th>\n<th>Time per Fold</th>\n<th>Total (5 folds)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a></td>\n<td>RTX 3090</td>\n<td>Intel Core i5-12400F (12) @ 4.4GHz</td>\n<td>32GB</td>\n<td>4-5 hours</td>\n<td>~24 hours</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a></td>\n<td>RTX 4090</td>\n<td>AMD Ryzen 9 7950X (32) @ 5.883GHz</td>\n<td>64GB</td>\n<td>~2 hours</td>\n<td>~10 hours</td>\n</tr>\n</tbody>\n</table>\n<h3>Inference</h3>\n<p>For inference, we sort DICOM slices by spatial position (SliceLocation → ImagePositionPatient → InstanceNumber), create 2.5D triplets for slices 1 to N-1<br>\nRun batch inference across both models, extract max confidence per class and use max(localization_confidences) as overall aneurysm presence probability<br>\nThe process can be observed in the gif below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F7a4a4d9d460b5a6b0cb6eb82a3c4b5ad%2Fyolo_predictions_1.2.826.0.1.3680043.8.498.10022688097731894079510930966432818105%20(2).gif?generation=1761065820372307&amp;alt=media\" alt=\"\"></p>\n<h3>Cross-Validation Scores</h3>\n<p>folds were stratified using <code>MultilabelStratifiedKFold</code> to ensure balanced distribution across:</p>\n<ul>\n<li>Aneurysm presence</li>\n<li>All 13 vessel locations</li>\n<li>Modality</li>\n</ul>\n<h3><strong>YOLO11m 2.5D Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.8184</td>\n<td>0.7549</td>\n<td>0.7867</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8114</td>\n<td>0.7897</td>\n<td>0.8005</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8134</td>\n<td>0.7827</td>\n<td>0.7981</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8162</td>\n<td>0.7872</td>\n<td>0.8017</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8540</td>\n<td>0.8281</td>\n<td>0.8410</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8227</strong></td>\n<td><strong>0.7885</strong></td>\n<td><strong>0.8056</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>EfficientNetV2-S 2.5D Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7978</td>\n<td>0.7790</td>\n<td>0.7884</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8159</td>\n<td>0.8269</td>\n<td>0.8214</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8155</td>\n<td>0.7921</td>\n<td>0.8038</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8153</td>\n<td>0.7907</td>\n<td>0.8030</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8499</td>\n<td>0.8560</td>\n<td>0.8529</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8189</strong></td>\n<td><strong>0.8090</strong></td>\n<td><strong>0.8139</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>Ensemble Results</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.8393</td>\n<td>0.7849</td>\n<td>0.8121</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.8330</td>\n<td>0.8291</td>\n<td>0.8310</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.8399</td>\n<td>0.8135</td>\n<td>0.8267</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.8370</td>\n<td>0.8093</td>\n<td>0.8232</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.8738</td>\n<td>0.8620</td>\n<td>0.8679</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.8446</strong></td>\n<td><strong>0.8198</strong></td>\n<td><strong>0.8322</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>What did not work:</h2>\n<p>In the initial stages of YOLO development <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> created an ensemble with 3xYolo11m and 0.69LB <a href=\"https://www.kaggle.com/code/yosukeyama/rsna2025-32ch-img-infer-lb-0-69-share\" target=\"_blank\">EfficientNetB2 public notebook</a>. That gave us a score of 0.78 LB and put us in the top 3 in early stages of competition</p>\n<p>After a bit of probing the leaderboard, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> found out that YOLO's aneurysm classification AUROC was very low (around ~0.58). Then we started to develop models that would complement YOLO and boost the aneurysm classification AUROC. (We later discovered that the poor results of early YOLO development were due to an insufficient amount of negative samples in training.)</p>\n<p>One of the class of models that we used was 2.5D models that just detects the presence of aneurysm in the given series with auxiliary segmentation head. These models were inspired by <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a>. These models in combination with yolo11m gave us a LB of 0.80. </p>\n<p>These models then became obsolete when we merged with  <a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a>, <a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a>.</p>\n<p>In the BYU competition, our best-performing backbone was <code>timm/convnextv2_base.fcmae_ft_in22k_in1k</code>. However, in the current competition, training with half precision caused the loss to quickly become NaN. Based on our investigation, this appears to be a common issue with ConvNeXt architectures. We hypothesize that if trained in full precision, this backbone could achieve performance comparable to YOLO11m and <code>timm/tf_efficientnetv2_s.in21k_ft_in1k</code>. Unfortunately, we were unable to verify this due to time constraints.</p>\n<p>Modifying the neck and head was also considered. From what we investigate with <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">awesome reverse engineer notebook</a> P3 features appeared to be the most utilized. We conducted a quick test on a single fold, but results on validation were more or less the same so we decided to move on.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7a4ca157151f82f0436dfdee766ec665%2Fimage%20(2).png?generation=1760534911466456&amp;alt=media\" alt=\"\"></p>\n<p>We also try to use our 3d segmentation result to <a href=\"https://www.kaggle.com/code/tom99763/rsna-iad-vessel-segmentation\" target=\"_blank\">filter out invalid location prediction from yolo</a>, but recall drops a lot.<br>\nMoreover, we tried extracting patches using YOLO and applying wavelet and log-polar transforms to the patches before predicting their classes; however, <a href=\"https://www.kaggle.com/code/tom99763/yolos-phw-mip-lp-rsna-i\" target=\"_blank\">this approach </a> did not yield good results.</p>\n<h2>EfficientV2s + 3D-CenterNet (Flayer)</h2>\n<p>This is a highly <strong>experimental project</strong>, made with a lot of help from AI. We started with the popular “LB 0.69 Public Notebook” , reconstructed the entire <strong>training pipeline and data preprocessing</strong> from it, and then incorporated a <strong>CenterNet-like mechanism to detect the center of the aneurysms</strong>.</p>\n<h3>Model architecture</h3>\n<p>We adopt a <strong>slice-wise 2D encoder + shallow 3D head</strong> design:</p>\n<ul>\n<li><strong>Encoder.</strong> A 2D ImageNet-pretrained <strong>EfficientNetV2-S</strong> (<code>tf_efficientnetv2_s.in21k_ft_in1k</code>) from <code>timm</code> in <em>features_only</em> mode with the <strong>penultimate</strong> feature map (<code>out_indices = (1,-2)</code>).</li>\n<li><strong>Auxiliary 2D head.</strong>The intermediate feature map (early stage) is fed into a lightweight 2D convolutional head to produce per-slice auxiliary logits, which encourage the encoder to capture vessel-relevant spatial cues before 3D aggregation. The auxiliary head is supervised with slice-level targets to stabilize early-layer learning.</li>\n<li><strong>2D→3D fusion.</strong> Each volume <code>(B, C=1, D, H, W)</code> is <strong>rearranged into <code>B·D</code> 2D slices</strong>, encoded independently, then <strong>reassembled</strong> to <code>(B, C_feat, D, H′, W′)</code> before the 3D head. If the backbone expects 3 channels, single-channel inputs are channel-repeated.</li>\n<li><strong>3D temporal head.</strong> A lightweight head (Conv3d-BN-ReLU ×2, <code>base_channels</code>) aggregates signals across depth.</li>\n<li><strong>Outputs.</strong> Two 1×1×1 conv heads:<ul>\n<li><strong>Heatmap head:</strong> <code>num_classes = 13</code> vessel locations (3D centerness maps)</li>\n<li><strong>Offset head:</strong> 3-channel sub-voxel offsets <code>(dz, dy, dx)</code></li></ul></li>\n</ul>\n<h3>Supervision targets</h3>\n<ul>\n<li>Place a <strong>3D Gaussian</strong> peak (stride-aware) at each annotated center on the class-specific heatmap.</li>\n<li>Store <strong>sub-voxel residual offsets</strong> at the nearest grid cell and mark them with an <strong>offset mask</strong>.</li>\n<li>Scale coordinates from the original series size to the current <code>(D, H, W)</code>; map to the output grid using <code>(stride_d=1, stride_h=stride_w=16)</code> so the output is <code>D × H/16 × W/16</code>.</li>\n<li>Default <strong>Gaussian σ = 0.2</strong>.</li>\n</ul>\n<h3>Segmentation masks</h3>\n<p>Vessel masks are produced in physical space by training a 3D DynUNet on volumes resampled to 0.7 mm isotropic spacing. The provided voxel-level annotations of 13 vessel classes are merged into a single binary vessel mask for training. The resulting 3D masks are then converted to slice-wise 2D masks to supervise the auxiliary head.</p>\n<h3>Losses</h3>\n<p>We train with four terms and sum them with weights:</p>\n<ul>\n<li><strong>CenterNet-style focal loss</strong> (weight = 1.0) on heatmaps (α=2, β=4).</li>\n<li><strong>L1 offset loss</strong> (weight = 1.0) (only at valid center cells).</li>\n<li><strong>BCE-with-logits classification loss</strong> (weight = 1.0) from series-level logits .</li>\n<li><strong>Auxiliary Dice &amp; BCE loss</strong> (weight = 0.5) on 2D auxiliary outputs when vessel segmentation masks are available; otherwise, it is set to zero, encouraging slice-wise feature consistency.</li>\n</ul>\n<h3>From voxelwise maps to series-level logits (for AUC &amp; ensembling)</h3>\n<p>Given heatmap logits <code>H ∈ ℝ^{B×13×D×H′×W′}</code>:</p>\n<ul>\n<li><p><strong>Per-class series logits:</strong> spatial <strong>max</strong> over <code>(D, H′, W′)</code> for each of the 13 classes.</p></li>\n<li><p><strong>Aneurysm Present (global) logit:</strong> the <strong>max over the 13</strong> class logits.</p>\n<p>These <strong>14 logits</strong> drive the BCE loss and are exported for meta-ensembling.</p></li>\n</ul>\n<h3>Data pipeline &amp; augmentations</h3>\n<ul>\n<li><strong>Input.</strong> Precomputed <code>.npy</code> volumes; <code>(D,H,W)</code> expanded to <code>(1,D,H,W)</code>— here D = 64.</li>\n<li><strong>Label scaling.</strong> Rescale coordinates to current volume size; apply stride mapping for target grids.</li>\n<li><strong>Augmentations.</strong> A <strong>shared 2D affine warp</strong> per volume (same transform for all slices): random rotation, scale, translation; vertical flip available (off by default).</li>\n<li><strong>Normalization.</strong> Albumentations normalization → tensor conversion; channel repeat to 3-ch if needed.</li>\n</ul>\n<h3>Optimization &amp; schedules</h3>\n<ul>\n<li><strong>Optimizer.</strong> AdamW (LR=2e-4, weight decay=1e-5), <strong>AMP</strong> mixed precision, <strong>grad clipping</strong> (=1.0).</li>\n<li><strong>LR schedule.</strong> Default <strong>CosineAnnealingLR</strong> (<code>T_max = epochs</code>, <code>eta_min = 1e-6</code>); <code>StepLR</code> / <code>ReduceLROnPlateau</code> also supported.</li>\n<li><strong>Checkpointing.</strong> Best model selected by <strong>mean AUC</strong> on validation.</li>\n</ul>\n<h3>Training configuration</h3>\n<ul>\n<li><strong>Epochs:</strong> 16</li>\n<li><strong>Batch size:</strong> 4 (train) / 2 (val)</li>\n<li><strong>accumulate:</strong>2</li>\n<li><strong>Workers:</strong> 16</li>\n<li><strong>Strides:</strong> depth <code>1</code>, lateral <code>16</code></li>\n<li><strong>Gaussian σ:</strong> 0.2</li>\n<li><strong>Backbone:</strong> EfficientNetV2-S, <code>pretrained=True</code>, not frozen by default</li>\n</ul>\n<h3>Cross-validation &amp; metrics</h3>\n<ul>\n<li>K-fold protocol driven by the training CSV split.</li>\n<li>Compute <strong>per-label ROC-AUC</strong>, plus <strong>mean</strong> and <strong>presence-weighted</strong> AUC each epoch.</li>\n<li>Scheduler stepping and best-by-mean-AUC saving per fold; final summary reports per-fold and average AUC.</li>\n</ul>\n<h3><strong>Flayer Results(without aux head)</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7762</td>\n<td>0.7747</td>\n<td>0.7754</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.7676</td>\n<td>0.8069</td>\n<td>0.7873</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.7508</td>\n<td>0.7892</td>\n<td>0.7700</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.7553</td>\n<td>0.7778</td>\n<td>0.7666</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.7734</td>\n<td>0.7758</td>\n<td>0.7746</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.7647</strong></td>\n<td><strong>0.7848</strong></td>\n<td><strong>0.7748</strong></td>\n</tr>\n</tbody>\n</table>\n<h3><strong>Flayer Results(with aux head)</strong></h3>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>loc_macro_auc</th>\n<th>cls_auc</th>\n<th>combined_mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold0</td>\n<td>0.7779</td>\n<td>0.7991</td>\n<td>0.7885</td>\n</tr>\n<tr>\n<td>fold1</td>\n<td>0.7828</td>\n<td>0.7987</td>\n<td>0.7908</td>\n</tr>\n<tr>\n<td>fold2</td>\n<td>0.7601</td>\n<td>0.8051</td>\n<td>0.7826</td>\n</tr>\n<tr>\n<td>fold3</td>\n<td>0.7683</td>\n<td>0.7826</td>\n<td>0.7755</td>\n</tr>\n<tr>\n<td>fold4</td>\n<td>0.7586</td>\n<td>0.7954</td>\n<td>0.7770</td>\n</tr>\n<tr>\n<td><strong>Average</strong></td>\n<td><strong>0.7695</strong></td>\n<td><strong>0.7962</strong></td>\n<td><strong>0.7829</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>Summary</h3>\n<p><strong>Flayer</strong> converts 2D EfficientNetV2-S slice features into 3D evidence maps with a light 3D head, supervises them using CenterNet-style heatmaps + offsets, and produces robust <strong>14-logit</strong> series-level predictions (including an aggregated <em>aneurysm present</em> logit) that plug cleanly into our stacked meta-classifier.</p>\n<h2>Meta Classifier</h2>\n<p>We designed a stacked ensemble architecture to integrate predictions from <strong>YOLO11m</strong>, <strong>YOLO11-EffV2s</strong>, and <strong>EffV2s-3D-CenterNet</strong> models. Specifically, each meta-classifier receives the concatenated predictions from all base models as input and produces final predictions for both <strong>aneurysm presence</strong> and <strong>location-specific detection</strong>. This setup enables the model to assess the presence of an aneurysm at a given location by considering predictions from all locations across all models, effectively capturing the ensemble’s collective reasoning about aneurysm presence.</p>\n<pre><code>\nmetadata = np.array([age, sex])\nX = np.concatenate([np.array([yolo11m_cls_preds[fold_id]]), yolo11m_loc_preds[fold_id],\n                                np.array([effv2s_cls_preds[fold_id]]), effv2s_loc_preds[fold_id],\n                                flayer_fold_preds[fold_id], metadata], axis=)[, :]\nlgb_pred = predict_prob_lgb(X, fold_id) \nxgb_pred = predict_prob_xgb(X, fold_id)\ncat_pred = predict_prob_cat(X, fold_id)\n</code></pre>\n<p>We experimented with two ensemble structures:</p>\n<ul>\n<li><strong>Series:</strong> A single YOLO prediction combined with one Flayer prediction is fed into the ensemble model to produce the final output. </li>\n</ul>\n<pre><code>( features, [() + flayer() + meta()])\n</code></pre>\n<ul>\n<li><strong>Parallel:</strong> Two predictions from two YOLOs and one predictions from Flayer are processed in parallel and jointly used by the ensemble model to generate the final prediction. </li>\n</ul>\n<pre><code>( features, [yolo11m() + ()+ () + ()])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2Fef14572c1840aec4e9d0ce115fe7adb3%2Fmeta.png?generation=1760538976614470&amp;alt=media\" alt=\"diagram-meta-models\"><br>\nThis comparison allows us to evaluate whether combining multiple feature streams in parallel provides additional complementary information beyond the sequential (series) configuration.</p>\n<p>We found that combining predictions from different YOLO backbones and flayer in parallel provides better improvement than series in the ensemble CV score. This suggests that the GBDT effectively leverages each model’s perception of aneurysm presence across locations to make accurate judgments. We were unable to complete the 5-fold submission due to issues encountered on the final day. Therefore, we present here our final selected result, which is based on the average of 2 folds.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Series</th>\n<th>Parallel</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5 folds CV (average)</td>\n<td>0.84</td>\n<td><strong>0.852</strong></td>\n</tr>\n<tr>\n<td>5 folds CV (nelder-mead optimized)</td>\n<td>0.841</td>\n<td><strong>0.858</strong></td>\n</tr>\n<tr>\n<td>public score (average; 2 folds)</td>\n<td>0.83</td>\n<td>0.83082</td>\n</tr>\n<tr>\n<td>private score (average; 2 folds)</td>\n<td>0.81</td>\n<td><strong>0.8230</strong></td>\n</tr>\n</tbody>\n</table>\n<p>We’ve also tried Keypoint-Based Dual-Graph Predictor in <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/bias-vibes-trade-off-20th-place-solution-keypoint-\" target=\"_blank\">Tom’s BYU approach</a> earlier and found it performed well in terms of ensemble CV performance. However, the YOLO-based feature extraction and preprocessing were too time-consuming and often led to timeouts. Therefore, we did not include this method in our final solution.</p>\n<h2>Codebase &amp; Submission</h2>\n<p>Submission: </p>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-classifier-final-ver?scriptVersionId=267890122\" target=\"_blank\"><strong>2x yolo + flayer + meta classifier (final ver)</strong></a></p></li>\n<li><p><a href=\"https://www.kaggle.com/code/tom99763/9th-place-solution-rsna-iad\" target=\"_blank\">kaggle models submission</a></p></li>\n</ul>\n<p>Meta classifier training: <a href=\"https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-training-final-ver\" target=\"_blank\"><strong>2x yolo + flayer meta training (final ver)</strong></a></p>\n<p>Github: <a href=\"https://github.com/tom99763/9th-place-solution-yolo-RSNA-IAD\" target=\"_blank\">9th-place-solution-yolo-RSNA-IAD</a></p>",
      "rawMarkdown": "## Acknowledgement\n\nFirst, we thank the competition hosts @evancalabrese, @shosys and every person involved from RSNA in data preparation and competition related process, and Kaggle staff for organizing this competition. \n\nBelow, we introduce the solution of team **Vibes and Genius Trade-Off** -- @tom99763, @sersasj, @iamparadox, @chihantsai, @atom1231!\n\n**TL;DR**\n\nOur solution combines three complementary approaches:\n\n- YOLO 2.5D with different backbones. Derived from [BYU competition](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly)\n- 3D CenterNet with 2D Effv2s extractor.\n- Meta-classifiers (LightGBM, XGBoost, CatBoost)\n\nThe final probabilities are obtained by averaging the outputs of the YOLO 2.5D models, 3D CenterNet with 2D Effv2s extractor and the three meta-classifiers.\n\nHere is the diagram of our approach.\n\n\nWe used two different variations of YOLO:\n\n1. **YOLOv11m** - The standard YOLO11 medium model from Ultralytcs\n2. **Custom YOLO with timm backbone** - YOLO architecture with `timm/tf_efficientnetv2_s.in21k_ft_in1k` as the backbone\n\nFor details about the YOLO+timm customization, please refer to the [BYU writeup](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly).\n\nWe treated each of the **13 vessel locations as separate bounding box classes**:\n\n- Left/Right Infraclinoid Internal Carotid Artery\n- Left/Right Supraclinoid Internal Carotid Artery\n- Left/Right Middle Cerebral Artery\n- Anterior Communicating Artery\n- Left/Right Anterior Cerebral Artery\n- Left/Right Posterior Communicating Artery\n- Basilar Tip\n- Other Posterior Circulation\n\n### **2.5D Strategy**\n\nImages were resized to 512×512, normalized with min-max and converted to 2.5D variation:\n\n`R (Red channel)   = slice i-1\nG (Green channel) = slice i\nB (Blue channel)  = slice i+1`\n\nThe figure below shows examples of slices in 2D vs 2.5D variation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fbc5e5f0a3d330bfc78d7961160b805f6%2F2D.png?generation=1760534875382926&alt=media)\n\nAs can be seen in the examples, the images vary significantly because **we didn't standardize Z spacing**.\n\nI (@sersasj) spent 1-2 weeks experimenting with Z-axis resize, but results were consistently worse. Perhaps I was doing something wrong, but didn't have time to investigate further.\n\nWe were very reluctant to try 2.5D without proper Z-spacing resize (in my mind didn’t seem a good idea - @sersasj). Nevertheless, the 2.5D approach even without Z-resampling improved results by 0.02+ in CV.\n\n### **Training Configuration**\n\nData Sampling Strategy:\n\nFor negative samples we took 10 evenly sampled slices per series.\n\nExample: For a 100-slice series → slices [1, 11, 22, 33, 44, 56, 67, 78, 89, 99]\n\nFor positive samples we used all slices containing annotations\n\n### **YOLO with timm/tf_efficientnetv2_s Backbone**\n\n```yaml\nbatch_size: 16\nepochs: 50\nmixup: 0.4\nmosaic: 0.4\ndrop_path_rate: 0.2 \ncls_loss: 1.0\noptimizer: AdamW\nmomentum: 0.9\nlearning_rate: auto  \n```\n\n### **YOLOv11m**\n\n```yaml\nbatch_size: 32\nepochs: 80\nmixup: 0.4\nmosaic: 0.4\ndroput: 0.3\ncls_loss: 1.0\noptimizer: AdamW\nmomentum: 0.9\nlearning_rate: auto  \n```\n\nWe also modified the fitness function to include auc metric and prioritize mAP@50.\n\n```python\nfitness = 0.5 × mAP@50 + 0.25 × mAP@50-95 + 0.25 × mAUC\n\n```\n\nDid this help? Maybe a little. In the [BYU competition](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly), we had already found that prioritizing mAP@50 over the Ultralytics standard 1.0×mAP@50-95 gave better results.\n\nAdding AUC seemed like a good idea at the time since it's the competition metric, but we didn't investigate thoroughly whether it significantly improved performance. It's possible the benefit was marginal, but we kept it for consistency.\n\n**Training Hardware and Time**\n\n| Hardware Setup | GPU | CPU | RAM | Time per Fold | Total (5 folds) |\n| --- | --- | --- | --- | --- | --- |\n| @sersasj | RTX 3090 | Intel Core i5-12400F (12) @ 4.4GHz | 32GB | 4-5 hours | ~24 hours |\n| @iamparadox | RTX 4090 | AMD Ryzen 9 7950X (32) @ 5.883GHz | 64GB | ~2 hours | ~10 hours |\n\n### Inference\n\nFor inference, we sort DICOM slices by spatial position (SliceLocation → ImagePositionPatient → InstanceNumber), create 2.5D triplets for slices 1 to N-1\nRun batch inference across both models, extract max confidence per class and use max(localization_confidences) as overall aneurysm presence probability\nThe process can be observed in the gif below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F7a4a4d9d460b5a6b0cb6eb82a3c4b5ad%2Fyolo_predictions_1.2.826.0.1.3680043.8.498.10022688097731894079510930966432818105%20(2).gif?generation=1761065820372307&alt=media)\n\n### Cross-Validation Scores\n\nfolds were stratified using `MultilabelStratifiedKFold` to ensure balanced distribution across:\n\n- Aneurysm presence\n- All 13 vessel locations\n- Modality\n\n### **YOLO11m 2.5D Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.8184 | 0.7549 | 0.7867 |\n| fold1 | 0.8114 | 0.7897 | 0.8005 |\n| fold2 | 0.8134 | 0.7827 | 0.7981 |\n| fold3 | 0.8162 | 0.7872 | 0.8017 |\n| fold4 | 0.8540 | 0.8281 | 0.8410 |\n| **Average** | **0.8227** | **0.7885** | **0.8056** |\n\n### **EfficientNetV2-S 2.5D Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7978 | 0.7790 | 0.7884 |\n| fold1 | 0.8159 | 0.8269 | 0.8214 |\n| fold2 | 0.8155 | 0.7921 | 0.8038 |\n| fold3 | 0.8153 | 0.7907 | 0.8030 |\n| fold4 | 0.8499 | 0.8560 | 0.8529 |\n| **Average** | **0.8189** | **0.8090** | **0.8139** |\n\n### **Ensemble Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.8393 | 0.7849 | 0.8121 |\n| fold1 | 0.8330 | 0.8291 | 0.8310 |\n| fold2 | 0.8399 | 0.8135 | 0.8267 |\n| fold3 | 0.8370 | 0.8093 | 0.8232 |\n| fold4 | 0.8738 | 0.8620 | 0.8679 |\n| **Average** | **0.8446** | **0.8198** | **0.8322** |\n\n## What did not work:\n\nIn the initial stages of YOLO development @sersasj created an ensemble with 3xYolo11m and 0.69LB [EfficientNetB2 public notebook](https://www.kaggle.com/code/yosukeyama/rsna2025-32ch-img-infer-lb-0-69-share). That gave us a score of 0.78 LB and put us in the top 3 in early stages of competition\n\nAfter a bit of probing the leaderboard, @iamparadox found out that YOLO's aneurysm classification AUROC was very low (around ~0.58). Then we started to develop models that would complement YOLO and boost the aneurysm classification AUROC. (We later discovered that the poor results of early YOLO development were due to an insufficient amount of negative samples in training.)\n\nOne of the class of models that we used was 2.5D models that just detects the presence of aneurysm in the given series with auxiliary segmentation head. These models were inspired by https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder. These models in combination with yolo11m gave us a LB of 0.80. \n\nThese models then became obsolete when we merged with  @chihantsai, @atom1231.\n\nIn the BYU competition, our best-performing backbone was `timm/convnextv2_base.fcmae_ft_in22k_in1k`. However, in the current competition, training with half precision caused the loss to quickly become NaN. Based on our investigation, this appears to be a common issue with ConvNeXt architectures. We hypothesize that if trained in full precision, this backbone could achieve performance comparable to YOLO11m and `timm/tf_efficientnetv2_s.in21k_ft_in1k`. Unfortunately, we were unable to verify this due to time constraints.\n\nModifying the neck and head was also considered. From what we investigate with @tatamikenn [awesome reverse engineer notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo) P3 features appeared to be the most utilized. We conducted a quick test on a single fold, but results on validation were more or less the same so we decided to move on.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7a4ca157151f82f0436dfdee766ec665%2Fimage%20(2).png?generation=1760534911466456&alt=media)\n\nWe also try to use our 3d segmentation result to [filter out invalid location prediction from yolo](https://www.kaggle.com/code/tom99763/rsna-iad-vessel-segmentation), but recall drops a lot.\nMoreover, we tried extracting patches using YOLO and applying wavelet and log-polar transforms to the patches before predicting their classes; however, [this approach ](https://www.kaggle.com/code/tom99763/yolos-phw-mip-lp-rsna-i) did not yield good results.\n\n## EfficientV2s + 3D-CenterNet (Flayer)\n\nThis is a highly **experimental project**, made with a lot of help from AI. We started with the popular “LB 0.69 Public Notebook” , reconstructed the entire **training pipeline and data preprocessing** from it, and then incorporated a **CenterNet-like mechanism to detect the center of the aneurysms**.\n\n### Model architecture\n\nWe adopt a **slice-wise 2D encoder + shallow 3D head** design:\n\n- **Encoder.** A 2D ImageNet-pretrained **EfficientNetV2-S** (`tf_efficientnetv2_s.in21k_ft_in1k`) from `timm` in *features_only* mode with the **penultimate** feature map (`out_indices = (1,-2)`).\n- **Auxiliary 2D head.**The intermediate feature map (early stage) is fed into a lightweight 2D convolutional head to produce per-slice auxiliary logits, which encourage the encoder to capture vessel-relevant spatial cues before 3D aggregation. The auxiliary head is supervised with slice-level targets to stabilize early-layer learning.\n- **2D→3D fusion.** Each volume `(B, C=1, D, H, W)` is **rearranged into `B·D` 2D slices**, encoded independently, then **reassembled** to `(B, C_feat, D, H′, W′)` before the 3D head. If the backbone expects 3 channels, single-channel inputs are channel-repeated.\n- **3D temporal head.** A lightweight head (Conv3d-BN-ReLU ×2, `base_channels`) aggregates signals across depth.\n- **Outputs.** Two 1×1×1 conv heads:\n    - **Heatmap head:** `num_classes = 13` vessel locations (3D centerness maps)\n    - **Offset head:** 3-channel sub-voxel offsets `(dz, dy, dx)`\n\n### Supervision targets\n\n- Place a **3D Gaussian** peak (stride-aware) at each annotated center on the class-specific heatmap.\n- Store **sub-voxel residual offsets** at the nearest grid cell and mark them with an **offset mask**.\n- Scale coordinates from the original series size to the current `(D, H, W)`; map to the output grid using `(stride_d=1, stride_h=stride_w=16)` so the output is `D × H/16 × W/16`.\n- Default **Gaussian σ = 0.2**.\n\n### Segmentation masks\nVessel masks are produced in physical space by training a 3D DynUNet on volumes resampled to 0.7 mm isotropic spacing. The provided voxel-level annotations of 13 vessel classes are merged into a single binary vessel mask for training. The resulting 3D masks are then converted to slice-wise 2D masks to supervise the auxiliary head.\n\n### Losses\n\nWe train with four terms and sum them with weights:\n\n- **CenterNet-style focal loss** (weight = 1.0) on heatmaps (α=2, β=4).\n- **L1 offset loss** (weight = 1.0) (only at valid center cells).\n- **BCE-with-logits classification loss** (weight = 1.0) from series-level logits .\n- **Auxiliary Dice & BCE loss** (weight = 0.5) on 2D auxiliary outputs when vessel segmentation masks are available; otherwise, it is set to zero, encouraging slice-wise feature consistency.\n\n### From voxelwise maps to series-level logits (for AUC & ensembling)\n\nGiven heatmap logits `H ∈ ℝ^{B×13×D×H′×W′}`:\n\n- **Per-class series logits:** spatial **max** over `(D, H′, W′)` for each of the 13 classes.\n- **Aneurysm Present (global) logit:** the **max over the 13** class logits.\n    \n    These **14 logits** drive the BCE loss and are exported for meta-ensembling.\n    \n\n### Data pipeline & augmentations\n\n- **Input.** Precomputed `.npy` volumes; `(D,H,W)` expanded to `(1,D,H,W)`— here D = 64.\n- **Label scaling.** Rescale coordinates to current volume size; apply stride mapping for target grids.\n- **Augmentations.** A **shared 2D affine warp** per volume (same transform for all slices): random rotation, scale, translation; vertical flip available (off by default).\n- **Normalization.** Albumentations normalization → tensor conversion; channel repeat to 3-ch if needed.\n\n### Optimization & schedules\n\n- **Optimizer.** AdamW (LR=2e-4, weight decay=1e-5), **AMP** mixed precision, **grad clipping** (=1.0).\n- **LR schedule.** Default **CosineAnnealingLR** (`T_max = epochs`, `eta_min = 1e-6`); `StepLR` / `ReduceLROnPlateau` also supported.\n- **Checkpointing.** Best model selected by **mean AUC** on validation.\n\n### Training configuration\n\n- **Epochs:** 16\n- **Batch size:** 4 (train) / 2 (val)\n- **accumulate:**2\n- **Workers:** 16\n- **Strides:** depth `1`, lateral `16`\n- **Gaussian σ:** 0.2\n- **Backbone:** EfficientNetV2-S, `pretrained=True`, not frozen by default\n\n### Cross-validation & metrics\n\n- K-fold protocol driven by the training CSV split.\n- Compute **per-label ROC-AUC**, plus **mean** and **presence-weighted** AUC each epoch.\n- Scheduler stepping and best-by-mean-AUC saving per fold; final summary reports per-fold and average AUC.\n\n### **Flayer Results(without aux head)**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7762 | 0.7747 | 0.7754 |\n| fold1 | 0.7676 | 0.8069 | 0.7873 |\n| fold2 | 0.7508 | 0.7892 | 0.7700 |\n| fold3 | 0.7553 | 0.7778 | 0.7666 |\n| fold4 | 0.7734 | 0.7758 | 0.7746 |\n| **Average** | **0.7647** | **0.7848** | **0.7748** |\n\n### **Flayer Results(with aux head)**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7779 | 0.7991 | 0.7885 |\n| fold1 | 0.7828 | 0.7987 | 0.7908 |\n| fold2 | 0.7601 | 0.8051 | 0.7826 |\n| fold3 | 0.7683 | 0.7826 | 0.7755 |\n| fold4 | 0.7586 | 0.7954 | 0.7770 |\n| **Average** | **0.7695** | **0.7962** | **0.7829** |\n\n\n\n### Summary\n\n**Flayer** converts 2D EfficientNetV2-S slice features into 3D evidence maps with a light 3D head, supervises them using CenterNet-style heatmaps + offsets, and produces robust **14-logit** series-level predictions (including an aggregated *aneurysm present* logit) that plug cleanly into our stacked meta-classifier.\n\n## Meta Classifier\n\nWe designed a stacked ensemble architecture to integrate predictions from **YOLO11m**, **YOLO11-EffV2s**, and **EffV2s-3D-CenterNet** models. Specifically, each meta-classifier receives the concatenated predictions from all base models as input and produces final predictions for both **aneurysm presence** and **location-specific detection**. This setup enables the model to assess the presence of an aneurysm at a given location by considering predictions from all locations across all models, effectively capturing the ensemble’s collective reasoning about aneurysm presence.\n\n```python\n#for each fold_id\nmetadata = np.array([age, sex])\nX = np.concatenate([np.array([yolo11m_cls_preds[fold_id]]), yolo11m_loc_preds[fold_id],\n                                np.array([effv2s_cls_preds[fold_id]]), effv2s_loc_preds[fold_id],\n                                flayer_fold_preds[fold_id], metadata], axis=0)[None, :]\nlgb_pred = predict_prob_lgb(X, fold_id) #(14, )\nxgb_pred = predict_prob_xgb(X, fold_id)\ncat_pred = predict_prob_cat(X, fold_id)\n```\n\nWe experimented with two ensemble structures:\n\n- **Series:** A single YOLO prediction combined with one Flayer prediction is fed into the ensemble model to produce the final output. \n```\n(30 features, [yolo(14) + flayer(14) + meta(2)])\n```\n- **Parallel:** Two predictions from two YOLOs and one predictions from Flayer are processed in parallel and jointly used by the ensemble model to generate the final prediction. \n```\n(44 features, [yolo11m(14) + yolo_effnets(14)+ flayer(14) + meta(2)])\n```\n![diagram-meta-models](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2Fef14572c1840aec4e9d0ce115fe7adb3%2Fmeta.png?generation=1760538976614470&alt=media)\nThis comparison allows us to evaluate whether combining multiple feature streams in parallel provides additional complementary information beyond the sequential (series) configuration.\n\nWe found that combining predictions from different YOLO backbones and flayer in parallel provides better improvement than series in the ensemble CV score. This suggests that the GBDT effectively leverages each model’s perception of aneurysm presence across locations to make accurate judgments. We were unable to complete the 5-fold submission due to issues encountered on the final day. Therefore, we present here our final selected result, which is based on the average of 2 folds.\n\n|  | Series | Parallel |\n| --- | --- | --- |\n| 5 folds CV (average) | 0.84 | **0.852** |\n| 5 folds CV (nelder-mead optimized) | 0.841 | **0.858** |\n| public score (average; 2 folds) | 0.83 | 0.83082 |\n| private score (average; 2 folds) | 0.81 | **0.8230** |\n\nWe’ve also tried Keypoint-Based Dual-Graph Predictor in [Tom’s BYU approach](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/bias-vibes-trade-off-20th-place-solution-keypoint-) earlier and found it performed well in terms of ensemble CV performance. However, the YOLO-based feature extraction and preprocessing were too time-consuming and often led to timeouts. Therefore, we did not include this method in our final solution.\n\n## Codebase & Submission\n\nSubmission: \n* [**2x yolo + flayer + meta classifier (final ver)**](https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-classifier-final-ver?scriptVersionId=267890122)\n\n* [kaggle models submission](https://www.kaggle.com/code/tom99763/9th-place-solution-rsna-iad)\n\nMeta classifier training: [**2x yolo + flayer meta training (final ver)**](https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-training-final-ver)\n\nGithub: [9th-place-solution-yolo-RSNA-IAD](https://github.com/tom99763/9th-place-solution-yolo-RSNA-IAD)",
      "votes": 47
    },
    {
      "id": 3304634,
      "postDate": "2025-10-21T01:24:42.317Z",
      "content": "<p>Flayer training code is updated in <a href=\"https://github.com/tom99763/9th-place-solution-RSNA-IAD\" target=\"_blank\">our github</a>, also publish our <a href=\"https://www.kaggle.com/models/tom99763/9th-place-models-rsna-iad\" target=\"_blank\">kaggle Model</a></p>",
      "rawMarkdown": "Flayer training code is updated in [our github](https://github.com/tom99763/9th-place-solution-RSNA-IAD), also publish our [kaggle Model](https://www.kaggle.com/models/tom99763/9th-place-models-rsna-iad)",
      "votes": 1
    },
    {
      "id": 3302673,
      "postDate": "2025-10-16T09:09:10.200Z",
      "content": "<p>I really like how kind this writing feels — congratulations!</p>",
      "rawMarkdown": "I really like how kind this writing feels — congratulations!"
    },
    {
      "id": 3302586,
      "postDate": "2025-10-16T05:24:13.120Z",
      "content": "<p>Hi, thanks for wonderful writeup. I have below questions.</p>\n<ol>\n<li>Did you guys considred trying out reorientation or resampling of the niftii volume in pre-processing pipeline?</li>\n<li>Any specific reason for usage of GBDT in the pipeline. If I am correct Age + Sex metadata is not there for test set. (Was it just matter of convinence?)</li>\n<li>Any motivation for 2D or 2.5D based solution? (Was it because of mixture of thick and thin slices).</li>\n<li>Did you guys had any attempt with pure E2E 3D based solution?</li>\n</ol>",
      "rawMarkdown": "Hi, thanks for wonderful writeup. I have below questions.\n\n1. Did you guys considred trying out reorientation or resampling of the niftii volume in pre-processing pipeline?\n2. Any specific reason for usage of GBDT in the pipeline. If I am correct Age + Sex metadata is not there for test set. (Was it just matter of convinence?)\n2. Any motivation for 2D or 2.5D based solution? (Was it because of mixture of thick and thin slices).\n3. Did you guys had any attempt with pure E2E 3D based solution?",
      "replies": [
        {
          "id": 3302596,
          "postDate": "2025-10-16T06:06:00.620Z",
          "content": "<blockquote>\n  <p>Did you guys had any attempt with pure E2E 3D based solution?</p>\n</blockquote>\n<p>The biggest issues encountered with a pure 3D solution are insufficient computational power (compute) and excessively long training times. The CenterNet within the solution is a compromised/truncated 3D solution designed to mitigate the constraints of compute and time.</p>",
          "rawMarkdown": "> Did you guys had any attempt with pure E2E 3D based solution?\n\nThe biggest issues encountered with a pure 3D solution are insufficient computational power (compute) and excessively long training times. The CenterNet within the solution is a compromised/truncated 3D solution designed to mitigate the constraints of compute and time.",
          "votes": 1
        },
        {
          "id": 3302638,
          "postDate": "2025-10-16T08:08:12.570Z",
          "content": "<pre><code> specific reason    GBDT  the pipeline.  I am correct Age + Sex metadata   there  test . (Was it just matter  convinence?)\n</code></pre>\n<p>The reason of using meta-classifiers is that we noticed when predicting a class, aggregating predictions from other classes can give a significant CV boost. Therefore, we believe the predicted probabilities of other classes can also serve as useful clues for identifying aneurysms in a specific locations.</p>\n<p>We noticed that the sample test data contain Age and Sex, but after submitting, with or without meta data the didn’t change Leaderboard (LB) score much, while the cross-validation (CV) score showed a noticeable difference.<br>\nTherefore, we decided to keep Age and Sex features, assuming that some data in hidden test set also includes them.</p>\n<pre><code>Any motivation for D .D solution? (Was it of mixture of thick thin slices).\n</code></pre>\n<p>In the first month, we found that the 3D classifier approach tended to overfit easily, as representing each voxel along the channel dimension made the model prone to overfitting. However, depth context is crucial for detecting irregular structural changes that could indicate an aneurysm, and slice-wise predictions not having expected result here. We found that  using 3 local adjancent depths as the channel dimension to represent each pixel outperforms other previous methods. This approach proved effective for YOLO in capturing anomaly locations based on slight depth structure changes and achieved better CV compared to YOLO models that takes processed single slices repeated along RGB channels.</p>",
          "rawMarkdown": "```\nAny specific reason for usage of GBDT in the pipeline. If I am correct Age + Sex metadata is not there for test set. (Was it just matter of convinence?)\n```\nThe reason of using meta-classifiers is that we noticed when predicting a class, aggregating predictions from other classes can give a significant CV boost. Therefore, we believe the predicted probabilities of other classes can also serve as useful clues for identifying aneurysms in a specific locations.\n\nWe noticed that the sample test data contain Age and Sex, but after submitting, with or without meta data the didn’t change Leaderboard (LB) score much, while the cross-validation (CV) score showed a noticeable difference.\nTherefore, we decided to keep Age and Sex features, assuming that some data in hidden test set also includes them.\n\n\n```\nAny motivation for 2D or 2.5D based solution? (Was it because of mixture of thick and thin slices).\n```\nIn the first month, we found that the 3D classifier approach tended to overfit easily, as representing each voxel along the channel dimension made the model prone to overfitting. However, depth context is crucial for detecting irregular structural changes that could indicate an aneurysm, and slice-wise predictions not having expected result here. We found that  using 3 local adjancent depths as the channel dimension to represent each pixel outperforms other previous methods. This approach proved effective for YOLO in capturing anomaly locations based on slight depth structure changes and achieved better CV compared to YOLO models that takes processed single slices repeated along RGB channels.",
          "votes": 1,
          "replies": [
            {
              "id": 3303104,
              "postDate": "2025-10-17T07:13:36.327Z",
              "content": "<p>Thanks for additional information. </p>",
              "rawMarkdown": "Thanks for additional information. "
            }
          ]
        }
      ]
    },
    {
      "id": 3302503,
      "postDate": "2025-10-16T01:10:49.220Z",
      "content": "<p>Congratulations, very well done!</p>",
      "rawMarkdown": "Congratulations, very well done!"
    }
  ],
  "comments": [
    {
      "id": 3304634,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-21T01:24:42.317000",
      "content": "<p>Flayer training code is updated in <a href=\"https://github.com/tom99763/9th-place-solution-RSNA-IAD\" target=\"_blank\">our github</a>, also publish our <a href=\"https://www.kaggle.com/models/tom99763/9th-place-models-rsna-iad\" target=\"_blank\">kaggle Model</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3302673,
      "author_name": "doheon114",
      "author_url": "",
      "post_date": "2025-10-16T09:09:10.200000",
      "content": "<p>I really like how kind this writing feels — congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302586,
      "author_name": "Ujjwal Pandey",
      "author_url": "",
      "post_date": "2025-10-16T05:24:13.120000",
      "content": "<p>Hi, thanks for wonderful writeup. I have below questions.</p>\n<ol>\n<li>Did you guys considred trying out reorientation or resampling of the niftii volume in pre-processing pipeline?</li>\n<li>Any specific reason for usage of GBDT in the pipeline. If I am correct Age + Sex metadata is not there for test set. (Was it just matter of convinence?)</li>\n<li>Any motivation for 2D or 2.5D based solution? (Was it because of mixture of thick and thin slices).</li>\n<li>Did you guys had any attempt with pure E2E 3D based solution?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 3302596,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "2025-10-16T06:06:00.620000",
          "content": "<blockquote>\n  <p>Did you guys had any attempt with pure E2E 3D based solution?</p>\n</blockquote>\n<p>The biggest issues encountered with a pure 3D solution are insufficient computational power (compute) and excessively long training times. The CenterNet within the solution is a compromised/truncated 3D solution designed to mitigate the constraints of compute and time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3302638,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-10-16T08:08:12.570000",
          "content": "<pre><code> specific reason    GBDT  the pipeline.  I am correct Age + Sex metadata   there  test . (Was it just matter  convinence?)\n</code></pre>\n<p>The reason of using meta-classifiers is that we noticed when predicting a class, aggregating predictions from other classes can give a significant CV boost. Therefore, we believe the predicted probabilities of other classes can also serve as useful clues for identifying aneurysms in a specific locations.</p>\n<p>We noticed that the sample test data contain Age and Sex, but after submitting, with or without meta data the didn’t change Leaderboard (LB) score much, while the cross-validation (CV) score showed a noticeable difference.<br>\nTherefore, we decided to keep Age and Sex features, assuming that some data in hidden test set also includes them.</p>\n<pre><code>Any motivation for D .D solution? (Was it of mixture of thick thin slices).\n</code></pre>\n<p>In the first month, we found that the 3D classifier approach tended to overfit easily, as representing each voxel along the channel dimension made the model prone to overfitting. However, depth context is crucial for detecting irregular structural changes that could indicate an aneurysm, and slice-wise predictions not having expected result here. We found that  using 3 local adjancent depths as the channel dimension to represent each pixel outperforms other previous methods. This approach proved effective for YOLO in capturing anomaly locations based on slight depth structure changes and achieved better CV compared to YOLO models that takes processed single slices repeated along RGB channels.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3303104,
              "author_name": "Ujjwal Pandey",
              "author_url": "",
              "post_date": "2025-10-17T07:13:36.327000",
              "content": "<p>Thanks for additional information. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302503,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-10-16T01:10:49.220000",
      "content": "<p>Congratulations, very well done!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3302294": "## Acknowledgement\n\nFirst, we thank the competition hosts @evancalabrese, @shosys and every person involved from RSNA in data preparation and competition related process, and Kaggle staff for organizing this competition. \n\nBelow, we introduce the solution of team **Vibes and Genius Trade-Off** -- @tom99763, @sersasj, @iamparadox, @chihantsai, @atom1231!\n\n**TL;DR**\n\nOur solution combines three complementary approaches:\n\n- YOLO 2.5D with different backbones. Derived from [BYU competition](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly)\n- 3D CenterNet with 2D Effv2s extractor.\n- Meta-classifiers (LightGBM, XGBoost, CatBoost)\n\nThe final probabilities are obtained by averaging the outputs of the YOLO 2.5D models, 3D CenterNet with 2D Effv2s extractor and the three meta-classifiers.\n\nHere is the diagram of our approach.\n\n\nWe used two different variations of YOLO:\n\n1. **YOLOv11m** - The standard YOLO11 medium model from Ultralytcs\n2. **Custom YOLO with timm backbone** - YOLO architecture with `timm/tf_efficientnetv2_s.in21k_ft_in1k` as the backbone\n\nFor details about the YOLO+timm customization, please refer to the [BYU writeup](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly).\n\nWe treated each of the **13 vessel locations as separate bounding box classes**:\n\n- Left/Right Infraclinoid Internal Carotid Artery\n- Left/Right Supraclinoid Internal Carotid Artery\n- Left/Right Middle Cerebral Artery\n- Anterior Communicating Artery\n- Left/Right Anterior Cerebral Artery\n- Left/Right Posterior Communicating Artery\n- Basilar Tip\n- Other Posterior Circulation\n\n### **2.5D Strategy**\n\nImages were resized to 512×512, normalized with min-max and converted to 2.5D variation:\n\n`R (Red channel)   = slice i-1\nG (Green channel) = slice i\nB (Blue channel)  = slice i+1`\n\nThe figure below shows examples of slices in 2D vs 2.5D variation:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2Fbc5e5f0a3d330bfc78d7961160b805f6%2F2D.png?generation=1760534875382926&alt=media)\n\nAs can be seen in the examples, the images vary significantly because **we didn't standardize Z spacing**.\n\nI (@sersasj) spent 1-2 weeks experimenting with Z-axis resize, but results were consistently worse. Perhaps I was doing something wrong, but didn't have time to investigate further.\n\nWe were very reluctant to try 2.5D without proper Z-spacing resize (in my mind didn’t seem a good idea - @sersasj). Nevertheless, the 2.5D approach even without Z-resampling improved results by 0.02+ in CV.\n\n### **Training Configuration**\n\nData Sampling Strategy:\n\nFor negative samples we took 10 evenly sampled slices per series.\n\nExample: For a 100-slice series → slices [1, 11, 22, 33, 44, 56, 67, 78, 89, 99]\n\nFor positive samples we used all slices containing annotations\n\n### **YOLO with timm/tf_efficientnetv2_s Backbone**\n\n```yaml\nbatch_size: 16\nepochs: 50\nmixup: 0.4\nmosaic: 0.4\ndrop_path_rate: 0.2 \ncls_loss: 1.0\noptimizer: AdamW\nmomentum: 0.9\nlearning_rate: auto  \n```\n\n### **YOLOv11m**\n\n```yaml\nbatch_size: 32\nepochs: 80\nmixup: 0.4\nmosaic: 0.4\ndroput: 0.3\ncls_loss: 1.0\noptimizer: AdamW\nmomentum: 0.9\nlearning_rate: auto  \n```\n\nWe also modified the fitness function to include auc metric and prioritize mAP@50.\n\n```python\nfitness = 0.5 × mAP@50 + 0.25 × mAP@50-95 + 0.25 × mAUC\n\n```\n\nDid this help? Maybe a little. In the [BYU competition](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly), we had already found that prioritizing mAP@50 over the Ultralytics standard 1.0×mAP@50-95 gave better results.\n\nAdding AUC seemed like a good idea at the time since it's the competition metric, but we didn't investigate thoroughly whether it significantly improved performance. It's possible the benefit was marginal, but we kept it for consistency.\n\n**Training Hardware and Time**\n\n| Hardware Setup | GPU | CPU | RAM | Time per Fold | Total (5 folds) |\n| --- | --- | --- | --- | --- | --- |\n| @sersasj | RTX 3090 | Intel Core i5-12400F (12) @ 4.4GHz | 32GB | 4-5 hours | ~24 hours |\n| @iamparadox | RTX 4090 | AMD Ryzen 9 7950X (32) @ 5.883GHz | 64GB | ~2 hours | ~10 hours |\n\n### Inference\n\nFor inference, we sort DICOM slices by spatial position (SliceLocation → ImagePositionPatient → InstanceNumber), create 2.5D triplets for slices 1 to N-1\nRun batch inference across both models, extract max confidence per class and use max(localization_confidences) as overall aneurysm presence probability\nThe process can be observed in the gif below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2F7a4a4d9d460b5a6b0cb6eb82a3c4b5ad%2Fyolo_predictions_1.2.826.0.1.3680043.8.498.10022688097731894079510930966432818105%20(2).gif?generation=1761065820372307&alt=media)\n\n### Cross-Validation Scores\n\nfolds were stratified using `MultilabelStratifiedKFold` to ensure balanced distribution across:\n\n- Aneurysm presence\n- All 13 vessel locations\n- Modality\n\n### **YOLO11m 2.5D Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.8184 | 0.7549 | 0.7867 |\n| fold1 | 0.8114 | 0.7897 | 0.8005 |\n| fold2 | 0.8134 | 0.7827 | 0.7981 |\n| fold3 | 0.8162 | 0.7872 | 0.8017 |\n| fold4 | 0.8540 | 0.8281 | 0.8410 |\n| **Average** | **0.8227** | **0.7885** | **0.8056** |\n\n### **EfficientNetV2-S 2.5D Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7978 | 0.7790 | 0.7884 |\n| fold1 | 0.8159 | 0.8269 | 0.8214 |\n| fold2 | 0.8155 | 0.7921 | 0.8038 |\n| fold3 | 0.8153 | 0.7907 | 0.8030 |\n| fold4 | 0.8499 | 0.8560 | 0.8529 |\n| **Average** | **0.8189** | **0.8090** | **0.8139** |\n\n### **Ensemble Results**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.8393 | 0.7849 | 0.8121 |\n| fold1 | 0.8330 | 0.8291 | 0.8310 |\n| fold2 | 0.8399 | 0.8135 | 0.8267 |\n| fold3 | 0.8370 | 0.8093 | 0.8232 |\n| fold4 | 0.8738 | 0.8620 | 0.8679 |\n| **Average** | **0.8446** | **0.8198** | **0.8322** |\n\n## What did not work:\n\nIn the initial stages of YOLO development @sersasj created an ensemble with 3xYolo11m and 0.69LB [EfficientNetB2 public notebook](https://www.kaggle.com/code/yosukeyama/rsna2025-32ch-img-infer-lb-0-69-share). That gave us a score of 0.78 LB and put us in the top 3 in early stages of competition\n\nAfter a bit of probing the leaderboard, @iamparadox found out that YOLO's aneurysm classification AUROC was very low (around ~0.58). Then we started to develop models that would complement YOLO and boost the aneurysm classification AUROC. (We later discovered that the poor results of early YOLO development were due to an insufficient amount of negative samples in training.)\n\nOne of the class of models that we used was 2.5D models that just detects the presence of aneurysm in the given series with auxiliary segmentation head. These models were inspired by https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder. These models in combination with yolo11m gave us a LB of 0.80. \n\nThese models then became obsolete when we merged with  @chihantsai, @atom1231.\n\nIn the BYU competition, our best-performing backbone was `timm/convnextv2_base.fcmae_ft_in22k_in1k`. However, in the current competition, training with half precision caused the loss to quickly become NaN. Based on our investigation, this appears to be a common issue with ConvNeXt architectures. We hypothesize that if trained in full precision, this backbone could achieve performance comparable to YOLO11m and `timm/tf_efficientnetv2_s.in21k_ft_in1k`. Unfortunately, we were unable to verify this due to time constraints.\n\nModifying the neck and head was also considered. From what we investigate with @tatamikenn [awesome reverse engineer notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo) P3 features appeared to be the most utilized. We conducted a quick test on a single fold, but results on validation were more or less the same so we decided to move on.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4310004%2F7a4ca157151f82f0436dfdee766ec665%2Fimage%20(2).png?generation=1760534911466456&alt=media)\n\nWe also try to use our 3d segmentation result to [filter out invalid location prediction from yolo](https://www.kaggle.com/code/tom99763/rsna-iad-vessel-segmentation), but recall drops a lot.\nMoreover, we tried extracting patches using YOLO and applying wavelet and log-polar transforms to the patches before predicting their classes; however, [this approach ](https://www.kaggle.com/code/tom99763/yolos-phw-mip-lp-rsna-i) did not yield good results.\n\n## EfficientV2s + 3D-CenterNet (Flayer)\n\nThis is a highly **experimental project**, made with a lot of help from AI. We started with the popular “LB 0.69 Public Notebook” , reconstructed the entire **training pipeline and data preprocessing** from it, and then incorporated a **CenterNet-like mechanism to detect the center of the aneurysms**.\n\n### Model architecture\n\nWe adopt a **slice-wise 2D encoder + shallow 3D head** design:\n\n- **Encoder.** A 2D ImageNet-pretrained **EfficientNetV2-S** (`tf_efficientnetv2_s.in21k_ft_in1k`) from `timm` in *features_only* mode with the **penultimate** feature map (`out_indices = (1,-2)`).\n- **Auxiliary 2D head.**The intermediate feature map (early stage) is fed into a lightweight 2D convolutional head to produce per-slice auxiliary logits, which encourage the encoder to capture vessel-relevant spatial cues before 3D aggregation. The auxiliary head is supervised with slice-level targets to stabilize early-layer learning.\n- **2D→3D fusion.** Each volume `(B, C=1, D, H, W)` is **rearranged into `B·D` 2D slices**, encoded independently, then **reassembled** to `(B, C_feat, D, H′, W′)` before the 3D head. If the backbone expects 3 channels, single-channel inputs are channel-repeated.\n- **3D temporal head.** A lightweight head (Conv3d-BN-ReLU ×2, `base_channels`) aggregates signals across depth.\n- **Outputs.** Two 1×1×1 conv heads:\n    - **Heatmap head:** `num_classes = 13` vessel locations (3D centerness maps)\n    - **Offset head:** 3-channel sub-voxel offsets `(dz, dy, dx)`\n\n### Supervision targets\n\n- Place a **3D Gaussian** peak (stride-aware) at each annotated center on the class-specific heatmap.\n- Store **sub-voxel residual offsets** at the nearest grid cell and mark them with an **offset mask**.\n- Scale coordinates from the original series size to the current `(D, H, W)`; map to the output grid using `(stride_d=1, stride_h=stride_w=16)` so the output is `D × H/16 × W/16`.\n- Default **Gaussian σ = 0.2**.\n\n### Segmentation masks\nVessel masks are produced in physical space by training a 3D DynUNet on volumes resampled to 0.7 mm isotropic spacing. The provided voxel-level annotations of 13 vessel classes are merged into a single binary vessel mask for training. The resulting 3D masks are then converted to slice-wise 2D masks to supervise the auxiliary head.\n\n### Losses\n\nWe train with four terms and sum them with weights:\n\n- **CenterNet-style focal loss** (weight = 1.0) on heatmaps (α=2, β=4).\n- **L1 offset loss** (weight = 1.0) (only at valid center cells).\n- **BCE-with-logits classification loss** (weight = 1.0) from series-level logits .\n- **Auxiliary Dice & BCE loss** (weight = 0.5) on 2D auxiliary outputs when vessel segmentation masks are available; otherwise, it is set to zero, encouraging slice-wise feature consistency.\n\n### From voxelwise maps to series-level logits (for AUC & ensembling)\n\nGiven heatmap logits `H ∈ ℝ^{B×13×D×H′×W′}`:\n\n- **Per-class series logits:** spatial **max** over `(D, H′, W′)` for each of the 13 classes.\n- **Aneurysm Present (global) logit:** the **max over the 13** class logits.\n    \n    These **14 logits** drive the BCE loss and are exported for meta-ensembling.\n    \n\n### Data pipeline & augmentations\n\n- **Input.** Precomputed `.npy` volumes; `(D,H,W)` expanded to `(1,D,H,W)`— here D = 64.\n- **Label scaling.** Rescale coordinates to current volume size; apply stride mapping for target grids.\n- **Augmentations.** A **shared 2D affine warp** per volume (same transform for all slices): random rotation, scale, translation; vertical flip available (off by default).\n- **Normalization.** Albumentations normalization → tensor conversion; channel repeat to 3-ch if needed.\n\n### Optimization & schedules\n\n- **Optimizer.** AdamW (LR=2e-4, weight decay=1e-5), **AMP** mixed precision, **grad clipping** (=1.0).\n- **LR schedule.** Default **CosineAnnealingLR** (`T_max = epochs`, `eta_min = 1e-6`); `StepLR` / `ReduceLROnPlateau` also supported.\n- **Checkpointing.** Best model selected by **mean AUC** on validation.\n\n### Training configuration\n\n- **Epochs:** 16\n- **Batch size:** 4 (train) / 2 (val)\n- **accumulate:**2\n- **Workers:** 16\n- **Strides:** depth `1`, lateral `16`\n- **Gaussian σ:** 0.2\n- **Backbone:** EfficientNetV2-S, `pretrained=True`, not frozen by default\n\n### Cross-validation & metrics\n\n- K-fold protocol driven by the training CSV split.\n- Compute **per-label ROC-AUC**, plus **mean** and **presence-weighted** AUC each epoch.\n- Scheduler stepping and best-by-mean-AUC saving per fold; final summary reports per-fold and average AUC.\n\n### **Flayer Results(without aux head)**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7762 | 0.7747 | 0.7754 |\n| fold1 | 0.7676 | 0.8069 | 0.7873 |\n| fold2 | 0.7508 | 0.7892 | 0.7700 |\n| fold3 | 0.7553 | 0.7778 | 0.7666 |\n| fold4 | 0.7734 | 0.7758 | 0.7746 |\n| **Average** | **0.7647** | **0.7848** | **0.7748** |\n\n### **Flayer Results(with aux head)**\n\n| Fold | loc_macro_auc | cls_auc | combined_mean |\n| --- | --- | --- | --- |\n| fold0 | 0.7779 | 0.7991 | 0.7885 |\n| fold1 | 0.7828 | 0.7987 | 0.7908 |\n| fold2 | 0.7601 | 0.8051 | 0.7826 |\n| fold3 | 0.7683 | 0.7826 | 0.7755 |\n| fold4 | 0.7586 | 0.7954 | 0.7770 |\n| **Average** | **0.7695** | **0.7962** | **0.7829** |\n\n\n\n### Summary\n\n**Flayer** converts 2D EfficientNetV2-S slice features into 3D evidence maps with a light 3D head, supervises them using CenterNet-style heatmaps + offsets, and produces robust **14-logit** series-level predictions (including an aggregated *aneurysm present* logit) that plug cleanly into our stacked meta-classifier.\n\n## Meta Classifier\n\nWe designed a stacked ensemble architecture to integrate predictions from **YOLO11m**, **YOLO11-EffV2s**, and **EffV2s-3D-CenterNet** models. Specifically, each meta-classifier receives the concatenated predictions from all base models as input and produces final predictions for both **aneurysm presence** and **location-specific detection**. This setup enables the model to assess the presence of an aneurysm at a given location by considering predictions from all locations across all models, effectively capturing the ensemble’s collective reasoning about aneurysm presence.\n\n```python\n#for each fold_id\nmetadata = np.array([age, sex])\nX = np.concatenate([np.array([yolo11m_cls_preds[fold_id]]), yolo11m_loc_preds[fold_id],\n                                np.array([effv2s_cls_preds[fold_id]]), effv2s_loc_preds[fold_id],\n                                flayer_fold_preds[fold_id], metadata], axis=0)[None, :]\nlgb_pred = predict_prob_lgb(X, fold_id) #(14, )\nxgb_pred = predict_prob_xgb(X, fold_id)\ncat_pred = predict_prob_cat(X, fold_id)\n```\n\nWe experimented with two ensemble structures:\n\n- **Series:** A single YOLO prediction combined with one Flayer prediction is fed into the ensemble model to produce the final output. \n```\n(30 features, [yolo(14) + flayer(14) + meta(2)])\n```\n- **Parallel:** Two predictions from two YOLOs and one predictions from Flayer are processed in parallel and jointly used by the ensemble model to generate the final prediction. \n```\n(44 features, [yolo11m(14) + yolo_effnets(14)+ flayer(14) + meta(2)])\n```\n![diagram-meta-models](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2221915%2Fef14572c1840aec4e9d0ce115fe7adb3%2Fmeta.png?generation=1760538976614470&alt=media)\nThis comparison allows us to evaluate whether combining multiple feature streams in parallel provides additional complementary information beyond the sequential (series) configuration.\n\nWe found that combining predictions from different YOLO backbones and flayer in parallel provides better improvement than series in the ensemble CV score. This suggests that the GBDT effectively leverages each model’s perception of aneurysm presence across locations to make accurate judgments. We were unable to complete the 5-fold submission due to issues encountered on the final day. Therefore, we present here our final selected result, which is based on the average of 2 folds.\n\n|  | Series | Parallel |\n| --- | --- | --- |\n| 5 folds CV (average) | 0.84 | **0.852** |\n| 5 folds CV (nelder-mead optimized) | 0.841 | **0.858** |\n| public score (average; 2 folds) | 0.83 | 0.83082 |\n| private score (average; 2 folds) | 0.81 | **0.8230** |\n\nWe’ve also tried Keypoint-Based Dual-Graph Predictor in [Tom’s BYU approach](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/bias-vibes-trade-off-20th-place-solution-keypoint-) earlier and found it performed well in terms of ensemble CV performance. However, the YOLO-based feature extraction and preprocessing were too time-consuming and often led to timeouts. Therefore, we did not include this method in our final solution.\n\n## Codebase & Submission\n\nSubmission: \n* [**2x yolo + flayer + meta classifier (final ver)**](https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-classifier-final-ver?scriptVersionId=267890122)\n\n* [kaggle models submission](https://www.kaggle.com/code/tom99763/9th-place-solution-rsna-iad)\n\nMeta classifier training: [**2x yolo + flayer meta training (final ver)**](https://www.kaggle.com/code/tom99763/2x-yolo-flayer-meta-training-final-ver)\n\nGithub: [9th-place-solution-yolo-RSNA-IAD](https://github.com/tom99763/9th-place-solution-yolo-RSNA-IAD)",
    "3304634": "Flayer training code is updated in [our github](https://github.com/tom99763/9th-place-solution-RSNA-IAD), also publish our [kaggle Model](https://www.kaggle.com/models/tom99763/9th-place-models-rsna-iad)",
    "3302673": "I really like how kind this writing feels — congratulations!",
    "3302586": "Hi, thanks for wonderful writeup. I have below questions.\n\n1. Did you guys considred trying out reorientation or resampling of the niftii volume in pre-processing pipeline?\n2. Any specific reason for usage of GBDT in the pipeline. If I am correct Age + Sex metadata is not there for test set. (Was it just matter of convinence?)\n2. Any motivation for 2D or 2.5D based solution? (Was it because of mixture of thick and thin slices).\n3. Did you guys had any attempt with pure E2E 3D based solution?",
    "3302503": "Congratulations, very well done!"
  }
}