{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Six ensembling ideas that failed, and the two that worked (0.873 → 0.919)\n\nPublic notebooks tend to show the thing that worked. This one shows the eight things I tried, because\nsix of them failed and the failures were more instructive than the successes — three of them looked\nlike clear wins on validation before turning out to be noise.\n\n**Where the numbers come from.** Six knee-MRI models, each trained separately and each scored on the\nleaderboard or on a held-out set, then blended:\n\n| member | architecture | notes |\n|---|---|---|\n| v2 | ResNet-34, sagittal fluid-sensitive only, mean-pooled slices | LB 0.873 |\n| v3 | ResNet-34, three planes, per-label plane-mixture head | LB 0.874 |\n| v4 | ConvNeXt-T, gated slice attention, EMA weights | LB 0.883 |\n| v5 | **RadImageNet** ResNet-50 (medically pretrained) | weakest member |\n| v6 / v7 | DINOv2-small, five series slots, geometric slice ordering | two folds |\n\n| submission | public LB |\n|---|---|\n| best single model (v4) | 0.883 |\n| 3-member rank mean | 0.898 |\n| 4-member (+RadImageNet) | 0.907 |\n| 5-member | 0.913 |\n| **6-member + per-target weights** | **0.919** |\n\nEverything below is reproducible here: the six members' validation predictions are attached as a\ndataset, so you can rerun every experiment on CPU in seconds without my model weights."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import itertools\nimport numpy as np\nfrom scipy.stats import rankdata\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n\nimport glob\npath = glob.glob('/kaggle/input/**/member_val_preds.npz', recursive=True)[0]\nz = np.load(path, allow_pickle=True)\nLABELS = [str(x) for x in z['labels']]\nY = z['Y'].astype(float)                       # radiologist labels, 58 annotated studies\nnames = [k for k in ['V2', 'V3', 'V4', 'V5', 'V6', 'V7'] if k in z.files]\nP = {k: z[k] for k in names}\n\ndef rank_norm(a):\n    return np.stack([rankdata(a[:, j]) / len(a) for j in range(a.shape[1])], 1)\n\nR = {k: rank_norm(v) for k, v in P.items()}\nRs = [R[k] for k in names]\n\ndef macro(pred, idx=None):\n    idx = slice(None) if idx is None else idx\n    out = []\n    for j in range(len(LABELS)):\n        y = (Y[idx, j] >= 0.5).astype(int)\n        if 0 < y.sum() < len(y):\n            out.append(roc_auc_score(y, pred[idx, j]))\n    return float(np.mean(out)) if out else np.nan\n\nprint(f'{len(Y)} annotated studies, {len(names)} members')\nfor k in names:\n    print(f'  {k}: {macro(R[k]):.4f}')\nprint(f'\\nequal rank mean: {macro(sum(Rs)/len(Rs)):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"## ✅ What worked #1 — rank mean rather than probability mean\n\nThe metric is ROC AUC, which reads only the *order* of scores. Averaging raw probabilities lets whichever\nmember happens to output the widest numeric range dominate the sum — a property the metric never asked\nabout. Rank-normalising each member per label first puts everyone on the same footing."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"prob_mean = np.mean([P[k] for k in names], axis=0)\nrank_mean = sum(Rs) / len(Rs)\nprint(f'probability mean : {macro(prob_mean):.4f}')\nprint(f'rank mean        : {macro(rank_mean):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"## ✅ What worked #2 — per-target complement weights\n\nEach member is stronger on some findings than others: the sagittal-only model reads ACL best, the\nmedically-pretrained one is the only member that reads Synovitis above 0.72. Giving every member double\nweight on its own specialty labels (declared in advance from training logs, then normalised per label)\nwas worth about +0.007 on validation and +0.004 on the leaderboard.\n\nThe important discipline: the specialty lists were written down **before** scoring the blend, not searched\nfor afterwards. The next section shows what happens when you search instead."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"SPECIALTY = {\n    'V2': ['ACL'],\n    'V3': [\"Baker's\", 'Lateral Meniscus', 'Lateral OA'],\n    'V4': ['Effusion'],\n    'V5': ['Synovitis', 'Fracture', 'Medial OA'],\n    'V6': ['Medial Meniscus', 'Contusion', 'Effusion'],\n    'V7': ['MCL', 'Contusion', 'Medial OA'],\n}\n\ndef complement_weights(subset=None):\n    subset = subset or names\n    W = np.ones((len(subset), len(LABELS)))\n    for i, k in enumerate(subset):\n        for lab in SPECIALTY[k]:\n            W[i, LABELS.index(lab)] += 1.0\n    return W / W.sum(0, keepdims=True)\n\ndef blend(subset, W):\n    return sum(R[k] * W[i][None, :] for i, k in enumerate(subset))\n\ncomp = blend(names, complement_weights())\nprint(f'equal weights      : {macro(rank_mean):.4f}')\nprint(f'complement weights : {macro(comp):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"## ❌ Failure 1 — per-label member dropping\n\nIf each member has weak labels, why not set its weight to **zero** there instead of merely down-weighting?\nOn validation this looked like the biggest win of the whole project."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"per = {k: {c: (roc_auc_score((Y[:, j] >= .5).astype(int), R[k][:, j])\n                if 0 < (Y[:, j] >= .5).sum() < len(Y) else 0.5)\n           for j, c in enumerate(LABELS)} for k in names}\n\ndef drop_weights(drop_n, per_table):\n    W = np.ones((len(names), len(LABELS)))\n    for j, c in enumerate(LABELS):\n        for k in sorted(names, key=lambda k: per_table[k][c])[:drop_n]:\n            W[names.index(k), j] = 0.0\n    for i, k in enumerate(names):\n        for lab in SPECIALTY[k]:\n            if W[i, LABELS.index(lab)] > 0:\n                W[i, LABELS.index(lab)] += 1.0\n    s = W.sum(0, keepdims=True); s[s == 0] = 1\n    return W / s\n\nfor d in (1, 2):\n    print(f'drop worst {d} per label: {macro(blend(names, drop_weights(d, per))):.4f}   <- looks great')\nprint(f'complement baseline    : {macro(comp):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"**+0.011 over the baseline.** That would have been the single largest gain of the project.\n\nIt is entirely an illusion. The member being dropped is chosen using the same 58 studies the blend is then\nscored on. Refit the choice on one half and score on the other, and the gain disappears:"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"rng = np.random.RandomState(0)\nn = len(Y)\nres = {'complement': [], 'drop1': [], 'drop2': []}\nfor _ in range(40):\n    p = rng.permutation(n); h = n // 2\n    for a, b in ((p[:h], p[h:]), (p[h:], p[:h])):\n        per_a = {k: {c: (roc_auc_score((Y[a, j] >= .5).astype(int), R[k][a, j])\n                         if 0 < (Y[a, j] >= .5).sum() < len(a) else 0.5)\n                     for j, c in enumerate(LABELS)} for k in names}\n        res['complement'].append(macro(comp, b))\n        res['drop1'].append(macro(blend(names, drop_weights(1, per_a)), b))\n        res['drop2'].append(macro(blend(names, drop_weights(2, per_a)), b))\n\nprint('honest out-of-fold (choice fitted on the other half):')\nfor k, v in res.items():\n    print(f'  {k:12} {np.nanmean(v):.4f} +/- {np.nanstd(v):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"In-sample **+0.011**, out-of-fold **+0.0003**. The entire gain was the selection itself."},{"cell_type":"markdown","metadata":{},"source":"## ❌ Failure 2 — searching for the best subset of members\n\nSame trap, one level up. With six members there are 57 subsets of size 3 to 6; the best of them scores\nnoticeably above the full set. A properly nested test — pick the subset on one half, score it on the\nuntouched half — shows that *selecting at all* costs you accuracy."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"subsets = [list(s) for r in range(3, len(names) + 1) for s in itertools.combinations(names, r)]\ncand = {'+'.join(s): blend(s, complement_weights(s)) for s in subsets}\n\nbest_name = max(cand, key=lambda k: macro(cand[k]))\nprint(f'best subset in-sample : {best_name}  {macro(cand[best_name]):.4f}')\nprint(f'all six               : {macro(comp):.4f}')\n\nsel, fixed = [], []\nrng = np.random.RandomState(1)\nfor _ in range(60):\n    p = rng.permutation(n); h = n // 2\n    for a, b in ((p[:h], p[h:]), (p[h:], p[:h])):\n        pick = max(cand.values(), key=lambda P_: macro(P_, a))     # choose on half a\n        sel.append(macro(pick, b))                                  # score on half b\n        fixed.append(macro(comp, b))\nprint(f'\\nnested: select-a-subset {np.nanmean(sel):.4f} | keep all six {np.nanmean(fixed):.4f}')\nprint(f'selecting costs {np.nanmean(sel)-np.nanmean(fixed):+.4f}')"},{"cell_type":"markdown","metadata":{},"source":"## ❌ Failures 3–6, briefly\n\n**Test-time augmentation.** Three crops per member, averaged. It genuinely helps each member individually\n(+0.003 to +0.005 each) — and moves the blend by **-0.001**. Rank-mean ensembling is already averaging\naway that variance, so TTA re-buys something you own, at three times the inference cost. Worth keeping for\na single-model submission; not for an ensemble.\n\n**Borrowing a correlated label.** Synovitis is the worst label for every model here (0.62-0.73) and\ncorrelates about 0.5 with Effusion, which is one of the best. Mixing a fraction of the Effusion rank into\nthe Synovitis prediction sounded free. Every mixing weight made it worse, monotonically."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"si, ei = LABELS.index('Synovitis'), LABELS.index('Effusion')\nfor w in (0.0, 0.1, 0.2, 0.3, 0.4):\n    q = comp.copy()\n    q[:, si] = (1 - w) * comp[:, si] + w * comp[:, ei]\n    syn = roc_auc_score((Y[:, si] >= .5).astype(int), q[:, si])\n    print(f'  borrow {w:.1f}: macro {macro(q):.4f}  synovitis {syn:.3f}')"},{"cell_type":"markdown","metadata":{},"source":"**Power-weighted ranks.** Sharpening or flattening the rank distribution before averaging (rank\\*\\*p).\np = 1 was best; every other exponent was worse.\n\n**A seventh member.** I trained a third DINOv2 fold. It was the *strongest single member* I had, and\nadding it moved the blend from 0.8987 to 0.8980 — nothing. Three folds of one architecture are near\nduplicates of each other: blending only the DINOv2 folds together scores 0.879, barely above the best one\nalone at 0.875. Meanwhile the **weakest** member in the whole set, RadImageNet at 0.843, was worth +0.009\nwhen it joined, because a medically pretrained encoder fails differently from ImageNet ones.\n\nThat is the single most useful thing I learned: **an ensemble is fed by disagreement, not by member quality.**"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"pairs = [(a, b) for a, b in itertools.combinations(names, 2)]\ncorr = [(np.mean([np.corrcoef(R[a][:, j], R[b][:, j])[0, 1] for j in range(len(LABELS))]), a, b)\n        for a, b in pairs]\ncorr.sort()\nprint('most different pairs (mean per-label rank correlation):')\nfor c, a, b in corr[:4]:\n    print(f'  {a}-{b}: {c:.3f}')\nprint('most similar pairs:')\nfor c, a, b in corr[-4:]:\n    print(f'  {a}-{b}: {c:.3f}')"},{"cell_type":"markdown","metadata":{},"source":"## Why validation lies here, in one plot\n\nFifty-eight annotated studies, 9 to 35 positives per label. Bootstrapping a *fixed* set of predictions\nshows how much a score can move by luck alone — and the interval is wider than the gap between most of\nthe things I was comparing."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"rng = np.random.RandomState(0)\nboot = []\nfor _ in range(2000):\n    idx = rng.randint(0, n, n)\n    boot.append(macro(comp, idx))\nlo, hi = np.percentile(boot, [2.5, 97.5])\n\nfig, ax = plt.subplots(figsize=(8, 4))\nax.hist(boot, bins=50, color='#4477AA')\nax.axvline(lo, color='#CC3311', ls='--'); ax.axvline(hi, color='#CC3311', ls='--')\nax.set_title(f'Same predictions, resampled studies — 95% CI width {hi-lo:.3f}')\nax.set_xlabel('macro AUC'); plt.tight_layout(); plt.show()\nprint(f'95% CI: {lo:.3f} to {hi:.3f}  (width {hi-lo:.3f})')\nprint('for comparison, the whole 0.873 -> 0.919 climb is 0.046')"},{"cell_type":"markdown","metadata":{},"source":"## What I would tell myself at the start\n\n1. **Blend by rank, not probability**, whenever the metric only reads order.\n2. **Add members that fail differently**, not members that score well. My weakest model was my most\n   valuable addition; my strongest addition was worthless.\n3. **Write weighting rules down before you score them.** Anything chosen by looking at a small validation\n   set will look good on that set and nowhere else.\n4. **Nested-validate any selection step.** Two of my three \"best ideas\" reversed sign under a proper test.\n5. **Bootstrap the validation metric once, at the start.** If the confidence interval is wider than the\n   effects you plan to chase, that set cannot referee your experiments — and mine could not.\n\nCorrections and counter-results welcome, especially from anyone who has made a learned stacker work on a\nvalidation set this small. I could not, and I would like to know what I was missing."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}