{"cells":[{"cell_type":"markdown","metadata":{},"source":"# 4,407 knee MRIs, 58 labels — the other 4,349 are in the reports\n\nThis competition hands you 4,407 knee MRI studies and asks for 12 binary findings per\nstudy. It also hands you a label table where **58 rows are filled in and 4,349 are\nempty**. That is not an oversight, and it is the single fact that should shape how you\nspend the next two months.\n\nThe supervision you are missing is not missing. It is sitting in the `Report` column as\nfree text — written by radiologists in **seven languages**, 178 of them in Greek script,\nwhich most language-detection one-liners will silently drop.\n\nThis notebook establishes the ground facts: what is labelled, what the reports look\nlike, what the series metadata actually contains (one of its three columns is a copy of\nanother), what a DICOM slice looks like, and how far a plain multilingual keyword\nmatcher gets you on the 58 studies where the answer is known — the floor any real\nreport-parsing method has to beat.\n\n*No model here. This is the map, not the route.*"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import os, re, glob, warnings\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nwarnings.filterwarnings('ignore')\n\n# Kaggle mounts competition data under /kaggle/input/<slug> on older notebooks and\n# /kaggle/input/competitions/<slug> on newer ones — find it instead of guessing\nP = glob.glob('/kaggle/input/**/train_series.csv', recursive=True)[0].rsplit('/', 1)[0]\nprint('data:', P)\n\nBLUE, ORANGE, GREY = '#3987e5', '#d95926', '#8a8a85'\nplt.rcParams.update({'figure.facecolor': 'white', 'axes.facecolor': 'white',\n                     'axes.spines.top': False, 'axes.spines.right': False,\n                     'axes.edgecolor': '#cccccc', 'font.size': 11,\n                     'axes.titlesize': 14, 'axes.titleweight': 'bold',\n                     'axes.titlelocation': 'left', 'axes.titlepad': 14})\n\ntrain  = pd.read_csv(f'{P}/train.csv')\nseries = pd.read_csv(f'{P}/train_series.csv')\ntest   = pd.read_csv(f'{P}/test.csv')\nsub    = pd.read_csv(f'{P}/sample_submission.csv')\n\nLABELS = [c for c in train.columns if c not in ('StudyInstanceUID', 'Report')]\nprint(f'train   {train.shape[0]:>6,} studies')\nprint(f'series  {series.shape[0]:>6,} series across {series.StudyInstanceUID.nunique():,} studies')\nprint(f'test    {test.shape[0]:>6,} studies   (public test is a 3-study stub)')\nprint(f'labels  {len(LABELS)}: {\", \".join(LABELS)}')"},{"cell_type":"markdown","metadata":{},"source":"## 1. The label table is 1.3% full\n\nEvery one of the 12 label columns is filled for exactly the same 58 studies. There is no\npartial labelling to exploit — a study is either fully annotated or not annotated at\nall."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"cov = train[LABELS].notna().sum()\nlabelled = train[train[LABELS].notna().any(axis=1)]\n\nprint(f'studies with at least one label : {len(labelled):>5,}')\nprint(f'studies with all twelve labels  : {train[LABELS].notna().all(axis=1).sum():>5,}')\nprint(f'distinct coverage counts        : {sorted(cov.unique())}   <- identical for every column')\nprint(f'labelled fraction               : {len(labelled) / len(train):.3%}')\n\nrates = labelled[LABELS].mean().sort_values()\nfig, ax = plt.subplots(figsize=(9, 5))\nax.barh(rates.index, rates.values, color=BLUE, height=0.62)\nfor k, v in rates.items():\n    ax.text(v + .008, k, f'{v:.0%}  ({int(v * len(labelled))}/58)', va='center',\n            fontsize=10, color='#444')\nax.set_xlim(0, rates.max() * 1.35)\nax.set_xticks([])\nax.tick_params(left=False)\nax.set_title('Positive rate in the 58 labelled studies')\nplt.show()"},{"cell_type":"markdown","metadata":{},"source":"Two things to carry forward. **Nothing is rare** — the least common finding still shows\nup in 9 of 58 studies, so this is not an imbalance problem. And **58 studies is far too\nfew to validate 12 binary heads**: at this size a single flipped study moves a positive\nrate by 1.7 points. Any local CV you build on these 58 alone will be noise."},{"cell_type":"markdown","metadata":{},"source":"## 2. The reports, and the language problem\n\nEach study carries a free-text radiology report — median ~980 characters. That text is\nwhere the other 4,349 label sets live. The catch is that it is not in one language.\n\nThe detector below is a deliberately crude stopword counter, because that is the point:\neven a crude detector separates six Latin-script languages plus Greek, and still leaves\n428 reports it cannot place."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"STOP = {\n    'English':    [' the ', ' and ', ' with ', ' no ', 'knee', 'findings', 'impression'],\n    'Spanish':    ['de la', ' del ', ' sin ', 'no hay', 'hallazgos', 'rodilla', 'menisco'],\n    'Turkish':    [' ve ', 'diz ', 'bulgu', 'düzey', 'izlenmekte', 'menisk'],\n    'German':     [' der ', ' und ', ' mit ', ' kein', 'knie', 'befund'],\n    'Dutch':      ['van de', ' met ', 'geen ', 'knie ', 'bevindingen', 'meniscus'],\n    'French':     [' de la ', ' avec ', ' sans ', 'genou', 'conclusion'],\n    'Portuguese': [' da ', ' com ', ' sem ', 'joelho', 'achados'],\n    'Greek':      ['ευρήματα', 'γόνατ', 'μηνίσκ', 'δεν ', 'εξέταση'],\n}\n\ndef detect(text):\n    t = ' ' + text.lower() + ' '\n    score = {k: sum(t.count(w) for w in ws) for k, ws in STOP.items()}\n    best = max(score, key=score.get)\n    return best if score[best] > 1 else 'unresolved'\n\ntrain['lang'] = train.Report.map(detect)\ncounts = train.lang.value_counts()\n\nfig, ax = plt.subplots(figsize=(9, 4.6))\ncol = [GREY if k == 'unresolved' else BLUE for k in counts.index]\nax.bar(counts.index, counts.values, color=col, width=.66)\nfor k, v in counts.items():\n    ax.text(k, v + 25, f'{v:,}', ha='center', fontsize=10, color='#444')\nax.set_ylim(0, counts.max() * 1.18)\nax.set_yticks([])\nax.tick_params(left=False)\nplt.setp(ax.get_xticklabels(), rotation=30, ha='right')\nax.set_title('Report language  (crude stopword detector)')\nplt.show()\n\nprint(counts.to_string())"},{"cell_type":"markdown","metadata":{},"source":"The `unresolved` bar is the honest part of this chart: those are reports too short,\ntoo abbreviation-heavy, or in a language not on the list. Treat that bar as your backlog,\nnot as noise — and note that the same detector puts several of the **labelled** 58 in\nthat bucket, so it is not a rare-corner problem.\n\nThree reports, three languages, so you can see what you are actually parsing:"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"for lg in ['English', 'Spanish', 'Greek']:\n    r = train[train.lang == lg].Report.iloc[1]\n    print(f'───── {lg} ' + '─' * 60)\n    print(r[:430].strip(), '…\\n')"},{"cell_type":"markdown","metadata":{},"source":"Same anatomy, three vocabularies, three negation grammars. Note that the Greek report\nis not an edge case in format — it has the ordinary TECHNIQUE / FINDINGS / IMPRESSION\nskeleton. It is only unusual to a pipeline that assumed Latin script."},{"cell_type":"markdown","metadata":{},"source":"## 3. Series metadata: three columns, two of them independent\n\n`train_series.csv` gives you `Fluid_Sensitive`, `Fat_Suppression` and\n`Anatomical_Plane` per series. The first two are the same column."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"ct = pd.crosstab(series.Fluid_Sensitive, series.Fat_Suppression)\nprint(ct.to_string())\nprint(f'\\noff-diagonal rows: {ct.values.sum() - np.trace(ct.values)} of {len(series):,}')\n\nper_study = series.groupby('StudyInstanceUID').size()\nfig, axes = plt.subplots(1, 2, figsize=(12, 4.2))\npl = series.Anatomical_Plane.value_counts()\naxes[0].bar(pl.index, pl.values, color=BLUE, width=.6)\nfor k, v in pl.items():\n    axes[0].text(k, v + 120, f'{v:,}', ha='center', fontsize=10, color='#444')\naxes[0].set_ylim(0, pl.max() * 1.15)\naxes[0].set_yticks([]); axes[0].tick_params(left=False)\naxes[0].set_title('Series by plane')\n\nh = per_study.value_counts().sort_index()\naxes[1].bar(h.index, h.values, color=ORANGE, width=.7)\naxes[1].set_xlabel('series in a study')\naxes[1].set_yticks([]); axes[1].tick_params(left=False)\naxes[1].set_title(f'Series per study  (median {int(per_study.median())}, max {per_study.max()})')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"`Fluid_Sensitive` and `Fat_Suppression` agree on all 24,371 series — zero\noff-diagonal. Encode one of them and drop the other; feeding both to a model just\nduplicates a feature.\n\nThe plane counts hide the more useful fact, which the next cell pulls out."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"planes = (series.assign(v=1)\n          .pivot_table(index='StudyInstanceUID', columns='Anatomical_Plane',\n                       values='v', aggfunc='sum').fillna(0).astype(int))\nprint('studies missing a plane entirely:')\nfor c in ['Sagittal', 'Coronal', 'Axial']:\n    print(f'  no {c:9s} {int((planes[c] == 0).sum()):>5,}')\n\nfl = series.groupby('StudyInstanceUID').Fluid_Sensitive.agg(['sum', 'size'])\nprint(f'\\nstudies with no fluid-sensitive series : {int((fl[\"sum\"] == 0).sum()):>5,}')\nprint(f'studies with no T1-type series        : '\n      f'{int((fl[\"sum\"] == fl[\"size\"]).sum()):>5,}')\n\nprint('\\nmost common (sagittal, coronal, axial) layouts:')\nprint(planes.groupby(['Sagittal', 'Coronal', 'Axial']).size()\n      .sort_values(ascending=False).head(5).to_string())"},{"cell_type":"markdown","metadata":{},"source":"**Every study has all three planes, and every study has both fluid-sensitive and\nnon-fluid-sensitive series.** Zero exceptions in 4,407 studies. That is unusually clean\nfor a medical imaging competition, and it means you can design a fixed three-plane,\ntwo-contrast input without a fallback path — the 40% of studies that are exactly\n2 sagittal + 2 coronal + 1 axial make a reasonable canonical layout to standardise on."},{"cell_type":"markdown","metadata":{},"source":"## 4. What a slice actually looks like\n\nPixel data is DICOM, one file per slice, nested `series/<study>/<series>/<slice>.dcm`.\nBelow: one study, one series per plane, four evenly spaced slices each."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import pydicom\n\nSTUDY = series.StudyInstanceUID.iloc[0]\nrows = series[series.StudyInstanceUID == STUDY].drop_duplicates('Anatomical_Plane')\n\nfig, axes = plt.subplots(len(rows), 4, figsize=(13, 3.3 * len(rows)))\nfor ax_row, (_, r) in zip(np.atleast_2d(axes), rows.iterrows()):\n    files = sorted(glob.glob(f'{P}/train_series/{STUDY}/{r.SeriesInstanceUID}/*.dcm'))\n    picks = np.linspace(0, len(files) - 1, 4).astype(int)\n    for ax, i in zip(ax_row, picks):\n        d = pydicom.dcmread(files[i])\n        ax.imshow(d.pixel_array, cmap='gray')\n        ax.set_axis_off()\n    ax_row[0].set_title(f'{r.Anatomical_Plane} · {len(files)} slices · '\n                        f'fluid-sensitive={r.Fluid_Sensitive}', fontsize=11, loc='left')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"### The tags worth reading\n\nSequence parameters are not decoration here — `Fluid_Sensitive` is a coarse summary of\nwhat `RepetitionTime` / `EchoTime` say precisely, and effusion is far easier to see on a\nfluid-sensitive sequence than on a T1."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def header(path):\n    d = pydicom.dcmread(path, stop_before_pixels=True)\n    g = lambda t, dflt='—': getattr(d, t, dflt)\n    return dict(rows=g('Rows'), cols=g('Columns'), thick=g('SliceThickness'),\n                spacing=str(g('PixelSpacing')), TR=g('RepetitionTime'),\n                TE=g('EchoTime'), tesla=g('MagneticFieldStrength'),\n                vendor=str(g('Manufacturer'))[:22], desc=str(g('SeriesDescription'))[:26])\n\nsample = series.sample(120, random_state=0)\nrecs = []\nfor _, r in sample.iterrows():\n    f = glob.glob(f'{P}/train_series/{r.StudyInstanceUID}/{r.SeriesInstanceUID}/*.dcm')\n    if f:\n        recs.append({**header(f[0]), 'plane': r.Anatomical_Plane,\n                     'fluid': r.Fluid_Sensitive})\nhdr = pd.DataFrame(recs)\nprint(hdr[['plane', 'desc', 'rows', 'cols', 'thick', 'TR', 'TE', 'tesla', 'vendor']]\n      .head(8).to_string(index=False))\nprint(f'\\nscanner vendors in a {len(hdr)}-series sample:')\nprint(hdr.vendor.value_counts().head().to_string())\nnum = lambda s: sorted(pd.to_numeric(s, errors='coerce').dropna().unique())\nprint(f'\\nfield strength : {num(hdr.tesla)}')\nprint(f'in-plane size  : {num(hdr[\"rows\"])}')\nprint(f'slice thickness: {num(hdr.thick)}')"},{"cell_type":"markdown","metadata":{},"source":"Resolution, vendor and field strength all vary across the corpus. Resize and intensity\n-normalise per series rather than assuming a global scale."},{"cell_type":"markdown","metadata":{},"source":"## 5. How far do keywords get you?\n\nThe obvious first move on 4,349 unlabelled reports is a keyword matcher. It is worth\nknowing exactly how well that works before reaching for anything bigger, so here it is,\nscored against the 58 studies where the truth is known.\n\nTwo design notes. The vocabulary is taken from a radiology glossary, **not tuned on the\n58** — otherwise the accuracy below would be meaningless. And negation is handled\nexplicitly: `no effusion`, `sin derrame`, `geen gewrichtsvocht` and Turkish\n`efüzyon izlenmedi` all have to cancel a hit, and Turkish puts its negation *after* the\nterm."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"NEG_L = (r'\\b(?:no|not|without|absent|geen|zonder|kein[e]?|ohne|sin|sans|senza|non|'\n         r'nessun[ao]?|δεν|χωρ[ιί]ς)\\b')\nNEG_R = r'\\b(?:izlenmedi|saptanmad[iı]|yok|g[oö]r[uü]lmedi|mevcut de[gğ]il)\\b'\n\nTERMS = {\n 'Effusion': r'effusion|joint fluid|derrame|gewrichtsvocht|erguss|[ée]panchement|'\n             r'versamento|ef[uü]zyon|s[iı]v[iı]|υγρ[οό]',\n \"Baker's\": r'baker|popliteal cyst|quiste popl[ií]teo|bakercyste|poplitea?le? cyste|'\n            r'bakerzyste|kyste popliet|baker kist',\n 'Fracture': r'fracture|fractura|fractuur|fraktur|frattura|k[iı]r[iı]k|avulsion',\n 'ACL': r'anterior cruciate|acl\\b|ligamento cruzado anterior|lca\\b|voorste kruisband|'\n        r'vkb\\b|vorderes kreuzband|ligament crois[ée] ant[ée]rieur|[oö]n [cç]apraz',\n}\nTEAR = (r'tear|torn|rupture|ruptur|rotura|ruptura|scheur|ruptuur|riss|d[ée]chirure|'\n        r'lesione|y[iı]rt[iı]k|rupt[uü]r')\n\ndef hit(text, pattern, need=None, window=90):\n    t = ' ' + re.sub(r'\\s+', ' ', text.lower()) + ' '\n    for m in re.finditer(pattern, t):\n        left, right = t[max(0, m.start() - window):m.start()], t[m.end():m.end() + window]\n        if re.search(NEG_L + r'\\W+(?:\\w+\\W+){0,4}$', left):      # negated before\n            continue\n        if re.search(r'^\\W*(?:\\w+\\W+){0,6}' + NEG_R, right):     # negated after\n            continue\n        if need and not re.search(need, left + ' ' + right):       # ligament needs a tear\n            continue\n        return 1\n    return 0\n\ndef extract(text):\n    out = {k: hit(text, p) for k, p in TERMS.items()}\n    out['ACL'] = hit(text, TERMS['ACL'], need=TEAR)\n    return out\n\npred = pd.DataFrame([extract(t) for t in labelled.Report])\nprint(f'{\"finding\":10s} {\"accuracy\":>9s} {\"sens\":>7s} {\"spec\":>7s}   n positive')\nfor c in pred.columns:\n    g, p = labelled[c].values, pred[c].values\n    print(f'{c:10s} {(g == p).mean():9.3f} {p[g == 1].mean():7.3f} '\n          f'{1 - p[g == 0].mean():7.3f}   {int(g.sum()):>2}/58')"},{"cell_type":"markdown","metadata":{},"source":"So: **0.86 on Baker's cyst, 0.66 on effusion.** Baker's works because it is a named\nobject that either appears in a report or does not. Effusion fails because it is graded —\n\"minimal joint fluid\", \"trace effusion\", \"physiological amount\" — and a binary keyword\nhas nowhere to put a grade. Fracture has 0.93 specificity against 0.44 sensitivity: when\nit fires it is right, but more than half the fractures are described by their appearance —\na subchondral defect ringed by marrow oedema — without the word *fracture* ever being\nwritten.\n\nThat is the shape of the whole problem. Keyword matching is a floor, and these numbers\nare the floor's height. Anything that parses these reports properly — an instruction-\ntuned model, per-language rules, or a translation pass first — should be measured\nagainst this table and not against zero.\n\nOne caution on the measurement itself: 58 studies means the standard error on any\naccuracy here is about ±6 points. Treat the column as an order of magnitude."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"full = pd.DataFrame([extract(t) for t in train.Report])\ncomp = pd.DataFrame({'keyword rate, all 4,407': full.mean(),\n                     'true rate, labelled 58': labelled[full.columns].mean()}).round(3)\nprint(comp.to_string())"},{"cell_type":"markdown","metadata":{},"source":"Baker's cyst lands at 0.217 across the full corpus against 0.207 in the labelled\nsample — the matcher reproduces the prevalence almost exactly. ACL lands at 0.195 against\n0.414, which says the labelled 58 are **not** a random draw from the corpus: they are\nenriched for the interesting cases. Do not calibrate a model's output prior on those 58."},{"cell_type":"markdown","metadata":{},"source":"## 6. A submission that scores\n\nThe public test set is a 3-study stub, so a submission here is a plumbing check rather\nthan a score. Column order and dtype are what break people."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"out = sub.copy()\nout[LABELS] = 0.5\nout.to_csv('submission.csv', index=False)\nprint(out.shape, '->', list(out.columns) == ['StudyInstanceUID'] + LABELS)\nout.head()"},{"cell_type":"markdown","metadata":{},"source":"## 7. What 58 studies can actually measure\n\nSection 1 said 58 studies is too few to validate twelve binary heads, and section 5 put a\n`+/- 6 points` caution under the keyword table. Both are hand-waves. This section replaces them\nwith a number, because the number turns out to govern how you should spend the next two months.\n\nThe question is not \"how noisy is one accuracy figure\" but **\"if my new method really is better,\nwhat is the chance this test set tells me so?\"** That is statistical power, and it is simulable:\ngive two scorers a known AUC gap, draw 58 studies with the real positive count for a finding,\ntake the paired bootstrap interval, and count how often it excludes zero."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"from scipy.stats import norm\n\nRNG = np.random.default_rng(7)\nN_STUDIES, NBOOT, TRIALS = len(labelled), 600, 400\nBASE = 0.672                      # the keyword floor measured in section 5\nPOS = labelled[LABELS].sum().astype(int).to_dict()\n\n\ndef fast_auc(y, s):\n    '''y, s: (k, N) -> (k,) AUC per row, via the rank-sum identity.'''\n    order = np.argsort(s, axis=1)\n    ranks = np.empty_like(order, dtype=np.float64)\n    np.put_along_axis(ranks, order, np.arange(1, s.shape[1] + 1, dtype=float), axis=1)\n    npos = y.sum(1); nneg = s.shape[1] - npos\n    out = ((ranks * y).sum(1) - npos * (npos + 1) / 2) / np.maximum(npos * nneg, 1)\n    out[(npos == 0) | (nneg == 0)] = np.nan\n    return out\n\n\ndef power(npos, gap, base=BASE):\n    '''P(the 95% paired bootstrap interval excludes zero) for a true improvement of `gap`.'''\n    y = np.zeros(N_STUDIES, int); y[:npos] = 1\n    mu_a, mu_b = norm.ppf(base + gap) * np.sqrt(2), norm.ppf(base) * np.sqrt(2)\n    hits = 0\n    for _ in range(TRIALS):\n        sa, sb = RNG.normal(mu_a * y, 1.0), RNG.normal(mu_b * y, 1.0)\n        idx = RNG.integers(0, N_STUDIES, (NBOOT, N_STUDIES))\n        yb = y[idx]\n        d = fast_auc(yb, sa[idx]) - fast_auc(yb, sb[idx])\n        d = d[~np.isnan(d)]\n        hits += bool(len(d) and np.quantile(d, .025) > 0)\n    return hits / TRIALS\n\n\nGAPS = [.05, .10, .15, .20, .25, .30]\nassert max(GAPS) + BASE < 1.0, 'the binormal separation is infinite at AUC 1'\n\npw = pd.DataFrame([{'finding': l, 'positives': int(POS[l]),\n                    **{f'+{g:.2f}': power(int(POS[l]), g) for g in GAPS}}\n                   for l in LABELS]).sort_values('positives')\nprint(f'Chance a real improvement over a {BASE:.3f} baseline is called significant, n=58\\n')\nprint(pw.round(2).to_string(index=False))\nmean_pw = pw[[c for c in pw.columns if c.startswith('+')]].mean()\nprint('\\nmean over the twelve findings:')\nprint(mean_pw.round(2).to_string())"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"fig, ax = plt.subplots(figsize=(8.6, 4.8))\ngaps = np.array([float(c) for c in mean_pw.index])\nax.plot(gaps, mean_pw.values, color=BLUE, lw=2.2, marker='o', ms=7, zorder=3)\nax.axhline(.8, color=ORANGE, lw=1.3, ls=(0, (4, 3)))\nax.text(.052, .815, '80% power', color=ORANGE, fontsize=10)\n\nxs = np.linspace(gaps.min(), gaps.max(), 200)\nmde = np.interp(.8, mean_pw.values, gaps)\nax.plot([mde, mde], [0, .8], color=ORANGE, lw=1.1, ls=(0, (2, 3)))\nax.text(mde + .006, .12, f'+{mde:.2f} AUC\\nbefore this test set\\ncan see it',\n        color=ORANGE, fontsize=10, va='bottom')\n\nfor g, p in zip(gaps, mean_pw.values):\n    ax.text(g, p + .035, f'{p:.0%}', ha='center', fontsize=9.5, color='#444')\nax.set_xlabel('true improvement over the keyword floor, in AUC')\nax.set_ylabel('chance the 58 studies detect it')\nax.set_ylim(0, 1.08); ax.set_yticks([0, .25, .5, .75, 1.0])\nax.set_yticklabels(['0', '25%', '50%', '75%', '100%'])\nax.set_title('58 studies cannot see an improvement smaller than about +0.24 AUC')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Read the curve at the size of improvement you are actually hoping for.\n\n**A +0.05 gain is detected 7% of the time.** If you tune a prompt, add a rule, or swap a\nbackbone and the truth is that you gained five points of AUC, this test set will almost\ncertainly tell you nothing happened.\n\n**80% power arrives at about +0.24 AUC.** That is an enormous effect -- the difference between\na coin flip and a usable classifier -- and it is the *smallest* thing 58 studies can reliably\nconfirm.\n\nThree consequences, and they are the practical content of this whole notebook:\n\n1. **A null result on the 58 is not evidence of no improvement.** It is the expected outcome for\n   any realistic improvement. Do not discard a method because these 58 studies shrugged at it.\n2. **Do not do model selection on them.** Choosing between ten variants by their score on 58\n   studies selects the luckiest variant, not the best one, and the gap between those two is\n   larger than the differences you are choosing between.\n3. **Spend them on catching breakage, not on ranking.** A pipeline that is silently wrong shows\n   up here immediately. A pipeline that is 4% better does not.\n\nThe companion notebook,\n[Weak labels for 12 knee findings](https://www.kaggle.com/code/nekkon/weak-labels-for-all-12-knee-mri-findings),\nis a direct illustration. Replacing the keyword matcher with a cross-lingual entailment model\nmoved medial meniscus by +0.32 AUC and that came back significant; it moved eight other findings\nby between +0.04 and +0.11 and every one of those came back as a tie. Reading this curve, that is\nexactly what should have happened whether or not those eight were real improvements."},{"cell_type":"markdown","metadata":{},"source":"## What this all adds up to\n\n1. **58 labelled studies out of 4,407.** Not an imbalance problem, a supervision problem.\n2. **The labels exist as prose**, in seven languages, with per-language negation.\n3. **A keyword floor of 0.66-0.86** depending on how nameable the finding is. Beat it\n   deliberately, and measure against it.\n4. **The 58 are not a random sample** - ACL prevalence is twice the corpus rate. Do not\n   use them to set priors, and do not trust a CV built on them alone.\n5. **`Fluid_Sensitive` == `Fat_Suppression`**, exactly, on all 24,371 series.\n6. **The imaging side is clean.** All three planes and both contrast types are present in\n   every study, so a fixed-shape input needs no fallback.\n7. **58 studies reach 80% power only at about +0.24 AUC.** Everything below that is invisible\n   to your local validation, including most real improvements. Use the 58 to catch breakage,\n   and get your signal from the 4,349 reports instead.\n\nIf this was useful, an upvote helps other people find it. Corrections in the comments are\nmore useful still - especially from anyone who reads Greek or Turkish radiology reports\nbetter than a regex does."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","nbconvert_exporter":"python","pygments_lexer":"ipython3"}},"nbformat":4,"nbformat_minor":5}