{"cells":[{"cell_type":"markdown","metadata":{},"source":"# 🦵 RSNA Knee Abnormality Detection — Deep EDA + Metric + CV Starter\n\n**The one number that decides this competition: only ~1.3% of training studies are labeled.**\nThe rest come with a free-text *radiology report* (in many languages) that you must mine for labels.\nSo this is really a **weak-supervision / report-mining** problem wearing an imaging-competition coat.\n\nThis notebook covers, end to end:\n1. The exact **metric** (macro-averaged ROC-AUC) — reproduced in code\n2. The **labeled-vs-report-only** split (the crux)\n3. **Label prevalence** & co-occurrence\n4. **Report** languages & lengths\n5. **Series / plane / sequence** structure (Sagittal·Coronal·Axial, fluid-sensitive, fat-sat)\n6. **Slices per series** distribution\n7. A **DICOM preview** across planes\n8. A **study-grouped, label-stratified CV** you can copy\n\n*If this helps your start, an upvote is very welcome 🙏 — questions in the comments.*","id":"cell00"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import os, glob, numpy as np, pandas as pd\nimport matplotlib.pyplot as plt\nplt.rcParams['figure.dpi']=120; plt.rcParams['axes.grid']=True\npd.set_option('display.width',120)\n\nCOMP='/kaggle/input/competitions/rsna-knee-abnormality-detection'\nif not os.path.isdir(COMP):\n    # fallback: locate the mounted competition folder\n    for r,d,_ in os.walk('/kaggle/input'):\n        if 'train_series' in d: COMP=r; break\nprint('competition root:', COMP)\nprint(sorted(os.listdir(COMP))[:10])\n\nLABELS=['ACL','MCL','Medial Meniscus','Lateral Meniscus','Medial OA','Lateral OA',\n        'PF OA','Effusion','Synovitis',\"Baker's\",'Contusion','Fracture']\ntrain=pd.read_csv(f'{COMP}/train.csv')\nseries=pd.read_csv(f'{COMP}/train_series.csv')\nprint('train.csv', train.shape, '| train_series.csv', series.shape)","id":"cell01"},{"cell_type":"markdown","metadata":{},"source":"## 1. The metric — macro-averaged ROC-AUC (per study)\n$$\\text{Score}=\\frac{1}{12}\\sum_{i=1}^{12}\\text{AUC}_i$$\nTwo consequences that should shape everything you do:\n- It is a **ranking** metric per column — only the *order* of your scores matters, so calibration\n  is irrelevant and **rank-averaging** is the natural way to ensemble.\n- Every label is weighted **equally (1/12)** regardless of how rare it is — so a rare finding is\n  worth just as much as a common one. Spend effort where the AUC *headroom* is largest.","id":"cell02"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"from sklearn.metrics import roc_auc_score\ndef macro_auc(y_true, y_pred):\n    y_true=np.asarray(y_true,float); y_pred=np.asarray(y_pred,float)\n    aucs=[]\n    for j in range(y_true.shape[1]):\n        col=y_true[:,j]\n        if len(np.unique(col[~np.isnan(col)]))>1:\n            m=~np.isnan(col); aucs.append(roc_auc_score(col[m],y_pred[m,j]))\n    return float(np.mean(aucs))\n# sanity: perfect=1, reversed=0, random~0.5\nrng=np.random.default_rng(0); y=rng.integers(0,2,(500,12)); y[0]=0; y[1]=1\nprint('perfect :', round(macro_auc(y,y.astype(float)),3))\nprint('reversed:', round(macro_auc(y,1-y.astype(float)),3))\nprint('random  :', round(macro_auc(y,rng.random((500,12))),3))","id":"cell03"},{"cell_type":"markdown","metadata":{},"source":"## 2. The crux — how many studies are actually labeled?\nEvery study has a **Report**. Only a tiny fraction carry the 12 ground-truth labels.","id":"cell04"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"has_label = train[LABELS].notna().all(axis=1)\nn_lab=int(has_label.sum()); n_tot=len(train)\nprint(f'Labeled studies : {n_lab} / {n_tot}  ({100*n_lab/n_tot:.1f}%)')\nprint(f'Report-only     : {n_tot-n_lab} / {n_tot}  ({100*(n_tot-n_lab)/n_tot:.1f}%)')\nprint('Studies with a Report:', int(train['Report'].notna().sum()))\n\nfig,ax=plt.subplots(figsize=(6,1.6))\nax.barh([0],[n_lab],color='#2a9d8f',label=f'labeled ({n_lab})')\nax.barh([0],[n_tot-n_lab],left=[n_lab],color='#e9c46a',label=f'report-only ({n_tot-n_lab})')\nax.set_yticks([]); ax.set_xlabel('studies'); ax.legend(loc='center'); ax.grid(False)\nax.set_title('Labeled vs report-only — the whole game is turning reports into labels')\nplt.show()","id":"cell05"},{"cell_type":"markdown","metadata":{},"source":"## 3. Label prevalence & co-occurrence (on the labeled subset)","id":"cell06"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"lab=train[has_label]\nprev=lab[LABELS].mean().sort_values()\nfig,ax=plt.subplots(figsize=(7,4))\nax.barh(prev.index, prev.values, color='#264653')\nax.set_xlabel('positive rate'); ax.set_title(f'Label prevalence (n={len(lab)} labeled studies)')\nfor i,v in enumerate(prev.values): ax.text(v+0.005,i,f'{v:.2f}',va='center',fontsize=8)\nplt.tight_layout(); plt.show()\n\nnpos=lab[LABELS].sum(axis=1)\nprint('Findings per labeled study: mean %.1f, median %.0f, max %d'%(npos.mean(),npos.median(),npos.max()))\nprint('=> multi-label: most knees have several findings at once.')","id":"cell07"},{"cell_type":"markdown","metadata":{},"source":"## 4. Reports — multilingual & variable length","id":"cell08"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"rl=train['Report'].astype(str).str.len()\nfig,ax=plt.subplots(figsize=(7,3))\nax.hist(rl.clip(upper=3500),bins=50,color='#e76f51')\nax.set_xlabel('report length (chars)'); ax.set_ylabel('# studies')\nax.set_title('Report length distribution'); plt.tight_layout(); plt.show()\nprint('length: median %d, p90 %d, max %d chars'%(rl.median(),rl.quantile(.9),rl.max()))\n# crude language hint: share of non-ASCII reports\nnonascii=train['Report'].astype(str).apply(lambda s: not s.isascii()).mean()\nprint(f'~{100*nonascii:.0f}% of reports contain non-ASCII characters (many languages present).')\nprint('\\nExample (first 300 chars):\\n', str(train.loc[has_label,\"Report\"].iloc[1])[:300])","id":"cell09"},{"cell_type":"markdown","metadata":{},"source":"## 5. Series structure — planes & sequence types\nEach study is several MRI **series** (acquisitions). The per-series metadata tells you the plane\nand whether the sequence is fluid-sensitive / fat-suppressed — useful to **route** pathologies\n(e.g. ACL & meniscus read best on sagittal; effusion on fluid-sensitive sequences).","id":"cell10"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"sps=series.groupby('StudyInstanceUID').size()\nprint('series per study: mean %.1f, median %.0f, range %d-%d'%(sps.mean(),sps.median(),sps.min(),sps.max()))\nfig,axes=plt.subplots(1,3,figsize=(12,3))\nseries['Anatomical_Plane'].value_counts().plot.bar(ax=axes[0],color='#2a9d8f',title='Anatomical_Plane')\nseries['Fluid_Sensitive'].value_counts().sort_index().plot.bar(ax=axes[1],color='#e9c46a',title='Fluid_Sensitive')\nseries['Fat_Suppression'].value_counts().sort_index().plot.bar(ax=axes[2],color='#e76f51',title='Fat_Suppression')\nfor a in axes: a.tick_params(axis='x',rotation=0)\nplt.tight_layout(); plt.show()\nprint('\\nPlane x Fluid_Sensitive:'); print(pd.crosstab(series.Anatomical_Plane, series.Fluid_Sensitive))","id":"cell11"},{"cell_type":"markdown","metadata":{},"source":"## 6. Slices per series","id":"cell12"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# count .dcm per series folder for a sample of studies (fast)\nsample_studies=train['StudyInstanceUID'].head(150)\ncnts=[]\nfor st in sample_studies:\n    for se in glob.glob(f'{COMP}/train_series/{st}/*'):\n        n=len(glob.glob(se+'/*.dcm'))\n        if n: cnts.append(n)\ncnts=np.array(cnts)\nfig,ax=plt.subplots(figsize=(7,3))\nax.hist(np.minimum(cnts,120),bins=40,color='#264653')\nax.set_xlabel('slices per series'); ax.set_ylabel('# series')\nax.set_title(f'Slices per series (sample of {len(cnts)} series): median {int(np.median(cnts))}')\nplt.tight_layout(); plt.show()","id":"cell13"},{"cell_type":"markdown","metadata":{},"source":"## 7. DICOM preview — one study across planes","id":"cell14"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import subprocess, sys\ntry:\n    import gdcm  # noqa\nexcept Exception:\n    subprocess.run([sys.executable,'-m','pip','install','-q','python-gdcm','pylibjpeg','pylibjpeg-libjpeg','pylibjpeg-openjpeg'],check=False)\nimport pydicom\ndef load(p):\n    d=pydicom.dcmread(p); a=d.pixel_array.astype(float)\n    lo,hi=np.percentile(a,1),np.percentile(a,99); a=np.clip((a-lo)/(hi-lo+1e-6),0,1)\n    if getattr(d,'PhotometricInterpretation','')=='MONOCHROME1': a=1-a\n    return a\nst=train['StudyInstanceUID'].iloc[0]\nsers=series[series.StudyInstanceUID==st]\nfig,axes=plt.subplots(1,min(4,len(sers)),figsize=(13,3.5))\naxes=np.atleast_1d(axes)\nfor ax,(_,row) in zip(axes,sers.iterrows()):\n    fs=sorted(glob.glob(f'{COMP}/train_series/{st}/{row.SeriesInstanceUID}/*.dcm'))\n    if not fs: continue\n    try: ax.imshow(load(fs[len(fs)//2]),cmap='gray')\n    except Exception as e: ax.set_title('decode err'); continue\n    ax.set_title(f'{row.Anatomical_Plane}\\nFS={row.Fluid_Sensitive}'); ax.axis('off')\nplt.suptitle('Mid-slice per series (one study)'); plt.tight_layout(); plt.show()","id":"cell15"},{"cell_type":"markdown","metadata":{},"source":"## 8. A study-grouped, label-stratified 5-fold CV\nThe unit scored is the **study** (there is no patient id linking studies), so plain multi-label\nstratified K-fold over studies both matches the test structure and avoids leakage. Guarantee both\nclasses appear per label per fold, or a column's AUC is undefined.","id":"cell16"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"def mlskf(Y,k=5,seed=42):\n    rng=np.random.default_rng(seed); n,L=Y.shape; folds=-np.ones(n,int)\n    need=(Y.sum(0)[None,:]/k).repeat(k,0); cap=np.full(k,n/k)\n    for lab in np.argsort(Y.sum(0)):\n        pos=[i for i in np.where(Y[:,lab]==1)[0] if folds[i]<0]; rng.shuffle(pos)\n        for i in pos:\n            b=np.where(need[:,lab]==need[:,lab].max())[0]\n            f=int(b[np.argmax(cap[b])]); folds[i]=f; need[f,lab]-=1; cap[f]-=1\n    for i in np.where(folds<0)[0]:\n        f=int(np.argmax(cap)); folds[i]=f; cap[f]-=1\n    return folds\n# demonstrate on the labeled subset\nY=lab[LABELS].astype(int).values; folds=mlskf(Y,5)\nchk=pd.DataFrame({l:[int(Y[folds==f,j].sum()) for f in range(5)] for j,l in enumerate(LABELS)},\n                 index=[f'fold{f}' for f in range(5)]).T\nprint('positives per label per fold (both-class check):'); print(chk)","id":"cell17"},{"cell_type":"markdown","metadata":{},"source":"## Takeaways\n- **Labels, not model, decide this competition** — 98.7% of studies are report-only, so the biggest\n  lever is how well you turn multilingual reports into labels.\n- The metric is **rank-only macro-AUC** — ensemble by rank-averaging; don't sweat calibration.\n- Rare labels count as much as common ones — chase **AUC headroom**.\n- Route pathologies to the right **plane / sequence**; aggregate slice→study thoughtfully.\n\nGood luck! ⭐ **Upvote if this saved you time** — happy to answer questions below.","id":"cell18"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}